EDBT 2026 Demo / reviewers in the wild / expert
Zhibin Wang 0004
dblp:67/1237-4
· DBLP profile ↗
37ranked-venue papers
0as first author
35since 2021 · last 2026
0000-0001-7618-7973ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 21 since 2021Artificial intelligence and machine learning · 21 · 20 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Boosting Multi-Modal Large Language Model With Enhanced Visual FeaturesabstractRecent advancements in computer vision (CV) and large language models (LLMs) have spurred significant interest in multi-modal large language models (MLLMs), which aim to integrate visual and textual modalities for enhanced understanding and generation tasks. While much of the existing research focuses on optimizing projectors and LLMs to improve MLLM performance, a critical question remains underexplored: Has the full potential of visual features in MLLMs been realized? To address this question, we identify two key limitations in current MLLM architectures and propose vMLLM, a vision-enhanced MLLM designed to fully leverage the capabilities of visual features. vMLLM introduces two novel components: the Multi-level Aggregation Module (MAM) and the Intra- and inter-modal Enhancement Module (IEM). The MAM aggregates multi-layer features from the vision encoder, capturing both high-level semantic information and low-level spatial details, thereby enriching the visual representation. The IEM enhances visual features through intra- and inter-modal interactions, effectively suppressing irrelevant information while amplifying task-relevant features, leading to more robust multimodal understanding. We conduct extensive experiments on multiple benchmarks, evaluating vMLLM across diverse settings, including different vision encoders, training dataset scales, and varying sizes of LLMs. Our results demonstrate that vMLLM consistently achieves significant performance improvements, validating its effectiveness in harnessing the potential of visual features. These findings highlight the importance of optimizing visual feature extraction and interaction mechanisms in MLLMs, paving the way for more advanced multimodal AI systems.. Weihuang Lin, Zhibin Wang 0004, Jiayi Ji, Xiaoshuai Sun, Chia-Wen Lin, Rongrong Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | MVPaint: Synchronized Multi-View Diffusion for Painting Anything 3DabstractTexturing is a crucial step in the 3D asset production workflow, which enhances the visual appeal and diversity of 3D assets. Despite recent advancements in Text-to-Texture (T2T) generation, existing methods often yield subpar results, primarily due to local discontinuities, inconsistencies across multiple views, and heavy dependence on UV unwrapping outcomes. To tackle these challenges, we propose a novel generation-refinement 3D texturing framework called MVPaint, which can generate high-resolution, seamless textures while emphasizing multi-view consistency. MVPaint mainly consists of three key modules. 1) Synchronized Multi-view Generation (SMG). Given a 3D mesh model, MVPaint first simultaneously generates multi-view images by employing a SMG model, which leads to coarse texturing results with unpainted parts due to missing observations. 2) Spatial-aware 3D Inpainting (S3I). To ensure complete 3D texturing, we introduce the S3I method, specifically designed to texture previously unobserved areas effectively. 3) UV Refinement (UVR). Furthermore, MVPaint employs a UVR module to improve the texture quality in the UV space, which first performs a UV-space Super-Resolution, followed by a Spatial-aware Seam-Smoothing algorithm for revising spatial texturing discontinuities caused by UV unwrapping. Moreover, we establish two T2T evaluation benchmarks: the Objaverse T2T benchmark and the GSO T2T benchmark, based on selected high-quality 3D meshes from the Objaverse dataset and the entire GSO dataset, respectively. Extensive experimental results demonstrate that MVPaint surpasses existing state-of-the-art methods. Notably, MVPaint could generate high-fidelity textures with minimal Janus issues and highly enhanced cross-view consistency. Juncheng Mu, Xianfang Zeng, Xin Chen 0040, Anqi Pang, Chi Zhang 0080, Zhibin Wang 0004, Gang Yu 0002, Ziwei Liu 0002, Liang Pan |
CVPR | 7 |
| 2025 | CR2PQ: Continuous Relative Rotary Positional Query for Dense Visual Representation LearningabstractDense visual contrastive learning (DRL) shows promise for learning localized information in dense prediction tasks, but struggles with establishing pixel/patch correspondence across different views (cross-contrasting). Existing methods primarily rely on self-contrasting the same view with variations, limiting input variance and hindering downstream performance. This paper delves into the mechanisms of self-contrasting and cross-contrasting, identifying the crux of the issue: transforming discrete positional embeddings to continuous representations. To address the correspondence problem, we propose a Continuous Relative Rotary Positional Query ({\mname}), enabling patch-level representation learning. Our extensive experiments on standard datasets demonstrate state-of-the-art (SOTA) results. Compared to the previous SOTA method (PQCL), our approach achieves significant improvements on COCO: with 300 epochs of pretraining, {\mname} obtains \textbf{3.4\%} mAP$^{bb}$ and \textbf{2.1\%} mAP$^{mk}$ improvements for detection and segmentation tasks, respectively. Furthermore, {\mname} exhibits faster convergence, achieving \textbf{10.4\%} mAP$^{bb}$ and \textbf{7.9\%} mAP$^{mk}$ improvements over SOTA with just 40 epochs of pretraining. Shaofeng Zhang, Qiang Zhou 0001, Sitong Wu, Haoru Tan, Zhibin Wang 0004, Jinfa Huang, Junchi Yan |
ICLR | 5 |
| 2025 | A Multimodal LLM for Chart Understanding and GenerationabstractMulti-modal large language models have demonstrated impressive performances on most vision-language tasks. However, the model generally lacks the understanding capabilities for specific domain data, particularly when it comes to interpreting chart figures. This is mainly due to the lack of relevant multi-modal instruction tuning datasets. In this article, we create a high-quality instruction-tuning dataset leveraging GPT-4. We develop a multi-step data generation process in which different steps are responsible for generating tabular data, creating chart figures, and designing instruction tuning data separately. Our method’s flexibility enables us to generate diverse, high-quality instruction-tuning data consistently and efficiently while maintaining a low resource expenditure. Additionally, it allows us to incorporate a wider variety of chart and task types not yet featured in existing datasets. Next, we introduce ChartLlama, a multi-modal large language model that we’ve trained using our created dataset. ChartLlama outperforms all prior methods in ChartQA, Chart-to-text, and Chart-extraction evaluation benchmarks. Additionally, ChartLlama significantly improves upon the baseline in our specially compiled chart dataset, which includes new chart and task types. The results of ChartLlama confirm the value and huge potential of our proposed data generation method in enhancing chart comprehension. Yucheng Han, Chi Zhang 0007, Xin Chen 0040, Fukun Yin, Xu Yang 0021, Zhibin Wang 0004, Gang Yu 0002, Hanwang Zhang |
IJCNN | 6 |
| 2025 | EasyOutPainter: One Step Image Outpainting With Both Continuous Multiple and ResolutionabstractImage outpainting aims to generate the content of an input sub-image outside its boundaries, which remains open for existing generative models. This paper explores image outpainting in three directions that have not been achieved in literature to our knowledge: outpainting 1) with continuous multiples (in contrast to the discrete ones by existing methods); 2) with arbitrary resolutions; and 3) in a single step (for any multiples and resolutions). The arbitrary multiple outpainting is achieved by utilizing randomly cropped views from the same image during training to capture arbitrary relative positional information. Specifically, by feeding one view and relative positional embeddings as queries, we can reconstruct another view. At inference, we generate images with arbitrary expansion multiples by inputting an anchor image and its corresponding positional embeddings. The continuous-resolution outpainting is achieved by introducing the multi-scale training strategy into generative models. Specifically, by disentangling the image resolution and the number of patches, it can generate images with arbitrary resolutions without post-processing. Meanwhile, we propose a query-based contrastive objective to make our method not rely on a pre-trained backbone network which is otherwise often required in peer methods. The comprehensive experimental results on public benchmarks show its superior performance over state-of-the-art approaches. Shaofeng Zhang, Qiang Zhou 0001, Zhibin Wang 0004, Hao Li 0030, Junchi Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Few-Shot Semantic Segmentation on Remote Sensing Images With Learnable PrototypeabstractDeep learning-based semantic segmentation has been the dominant solution to quickly capture regions of interest (ROIs) in remote sensing images. However, the annotation and training cost of a fully-supervised segmentation model is often too high due to the requirement for elaborate masks. Additionally, trained models are limited to recognize only those classes defined in the training set. This has led to increased interest in how to cheaply adapt learned knowledge to new unseen objects. In this paper, we propose a meta-learning-based few-shot method called Learnable Prototype Few-Shot Segmentation (LPFS) to quickly adapt models to previously unseen geographic categories with only a few support examples of remote sensing images. Specifically, we first build a learnable prototype module based on variational auto-encoder (VAE) to eliminate inter-class ambiguity and extract high-level semantic prototypes from the support set effectively. We then design a global-attention correlation map to achieve low-level structural feature alignment between the support and query images. Additionally, we introduce a base learner to alleviate the bias caused by the meta-learning network on base classes. The extensive experiments on the public few-shot segmentation benchmark iSAID-5idemonstrate that our method sets a new strong baseline for few-shot semantic segmentation on remote sensing images. Jing Wang 0224, Yuang Liu, Qiang Zhou 0001, Zhibin Wang 0004, Fan Wang 0019 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Paint3D: Paint Anything 3D With Lighting-Less Texture Diffusion ModelsabstractThis paper presents Paint3D, a novel coarse-to-fine generative framework that is capable of producing high-resolution, lighting-less, and diverse 2K UV texture maps for untextured 3D meshes conditioned on text or image inputs. The key challenge addressed is generating high-quality textures without embedded illumination information, which allows the textures to be re-lighted or re-edited within modern graphics pipelines. To achieve this, our method first leverages a pre-trained depth-aware 2D diffusion model to generate view-conditional images and per-form multi-view texture fusion, producing an initial coarse texture map. However, as 2D models cannot fully repre-sent 3D shapes and disable lighting effects, the coarse texture map exhibits incomplete areas and illumination artifacts. To resolve this, we train separate UV Inpainting and UVHD diffusion models specialized for the shape-aware re-finement of incomplete areas and the removal of illumination artifacts. Through this coarse-to-fine process, Paint3D can produce high-quality 2K UV textures that maintain se-mantic consistency while being lighting-less, significantly advancing the state-of-the-art in texturing 3D objects. Xianfang Zeng, Xin Chen 0040, Zhongqi Qi, Wen Liu 0003, Zibo Zhao 0001, Zhibin Wang 0004, Yong Liu 0007, Gang Yu 0002 |
CVPR | 6 |
| 2024 | Continuous-Multiple Image Outpainting in One-Step via Positional Query and A Diffusion-based ApproachabstractImage outpainting aims to generate the content of an input sub-image beyond its original boundaries. It is an important task in content generation yet remains an open problem for generative models. This paper pushes the technical frontier of image outpainting in two directions that have not been resolved in literature: 1) outpainting with arbitrary and continuous multiples (without restriction), and 2) outpainting in a single step (even for large expansion multiples). Moreover, we develop a method that does not depend on a pre-trained backbone network, which is in contrast commonly required by the previous SOTA outpainting methods. The arbitrary multiple outpainting is achieved by utilizing randomly cropped views from the same image during training to capture arbitrary relative positional information. Specifically, by feeding one view and positional embeddings as queries, we can reconstruct another view. At inference, we generate images with arbitrary expansion multiples by inputting an anchor image and its corresponding positional embeddings. The one-step outpainting ability here is particularly noteworthy in contrast to previous methods that need to be performed for $N$ times to obtain a final multiple which is $N$ times of its basic and fixed multiple. We evaluate the proposed approach (called PQDiff as we adopt a diffusion-based generator as our embodiment, under our proposed \textbf{P}ositional \textbf{Q}uery scheme) on public benchmarks, demonstrating its superior performance over state-of-the-art approaches. Specifically, PQDiff achieves state-of-the-art FID scores on the Scenery (\textbf{21.512}), Building Facades (\textbf{25.310}), and WikiArts (\textbf{36.212}) datasets. Furthermore, under the 2.25x, 5x and 11.7x outpainting settings, PQDiff only takes \textbf{40.6\%}, \textbf{20.3\%} and \textbf{10.2\%} of the time of the benchmark state-of-the-art (SOTA) method. Shaofeng Zhang, Jinfa Huang, Qiang Zhou 0001, Zhibin Wang 0004, Fan Wang 0019, Jiebo Luo 0001, Junchi Yan |
ICLR | 4 |
| 2024 | Multimodal LLM Enhanced Cross-lingual Cross-modal RetrievalabstractCross-lingual cross-modal retrieval (CCR) aims to retrieve visually relevant content based on non-English queries, without relying on human-labeled cross-modal data pairs during training. One popular approach involves utilizing machine translation (MT) to create pseudo-parallel data pairs, establishing correspondence between visual and non-English textual data. However, aligning their representations poses challenges due to the significant semantic gap between vision and text, as well as the lower quality of non-English representations caused by pre-trained encoders and data noise. To overcome these challenges, we propose LECCR, a novel solution that incorporates the multi-modal large language model (MLLM) to improve the alignment between visual and non-English representations. Specifically, we first employ MLLM to generate detailed visual content descriptions and aggregate them into multi-view semantic slots that encapsulate different semantics. Then, we take these semantic slots as internal features and leverage them to interact with the visual features. By doing so, we enhance the semantic information within the visual features, narrowing the semantic gap between modalities and generating local visual semantics for subsequent multi-level matching. Additionally, to further enhance the alignment between visual and non-English features, we introduce softened matching under English guidance. This approach provides more comprehensive and reliable inter-modal correspondences between visual and non-English features. Extensive experiments on four CCR benchmarks, i.e., Multi30K, MSCOCO, VATEX, and MSR-VTT-CN, demonstrate the effectiveness of our proposed method. Code: https://github.com/LiJiaBei-7/leccr. Le Wang 0003, Qiang Zhou 0001, Zhibin Wang 0004, Hao Li 0030, Gang Hua 0001, Wei Tang 0016 |
ACM Multimedia | 4 |
| 2024 | I2EBench: A Comprehensive Benchmark for Instruction-based Image EditingabstractSignificant progress has been made in the field of Instruction-based Image Editing (IIE). However, evaluating these models poses a significant challenge. A crucial requirement in this field is the establishment of a comprehensive evaluation benchmark for accurately assessing editing results and providing valuable insights for its further development. In response to this need, we propose I2EBench, a comprehensive benchmark designed to automatically evaluate the quality of edited images produced by IIE models from multiple dimensions. I2EBench consists of 2,000+ images for editing, along with 4,000+ corresponding original and diverse instructions. It offers three distinctive characteristics: 1) Comprehensive Evaluation Dimensions: I2EBench comprises 16 evaluation dimensions that cover both high-level and low-level aspects, providing a comprehensive assessment of each IIE model. 2) Human Perception Alignment: To ensure the alignment of our benchmark with human perception, we conducted an extensive user study for each evaluation dimension. 3) Valuable Research Insights: By analyzing the advantages and disadvantages of existing IIE models across the 16 dimensions, we offer valuable research insights to guide future development in the field. We will open-source I2EBench, including all instructions, input images, human annotations, edited images from all evaluated methods, and a simple script for evaluating the results from new IIE models. The code, dataset, and generated images from all IIE models are provided in GitHub: https://github.com/cocoshe/I2EBench. Jiayi Ji, Ke Ye, Weihuang Lin, Zhibin Wang 0004, Yonghan Zheng, Qiang Zhou 0001, Xiaoshuai Sun, Rongrong Ji |
NeurIPS | 5 |
| 2024 | Dynamic Token-Pass Transformers for Semantic SegmentationabstractVision transformers (ViT) usually extract features via forwarding all the tokens in the self-attention layers from top to toe. In this paper, we introduce dynamic token-pass vision transformers (DoViT) for semantic segmentation, which can adaptively reduce the inference cost for images with different complexity. DoViT gradually stops partial easy tokens from self-attention calculation and keeps the hard tokens forwarding until meeting the stopping criteria. We employ lightweight auxiliary heads to make the token-pass decision and divide the tokens into keeping/stopping parts. With a token separate calculation, the self-attention layers are speeded up with sparse tokens and still work friendly with hardware. A token reconstruction module is built to collect and reset the grouped tokens to their original position in the sequence, which is necessary to predict correct semantic masks. We conduct extensive experiments on two common semantic segmentation tasks, and demonstrate that our method greatly reduces about 40% ∼ 60% FLOPs and the drop of mIoU is within 0.8% for various segmentation transformers. The throughput and inference speed of ViT-L/B are increased to more than 2× on Cityscapes. Code is available at https://github.com/FLHonker/DoViT-code. Yuang Liu, Qiang Zhou 0001, Jing Wang 0224, Zhibin Wang 0004, Fan Wang 0019, Jun Wang 0006, Wei Zhang 0056 |
WACV | 4 |
| 2024 | PolyRoad: Polyline Transformer for Topological Road-Boundary DetectionabstractTopological road-boundary detection using remote sensing imagery plays a critical role in creating high-definition (HD) maps and enabling autonomous driving. Previous approaches follow an iterative graph-growing paradigm for road-boundary extraction, where road boundaries are predicted vertex by vertex and instance by instance to output a graph, resulting in limitations of low inference speed. In this work, we formulate the road boundaries as polylines instead of a graph and propose a novel polyline transformer for topological road-boundary detection, termed PolyRoad. PolyRoad is built on the transformer architecture and is capable of detecting all road boundaries in parallel, which greatly improves the training and inference speed compared with the graph-based methods. To perform bipartite matching between the ground truth and predicted polylines, we develop a polyline matching cost to measure the distance, considering the order of open and closed polylines. In addition, we propose three different losses for supervising polyline learning: the order-oriented$L1$loss, direction loss, and mask loss. The order-oriented$L1$loss provides the point-level supervision to constrain the absolute position of each point of the road-boundary polylines. The direction loss provides the direction-level supervision to constrain the geometry shape of the predicted polylines by supervising the relative position of adjacent points. The mask loss provides the pixel-level supervision of the predicted polylines by converting the vector-format polylines into raster-format binary masks. Comprehensive experiments are conducted on the Topo-boundary dataset. Quantitative and qualitative results show that PolyRoad achieves superior performance than prior methods in both pixel-level and geometry-level metrics. More notably, PolyRoad achieves$3.37 \times $and$22.85 \times $faster inference speeds than Enhanced-iCurb and VecRoad, respectively. Zhibin Wang 0004, Zhou Huang 0002, Yu Liu 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Inter-Class and Inter-Domain Semantic Augmentation for Domain GeneralizationabstractThe domain generalization approach seeks to develop a universal model that performs well on unknown target domains with the aid of diverse source domains. Data augmentation has proven to be an effective method to enhance domain generalization in computer vision. Recently, semantic-level based data augmentation has yielded remarkable results. However, these methods focus on sampling semantic directions on feature space from intra-class and intra-domain, limiting the diversity of the source domain. To address this issue, we propose a novel approach called Inter-Class and Inter-Domain Semantic Augmentation (CDSA) for domain generalization. We first introduce a sampling-based method called CrossSmooth to obtain semantic directions from inter-class. Then, CrossVariance obtains the styles of different domains by sampling semantic directions. Our experiments on four well-known domain generalization benchmark datasets (Digits-DG, PACS, Office-Home, and DomainNet) demonstrate the effectiveness of our approach. We also validate our approach on commonly-used semantic segmentation datasets, namely GTAV, SYNTHIA, Cityscapes, Mapillary, and BDDS which also show significant improvements. Mengzhu Wang, Yuehua Liu, Jianlong Yuan, Shanshan Wang 0008, Zhibin Wang 0004, Wei Wang 0335 |
IEEE Trans. Image Process. | 5 |
| 2023 | SwinRDM: Integrate SwinRNN with Diffusion Model towards High-Resolution and High-Quality Weather ForecastingabstractData-driven medium-range weather forecasting has attracted much attention in recent years. However, the forecasting accuracy at high resolution is unsatisfactory currently. Pursuing high-resolution and high-quality weather forecasting, we develop a data-driven model SwinRDM which integrates an improved version of SwinRNN with a diffusion model. SwinRDM performs predictions at 0.25-degree resolution and achieves superior forecasting accuracy to IFS (Integrated Forecast System), the state-of-the-art operational NWP model, on representative atmospheric variables including 500 hPa geopotential (Z500), 850 hPa temperature (T850), 2-m temperature (T2M), and total precipitation (TP), at lead times of up to 5 days. We propose to leverage a two-step strategy to achieve high-resolution predictions at 0.25-degree considering the trade-off between computation memory and forecasting accuracy. Recurrent predictions for future atmospheric fields are firstly performed at 1.40625-degree resolution, and then a diffusion-based super-resolution model is leveraged to recover the high spatial resolution and finer-scale atmospheric details. SwinRDM pushes forward the performance and potential of data-driven models for a large margin towards operational applications. Zhibin Wang 0004, Fan Wang 0019 |
AAAI | 4 |
| 2023 | Point-Teaching: Weakly Semi-supervised Object Detection with Point AnnotationsabstractPoint annotations are considerably more time-efficient than bounding box annotations. However, how to use cheap point annotations to boost the performance of semi-supervised object detection is still an open question. In this work, we present Point-Teaching, a weakly- and semi-supervised object detection framework to fully utilize the point annotations. Specifically, we propose a Hungarian-based point-matching method to generate pseudo labels for point-annotated images. We further propose multiple instance learning (MIL) approaches at the level of images and points to supervise the object detector with point annotations. Finally, we propose a simple data augmentation, named Point-Guided Copy-Paste, to reduce the impact of those unmatched points. Experiments demonstrate the effectiveness of our method on a few datasets and various data regimes. In particular, Point-Teaching outperforms the previous best method Group R-CNN by 3.1 AP with 5% fully labeled data and 2.3 AP with 30% fully labeled data on the MS COCO dataset. We believe that our proposed framework can largely lower the bar of learning accurate object detectors and pave the way for its broader applications. The code is available at https://github.com/YongtaoGe/Point-Teaching. Yongtao Ge, Qiang Zhou 0001, Chunhua Shen, Zhibin Wang 0004, Hao Li 0030 |
AAAI | 5 |
| 2023 | Frequency Domain Disentanglement for Arbitrary Neural Style TransferabstractArbitrary neural style transfer has been a popular research topic due to its rich application scenarios. Effective disentanglement of content and style is the critical factor for synthesizing an image with arbitrary style. The existing methods focus on disentangling feature representations of content and style in the spatial domain where the content and style components are innately entangled and difficult to be disentangled clearly. Therefore, these methods always suffer from low-quality results because of the sub-optimal disentanglement. To address such a challenge, this paper proposes the frequency mixer (FreMixer) module that disentangles and re-entangles the frequency spectrum of content and style components in the frequency domain. Since content and style components have different frequency-domain characteristics (frequency bands and frequency patterns), the FreMixer could well disentangle these two components. Based on the FreMixer module, we design a novel Frequency Domain Disentanglement (FDD) framework for arbitrary neural style transfer. Qualitative and quantitative experiments verify that the proposed method can render better stylized results compared to the state-of-the-art methods. Hao Luo 0004, Pichao Wang, Zhibin Wang 0004, Shang Liu 0002, Fan Wang 0019 |
AAAI | 4 |
| 2023 | Efficient Mask Correction for Click-Based Interactive Image SegmentationabstractThe goal of click-based interactive image segmentation is to extract target masks with the input of positive/negative clicks. Every time a new click is placed, existing methods run the whole segmentation network to obtain a corrected mask, which is inefficient since several clicks may be needed to reach satisfactory accuracy. To this end, we propose an efficient method to correct the mask with a lightweight mask correction network. The whole network remains a low computational cost from the second click, even if we have a large backbone. However, a simple correction network with limited capacity is not likely to achieve comparable performance with a classic segmentation network. Thus, we propose a click-guided self-attention module and a click-guided correlation module to effectively exploits the click information to boost performance. First, several tem-plates are selected based on the semantic similarity with click features. Then the self-attention module propagates the template information to other pixels, while the correlation module directly uses the templates to obtain target out-lines. With the efficient architecture and two click-guided modules, our method shows preferable performance and efficiency compared to existing methods. The code will be released at https://github.com/feiaxyt/EMC-Click. Jianlong Yuan, Zhibin Wang 0004, Fan Wang 0019 |
CVPR | 3 |
| 2023 | Foundation Model Drives Weakly Incremental Learning for Semantic SegmentationabstractModern incremental learning for semantic segmentation methods usually learn new categories based on dense annotations. Although achieve promising results, pixel-by-pixel labeling is costly and time-consuming. Weakly incremental learning for semantic segmentation (WILSS) is a novel and attractive task, which aims at learning to segment new classes from cheap and widely available image-level labels. Despite the comparable results, the image-level labels can not provide details to locate each segment, which limits the performance of WILSS. This inspires us to think how to improve and effectively utilize the supervision of new classes given image-level labels while avoiding forgetting old ones. In this work, we propose a novel and data-efficient frame-work for WILSS, named FMWISS. Specifically, we propose pre-training based co-segmentation to distill the knowledge of complementary foundation models for generating dense pseudo labels. We further optimize the noisy pseudo masks with a teacher-student architecture, where a plug-in teacher is optimized with a proposed dense contrastive loss. Moreover, we introduce memory-based copy-paste augmentation to improve the catastrophic forgetting problem of old classes. Extensive experiments on Pascal VOC and COCO datasets demonstrate the superior performance of our framework, e.g., FMWISS achieves 70.7% and 73.3% in the 15–5 VOC setting, outperforming the state-of-the-art method by 3.4% and 6.1%, respectively. Chaohui Yu, Qiang Zhou 0001, Jingliang Li, Jianlong Yuan, Zhibin Wang 0004, Fan Wang 0019 |
CVPR | 5 |
| 2023 | D2Q-DETR: Decoupling and Dynamic Queries for Oriented Object Detection with TransformersabstractDespite the promising results, existing oriented object detection methods usually involve heuristically designed rules, e.g., RRoI generation, rotated NMS. In this paper, we propose an end-to-end framework for oriented object detection, which simplifies the model pipeline and obtains superior performance. Our framework is based on DETR, with the box regression head replaced with a points prediction head. The learning of points is more flexible, and the distribution of points can reflect the angle and size of the target rotated box. We further propose to decouple the query features into classification and regression features, which significantly improves the model precision. Aerial images usually contain thousands of instances. To better balance model precision and efficiency, we propose a novel dynamic query design, which reduces the number of object queries in stacked decoder layers without sacrificing model performance. Finally, we rethink the label assignment strategy of existing DETR-like detectors and propose an effective label re-assignment strategy for improved performance. We name our method D2Q-DETR. Experiments on the largest and challenging DOTA-v1.0 and DOTA-v1.5 datasets show that D2Q-DETR outperforms existing NMS-based and NMS-free oriented object detection methods and achieves the new state-of-the-art. Qiang Zhou 0001, Chaohui Yu, Zhibin Wang 0004, Fan Wang 0019 |
ICASSP | 3 |
| 2023 | Robust Geometry-Preserving Depth Estimation Using Differentiable RenderingabstractIn this study, we address the challenge of 3D scene structure recovery from monocular depth estimation. While traditional depth estimation methods leverage labeled datasets to directly predict absolute depth, recent advancements advocate for mix-dataset training, enhancing generalization across diverse scenes. However, such mixed dataset training yields depth predictions only up to an unknown scale and shift, hindering accurate 3D reconstructions. Existing solutions necessitate extra 3D datasets or geometry-complete depth annotations, constraints that limit their versatility. In this paper, we propose a learning framework that trains models to predict geometry-preserving depth without requiring extra data or annotations. To produce realistic 3D structures, we render novel views of the reconstructed scenes and design loss functions to promote depth estimation consistency across different views. Comprehensive experiments underscore our framework’s superior generalization capabilities, surpassing existing state-of-the-art methods on several benchmark datasets without leveraging extra training information. Moreover, our innovative loss functions empower the model to autonomously recover domain-specific scale-and-shift coefficients using solely unlabeled images. Chi Zhang 0007, Wei Yin 0006, Gang Yu 0002, Zhibin Wang 0004, Tao Chen 0003, Joey Tianyi Zhou, Chunhua Shen |
ICCV | 4 |
| 2023 | LMSeg: Language-guided Multi-dataset Segmentation
Qiang Zhou 0001, Yuang Liu, Chaohui Yu, Jingliang Li, Zhibin Wang 0004, Fan Wang 0019 |
ICLR | 5 |
| 2023 | Patch-level Contrastive Learning via Positional Query for Visual Pre-trainingabstractDense contrastive learning (DCL) has been recently explored for learning localized information for dense prediction tasks (e.g., detection and segmentation). It still suffers the difficulty of mining pixels/patches correspondence between two views. A simple way is inputting the same view twice and aligning the pixel/patch representation. However, it would reduce the variance of inputs, and hurts the performance. We propose a plug-in method PQCL (Positional Query for patch-level Contrastive Learning), which allows performing patch-level contrasts between two views with exact patch correspondence. Besides, by using positional queries, PQCL increases the variance of inputs, to enhance training. We apply PQCL to popular transformer-based CL frameworks (DINO and iBOT, and evaluate them on classification, detection and segmentation tasks, where our method obtains stable improvements, especially for dense tasks. It achieves new state-of-the-art in most settings. Code is available at https://github.com/Sherrylone/Query_Contrastive. Shaofeng Zhang, Qiang Zhou 0001, Zhibin Wang 0004, Fan Wang 0019, Junchi Yan |
ICML | 3 |
| 2023 | UniNeXt: Exploring A Unified Architecture for Vision RecognitionabstractVision Transformers have shown great potential in computer vision tasks. Most recent works have focused on elaborating the spatial token mixer for performance gains. However, we observe that a well-designed general architecture can significantly improve the performance of the entire backbone, regardless of which spatial token mixer is equipped. In this paper, we propose UniNeXt, an improved general architecture for the vision backbone. To verify its effectiveness, we instantiate the spatial token mixer with various typical and modern designs, including both convolution and attention modules. Compared with the architecture in which they are first proposed, our UniNeXt architecture can steadily boost the performance of all the spatial token mixers, and narrows the performance gap among them. Surprisingly, our UniNeXt equipped with naive local window attention even outperforms the previous state-of-the-art. Interestingly, the ranking of these spatial token mixers also changes under our UniNeXt, suggesting that an excellent spatial token mixer may be stifled due to a suboptimal general architecture, which further shows the importance of the study on the general architecture of vision backbone. Code is available at UniNeXt. Fangjian Lin, Jianlong Yuan, Sitong Wu, Fan Wang 0019, Zhibin Wang 0004 |
ACM Multimedia | 5 |
| 2023 | Mixture-of-Experts Learner for Single Long-Tailed Domain GeneralizationabstractDomain generalization (DG) refers to the task of training a model on multiple source domains and test it on a different target domain with different distribution. In this paper, we address a more challenging and realistic scenario known as Single Long-Tailed Domain Generalization, where only one source domain is available and the minority class in this domain has an abundance of instances in other domains. To tackle this task, we propose a novel approach called Mixture-of-Experts Learner for Single Long-Tailed Domain Generalization (MoEL), which comprises two key strategies. The first strategy is a simple yet effective data augmentation technique that leverages saliency maps to identify important regions on the original images and preserves these regions during augmentation. The second strategy is a new skill-diverse expert learning approach that trains multiple experts from a single long-tailed source domain and leverages mutual learning to aggregate their learned knowledge for the unknown target domain. We evaluate our method on various benchmark datasets, including Digits-DG, CIFAR-10-C, PACS, and DomainNet, and demonstrate its superior performance compared to previous single domain generalization methods. Additionally, the ablation study is also conducted to illustrate the inner workings of our approach. Mengzhu Wang, Jianlong Yuan, Zhibin Wang 0004 |
ACM Multimedia | 3 |
| 2023 | Points-to-3D: Bridging the Gap between Sparse Points and Shape-Controllable Text-to-3D GenerationabstractText-to-3D generation has recently garnered significant attention, fueled by 2D diffusion models trained on billions of image-text pairs. Existing methods primarily rely on score distillation to leverage the 2D diffusion priors to supervise the generation of 3D models, e.g., NeRF. However, score distillation is prone to suffer the view inconsistency problem, and implicit NeRF modeling can also lead to an arbitrary shape, thus leading to less realistic and uncontrollable 3D generation. In this work, we propose a flexible framework of Points-to-3D to bridge the gap between sparse yet freely available 3D points and realistic shape-controllable 3D generation by distilling the knowledge from both 2D and 3D diffusion models. The core idea of Points-to-3D is to introduce controllable sparse 3D points to guide the text-to-3D generation. Specifically, we use the sparse point cloud generated from the 3D diffusion model, Point-E, as the geometric prior, conditioned on a single reference image. To better utilize the sparse 3D points, we propose an efficient point cloud guidance loss to adaptively drive the NeRF's geometry to align with the shape of the sparse 3D points. In addition to controlling the geometry, we propose to optimize the NeRF for a more view-consistent appearance. To be specific, we perform score distillation to the publicly available 2D image diffusion model ControlNet, conditioned on text as well as depth map of the learned compact geometry. Qualitative and quantitative comparisons demonstrate that Points-to-3D improves view consistency and achieves good shape controllability for text-to-3D generation. Points-to-3D provides users with a new way to improve and control text-to-3D generation. Chaohui Yu, Qiang Zhou 0001, Jingliang Li, Zhe Zhang 0049, Zhibin Wang 0004, Fan Wang 0019 |
ACM Multimedia | 5 |
| 2023 | Semi-supervised Semantic Segmentation with Mutual Knowledge DistillationabstractConsistency regularization has been widely studied in recent semi- supervised semantic segmentation methods, and promising per- formance has been achieved. In this work, we propose a new con- sistency regularization framework, termed mutual knowledge dis- tillation (MKD), combined with data and feature augmentation. We introduce two auxiliary mean-teacher models based on consis- tency regularization. More specifically, we use the pseudo-labels generated by a mean teacher to supervise the student network to achieve a mutual knowledge distillation between the two branches. In addition to using image-level strong and weak augmentation, we also discuss feature augmentation. This involves considering various sources of knowledge to distill the student network. Thus, we can significantly increase the diversity of the training samples. Experiments on public benchmarks show that our framework out- performs previous state-of-the-art (SOTA) methods under various semi-supervised settings. Code is available at https://github.com/jianlong-yuan/semi-mmseg. Jianlong Yuan, Jinchao Ge, Zhibin Wang 0004, Yifan Liu 0001 |
ACM Multimedia | 3 |
| 2023 | Data Pruning via Moving-one-Sample-outabstractIn this paper, we propose a novel data-pruning approach called moving-one-sample-out (MoSo), which aims to identify and remove the least informative samples from the training set. The core insight behind MoSo is to determine the importance of each sample by assessing its impact on the optimal empirical risk. This is achieved by measuring the extent to which the empirical risk changes when a particular sample is excluded from the training set. Instead of using the computationally expensive leaving-one-out-retraining procedure, we propose an efficient first-order approximator that only requires gradient information from different training stages. The key idea behind our approximation is that samples with gradients that are consistently aligned with the average gradient of the training set are more informative and should receive higher scores, which could be intuitively understood as follows: if the gradient from a specific sample is consistent with the average gradient vector, it implies that optimizing the network using the sample will yield a similar effect on all remaining samples.
Experimental results demonstrate that MoSo effectively mitigates severe performance degradation at high pruning ratios and achieves satisfactory performance across various settings. Experimental results demonstrate that MoSo effectively mitigates severe performance degradation at high pruning ratios and outperforms state-of-the-art methods by a large margin across various settings. Haoru Tan, Sitong Wu, Yukang Chen, Zhibin Wang 0004, Fan Wang 0019, Xiaojuan Qi 0001 |
NeurIPS | 5 |
| 2023 | An M-Nary SAR Image Change Detection Based on GAN Architecture SearchabstractChange detection (CD) in synthetic aperture radar (SAR) images aims to detect changed areas by considering the changes in backscattering coefficients. However, the changes can be further divided into positive and negative changes in terms of the increase or decrease of backscattering coefficient, so the CD task can be divided into binary and ternary according to the number of existent categories. This paper introduces an M-nary (binary or ternary) SAR change detection procedure based on the generative adversarial network (GAN) and neural architecture search (NAS) strategy to detect which changes exist in the SAR image-pair and design specialized classifiers for both binary and ternary CD. First, a difference image generation approach based on the salient changed region extraction and neighborhood information is designed for a robust difference representation on the M-nary CD. Due to the further subdivision of changes, the insufficiency of labeled data presents the M-nary change detection with a dilemma. Concerning the lack of labeled information, this paper presents a labeled sample generation strategy based on the GAN architecture search to supplement sample data. Since GAN training is inherently unstable, NAS provides an effective means of searching GAN architecture automatically and ameliorates the reliability of generated samples. During the architecture search procedure, a double-phase evolutionary search strategy is introduced to further improve the stability of GAN training. The experimental results with theoretical analysis prove the validity, robustness, and potential of our method in synthetic as well as real SAR datasets. Maoguo Gong, Tianqi Gao, Mingyang Zhang 0002, Wei Li 0032, Zhibin Wang 0004, Dezhong Li |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Land Use and Land Cover Mapping in China Using Multimodal Fine-Grained Dual NetworkabstractWith the advancement of geo-systems and the increased availability of satellite data, a plethora of Land-Use and Land-Cover (LULC) products have been developed. The existing LULC products primarily relied on time-series imagery to classify land by pixel-based classifiers, allowing for local analysis and accurate boundary detection. However, the advent of deep learning has shifted towards the use of patch-based CNN models for generating land cover maps. In this paper, (1) we create a training dataset for China using a voting strategy based on three off-the-shelf available LULC products, avoiding the labor-intensive manual annotation. (2) We design a novel CNN-based model for LULC task, called Multi-modal Fine-grained Dual Network (dubbed as Dual-Net), which takes dual-date images to generate final maps, and reduces the need for gap-free temporal sequences or separate cloud detection. To leverage the correlation between location, date, and category, we embed multi-modal information (dates and geo-locations) to the model. Further, by incorporating low-level constraints and using pseudo-label refinement, we continually improve the performance and achieve more refined segmentation. (3) Due to the lack of a suitable validation dataset for China, we create a new validation dataset called China Sentinel2 Validation Dataset (CSVD) by manually annotating 733 finely labeled images of 1024 × 1024 pixels of China-specific Sentinel2 data. (4) Extensive experiments demonstrate that our model outperforms existing LULC products and produces more fine-grained segmentation results comparable to other patch-based products. Finally, we release annual LULC maps for China in 2020-2022 and also make our model accessible online for real-time results export. Shang Liu 0002, Yixuan Zhu, Zhibin Wang 0004, Mingyang Yang, Fan Wang 0019 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Poseur: Direct Human Pose Regression with Transformers
Weian Mao, Yongtao Ge, Chunhua Shen, Zhi Tian, Zhibin Wang 0004, Anton van den Hengel |
ECCV (6) | 6 |
| 2022 | Semantic Data Augmentation based Distance Metric Learning for Domain GeneralizationabstractDomain generalization (DG) aims to learn a model on one or more different but related source domains that could be generalized into an unseen target domain. Existing DG methods try to prompt the diversity of source domains for the model's generalization ability, while they may have to introduce auxiliary networks or striking computational costs. On the contrary, this work applies the implicit semantic augmentation in feature space to capture the diversity of source domains. Concretely, an additional loss function of distance metric learning (DML) is included to optimize the local geometry of data distribution. Besides, the logits from cross entropy loss with infinite augmentations is adopted as input features for the DML loss in lieu of the deep features. We also provide a theoretical analysis to show that the logits can approximate the distances defined on original features well. Further, we provide an in-depth analysis of the mechanism and rational behind our approach, which gives us a better understanding of why leverage logits in lieu of features can help domain generalization. The proposed DML loss with the implicit augmentation is incorporated into a recent DG method, that is, Fourier Augmented Co-Teacher framework (FACT). Meanwhile, our method also can be easily plugged into various DG methods. Extensive experiments on three benchmarks (Digits-DG, PACS and Office-Home) have demonstrated that the proposed method is able to achieve the state-of-the-art performance. Mengzhu Wang, Jianlong Yuan, Qi Qian 0001, Zhibin Wang 0004, Hao Li 0030 |
ACM Multimedia | 4 |
| 2022 | MimCo: Masked Image Modeling Pre-training with Contrastive TeacherabstractRecent masked image modeling (MIM) has received much attention in self-supervised learning (SSL), which requires the target model to recover the masked part of the input image. Although MIM-based pre-training methods achieve new state-of-the-art performance when transferred to many downstream tasks, the visualizations show that the learned representations are less separable, especially compared to those based on contrastive learning pre-training. This inspires us to think whether the linear separability of MIM pre-trained representation can be further improved, thereby improving the pre-training performance. Since MIM and contrastive learning tend to utilize different data augmentations and training strategies, combining these two pretext tasks is not trivial. In this work, we propose a novel and flexible pre-training framework, named MimCo, which combines MIM and contrastive learning through two-stage pre-training. Specifically, MimCo takes a pre-trained contrastive learning model as the teacher model and is pre-trained with two types of learning targets: patch-level and image-level reconstruction losses. Qiang Zhou 0001, Chaohui Yu, Hao Luo 0004, Zhibin Wang 0004, Hao Li 0030 |
ACM Multimedia | 4 |
| 2022 | Contrastive Haze-Aware Learning for Dynamic Remote Sensing Image DehazingabstractImage dehazing methods aim to recover a clear image from its hazy counterpart. While various dehazing methods have been proposed, their performance on real-world remote sensing (RS) images remains unsatisfying. A key reason is that the complex weather and imaging conditions (e.g., large fields of view) cause the haze condition to dramatically change in different images, while most existing methods fail to flexibly adapt their dehazing model to the specific haze condition in each image. To mitigate this problem, we present a contrastive haze-aware learning based dynamic dehazing method which demonstrates two aspects of advantage. On one hand, a contrastive clustering scheme is utilized to learn the image-wise haze representation using a set of real-world hazy images in an unsupervised manner, which enables identifying and discriminating the specific haze condition in each given hazy image. On the other hand, with the learned haze representation, a parameter generator can produce haze-aware parameters to dynamically construct a dehazing model for the given hazy image, which empowers us to adaptively dehaze the image based on its specific haze condition and thus improves the generalization ability. In addition, a new contrastive loss defined based on the learned haze representation is further utilized for model training and leads to better performance. To demonstrate the effectiveness of the proposed method, we evaluate it on two benchmark RS image datasets including various real-world hazy images, and observe obviously superiority over other state-of-the-art competitors. Jiangtao Nie, Wei Wei 0008, Lei Zhang 0054, Jianlong Yuan, Zhibin Wang 0004, Hao Li 0030 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Instant-Teaching: An End-to-End Semi-Supervised Object Detection FrameworkabstractSupervised learning based object detection frameworks demand plenty of laborious manual annotations, which may not be practical in real applications. Semi-supervised object detection (SSOD) can effectively leverage unlabeled data to improve the model performance, which is of great significance for the application of object detection models. In this paper, we revisit SSOD and propose Instant-Teaching, a completely end-to-end and effective SSOD framework, which uses instant pseudo labeling with extended weak-strong data augmentations for teaching during each training iteration. To alleviate the confirmation bias problem and improve the quality of pseudo annotations, we further propose a co-rectify scheme based on Instant-Teaching, denoted as Instant-Teaching∗. Extensive experiments on both MS-COCO and PASCAL VOC datasets substantiate the superiority of our framework. Specifically, our method surpasses state-of-the-art methods by 4.2 mAP on MS-COCO when using 2% labeled data. Even with full supervised information of MS-COCO, the proposed method still outperforms state-of-the-art methods by about 1.0 mAP. On PASCAL VOC, we can achieve more than 5 mAP improvement by applying VOC07 as labeled data and VOC12 as unlabeled data. Qiang Zhou 0001, Chaohui Yu, Zhibin Wang 0004, Qi Qian 0001, Hao Li 0030 |
CVPR | 3 |
| 2021 | A Simple Baseline for Semi-supervised Semantic Segmentation with Strong Data Augmentation*abstractRecently, significant progress has been made on semantic segmentation. However, the success of supervised semantic segmentation typically relies on a large amount of labeled data, which is time-consuming and costly to obtain. Inspired by the success of semi-supervised learning methods for image classification, here we propose a simple yet effective semi-supervised learning framework for semantic segmentation. We demonstrate that the devil is in the details: a set of simple designs and training techniques can collectively improve the performance of semi-supervised semantic segmentation significantly. Previous works [3], [25] fail to effectively employ strong augmentation in pseudo-label learning, as the large distribution disparity caused by strong augmentation harms the batch nor-malization statistics. We design a new batch normalization, namely distribution-specific batch normalization (DSBN) to address this problem and show the importance of strong augmentation for semantic segmentation. Moreover, we design a self-correction loss, which is effective in terms of noise resistance. We conduct a series of ablation studies to show the effectiveness of each component. Our method achieves state-of-the-art results in the semi-supervised settings on the Cityscapes and Pascal VOC datasets. Jianlong Yuan, Yifan Liu 0001, Chunhua Shen, Zhibin Wang 0004, Hao Li 0030 |
ICCV | 4 |
| 2020 | Unsupervised Style Transfer via Dualgan for Cross-Domain Aerial Image ClassificationabstractDue to its wide applications, aerial image classification, which is also called semantic segmentation of aerial imagery, attracts increasing research interest in recent years. Until now, deep semantic segmentation network (DSSN) has been widely adopted to address aerial image classification and achieves tremendous success. However, the superior performance of DSSN highly depends on massive targeted data with labels. When DSSN is trained on data from the source domain but tested on data from the target domain, the performance of DSSN is often very limited due to the data shift between source and target domains. To alleviate the disadvantage influence of data shift, this paper proposes a domain adaptation approach via unsupervised style transfer to cope with cross-domain aerial image classification. More specifically, this paper innovatively recommends DualGAN to conduct unsupervised style transfer for mapping aerial images in the source domain to the target domain. The mapped aerial imagery with labels is adopted to train DSSN, which is further used to classify aerial imagery in the target domain. To verify the validity of the presented approach, we give two cross-domain experimental settings including: (I) variation of geographic location; (II) variation of both geographic location and imaging mode. Extensive experiments under two typical cross-domain settings show that our proposed method can obviously outperform the state-of-the-art methods. Yansheng Li 0001, Te Shi 0001, Wei Chen 0089, Yongjun Zhang 0002, Zhibin Wang 0004, Hao Li 0030 |
IGARSS | 5 |
| 2017 | The Opensesame NIST 2016 Speaker Recognition Evaluation System
Qi Qian 0001, Zhibin Wang 0004, Qingen Zhao, Tianzhou Wang, Hao Li 0030, Shenghuo Zhu, Rong Jin 0001, Tuo Zhao |
INTERSPEECH | 3 |