EDBT 2026 Demo / reviewers in the wild / expert
Wei Chen 0089
dblp:181/2832-89
· DBLP profile ↗
15ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0001-9937-4122ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | REST: Holistic Learning for End-to-End Semantic Segmentation of Whole-Scene Remote Sensing ImageryabstractSemantic segmentation of remote sensing imagery (RSI) is a fundamental task that aims at assigning a category label to each pixel. To pursue precise segmentation with one or more fine-grained categories, semantic segmentation often requires holistic segmentation of whole-scene RSI (WRI), which is normally characterized by a large size. However, conventional deep learning methods struggle to handle holistic segmentation of WRI due to the memory limitations of the graphics processing unit (GPU), thus requiring to adopt suboptimal strategies such as cropping or fusion, which result in performance degradation. Here, we introduce the Robust End-to-end semantic Segmentation architecture for whole-scene remoTe sensing imagery (REST). REST is the first intrinsically endtoend framework for truly holistic segmentation of WRI, supporting a wide range of encoders and decoders in a plugandplay fashion. It enables seamless integration with mainstream semantic segmentation methods, and even more advanced foundation models. Specifically, we propose a novel spatial parallel interaction mechanism (SPIM) within REST to overcome GPU memory constraints and achieve global context awareness. Unlike traditional parallel methods, SPIM enables REST to process a WRI effectively and efficiently by combining parallel computation with a divideandconquer strategy. Both theoretical analysis and experiments demonstrate that REST attains nearlinear throughput scalability as additional GPUs are employed. Extensive experiments demonstrate that REST consistently outperforms existing cropping-based and fusion-based methods across a variety of scenarios, ranging from single-class to multi-class segmentation, from multispectral to hyperspectral imagery, and from satellite to drone platforms. The robustness and versatility of REST are expected to offer a promising solution for the holistic segmentation of WRI, with the potential for further extension to large-size medical imagery segmentation. Wei Chen 0089, Lorenzo Bruzzone, Bo Dang 0002, Yuan Gao 0015, Youming Deng, Jin-Gang Yu, Liangqi Yuan, Yansheng Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Full-Scope Vectorization of Geographical Elements from Large-Size Remote Sensing ImageryabstractLarge-size very-high-resolution (VHR) remote sensing imagery has emerged as a critical data source for high-precision vector mapping of multi-scale geographical elements such as building, water, road and etc. When dealing with the large-size image, due to the limited memory of GPU, the deep learning-based vector mapping methods often employ the sliding block strategy. This inevitably leads to the degenerated performance because of the stitching difficulty of the sliding blocks' vector mapping results. Therefore, it is necessary to conduct full-scope vector mapping via mining the consistent cue in large-size remote sensing imagery. To this end, this paper presents a novel global context-aware local point optimization method. To leverage the global context, this paper proposes a novel pyramid fusion network (PFNet) to conduct semantic segmentation of the large-size image in an end-to-end manner. Under the constraint of the global semantic segmentation result, a new inflection-point perception network (IPNet) is proposed to generate a set of stable points to depict the boundary of each element. Extensive experiments on building, water and road datasets, where each image has over 100 million pixels, show that our method obviously outperforms the existing methods. Yansheng Li 0001, Wanchun Li, Bo Dang 0002, Yu Wang 0222, Wei Chen 0089, Bingnan Yang, Yongjun Zhang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Tuning-Free Long Video Generation via Global-Local Collaborative DiffusionabstractCreating high-fidelity, coherent long videos is a sought-after aspiration. While recent video diffusion models have shown promising potential, they still grapple with spatiotemporal inconsistencies and high computational resource demands. We propose Global-Local Collaborative Diffusion (GLC-Diffusion), a tuning-free method for long video generation. It models the long video denoising process by establishing denoising trajectories through Global-Local Collaborative Denoising (GLCD) to ensure overall content consistency and temporal coherence between frames. Additionally, we introduce a Noise Reinitialization strategy which combines local noise shuffling with frequency fusion to improve global content consistency and visual diversity. Further, we propose a Video Motion Consistency Refinement (VMCR) module that computes the gradient of pixel-wise and frequency-wise losses to enhance visual consistency and temporal smoothness. Extensive experiments, including quantitative and qualitative evaluations on videos of varying lengths (e.g., 3× and 6× longer), demonstrate that our method effectively integrates with existing video diffusion models, producing coherent, high-fidelity long videos superior to previous approaches. Yongjia Ma, Junlin Chen, Donglin Di, Qi Xie 0009, Lei Fan 0007, Wei Chen 0089, Na Zhao 0004, Xun Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | GRPose: Learning Graph Relations for Human Image Generation with Pose PriorsabstractRecent methods using diffusion models have made significant progress in human image generation with various control signals such as pose priors. However, existing efforts are still struggling to generate high-quality images with consistent pose alignment, resulting in unsatisfactory output. In this paper, we propose a framework that delves into the graph relations of pose priors to provide control information for human image generation. The main idea is to establish a graph topological structure between the pose priors and latent representation of diffusion models to capture the intrinsic associations between different pose parts. A Progressive Graph Integrator (PGI) is designed to learn the spatial relationships of the pose priors with the graph structure, adopting a hierarchical strategy within an Adapter to gradually propagate information across different pose parts. Besides, a pose perception loss is introduced based on a pretrained pose estimation network to minimize the pose differences. Extensive qualitative and quantitative experiments conducted on the Human-Art and LAION-Human datasets clearly demonstrate that our model can achieve significant performance improvement over the latest benchmark models. Xiangchen Yin, Donglin Di, Lei Fan 0007, Hao Li 0030, Wei Chen 0089, Gouxiao Fei, Yang Song 0001, Xiao Sun 0003, Xun Yang 0001 |
AAAI | 5 |
| 2025 | DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation
Donglin Di, Wenzhang Sun, Yongjia Ma, Hao Li 0030, Wei Chen 0089, Lei Fan 0007, Tonghua Su, Xun Yang 0001 |
ICCV | 6 |
| 2025 | HoliTracer: Holistic Vectorization of Geographic Objects from Large-Size Remote Sensing ImageryabstractWith the increasing resolution of remote sensing imagery (RSI), large-size RSI has emerged as a vital data source for high-precision vector mapping of geographic objects. Existing methods are typically constrained to processing small image patches, which often leads to the loss of contextual information and produces fragmented vector outputs. To address these, this paper introduces HoliTracer, the first framework designed to holistically extract vectorized geographic objects from large-size RSI. In HoliTracer, we enhance segmentation of large-size RSI using the Context Attention Net (CAN), which employs a local-to-global attention mechanism to capture contextual dependencies. Furthermore, we achieve holistic vectorization through a robust pipeline that leverages the Mask Contour Reformer (MCR) to reconstruct polygons and the Polygon Sequence Tracer (PST) to trace vertices. Extensive experiments on large-size RSI datasets, including buildings, water bodies, and roads, demonstrate that HoliTracer outperforms state-of-the-art methods. Our code and data are available in https://github.com/vvangfaye/HoliTracer. Yu Wang 0222, Bo Dang 0002, Wanchun Li, Wei Chen 0089, Yansheng Li 0001 |
ICCV | 4 |
| 2025 | QR-LoRA: Efficient and Disentangled Fine-Tuning via QR Decomposition for Customized GenerationabstractExisting text-to-image models often rely on parameter fine-tuning techniques such as Low-Rank Adaptation (LoRA) to customize visual attributes. However, when combining multiple LoRA models for content-style fusion tasks, unstructured modifications of weight matrices often lead to undesired feature entanglement between content and style attributes. We propose QR-LoRA, a novel fine-tuning framework leveraging QR decomposition for structured parameter updates that effectively separate visual attributes. Our key insight is that the orthogonal Q matrix naturally minimizes interference between different visual features, while the upper triangular R matrix efficiently encodes attribute-specific transformations. Our approach fixes both Q and R matrices while only training an additional task-specific $ΔR$ matrix. This structured design reduces trainable parameters to half of conventional LoRA methods and supports effective merging of multiple adaptations without cross-contamination due to the strong disentanglement properties between $ΔR$ matrices. Experiments demonstrate that QR-LoRA achieves superior disentanglement in content-style fusion tasks, establishing a new paradigm for parameter-efficient, disentangled fine-tuning in generative models. The project page is available at: https://luna-ai-lab.github.io/QR-LoRA/. Yongjia Ma, Donglin Di, Jianxun Cui, Hao Li 0030, Wei Chen 0089, Xun Yang 0001, Wangmeng Zuo |
ICCV | 6 |
| 2025 | PRISM: A Benchmark for Unveiling Cross-modal Knowledge Inconsistency in Large Vision-Language ModelsabstractRecent advances in Large Vision-Language Models (LVLMs) have unearthed boosted performance of multi-modal understanding. In this paper, however, we for the first time uncover a critically under-explored challenge persisting in this trend, that LVLMs unfortunately exhibit cross-modal knowledge inconsistencies. Cross-modal knowledge inconsistency refers to the tendency of providing semantically inconsistent responses to contexts that are semantically equivalent but expressed in different modalities. In real-world applications, users can rely on either text or image to express their ideas. Inconsistent responses across modalities can confuse users, challenging the reliabilities of LVLMs in practice. Therefore, we argue that evaluating performance on either multi-modal or text-only task is insufficient; and waiving the mentioned cross-modal knowledge inconsistency is crucial. The paper proposes PRISM, the first-ever benchmark for measuring the inconsistency, and the corresponding evaluation metric Know-Inc. PRISM covers commonsense, encyclopedia, and mathematics knowledge, with manually-screened samples of semantic alignment. From the evaluation results of up to 27 LVLMs with diverse structures, we conclude that: 1) LVLMs show a preference for textual input, 2) there is a correlation between inconsistency and accuracy, and 3) the inconsistency is more prominent in encyclopedia knowledge. These findings can shed light on further optimization and development of LVLMs. Weinan Zhang 0003, Donglin Di, Wei Chen 0089, Ting Liu 0001 |
ACM Multimedia | 7 |
| 2025 | Hyper-3DG: Text-to-3D Gaussian Generation via Hypergraph
Donglin Di, Chaofan Luo, Zhou Xue, Wei Chen 0089, Xun Yang 0001, Yue Gao 0002 |
Int. J. Comput. Vis. | 5 |
| 2025 | Randomness-Restricted Diffusion Model for Ocular Surface Structure SegmentationabstractOcular surface diseases affect a significant portion of the population worldwide. Accurate segmentation and quantification of different ocular surface structures are crucial for the understanding of these diseases and clinical decision-making. However, the automated segmentation of the ocular surface structure is relatively unexplored and faces several challenges. Ocular surface structure boundaries are often inconspicuous and obscured by glare from reflections. In addition, the segmentation of different ocular structures always requires training of multiple individual models. Thus, developing a one-model-fits-all segmentation approach is desirable. In this paper, we introduce a randomness-restricted diffusion model for multiple ocular surface structure segmentation. First, a time-controlled fusion-attention module (TFM) is proposed to dynamically adjust the information flow within the diffusion model, based on the temporal relationships between the network's input and time. TFM enables the network to effectively utilize image features to constrain the randomness of the generation process. We further propose a low-frequency consistency filter and a new loss to alleviate model uncertainty and error accumulation caused by the multi-step denoising process. Extensive experiments have shown that our approach can segment seven different ocular surface structures. Our method performs better than both dedicated ocular surface segmentation methods and general medical image segmentation methods. We further validated the proposed method over two clinical datasets, and the results demonstrated that it is beneficial to clinical applications, such as the meibomian gland dysfunction grading and aqueous deficient dry eye diagnosis. Huaying Hao, Yifan Zhao 0001, Yanda Meng, Jiang Liu 0001, Yalin Zheng, Wei Chen 0089, Yitian Zhao |
IEEE Trans. Medical Imaging | 8 |
| 2025 | TrAME: Trajectory-Anchored Multi-View Editing for Text-Guided 3D Gaussian ManipulationabstractDespite significant strides in the field of 3D scene editing, current methods encounter substantial challenge, particularly in preserving 3D consistency during the multi-view editing process. To tackle this challenge, we propose a progressive 3D editing strategy that ensures multi-view consistency via a Trajectory-Anchored Scheme (TAS) with a dual-branch editing mechanism. Specifically, TAS facilitates a tightly coupled iterative process between 2D view editing and 3D updating, preventing error accumulation yielded from the text-to-image process. Additionally, we explore the connection between optimization-based methods and reconstruction-based methods, offering a unified perspective for selecting superior design choices, supporting the rationale behind the designed TAS. We further present a tuning-free View-Consistent Attention Control (VCAC) module that leverages cross-view semantic and geometric reference from the source branch to yield aligned views from the target branch during the editing of 2D views. To validate the effectiveness of our method, we analyze 2D examples to demonstrate the improved consistency with the VCAC module. Extensive quantitative and qualitative results in text-guided 3D scene editing clearly indicate that our method can achieve superior editing quality compared with state-of-the-art 3D scene editing methods. Our project site is athttps://fkcptlst.github.io/TrAME/ Chaofan Luo, Donglin Di, Xun Yang 0001, Yongjia Ma, Zhou Xue, Wei Chen 0089, Xiaofei Gou, Yebin Liu |
IEEE Trans. Multim. | 6 |
| 2023 | MFVNet: a deep adaptive fusion network with multiple field-of-views for remote sensing image semantic segmentation
Yansheng Li 0001, Wei Chen 0089, Xin Huang 0002, Zhi Gao 0005, Tao He 0002, Yongjun Zhang 0002 |
Sci. China Inf. Sci. | 2 |
| 2021 | Emotional dialog generation via multiple classifiers based on a generative adversarial networkabstractHuman-machine dialog generation is an essential topic of research in the field of natural language processing. Generating high-quality, diverse, fluent, and emotional conversation is a challenging task. Based on continuing advancements in artificial intelligence and deep learning, new methods have come to the forefront in recent times. In particular, the end-to-end neural network model provides an extensible conversation generation framework that has the potential to enable machines to understand semantics and automatically generate responses. However, neural network models come with their own set of questions and challenges. The basic conversational model framework tends to produce universal, meaningless, and relatively "safe" answers. Based on generative adversarial networks (GANs), a new emotional dialog generation framework called EMC-GAN is proposed in this study to address the task of emotional dialog generation. The proposed model comprises a generative and three discriminative models. The generator is based on the basic sequence-to-sequence (Seq2Seq) dialog generation model, and the aggregate discriminative model for the overall framework consists of a basic discriminative model, an emotion discriminative model, and a fluency discriminative model. The basic discriminative model distinguishes generated fake sentences from real sentences in the training corpus. The emotion discriminative model evaluates whether the emotion conveyed via the generated dialog agrees with a pre-specified emotion, and directs the generative model to generate dialogs that correspond to the category of the pre-specified emotion. Finally, the fluency discriminative model assigns a score to the fluency of the generated dialog and guides the generator to produce more fluent sentences. Based on the experimental results, this study confirms the superiority of the proposed model over similar existing models with respect to emotional accuracy, fluency, and consistency. The proposed EMC-GAN model is capable of generating consistent, smooth, and fluent dialog that conveys pre-specified emotions, and exhibits better performance with respect to emotional accuracy, consistency, and fluency compared to its competitors. Wei Chen 0089, Xinmiao Chen, Xiao Sun 0003 |
Virtual Real. Intell. Hardw. | 1 |
| 2020 | Deep Networks Under Block-Level Supervision for Pixel-Level Cloud Detection in Multi-Spectral Satellite ImageryabstractCloud cover hinders the usability of optical remote sensing imagery. Existing cloud detection methods either require hand-crafted features or utilize deep networks. Generally, deep networks perform better than hand-crafted features. However, deep networks for cloud detection need massive and expensive pixel-level annotation labels. To alleviate that, this paper proposes a weakly supervised deep learning-based cloud detection method using only block-level labels, with a new global convolutional pooling operation and a local pooling pruning strategy to improve the performance. For evaluating, we collect a training dataset containing over 160,000 image blocks with block-level labels and a testing dataset including ten large image scenes with pixel-level labels. Even under extremely weak supervision, our method performed well with the average overall accuracy reached 97.2 %. Experiments demonstrate that our proposed method obviously outperforms the state-of-the-art methods. Wei Chen 0089, Yansheng Li 0001, Yongjun Zhang 0002, Xiaolong Hao |
IGARSS | 1 |
| 2020 | Unsupervised Style Transfer via Dualgan for Cross-Domain Aerial Image ClassificationabstractDue to its wide applications, aerial image classification, which is also called semantic segmentation of aerial imagery, attracts increasing research interest in recent years. Until now, deep semantic segmentation network (DSSN) has been widely adopted to address aerial image classification and achieves tremendous success. However, the superior performance of DSSN highly depends on massive targeted data with labels. When DSSN is trained on data from the source domain but tested on data from the target domain, the performance of DSSN is often very limited due to the data shift between source and target domains. To alleviate the disadvantage influence of data shift, this paper proposes a domain adaptation approach via unsupervised style transfer to cope with cross-domain aerial image classification. More specifically, this paper innovatively recommends DualGAN to conduct unsupervised style transfer for mapping aerial images in the source domain to the target domain. The mapped aerial imagery with labels is adopted to train DSSN, which is further used to classify aerial imagery in the target domain. To verify the validity of the presented approach, we give two cross-domain experimental settings including: (I) variation of geographic location; (II) variation of both geographic location and imaging mode. Extensive experiments under two typical cross-domain settings show that our proposed method can obviously outperform the state-of-the-art methods. Yansheng Li 0001, Te Shi 0001, Wei Chen 0089, Yongjun Zhang 0002, Zhibin Wang 0004, Hao Li 0030 |
IGARSS | 3 |