VLDB 2026 Research / reviewers in the wild / expert
Wei Dong 0010
dblp:92/748-10
· DBLP profile ↗
27ranked-venue papers
9as first author
26since 2021 · last 2026
0000-0003-0263-3584ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 8 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 13 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seeing Beyond Illusion: Generalized and Efficient Mirror DetectionabstractReflective imaging enables the mirror imagings and physical entities to possess identical attributes, e.g., color and shape. Current mirror detection (MD) methods primarily rely on designing functional components to establish the correlation and disparities between the imagings and entities, thereby identifying the mirror regions. However, the exploration of extended scenes with dynamic content changes is rarely investigated. Therefore, we propose the MirrorSAM designed for MD based on the Segment Anything Model (SAM). Specifically, due to the varying reflections produced by mirrors in different positions and the complex visual space that interferes with localization, we design the hierarchical mixture of direction experts (HMDE) in the low-rank space to reduce biases towards entities in SAM and dynamically adjust experts based on the input scene. We observe differences in depth between mirrors and adjacent areas, and propose the depth token calibration (DTC), which introduces a learnable depth token to generate the depth map and serve as an error correction factor. We further formulate the selective pixel-prototype contrastive (SPPC) loss, selecting partially confusable samples to promote the decoupling of mirror and non-mirror representations. Extensive experiments conducted on four mirror benchmarks and two settings demonstrate that our approach surpasses state-of-the-art methods with few trainable parameters and FLOPs. We further extend to four transparent surface benchmarks to validate generalization. Mingfeng Zha, Guoqing Wang 0001, Tianyu Li 0003, Wei Dong 0010, Peng Wang 0023, Yang Yang 0002 |
AAAI | 4 |
| 2026 | Quantum-resistant blockchain architecture for secure vehicular networks: A ML-KEM-enabled approach with PoA and PoP consensus
Junsheng Wu, Weigang Li 0005, Zhijun Lin, Wei Dong 0010, Ghulam Mohiuddin |
Future Gener. Comput. Syst. | 7 |
| 2026 | Boosting HDR Image Reconstruction via Semantic Knowledge TransferabstractRecovering High Dynamic Range (HDR) images from multiple Standard Dynamic Range (SDR) images becomes challenging when the SDR images exhibit noticeable degradation and missing content. Leveraging scene-specific semantic priors offers a promising solution for restoring heavily degraded regions. However, these priors are typically extracted from sRGB SDR images, the domain/format gap poses a significant challenge when applying it to HDR imaging. To address this issue, we propose a general framework that transfers semantic knowledge derived from SDR domain via self-distillation to boost existing HDR reconstruction. Specifically, the proposed framework first introduces the Semantic Priors Guided Reconstruction Model (SPGRM), which leverages SDR image semantic knowledge to address ill-posed problems in the initial HDR reconstruction results. Subsequently, we leverage a self-distillation mechanism that constrains the color and content information with semantic knowledge, aligning the external outputs between the baseline and SPGRM. Furthermore, to transfer the semantic knowledge of the internal features, we utilize a Semantic Knowledge Alignment Module (SKAM) to fill the missing semantic contents with the complementary masks. Extensive experiments demonstrate that our framework significantly boosts HDR imaging quality for existing methods without altering the network architecture. Tao Hu 0013, Longyao Wu, Wei Dong 0010, Peng Wu 0015, Jinqiu Sun, Xiaogang Xu 0002, Qingsen Yan, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Ghost-Free HDR Imaging via Latent Low-Frequency Priors and Deformable Attention AlignmentabstractRecovering ghost-free High Dynamic Range (HDR) images from multiple Low Dynamic Range (LDR) images becomes challenging when the LDR images exhibit saturation and significant motion. Recent Diffusion Models (DMs) have been introduced in HDR imaging field, showing promising performance, particularly in achieving visually perceptible better results compared to previous DNN-based methods. However, DMs require extensive iterations with large models to estimate entire images, resulting in inefficiency that hinders their practical application. To address this challenge, we propose the Low-Frequency aware Diffusion (LF-Diff) model for ghost-free HDR imaging. The key idea of LF-Diff is implementing the DMs in a highly compacted latent space and integrating it into a regression-based model to enhance the details of reconstructed images. Specifically, as low-frequency information is closely related to human visual perception we propose to utilize DMs to create compact low-frequency priors for the reconstruction process. These priors are integrated into a carefully designed Dynamic HDR Reconstruction Network (DHRNet), which employs a regression-based approach to produce high-quality HDR images. Furthermore, we introduce the Attention-guided Deformable Alignment Module (ADAM) that utilizes correlation-driven feature matching to learn deformable receptive fields for self-attention, enabling efficient pre-alignment of LDR images by focusing on salient regions. Extensive experiments on synthetic and real-world benchmark datasets demonstrate that our LF-Diff performs favorably against several state-of-the-art methods and is $10\times $ faster than previous DM-based methods. Tao Hu 0013, Qingsen Yan, Wei Dong 0010, Peng Wu 0015, Yuankai Qi, Weisi Lin, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | TG-LLaVA: Text Guided LLaVA via Learnable Latent EmbeddingsabstractCurrently, inspired by the success of vision-language models (VLMs), an increasing number of researchers are focusing on improving VLMs and have achieved promising results. However, most existing methods concentrate on optimizing the connector and enhancing the language model component, while neglecting improvements to the vision encoder itself. In contrast, we propose Text Guided LLaVA (TG-LLaVA) in this paper, which optimizes VLMs by guiding the vision encoder with text, offering a new and orthogonal optimization direction. Specifically, inspired by the purpose-driven logic inherent in human behavior, we use learnable latent embeddings as a bridge to analyze textual instruction and add the analysis results to the vision encoder as guidance, refining it. Subsequently, another set of latent embeddings extracts additional detailed text-guided information from high-resolution local patches as auxiliary information. Finally, with the guidance of text, the vision encoder can extract text-related features, similar to how humans focus on the most relevant parts of an image when considering a question. This results in generating better answers. Experiments on various datasets validate the effectiveness of the proposed method. Remarkably, without the need for additional training data, our proposed method can bring more benefits to the baseline (LLaVA-1.5) compared with other concurrent methods. Furthermore, the proposed method consistently brings improvement in different settings. Dawei Yan 0001, Hao Chen 0041, Weihua Luo, Wei Dong 0010, Qingsen Yan, Haokui Zhang, Chunhua Shen |
AAAI | 7 |
| 2025 | HVI: A New Color Space for Low-light Image EnhancementabstractLow-Light Image Enhancement (LLIE) is a crucial computer vision task that aims to restore detailed visual information from corrupted low-light images. Many existing LLIE methods are based on standard RGB (sRGB) space, which often produce color bias and brightness artifacts due to inherent high color sensitivity in sRGB. While converting the images using Hue, Saturation and Value (HSV) color space helps resolve the brightness issue, it introduces significant red and black noise artifacts. To address this issue, we propose a new color space for LLIE, namely Horizontal/Vertical-Intensity (HVI), defined by polarized HS maps and learnable intensity. The former enforces small distances for red coordinates to remove the red artifacts, while the latter compresses the low-light regions to remove the black artifacts. To fully leverage the chromatic and intensity information, a novel Color and Intensity Decoupling Network (CIDNet) is further introduced to learn accurate photometric mapping function under different lighting conditions in the HVI space. Comprehensive results from benchmark and ablation experiments show that the proposed HVI color space with CIDNet outperforms the state-of-the-art methods on 10 datasets. The code is available at https://github.com/Fediory/HVI-CIDNet. Qingsen Yan, Yixu Feng, Guansong Pang, Kangbiao Shi, Peng Wu 0015, Wei Dong 0010, Jinqiu Sun, Yanning Zhang 0001 |
CVPR | 7 |
| 2025 | Efficient Adaptation of Pre-Trained Vision Transformer Underpinned by Approximately Orthogonal Fine-Tuning StrategyabstractA prevalent approach in Parameter-Efficient Fine-Tuning (PEFT) of pre-trained Vision Transformers (ViT) involves freezing the majority of the backbone parameters and solely learning low-rank adaptation weight matrices to accommodate downstream tasks. These low-rank matrices are commonly derived through the multiplication structure of down-projection and up-projection matrices, exemplified by methods such as LoRA and Adapter. In this work, we observe an approximate orthogonality among any two row or column vectors within any weight matrix of the backbone parameters; however, this property is absent in the vectors of the down/up-projection matrices. Approximate orthogonality implies a reduction in the upper bound of the model's generalization error, signifying that the model possesses enhanced generalization capability. If the fine-tuned down/up-projection matrices were to exhibit this same property as the pre-trained backbone matrices, could the generalization capability of fine-tuned ViTs be further augmented? To address this question, we propose an Approximately Orthogonal Fine-Tuning (AOFT) strategy for representing the low-rank weight matrices. This strategy employs a single learnable vector to generate a set of approximately orthogonal vectors, which form the down/up-projection matrices, thereby aligning the properties of these matrices with those of the backbone. Extensive experimental results demonstrate that our method achieves competitive performance across a range of downstream image classification tasks, confirming the efficacy of the enhanced generalization capability embedded in the down/up-projection matrices. Yiting Yang, Qingsen Yan, Haokui Zhang, Wei Dong 0010, Guoqing Wang 0001, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
ICCV | 6 |
| 2025 | SUNet: A Semantic-Driven Framework for Universal Image EnhancementabstractDeep learning has significantly improved image quality in enhancement and retouching tasks. However, current methods, such as HDRNet, CSRNet, and 3D LUT, primarily rely on low-level visual features and lack in-depth utilization of image semantic information, resulting in global average enhancement, color inconsistencies, and loss of brightness details. Noise and blur in images intensify the blending of distinct features and contribute to information degradation, making it more challenging to extract semantic details and thereby limiting the recovery performance. In order to solve the above problem, this paper proposes SUNet, an image restoration model enhanced with semantic information. By incorporating fine-grained semantic information into UNet and employing orthogonal and decoupled feature representations, SUNet significantly improves restoration performance without relying on specific segmentation annotations. The introduction of semantic information enables the model to differentiate between different regions in the image distinctly, making feature extraction more targeted and progressively reducing feature coupling. This enhances the model’s ability to represent semantically relevant features while avoiding interference from blurry or noisy features. Our contribution effectively bridges the gap between global enhancement techniques and the need for local semantic accuracy, laying the foundation for more sophisticated image enhancement methods. Experimental evaluations on public datasets demonstrate that our approach outperforms state-of-the-art methods. Dawei Yan 0001, Ghulam Mohiuddin, Marcin Wozniak, Wei Dong 0010 |
IJCNN | 9 |
| 2025 | CLIP-guided continual novel class discovery
Qingsen Yan, Yiting Yang, Yutong Dai 0001, Katarzyna Wiltos, Marcin Wozniak, Wei Dong 0010, Yanning Zhang 0001 |
Knowl. Based Syst. | 7 |
| 2025 | Efficient Image Enhancement With a Diffusion-Based Frequency PriorabstractDue to the lack of appropriate priors, generating the content of dark regions remains a challenge in low-light image enhancement tasks. Currently, diffusion models employ robust image generation capabilities for enhancing low-light images. However, diffusion models require multiple iterations at the image feature level to generate details and content, which limits the speed. Moreover, the diffusion-based methods tend to generate unexpected artifacts in the degraded regions. To address these issues, we propose a Frequency Priors-guided Image Enhancement (FPIE) network, including a frequency prior generation network and an image restoration network. FPIE significantly accelerates inference by learning abstract prior with frequency domain constraints. Concretely, to learn compacted priors at the frequency domain, we introduce a joint training approach for the prior generation and restoration models to constrain the distribution of priors. Furthermore, to better utilize frequency-domain features for enhancing the network’s generation capabilities, a wavelet-based transformer block is introduced to produce intricate details and avoid the artifacts of the output. Extensive experimental results on the commonly used benchmarks demonstrate that our approach achieves state-of-the-art performances and well generalization to real-world images. Qingsen Yan, Tao Hu 0013, Peng Wu 0015, Duwei Dai, Shuhang Gu, Wei Dong 0010, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Low-Rank Rescaled Vision Transformer Fine-Tuning: A Residual Design ApproachabstractParameter-efficient fine-tuning for pre-trained Vision Transformers aims to adeptly tailor a model to downstream tasks by learning a minimal set of new adaptation parameters while preserving the frozen majority of pre-trained parameters. Striking a balance between retaining the generalizable representation capacity of the pre-trained model and acquiring task-specific features poses a key challenge. Currently, there is a lack of focus on guiding this delicate trade-off. In this study, we approach the problem from the perspective of Singular Value Decomposition (SVD) of pre-trained parameter matrices, providing insights into the tuning dynamics of existing methods. Building upon this understanding, we propose a Residual-based Low-Rank Rescaling (RLRR) fine-tuning strategy. This strategy not only enhances flexibility in parameter tuning but also ensures that new parameters do not deviate excessively from the pre-trained model through a residual design. Extensive experiments demonstrate that our method achieves competitive performance across various downstream image classification tasks, all while maintaining comparable new parameters. We believe this work takes a step forward in offering a unified perspective for interpreting existing methods and serves as motivation for the development of new approaches that move closer to effectively considering the crucial trade-off mentioned above. Our code is available at https://github.com/zstarN70/RLRR.git. Wei Dong 0010, Bihui Chen, Dawei Yan 0001, Zhijun Lin, Qingsen Yan, Peng Wang 0023, Yang Yang 0002 |
CVPR | 1 |
| 2024 | Efficient Adaptation of Pre-trained Vision Transformer via Householder TransformationabstractA common strategy for Parameter-Efficient Fine-Tuning (PEFT) of pre-trained Vision Transformers (ViTs) involves adapting the model to downstream tasks by learning a low-rank adaptation matrix. This matrix is decomposed into a product of down-projection and up-projection matrices, with the bottleneck dimensionality being crucial for reducing the number of learnable parameters, as exemplified by prevalent methods like LoRA and Adapter. However, these low-rank strategies typically employ a fixed bottleneck dimensionality, which limits their flexibility in handling layer-wise variations. To address this limitation, we propose a novel PEFT approach inspired by Singular Value Decomposition (SVD) for representing the adaptation matrix. SVD decomposes a matrix into the product of a left unitary matrix, a diagonal matrix of scaling values, and a right unitary matrix. We utilize Householder transformations to construct orthogonal matrices that efficiently mimic the unitary matrices, requiring only a vector. The diagonal values are learned in a layer-wise manner, allowing them to flexibly capture the unique properties of each layer. This approach enables the generation of adaptation matrices with varying ranks across different layers, providing greater flexibility in adapting pre-trained models. Experiments on standard downstream vision tasks demonstrate that our method achieves promising fine-tuning performance. Wei Dong 0010, Yiting Yang, Zhijun Lin, Qingsen Yan, Haokui Zhang, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
NeurIPS | 1 |
| 2024 | SAMT-generator: A second-attention for image captioning based on multi-stage transformer network
Xiaobao Yang 0001, Yang Yang 0002, Sugang Ma, Wei Dong 0010, Marcin Wozniak |
Neurocomputing | 5 |
| 2024 | Dynamic center point learning for multiple object tracking under Severe occlusions
Yaoqi Hu, Axi Niu, Jinqiu Sun, Yu Zhu 0004, Qingsen Yan, Wei Dong 0010, Marcin Wozniak, Yanning Zhang 0001 |
Knowl. Based Syst. | 6 |
| 2024 | Hierarchical aggregation perceptual pipeline for tactical intention recognition
Ying Li 0055, Junsheng Wu, Weigang Li 0005, Wei Dong 0010, Aiqing Fang |
Multim. Tools Appl. | 4 |
| 2024 | Self-Supervised Node Representation Learning via Node-to-Neighbourhood AlignmentabstractSelf-supervised node representation learning aims to learn node representations from unlabelled graphs that rival the supervised counterparts. The key towards learning informative node representations lies in how to effectively gain contextual information from the graph structure. In this work, we present simple-yet-effective self-supervised node representation learning via aligning the hidden representations of nodes and their neighbourhood. Our first idea achieves such node-to-neighbourhood alignment by directly maximizing the mutual information between their representations, which, we prove theoretically, plays the role of graph smoothing. Our framework is optimized via a surrogate contrastive loss and a Topology-Aware Positive Sampling (TAPS) strategy is proposed to sample positives by considering the structural dependencies between nodes, which enables offline positive selection. Considering the excessive memory overheads of contrastive learning, we further propose a negative-free solution, where the main contribution is a Graph Signal Decorrelation (GSD) constraint to avoid representation collapse and over-smoothing. The GSD constraint unifies some of the existing constraints and can be used to derive new implementations to combat representation collapse. By applying our methods on top of simple MLP-based node representation encoders, we learn node representations that achieve promising node classification performance on a set of graph-structured datasets from small- to large-scale. Wei Dong 0010, Dawei Yan 0001, Peng Wang 0023 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | KGSR: A kernel guided network for real-world blind super-resolution
Qingsen Yan, Axi Niu, Wei Dong 0010, Marcin Wozniak, Yanning Zhang 0001 |
Pattern Recognit. | 4 |
| 2024 | Uncertainty estimation in HDR imaging with Bayesian neural networks
Qingsen Yan, Haishen Wang, Yuhang Liu 0002, Wei Dong 0010, Marcin Wozniak, Yanning Zhang 0001 |
Pattern Recognit. | 5 |
| 2024 | Toward High-Quality HDR Deghosting With Conditional Diffusion ModelsabstractHigh Dynamic Range (HDR) images can be recovered from several Low Dynamic Range (LDR) images by existing Deep Neural Networks (DNNs) techniques. Despite the remarkable progress, DNN-based methods still generate ghosting artifacts when LDR images have saturation and large motion, which hinders potential applications in real-world scenarios. To address this challenge, we formulate the HDR deghosting problem as an image generation that leverages LDR features as the diffusion model’s condition, consisting of the feature condition generator and the noise predictor. Feature condition generator employs attention and Domain Feature Alignment (DFA) layer to transform the intermediate features to avoid ghosting artifacts. With the learned features as conditions, the noise predictor leverages a stochastic iterative denoising process for diffusion models to generate an HDR image by steering the sampling process. Furthermore, to mitigate semantic confusion caused by the saturation problem of LDR images, we design a sliding window noise estimator to sample smooth noise in a patch-based manner. In addition, an image space loss is proposed to avoid the color distortion of the estimated HDR results. We empirically evaluate our model on benchmark datasets for HDR imaging. The results demonstrate that our approach achieves state-of-the-art performances and well generalization to real-world images. Qingsen Yan, Tao Hu 0013, Hao Tang 0005, Yu Zhu 0004, Wei Dong 0010, Luc Van Gool, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Efficient Adaptation of Large Vision Transformer via Adapter Re-ComposingabstractThe advent of high-capacity pre-trained models has revolutionized problem-solving in computer vision, shifting the focus from training task-specific models to adapting pre-trained models. Consequently, effectively adapting large pre-trained models to downstream tasks in an efficient manner has become a prominent research area. Existing solutions primarily concentrate on designing lightweight adapters and their interaction with pre-trained models, with the goal of minimizing the number of parameters requiring updates. In this study, we propose a novel Adapter Re-Composing (ARC) strategy that addresses efficient pre-trained model adaptation from a fresh perspective. Our approach considers the reusability of adaptation parameters and introduces a parameter-sharing scheme. Specifically, we leverage symmetric down-/up-projections to construct bottleneck operations, which are shared across layers. By learning low-dimensional re-scaling coefficients, we can effectively re-compose layer-adaptive adapters. This parameter-sharing strategy in adapter design allows us to further reduce the number of new parameters while maintaining satisfactory performance, thereby offering a promising approach to compress the adaptation cost. We conduct experiments on 24 downstream image classification tasks using various Vision Transformer variants to evaluate our method. The results demonstrate that our approach achieves compelling transfer learning performance with a reduced parameter count. Our code is available at https://github.com/DavidYanAnDe/ARC. Wei Dong 0010, Dawei Yan 0001, Zhijun Lin, Peng Wang 0023 |
NeurIPS | 1 |
| 2023 | Denoising Aggregation of Graph Neural Networks by Using Principal Component AnalysisabstractTo avoid the overfitting phenomenon that appeared in performing graph neural networks (GNNs) on test examples, the feature encoding scheme of GNNs usually introduces the dropout procedure. However, after learning latent node representations under this scheme, Gaussian noise produced by the dropout operation is inevitably transmitted into the next neighborhood aggregation step, which necessarily hampers the unbiased aggregation ability of GNN models. To address this issue, in this article, we present a novel aggregator, denoising aggregation (DNAG), which utilizes principal component analysis (PCA) to preserve the aggregated real signals from neighboring features and simultaneously filter out the Gaussian noise. The idea is different from using PCA on traditional applications to reduce the feature dimension. We regard PCA as an aggregator to compress the neighboring node features to have better expressive denoising power. We propose new training architectures to simplify the intensive computation of PCA in DNAG. Numerical experiments show the apparent superiority of the proposed DNAG models in gaining more denoising capability and achieving the state of the art for a set of predictive tasks on several graph-structured datasets. Wei Dong 0010, Marcin Wozniak, Junsheng Wu, Weigang Li 0005, Zongwen Bai |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | SCOAD: Single-Frame Click Supervision for Online Action Detection
Dawei Yan 0001, Wei Dong 0010, Qingsen Yan |
ACCV (4) | 4 |
| 2022 | DualBLN: Dual Branch LUT-Aware Network for Real-Time Image Retouching
Chengzhe Lu, Dawei Yan 0001, Wei Dong 0010, Qingsen Yan |
ACCV (3) | 4 |
| 2022 | Node Representation Learning in Graph via Node-to-Neighbourhood Mutual Information MaximizationabstractThe key towards learning informative node representations in graphs lies in how to gain contextual information from the neighbourhood. In this work, we present a simple-yet-effective self-supervised node representation learning strategy via directly maximizing the mutual information between the hidden representations of nodes and their neighbourhood, which can be theoretically justified by its link to graph smoothing. Following InfoNCE, our framework is optimized via a surrogate contrastive loss, where the positive selection underpins the quality and efficiency of rep-resentation learning. To this end, we propose a topology-aware positive sampling strategy, which samples positives from the neighbourhood by considering the structural dependencies between nodes and thus enables positive selection upfront. In the extreme case when only one positive is sampled, we fully avoid expensive neighbourhood aggregation. Our methods achieve promising performance on various node classification datasets. It is also worth mentioning by applying our loss function to MLP based node encoders, our methods can be orders of faster than existing solutions. Our codes and supplementary materials are available at https://github.com/dongwei156/n2n. Wei Dong 0010, Junsheng Wu, ZongYuan Ge, Peng Wang 0023 |
CVPR | 1 |
| 2022 | Improving performance and efficiency of Graph Neural Networks by injective aggregation
Wei Dong 0010, Junsheng Wu, Xinwan Zhang, Zongwen Bai, Peng Wang 0023, Marcin Wozniak |
Knowl. Based Syst. | 1 |
| 2021 | MobileGCN applied to low-dimensional node feature learning
Wei Dong 0010, Junsheng Wu, Zongwen Bai, Yaoqi Hu, Weigang Li 0005, Marcin Wozniak |
Pattern Recognit. | 1 |
| 2020 | Design of affinity-aware encoding by embedding graph centrality for graph classification
Wei Dong 0010, Junsheng Wu, Zongwen Bai, Weigang Li 0005 |
Neurocomputing | 1 |