VLDB 2026 Research / reviewers in the wild / expert
Dafeng Zhang
dblp:29/8349
· DBLP profile ↗
10ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Noise Resilience in Face Clustering via Sparse Differential TransformerabstractThe method used to measure relationships between face embeddings plays a crucial role in determining the performance of face clustering. Existing methods employ the Jaccard similarity coefficient instead of the traditional cosine distance to enhance the measurement accuracy. However, these methods introduce an excessive number of irrelevant nodes, producing Jaccard coefficients with limited discriminative power and adversely affecting clustering performance. To address this issue, we propose a prediction-driven Top-K Jaccard similarity coefficient that enhances the purity of neighboring nodes, thereby improving the reliability of similarity measurements. Nevertheless, accurately predicting the optimal number of neighbors (Top-K) remains challenging, leading to suboptimal clustering results. To overcome this limitation, we develop a Transformer-based prediction model that examines the relationships between the central node and its neighboring nodes near the Top-K to further enhance the reliability of similarity estimation. However, vanilla Transformer, when applied to predict relationships between nodes, often introduces noise due to their overemphasis on irrelevant feature relationships. To address these challenges, we propose a Sparse Differential Transformer (SDT), instead of the vanilla Transformer, to eliminate noise and enhance the model's anti-noise capabilities. Extensive experiments on multiple datasets, such as MS-Celeb-1M, demonstrate that our approach achieves state-of-the-art (SOTA) performance, outperforming existing methods and providing a more robust solution for face clustering. Dafeng Zhang, Yongqi Song, Shizhuo Liu |
AAAI | 1 |
| 2025 | Find Details in Long Videos: Tower-of-Thoughts and Self-Retrieval Augmented Generation for Video UnderstandingabstractThe Large Vision-Language Model (LVLM) has achieved impressive performance in the field of visual-language understanding. However, its ability to understand longer videos is still limited due to the length and information diversity of multi-modal videos. Moreover, accurately matching detailed content within videos remains an open research problem. We design a new framework for LVLM inference, "Tower of Thoughts" (ToT), which extends the "Chain-of-Thought" (CoT) approach to the visual domain and constructs the high-dimensional semantics of the complete videos from the bottom up. Meanwhile, to achieve question-answering for video details within the constraints of the restricted context window, we propose a method of self-retrieval augmented generation (SRAG), which makes it possible to obtain details from long videos by storing and accessing video text as dense vectors in non-parametric memory. The solution of combining the ToT with SRAG enables our model to have cross-modal high-density semantic fusion and comprehensive and accurate generation capabilities, thereby achieving rationalized video answers. Experiments on public benchmarks demonstrate the effectiveness of our proposed method. In addition, we also conducted experiments on multi-modal long videos in the open world and achieved remarkable outcomes. These results provide new perspectives and technical routes for the future development of visual language models. Tong Yue, Mingrui Xiao, Dafeng Zhang, Yali Li 0001, Shengjin Wang |
ICASSP | 3 |
| 2025 | MambaNext: An Enhanced Backbone Network with Focus Linear AttentionabstractIn response to the limitations of current linear attention models, such as Vision Mamba, which fail to mimic the human visual system’s ability to focus on objects and then shift attention to the surrounding context when ambiguity arises, we introduce the MambaNext model. This novel backbone network incorporates two key innovations: the Focus Linear Attention Module (FLAM) and the Star Fusion strategy. FLAM is designed to enhance object recognition by emulating the human visual system’s focused center, thereby reducing the interference from background elements. On the other hand, Star Fusion acts as a unique operation that utilizes global information as an attention map to guide local information toward more relevant features. Additionally, it also implicitly increases the feature dimensions to improve linear separability. The experimental results demonstrate that our MambaNext has achieved state-of-the-art performance across multiple computer vision tasks including classification, detection, and segmentation, outperforming existing Vision Mamba methods. Dafeng Zhang, Shizhuo Liu |
ICASSP | 1 |
| 2025 | From Cradle to Cane: A Two-Pass Framework for High-Fidelity Lifespan Face AgingabstractFace aging has become a crucial task in computer vision, with applications ranging from entertainment to healthcare. However, existing methods struggle with achieving a realistic and seamless transformation across the entire lifespan, especially when handling large age gaps or extreme head poses. The core challenge lies in balancing $age\ accuracy$ and $identity\ preservation$—what we refer to as the $Age\text{-}ID\ trade\text{-}off$. Most prior methods either prioritize age transformation at the expense of identity consistency or vice versa. In this work, we address this issue by proposing a $two\text{-}pass$ face aging framework, named $Cradle2Cane$, based on few-step text-to-image (T2I) diffusion models. The first pass focuses on solving $age\ accuracy$ by introducing an adaptive noise injection ($AdaNI$) mechanism. This mechanism is guided by including prompt descriptions of age and gender for the given person as the textual condition.
Also, by adjusting the noise level, we can control the strength of aging while allowing more flexibility in transforming the face.
However, identity preservation is weakly ensured here to facilitate stronger age transformations.
In the second pass, we enhance $identity\ preservation$ while maintaining age-specific features by conditioning the model on two identity-aware embeddings ($IDEmb$): $SVR\text{-}ArcFace$ and $Rotate\text{-}CLIP$. This pass allows for denoising the transformed image from the first pass, ensuring stronger identity preservation without compromising the aging accuracy.
Both passes are $jointly\ trained\ in\ an\ end\text{-}to\text{-}end\ way\$. Extensive experiments on the CelebA-HQ test dataset, evaluated through Face++ and Qwen-VL protocols, show that our $Cradle2Cane$ outperforms existing face aging methods in age accuracy and identity consistency.
Additionally, $Cradle2Cane$ demonstrates superior robustness when applied to in-the-wild human face images, where prior methods often fail. This significantly broadens its applicability to more diverse and unconstrained real-world scenarios. Code is available at https://github.com/byliutao/Cradle2Cane. Dafeng Zhang, Gengchen Li, Shizhuo Liu, Yongqi Song, Senmao Li, Shiqi Yang 0002, Boqian Li, Kai Wang 0060, Yaxing Wang |
NeurIPS | 2 |
| 2025 | DIFFSSR: Stereo Image Super-resolution Using Differential TransformerabstractIn the field of computer vision, the task of stereo image super-resolution (StereoSR) has garnered significant attention due to its potential applications in augmented reality, virtual reality, and autonomous driving. Traditional Transformer-based models, while powerful, often suffer from attention noise, leading to suboptimal reconstruction issues in super-resolved images. This paper introduces DIFFSSR, a novel neural network architecture designed to address these challenges. We introduce the Diff Cross Attention Block (DCAB) and the Sliding Stereo Cross-Attention Module (SSCAM) to enhance feature integration and mitigate the impact of attention noise. The DCAB differentiates between relevant and irrelevant context, amplifying attention to important features and canceling out noise. The SSCAM, with its sliding window mechanism and disparity-based attention, adapts to local variations in stereo images, preserving details, and addressing the performance degradation due to misalignment of horizontal epipolar lines in stereo images. Extensive experiments on benchmark datasets demonstrate that DIFFSSR outperforms state-of-the-art methods, including NAFSSR and SwinFIRSSR, in terms of both quantitative metrics and visual quality. Dafeng Zhang |
NeurIPS | 1 |
| 2023 | Context-Aware Face Clustering with Graph Convolutional NetworksabstractFace clustering is a necessary tool in the field of face-related algorithm research, which is widely used in album management and unlabeled data management. Recent works which use Graph Convolution Network (GCN) to extract the global features have achieved impressive results in the face clustering task. However, these works have a main drawback that they ignore the influence of the local features. In this paper, we propose a Context-Aware Graph Convolutional Network (CAGCN) to explicitly consider both the global and local information. We also propose a deduplication algorithm based on the Jaccard Similarity to improve the efficiency of face clustering. Experiments show that the proposed method can improve the integrity of the feature representations and the robustness of the clustering algorithm. We applied our algorithm on three popular large-scale benchmarks and achieved state-of-the-art performance comparing to the existing methods. Dafeng Zhang, Jiangbo Guo, Zhezhu Jin |
ICASSP | 1 |
| 2023 | MRNET: Multi-Refinement Network for Dual-Pixel Images Defocus DeblurringabstractDefocus blurring is an inevitable phenomenon in cameras. Though many methods have been proposed, the problem is still challenging because of their low deblurring performance and long processing time. To solve this problem, we propose an efficient Multi-Refinement Network (MRNet) for dual-pixel images defocus deblurring. The MRNet contains two core modules that are alignment module and reconstruction module, respectively. We design a Siamese Pyramid Network (SPN) as alignment module to alleviate the misalignment problem of left and right views. At the same time, a Multi-Scale Residuals Group Module (MSRGM) is proposed in the reconstruction module, which can extract and fuse features from different scales to obtain better deblurring performance. Specifically, the reconstruction module is composed of multiple MSRGM modules, and each MSRGM is a refinement of the previous one, which is our leitmotif - Multi-Refinement. Experimental results on the popular benchmarks show that the proposed method can significantly improve the performance of defocus deblurring. Dafeng Zhang, Zhezhu Jin |
ICASSP | 1 |
| 2023 | CLIP-Cluster: CLIP-Guided Attribute Hallucination for Face ClusteringabstractOne of the most important yet rarely studied challenges for supervised face clustering is the large intra-class variance caused by different face attributes such as age, pose, and expression. Images of the same identity but with different face attributes usually tend to be clustered into different sub-clusters. For the first time, we proposed an attribute hallucination framework named CLIP-Cluster to address this issue, which first hallucinates multiple representations for different attributes with the powerful CLIP model and then pools them by learning neighbor-adaptive attention. Specifically, CLIP-Cluster first introduces a text-driven attribute hallucination module, which allows one to use natural language as the interface to hallucinate novel attributes for a given face image based on the well-aligned image-language CLIP space. Furthermore, we develop a neighbor-aware proxy generator that fuses the features describing various attributes into a proxy feature to build a bridge among different sub-clusters and reduce the intra-class variance. The proxy feature is generated by adaptively attending to the hallucinated visual features and the source one based on the local neighbor information. On this basis, a graph built with the proxy representations is used for subsequent clustering operations. Extensive experiments show our proposed approach outperforms state-of-the-art face clustering methods with high inference efficiency. Shuai Shen, Wanhua Li 0001, Dafeng Zhang, Zhezhu Jin, Jie Zhou 0001, Jiwen Lu |
ICCV | 4 |
| 2023 | Graph Convolutional Network-Based Rumor Blocking on Social NetworksabstractMisinformation and rumors can spread rapidly and widely through online social networks, seriously endangering social stability. Therefore, rumor blocking on social networks has become a hot research topic. In the existing research, when users receive two opposing opinions, they tend to believe the one arrives first. In this article, we argue that users will dialectically trust the information based on their own opinions rather than the rule of first-come-first-listen. We propose a confidence-based opinion adoption (CBOA) model, which considers the opinion and confidence according to the traditional linear threshold (LT) model. Based on this model, we propose the directed graph convolutional network (DGCN) method to select the$k$most influential positive cascade nodes to suppress the propagation of rumors. Finally, we verify our method on four real network datasets. The experimental results show that our method can sufficiently suppress the propagation of rumors and obtains smaller number of rumor nodes than the baseline algorithms. Qiang He 0002, Dafeng Zhang, Xingwei Wang 0001, Lianbo Ma 0004, Min Huang 0001 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2022 | Dynamic Multi-Scale Network for Dual-Pixel Images Defocus Deblurring with TransformerabstractRecent works achieve excellent results in dual-pixel defo-cus deblurring task by using convolutional neural network (CNN), while the scarcity of data limits the exploration and attempt of vision transformer in this task. In this paper, we propose a dynamic multi-scale network, named DMT-Net, for dual-pixel images defocus deblurring. In DMTNet, the feature extraction module is composed of several vision transformer blocks, which uses its powerful feature extraction capability to obtain robust features. The reconstruction module is composed of several Dynamic Multi-scale Sub-reconstruction Module (DMSSRM). DMSSRM restores images by adaptively assigning weights to features from differ-ent scales according to the blur distribution and content in-formation of the input images. DMTNet combines the ad-vantages of transformer and CNN, in which the vision trans-former improves the performance ceiling of CNN, and the inductive bias of CNN enables transformer to extract more robust features without relying on a large amount of data. Experimental results on the popular benchmarks demonstrate that our DMTNet significantly outperforms state-of-the-art methods. Dafeng Zhang |
ICME | 1 |