EDBT 2026 Demo / reviewers in the wild / expert
Xinxin Zhang 0004
dblp:68/1939-4
· DBLP profile ↗
25ranked-venue papers
10as first author
19since 2021 · last 2026
0000-0001-6069-5391ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 10 first-author · 12 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploiting dynamic spatio-temporal correlations for origin-destination demand predictionabstractAccurate Origin-Destination (OD) demand prediction is fundamental to intelligent transportation systems (ITS), enabling real-time traffic management, dynamic vehicle dispatch, and efficient resource allocation in urban environments. However, OD demand exhibits complex, dynamic, and highly coupled spatio-temporal patterns that remain challenging for existing models. We propose a novel Dynamic Spatio-Temporal Correlation Network (DSTCN) for OD demand forecasting. DSTCN features three key components: (1) a bidirectional demand trend modeling module (Glstm2D) that jointly learns demand evolution from both origin and destination perspectives; (2) a Transformer-based spatial similarity module (Simformer) to dynamically extract and integrate inter-regional correlations across the OD matrix; and (3) a temporal fusion and modeling module (FF-TM) that combines and processes multi-source spatio-temporal features for next-step prediction. Extensive experiments on three large-scale real-world datasets (NYC-TOD2018, NYC-TOD2019, and HZMetro) demonstrate that DSTCN consistently outperforms state-of-the-art baselines across diverse urban scenarios. Yongshun Gong, Piao Yu, Xu Zhang 0039, Xinxin Zhang 0004, Xiushan Nie, Haoliang Sun |
Expert Syst. Appl. | 4 |
| 2025 | Disparity-Guided Cross-View Transformer For Stereo Image Super-ResolutionabstractAlthough transformer-based methods excel in stereo image super-resolution, the full potential of the distinctive, complementary information inherent in stereo images has not been fully utilized. We propose a Disparity-Guided Cross-View Transformer (DCT) to extract features across dimensions and views, achieving a more comprehensive feature representation. The proposed method introduces mutual attention within the transformer architecture, establishing the difference between left and right views through cross-view interaction. The proposed algorithm effectively harnesses the complementary information present in stereo image pairs, enhancing the restoration performance. Furthermore, we propose a disparity-guided cross-modal residual fusion module that leverages disparity information as prior knowledge to substantially improve image reconstruction. This module significantly complements the missing information in stereo images, enabling the network to comprehend more effectively and reconstruct the image content with greater accuracy. Extensive experimental results and ablation studies demonstrate the effectiveness of our method. Bingting Li, Wenjing Shang, Yongshun Gong, Qiangchang Wang, Xinxin Zhang 0004, Yilong Yin |
ICASSP | 5 |
| 2025 | LOFI: Harnessing Attention Dynamics for Facial Expression Recognition with Noisy LabelsabstractFacial expression recognition (FER) faces unique challenges from expression ambiguity and noisy labels, degrading performance in real-world applications. While leveraging attention, existing methods frequently neglect attention dynamic mechanism of dispersion followed by focus and the spatially structural knowledge essential for effectively guiding this dynamic dispersion of attention. To address this, we propose the Last fOcus First dIsperse (LOFI), which harnesses attention dynamics and spatial structure information dynamically refining focus during classification to mitigate the impact of noise labels. LOFI comprises two modules: Spatial Keypoint-enhanced Fused Attention (SKFA), which disperses focus on subtle, critical features near facial landmarks, and Hybrid Consistency-Calibrated Loss (HCCL), which employs consistency and re-weighting strategies focusing attention to boost performance. The synergy between these modules enables LOFI to adapt to various noise levels and challenging classes. Extensive experiments demonstrate that LOFI outperforms existing state-of-the-art (SOTA) methods in noisy FER, offering a robust solution for real-world applications. Qiangchang Wang, Xinxin Zhang 0004, Yilong Yin |
ICASSP | 3 |
| 2025 | Improving Generalization in Meta-Learning via Meta-Gradient AugmentationabstractMeta-learning methods typically follow a two-loop framework, where each loop potentially suffers from notorious overfitting, hindering rapid adaptation and generalization to new tasks. Existing methods address this by enhancing the mutual-exclusivity or diversity of training samples, but these data manipulation strategies are data-dependent and insufficiently flexible. This work proposes a data-independent Meta-Gradient Augmentation (MGAug) method from the perspective of gradient regularization. The key idea is first to break the rote memories by network pruning to address memorization overfitting in the inner loop, then use the gradients of pruned sub-networks to augment meta-gradients, alleviating overfitting in the outer loop. Specifically, we explore three pruning strategies, including random width pruning, random parameter pruning, and a newly proposed catfish pruning that measures a Meta-Memorization Carrying Amount (MMCA) score for each parameter and prunes high-score ones to break rote memories. The proposed MGAug is theoretically guaranteed by the generalization bound from the PAC-Bayes framework. Extensive experiments on multiple few-shot learning benchmarks validate MGAug's effectiveness and significant improvement over various meta-baselines. Ren Wang 0011, Haoliang Sun, Yuxiu Lin, Xinxin Zhang 0004, Yilong Yin |
IJCAI | 4 |
| 2025 | Enhancing origin-destination flow prediction via bi-directional spatio-temporal inference and interconnected feature evolution
Piao Yu, Xu Zhang 0039, Yongshun Gong, Jian Zhang 0002, Haoliang Sun, Junjie Zhang 0002, Xinxin Zhang 0004, Yilong Yin |
Expert Syst. Appl. | 7 |
| 2025 | Diverse Information Aggregation with Adaptive Graph Construction and prompts for deepfake detection
Zhenhua Bai, Qiangchang Wang, Lu Yang 0005, Xinxin Zhang 0004, Yanbo Gao, Yilong Yin |
Image Vis. Comput. | 4 |
| 2025 | PartialST: partial spatial-temporal learning for urban flow prediction
Xinxin Zhang 0004, Yongshun Gong |
Neural Comput. Appl. | 3 |
| 2025 | Learning Better Video Query With SAM for Video Instance SegmentationabstractRecently, Transformer-based offline video instance segmentation (VIS) solutions have made significant progress by decomposing the whole task into global segmentation map generation and instance discrimination. We argue that the quality of video queries that represent all instances in a video clip is crucial for offline VIS methods. Existing methods typically interact video queries with dense spatio-temporal features, resulting in significant computational complexity and redundant information. Thus, we propose a novel video instance segmentation framework, LBVQ, dedicated to learning better video queries. Specifically, we first obtain the frame queries for each frame independently without any complex inter-frame spatial-temporal association operations. Secondly, we propose an adaptive query initialization module (AQI), which adaptively integrates frame queries to initialize video queries instead of traditional random initialization strategies. This initialization method preserves rich instance clues and accelerates the optimization of the whole model. Finally, to enhance the quality of video queries, we propose a query propagation module (QPM) that captures relevant instance information in frame queries frame by frame, greatly improving the model’s understanding of long videos. By learning higher quality video queries, LBVQ achieves the state-of-the-art on VIS benchmarks with a ResNet-50 backbone: 52.2 AP, 44.8 AP on YouTube-VIS 2019 & 2021. Moreover, LBVQ achieves 39.7 AP on YouTube-VIS 2022 and 22.2 AP on OVIS, demonstrating superior potential for long videos. To further improve the quality of segmentation masks, a large-scale pretrained SAM is employed to refine the segmentation results. Code is available at https://github.com/fanghaook/LBVQ. Hao Fang 0010, Xiaofei Zhou 0003, Xinxin Zhang 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Fine-Grained Urban Flow Inference with Dynamic Multi-scale Representation Learning
Shilu Yuan, Wei Liu 0007, Xinxin Zhang 0004, Meng Chen 0003, Junjie Zhang 0002, Yongshun Gong |
DASFAA (2) | 4 |
| 2024 | Unified Embedding Alignment for Open-Vocabulary Video Instance Segmentation
Hao Fang 0010, Peng Wu 0014, Yawei Li 0001, Xinxin Zhang 0004, Xiankai Lu |
ECCV (70) | 4 |
| 2024 | Two-phase Parametric Registration for Retinal ImagesabstractWe propose a two-phase parametric registration algorithm for retinal images. Our algorithm focuses on dealing with the geometric transformation and the intensity transformation in the retinal image registration problem. In the first phase, we efficiently detect only one pair of feature points in the source and the target retinal images to estimate a translation transformation and get a warped source image. In the second phase, we estimate both the intensity and the geometric transformations between the target image and the warped source image by fitting parametric expressions. The displacement field is generated by a super fast and accurate coarse-to-fine elastic registration algorithm—local all-pass filters algorithm (LAP). At each iteration of the LAP, the elastic displacement field and the intensity difference take turns being fitted by two different low-order polynomial functions. The fitting steps are performed by solving linear systems of equations efficiently. Experiments on real retinal image datasets demonstrated the high accuracy and computational efficiency of the proposed retinal image registration method. Xinxin Zhang 0004, Xiankai Lu, Jizhou Li, Yongshun Gong, Qiangchang Wang, Yilong Yin |
ICME | 1 |
| 2024 | Cascaded Cross-modal Alignment for Visible-Infrared Person Re-Identification
Qiangchang Wang, Xinxin Zhang 0004, Yilong Yin |
Knowl. Based Syst. | 4 |
| 2024 | Spatio-Temporal Multi-Image Reflection RemovalabstractIn this letter, we propose a precise algorithm to eliminate reflections from two images by utilizing temporal and spatial priors. For the temporal prior, we compute the motion information between reflection layers in the two input reflection-contaminated images. Different from numerous popular multi-image reflection removal methods, our proposed algorithm does not assume that two input images are captured under similar lighting conditions and the same camera settings. Furthermore, the proposed algorithm is robust to the difference between the two reflection layers, such as moving objects and different reflections. For the spatial term, a sparsity gradient regularization is adopted to enforce the spatial smoothness of transmission layers and reflection layers. Importantly, the proposed algorithm does not rely on additional training data or high-performance computing devices. Experimental results on both synthetic images and real-world photographs demonstrate that the proposed algorithm achieves State-of-the-Art performance. Xinxin Zhang 0004, Wenjing Shang, Qiangchang Wang, Yongshun Gong |
IEEE Signal Process. Lett. | 1 |
| 2024 | Focusing on Subtle Differences: A Feature Disentanglement Model for Series Photo SelectionabstractNowadays, capturing cherished moments results in an abundance of photos, which necessitates the selection of the finest one from a pool of akin images—a process both intricate and time-intensive. Thus, series photo selection (SPS) techniques have been developed to recommend the optimal moment from nearly identical photos through the use of aesthetic quality assessment. However, addressing SPS proves demanding due to the subtle nuances within such imagery. Existing approaches predominantly rely on diverse feature types (e.g., color, layout, generic features) extracted from original images to discern the qualified shot, yet they disregard disentangling generality and specificity at the feature level. This study aims to detect subtle aesthetic distinctions among akin photos. We propose a feature separation model that captures all label-relevant information through an encoder. We introduce Information Bottleneck (IB) learning to obtain non-redundant representations of image pairs and filter out noise information from the representations. Our model segregates image features into shared and specific attributes by employing feature constraints to boost mutual information across images and guide meaningful information within individual images. This process filters out extraneous data within individual images, thus significantly enhancing the representation of similar image pairs. Extensive experiments on the Phototriage dataset show that our model can accentuate subtle disparities and achieve superior results when compared to alternative methods. Yongshun Gong, Xinxin Zhang 0004, Jian Zhang 0002, Yilong Yin |
IEEE Trans. Multim. | 4 |
| 2023 | Mask- and Contrast-Enhanced Spatio-Temporal Learning for Urban Flow PredictionabstractAs a critical mission of intelligent transportation systems, urban flow prediction (UFP) benefits in many city services including trip planning, congestion control, and public safety. Despite the achievements of previous studies, limited efforts have been observed on simultaneous investigation of the heterogeneity in both space and time aspects. That is, regional correlations would be variable at different timestamps. In this paper, we propose a spatio-temporal learning framework with mask and contrast enhancements to capture spatio-temporal variabilities among city regions. We devise a mask-enhanced pre-training task to learn latent correlations across the spatial and temporal dimensions, and then a graph-based method is developed to extract the significance of regions by using the inter-regional attention weights. To further acquire contrastive correlations of regions, we elaborate a pre-trained contrastive learning task with the global-local cross-attention mechanism. Thereafter, two well-trained encoders have strong capability to capture latent spatio-temporal representations for the flow forecasting with time-varying. Extensive experiments conducted on real-world urban flow datasets demonstrate that our method compares favorably with other state-of-the-art models. Xu Zhang 0039, Yongshun Gong, Xinxin Zhang 0004, Chengqi Zhang, Xiangjun Dong 0001 |
CIKM | 3 |
| 2023 | Personalized Single Image Reflection Removal Network through Adaptive Cascade RefinementabstractIn this paper, we aim to restore a reflection-free image from a single reflection-contaminated image captured through the glass. Many deep-learning-based methods attempt to solve the challenging problem by utilizing a uniform model obtained from training data for all test images. Hence, the distinctive characteristics of the test images are not considered. Besides, several methods use a cascade structure in image restoration to refine the results. But they blindly cascade modules with the same weights, improving the model's performance only to a certain extent. To address these problems, we propose a personalized single-image reflection removal network through adaptive cascade refinement (PNACR) based on meta-learning and self-supervised learning. While meta-learning can rapidly adapt to a new task with a few samples, PNACR can remove reflections of a new image with its distinctive characteristics learned by self-supervised learning. Furthermore, the proposed adaptive cascade model can adjust the weights of the model at the next iteration according to the output of the model at the current iteration, significantly improving the model's performance. Hence, the proposed model can learn information from both external training data and the new input image to provide a personalized reflection removal model for each new input image. Extensive comparison and ablation experiments on publicly available datasets demonstrate the validity of the proposed method in quantitative evaluation metrics and qualitative visualization. Mengyi Wang 0001, Xinxin Zhang 0004, Yongshun Gong, Yilong Yin |
ACM Multimedia | 2 |
| 2023 | Single Image Reflection Removal Based on Dark Channel Sparsity PriorabstractThe major task of reflection removal methods is to restore a reflection-free image from a reflection-contaminated image taken through glass. We propose an algorithm to remove reflections from a single image by means of the$l_{0}$-regularized dark channel sparsity prior and an$l_{0}$gradient sparsity prior. In addition, we analyze the difference between the dark channel map in the reflection-contaminated image and the reflection-free image empirically and mathematically. Moreover, a new data fidelity term is introduced to handle strong reflections and preserve high-frequency details in the recovered transmission image. Different from the model used in most state-of-the-art methods, our reflection removal model does not rely on the assumption of out-of-focus objects in the reflection layer. Quantitative evaluation on several publicly available real-world image datasets including ground-truth demonstrates the high accuracy of our algorithm. Qualitative evaluation of extensive experimental results on real-world images shows the competitive performance of the proposed method compared with the state-of-the-art reflection removal methods. Xinxin Zhang 0004, Kaixin Xing, Da Chen 0002, Yilong Yin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Geodesic Paths for Image Segmentation With Implicit Region-Based Homogeneity EnhancementabstractMinimal paths are regarded as a powerful and efficient tool for boundary detection and image segmentation due to its global optimality and the well-established numerical solutions such as fast marching method. In this paper, we introduce a flexible interactive image segmentation model based on the Eikonal partial differential equation (PDE) framework in conjunction with region-based homogeneity enhancement. A key ingredient in the introduced model is the construction of local geodesic metrics, which are capable of integrating anisotropic and asymmetric edge features, implicit region-based homogeneity features and/or curvature regularization. The incorporation of the region-based homogeneity features into the metrics considered relies on an implicit representation of these features, which is one of the contributions of this work. Moreover, we also introduce a way to build simple closed contours as the concatenation of two disjoint open curves. Experimental results prove that the proposed model indeed outperforms state-of-the-art minimal paths-based image segmentation approaches. Da Chen 0002, Xinxin Zhang 0004, Minglei Shu, Laurent D. Cohen |
IEEE Trans. Image Process. | 3 |
| 2021 | Handling Outliers by Robust M-Estimation in Blind Image DeblurringabstractThe major task of traditional motion deblurring methods is to estimate the blur kernel and restore the latent image. In low-light conditions, the pointolite is likely to produce saturated light streaks in captured blurred images. The light streaks are usually double-edged swords—outliers to the deconvolution, but a cue to kernel estimation. In this paper, we propose a novel blind motion deblurring method for blurred images including light streaks. The main idea is to model the non-linear blur caused by outliers as the Huber's M-estimation in blind deconvolution and take the shape of the light streak as a cue to estimate the blur kernel. Specifically, the optimal light streak patch is selected automatically according to the characteristics of light streaks and the blur kernel. This simple yet effective selection strategy solves the problems of false detection of candidate light streaks and optimal light streak in existing methods. Then, the optimal light streak patch is parameterized as a prior and is combined with other regularizers to estimate the blur kernel. Compared with the state-of-the-art kernel estimation methods, the proposed algorithm reduces the influence of outliers on deconvolution and utilizes more information. Thus, the restored image is more accurate. Experimental results on both synthetic and real images demonstrate the high accuracy of our algorithm. Xinxin Zhang 0004, Ronggang Wang, Da Chen 0002, Yang Zhao 0002, Wen Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | All-Pass Parametric Image RegistrationabstractImage registration is a required step in many practical applications that involve the acquisition of multiple related images. In this paper, we propose a methodology to deal with both the geometric and intensity transformations in the image registration problem. The main idea is to modify an accurate and fast elastic registration algorithm (Local All-Pass-LAP) so that it returns a parametric displacement field, and to estimate the intensity changes by fitting another parametric expression. Although we demonstrate the methodology using a low-order parametric model, our approach is highly flexible and easily allows substantially richer parametrisations, while requiring only limited extra computation cost. In addition, we propose two novel quantitative criteria to evaluate the accuracy of the alignment of two images ("salience correlation") and the number of degrees of freedom ("parsimony") of a displacement field, respectively. Experimental results on both synthetic and real images demonstrate the high accuracy and computational efficiency of our methodology. Furthermore, we demonstrate that the resulting displacement fields are more parsimonious than the ones obtained in other state-of-the-art image registration approaches. Xinxin Zhang 0004, Christopher Gilliam, Thierry Blu |
IEEE Trans. Image Process. | 1 |
| 2019 | Parametric Registration for Mobile Phone ImagesabstractImage registration is a significant step in a wide range of practical applications and it is a fundamental problem in various computer vision tasks. In this paper, we propose a highly accurate and fast parametric registration method for mobile phone photos. The proposed algorithm is based on a fast and accurate elastic registration algorithm, the Local All-Pass (LAP) algorithm, which performs in a coarse-to-fine manner. At each iteration, the LAP displacement field is fitted by a parametric model. Thus the image registration problem is equivalent to finding a few parameters to describe the displacement field. The fitting step can be performed very efficiently by solving a linear system of equations. In terms of the fitting model, it is easy to change the type of models to do the parametric fitting for specific applications. Experimental results on both synthetic and real images demonstrate the high accuracy and computational efficiency of the proposed algorithm. Xinxin Zhang 0004, Christopher Gilliam, Thierry Blu |
ICIP | 1 |
| 2017 | Iterative fitting after elastic registration: An efficient strategy for accurate estimation of parametric deformationsabstractWe propose an efficient method for image registration based on iteratively fitting a parametric model to the output of an elastic registration. It combines the flexibility of elastic registration — able to estimate complex deformations — with the robustness of parametric registration — able to estimate very large displacement. Our approach is made feasible by using the recent Local All-Pass (LAP) algorithm; a fast and accurate filter-based method for estimating the local deformation between two images. Moreover, at each iteration we fit a linear parametric model to the local deformation which is equivalent to solving a linear system of equations (very fast and efficient). We use a quadratic polynomial model however the framework can easily be extended to more complicated models. The significant advantage of the proposed method is its robustness to model mis-match (e.g. noise and blurring). Experimental results on synthetic images and real images demonstrate that the proposed algorithm is highly accurate and outperforms a selection of image registration approaches. Xinxin Zhang 0004, Christopher Gilliam, Thierry Blu |
ICIP | 1 |
| 2017 | Iterative fitting after elastic registration: An efficient strategy for accurate estimation of parametric deformationsabstractVideo phylogeny research about joint analysis of correlated video sequences has shown the possibility of developing interesting forensic applications. As an example, it is possible to study the provenance of near-duplicate (ND) video sequences, i.e., videos generated from the same original one through content preserving transformations. To perform this kind of analysis, accurate detection of ND videos is paramount. In this paper, we propose an algorithm for ND video detection and clustering in a challenging setup. Specifically, we analyze a scenario in which many videos, depicting the same event, are recorded by different users. This situation is critical as non-ND videos acquired from very close viewpoints run the risk of being incorrectly detected as ND. The proposed approach leverages on robust hashing properties and the concept of sensor noise traces. Xinxin Zhang 0004, Christopher Gilliam, Thierry Blu |
ICIP | 1 |
| 2016 | Spatially variant defocus blur map estimation and deblurring from a single image
Xinxin Zhang 0004, Ronggang Wang, Xiubao Jiang, Wenmin Wang 0001, Wen Gao 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2015 | Image deblurring using robust sparsity priorsabstractIn this paper, we propose a robust method to remove motion blur from a single photograph. We find that an inaccurate kernel and an unreliable final latent image reconstruction method are two main factors leading to low-quality restored images. To improve image quality, we do the following technical contributions. For robust blur kernel estimation, first, an edge mask and a smooth constraint are used to provide reliable intermediate latent images for salient structure extraction; second, we adopt an effective salient structure selection method to remove detrimental edges for kernel estimation; third, we use a gradient sparsity prior to remove kernel noise and ensure the continuity of blur kernels. For final latent image reconstruction, we combine the merits of both the TV-l2model and the hyper-Laplacian model to preserve tiny details and eliminate noise. Experimental results on synthetically blurred images and real photographs demonstrate that the proposed algorithm performs better than state-of-the-art approaches. Xinxin Zhang 0004, Ronggang Wang, Yonghong Tian 0001, Wenmin Wang 0001, Wen Gao 0001 |
ICIP | 1 |