EDBT 2026 Demo / reviewers in the wild / expert
Xue Zhang 0008
dblp:29/2362-8
· DBLP profile ↗
18ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0002-6579-7845ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pose-Guided Multi-Cue Explicit Query Construction for Disambiguating Human-Object InteractionsabstractHuman-Object Interaction (HOI) detection remains challenging due to the semantic ambiguity of interaction categories and the limited discriminability of their feature representations. Existing approaches often improve recognition by employing sophisticated models or auxiliary textual annotations. While effective in certain gains, these solutions incur additional computational or annotation costs and struggle to capture intrinsic interaction regularities. To address these issues, we propose Pose-Guided Multi-Cue Explicit Query Construction (PM-EQC), a unified Transformer-based framework that builds upon collaborative modeling of appearance, spatial, and pose cues for discriminative interaction reasoning. At its core, the Collaborative Multi-Cue Query Constructor (CM-CQC) jointly models dependencies among visual cues to generate explicit query embeddings. CM-CQC further incorporates a hierarchical pose contextualization mechanism: global body configurations adaptively guide attention to local critical joints, yielding fine-grained pose embeddings and more precise interaction disambiguation. Owing to its modular design, PM-EQC integrates seamlessly with diverse backbones and benefits from their advances. Extensive experiments on PhysLab, HICO-DET, and V-COCO datasets demonstrate that PM-EQC achieves state-of-the-art performance, and the code is publicly available at https://github.com/ZMHSDUST/ PM-EQC. Minghao Zou, Qingtian Zeng, Xue Zhang 0008, Guiyuan Yuan, Xiaoshuai Hao, Jun Liu 0036, Wei Zhou 0021 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | CRViT: Vision transformer advanced by causality and inductive bias for image recognition
Faming Lu, Kunhao Jia, Xue Zhang 0008 |
Appl. Intell. | 3 |
| 2024 | Soft Image Segmentation Using Gradient Graph Laplacian RegularizerabstractWe revisit the well-studied image segmentation problem from a soft labeling perspective: instead of estimating integer labels per pixel indicating a finite set of classes, each pixel is assigned a real number that conveys the level of uncertainty in the estimated class label. Soft labels are useful, for example, for subsequent human editing or composition. Specifically, given a set of pre-computed super-pixel labels and feature vectors per pixel, we formulate a convex optimization objective regularized by signal-dependent gradient graph Laplacian regularizers (GGLR), which promotes piecewise planar (PWP) signal reconstruction. Unlike a previous well-known soft segmentation scheme that requires expensive computation of the first 100 eigenvectors, our optimization can be solved efficiently in linear time via conjugate gradient (CG). Experimental results show that our method produces satisfactory soft labels per pixel for images in two public datasets at a reduced computation cost compared to the previous soft segmentation scheme. Fei Chen 0012, Gene Cheung, Xue Zhang 0008 |
ICASSP | 3 |
| 2024 | Graph-Enhanced Hybrid Sampling for Multi-Armed Bandit RecommendationabstractGraph-based multi-armed bandit algorithms utilize the relationship between users to select the best item to recommend for maximal reward, which is decided by items’ features and un-known users’ preferences. Therefore, the precise estimation of users’ preferences is fairly important and indispensable for bandit sampling, though it is not the ultimate target. However, existing algorithms generally neglect this crucial point and utilize reward maximization as objective in the first beginning, using inaccurate estimation as input, which deteriorates the performance from a long-term perspective. In this paper, we will propose one hybrid sampling framework for bandit selection, which at first purely focuses on the performance of estimation and then on the performance of reward maximization. Specifically, we propose an ‘unsupervised’ bandit selection objective to minimize expected estimation error, which doesn’t take users’ preferences as input and suppresses an approximate upper-bound of cumulative regret. Then, we design a low-complexity selection algorithm to optimize this formulated problem with simple multiplications between items’ features and users’ graphical relations. Subsequently, for reward maximization, we cascade one graph-based algorithm to find the following bandits on the basis of our proposed warm-starts. Extensive experiments on different graphs indicate that our proposed hybrid framework is substantially better than existing popular methods in terms of recommendation performance. Taihao Li, Wuyue Zhang, Xue Zhang 0008, Cheng Yang 0003 |
ICASSP | 4 |
| 2024 | Weakly-Supervised Action Learning in Procedural Task Videos via Process Knowledge DecompositionabstractAction learning is a research area that aims to recognize the action category of each frame in the video. Context information is crucial for learning actions, but most existing methods face two challenges in exploiting this information: 1) They apply global attention to aggregate global features for action representation, resulting in inefficiency and redundancy. 2) They impose implicit action constraints to regularize the action distribution, leading to subjectivity, interpretability issues, and optimization difficulties. To address these challenges, we propose an end-to-end weakly-supervised Action Learning framework with Process Knowledge Decomposition (AL-PKD), which leverages the intrinsic characteristics of procedural task videos. To enhance the effectiveness and adaptability of context aggregation, we first design the TEAL-Net action recognition network. Specifically, the TEAL-Net accounts for the diverse neighbor distributions of action nodes across categories and collects local neighborhood features with different receptive fields through feature pyramids, improving the accuracy and efficiency of action representation. Moreover, to overcome the drawbacks of implicit constraint strategies, we next employ process mining techniques to extract three types of explicit action pair constraints: sequentiality, concurrency, and selectivity. These constraints guide the model’s predictions and improve the interpretability of the learning process. Finally, we use the Viterbi algorithm to dynamically infer the optimal action boundaries based on the frame-level predictions, which helps to eliminate local misclassifications. Experiments on three datasets of Breakfast, CrossTask, and PEVD demonstrate that our method achieves state-of-the-art performance. Minghao Zou, Qingtian Zeng, Xue Zhang 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Revisit Sampling Theory of Bandlimited Graph Signals: One Bridge Between GSP and DSPabstractSampling of bandlimited (BL) graph signals is one fundamental problem in graph signal processing (GSP), whose underlying kernel is an irregular graph rather than regular 1-D time-series kernel in classical discrete signal processing (DSP). Though there were amounts of sampling objectives and algorithms proposed for BL graph signals, the essential relationship between those sampling objectives in GSP and Nyquist sampling theorem in DSP is still undiscovered. In this paper, we bridge this gap by revisiting sampling theory in GSP thoroughly. In specific, we first figure out that the minimal sample size used in GSP for unique recovery can derive exact Nyquist sampling frequency in DSP when eigenvector matrix is discrete Fourier transform (DFT) matrix. Then, we propose a graph sampling objective for directed cyclic graph as one bridge, which leads to uniform sampling pattern in DSP and has the same optimal solution as other popular graph sampling objectives in GSP, thus connecting GSP and DSP closely via sampling. Finally, we present simulations to demonstrate potential applications inspired by our study, such as fast reconstruction of BL graph signals. Taihao Li, Xue Zhang 0008 |
ICASSP | 3 |
| 2023 | Hierarchical Class Level Attribute Guided Generative Meta Learning for Pest Image Zero-shot LearningabstractExisting pest image classification models require a large number of labeled training images. However, labels for most pest images in the real world do not exist. Therefore, the zero-shot learning method based on generative meta-learning provides an effective solution, which first uses attributes to transfer knowledge from seen classes to unseen classes, and then synthesizes the features of unseen classes. We observe that seen and unseen classes share the same high-level attributes, which can be used to learn a shared set of optimal parameters for seen and unseen classes. Therefore, we propose a novel Hierarchical Class level Attribute guided Generative meta model for pest image Zero-shot Learning (HCAG-ZSL). HCAG-ZSL uses the pre-built Taxonomic Attribute Tree to get the high-level attributes corresponding to the class attributes. These attributes are then fed into a well-designed generator to generate visual features. Extensive experiments show that the proposed model outperforms state-of-the-art generative meta models. Shansong Wang, Qingtian Zeng, Weijian Ni, Xue Zhang 0008, Cheng Cheng 0018 |
ICME | 4 |
| 2023 | Multi-modal pseudo-information guided unsupervised deep metric learning for agricultural pest images
Shansong Wang, Qingtian Zeng, Xue Zhang 0008, Weijian Ni, Cheng Cheng 0018 |
Inf. Sci. | 3 |
| 2022 | Graph-Based Depth Denoising & Dequantization for Point Cloud EnhancementabstractA 3D point cloud is typically constructed from depth measurements acquired by sensors at one or more viewpoints. The measurements suffer from both quantization and noise corruption. To improve quality, previous works denoise a point cloud a posteriori after projecting the imperfect depth data onto 3D space. Instead, we enhance depth measurements directly on the sensed images a priori, before synthesizing a 3D point cloud. By enhancing near the physical sensing process, we tailor our optimization to our depth formation model before subsequent processing steps that obscure measurement errors. Specifically, we model depth formation as a combined process of signal-dependent noise addition and non-uniform log-based quantization. The designed model is validated (with parameters fitted) using collected empirical data from a representative depth sensor. To enhance each pixel row in a depth image, we first encode intra-view similarities between available row pixels as edge weights via feature graph learning. We next establish inter-view similarities with another rectified depth image via viewpoint mapping and sparse linear interpolation. This leads to a maximum a posteriori (MAP) graph filtering objective that is convex and differentiable. We minimize the objective efficiently using accelerated gradient descent (AGD), where the optimal step size is approximated via Gershgorin circle theorem (GCT). Experiments show that our method significantly outperformed recent point cloud denoising schemes and state-of-the-art image denoising schemes in two established point cloud quality metrics. Xue Zhang 0008, Gene Cheung, Jiahao Pang, Yash Sanghvi, Abhiram Gnanasambandam, Stanley H. Chan |
IEEE Trans. Image Process. | 1 |
| 2021 | Fast & Robust Image Interpolation Using Gradient Graph Laplacian RegularizerabstractIn the graph signal processing (GSP) literature, it has been shown that signal-dependent graph Laplacian regularizer (GLR) can efficiently promote piecewise constant (PWC) signal reconstruction for various image restoration tasks. However, for planar image patches, like total variation (TV), GLR may suffer from the well-known “staircase” effect. To remedy this problem, we generalize GLR to gradient graph Laplacian regularizer (GGLR) that provably promotes piecewise planar (PWP) signal reconstruction for the image interpolation problem—a 2D grid with random missing pixels that requires completion. Specifically, we first construct two higher-order gradient graphs to connect local horizontal and vertical gradients. Each local gradient is estimated using structure tensor, which is robust using known pixels in a small neighborhood, mitigating the problem of larger noise variance when computing gradient of gradients. Moreover, unlike total generalized variation (TGV), GGLR retains the quadratic form of GLR, leading to an unconstrained quadratic programming (QP) problem per iteration that can be solved quickly using conjugate gradient (CG). We derive the means-square-error minimizing weight parameter for GGLR, trading off bias and variance of the signal estimate. Experiments show that GGLR outperformed competing schemes in interpolation quality for severely damaged images at a reduced complexity. Fei Chen 0012, Gene Cheung, Xue Zhang 0008 |
ICIP | 3 |
| 2021 | Graph Learning Based Head Movement Prediction for Interactive 360 Video StreamingabstractUltra-high definition (UHD) 360 videos encoded in fine quality are typically too large to stream in its entirety over bandwidth (BW)-constrained networks. One popular approach is to interactively extract and send a spatial sub-region corresponding to a viewer's current field-of-view (FoV) in a head-mounted display (HMD) for more BW-efficient streaming. Due to the non-negligible round-trip-time (RTT) delay between server and client, accurate head movement prediction foretelling a viewer's future FoVs is essential. In this paper, we cast the head movement prediction task as a sparse directed graph learning problem: three sources of relevant information-collected viewers' head movement traces, a 360 image saliency map, and a biological human head model-are distilled into a view transition Markov model. Specifically, we formulate a constrained maximum a posteriori (MAP) problem with likelihood and prior terms defined using the three information sources. We solve the MAP problem alternately using a hybrid iterative reweighted least square (IRLS) and Frank-Wolfe (FW) optimization strategy. In each FW iteration, a linear program (LP) is solved, whose runtime is reduced thanks to warm start initialization. Having estimated a Markov model from data, we employ it to optimize a tile-based 360 video streaming system. Extensive experiments show that our head movement prediction scheme noticeably outperformed existing proposals, and our optimized tile-based streaming scheme outperformed competitors in rate-distortion performance. Xue Zhang 0008, Gene Cheung, Yao Zhao 0001, Patrick Le Callet, Chunyu Lin, Jack Z. G. Tan |
IEEE Trans. Image Process. | 1 |
| 2020 | Sparse Directed Graph Learning for Head Movement Prediction in 360 Video StreamingabstractHigh-definition 360 videos encoded in fine quality are typically too large in size to stream in its entirety over bandwidth (BW)-constrained networks. One popular remedy is to interactively extract and send a spatial sub-region corresponding to a viewer's current field-of-view (FoV) in a head-mounted display (HMD) for more BW-efficient streaming. Due to the non-negligible round-trip-time (RTT) delay between server and client, accurate head movement prediction that foretells a viewer's future FoVs is essential. Existing approaches are either overly simplistic in modelling and predict poorly when RTT is large, or are over-reliant on data-driven learning, resulting in inflexible models that are not robust to RTT heterogeneity. In this paper, we cast the head movement prediction task as a sparse directed graph learning problem, where three sources of relevant information-a 360 image saliency map, collected viewers' head movement traces, and a biological head rotation model-are aggregated into a unified Markov model. Specifically, we formulate a constrained optimization problem to minimize an l2-norm fidelity term and a sparsity term, corresponding to trace data / saliency consistency and a sparse graph model prior respectively. We solve the problem alternately using a hybrid iterative reweighted least square (IRLS) and Frank-Wolfe optimization strategy. Extensive experiments show that our head movement prediction scheme noticeably outperforms existing proposals across a wide range of RTTs. Xue Zhang 0008, Gene Cheung, Patrick Le Callet, Jack Z. G. Tan |
ICASSP | 1 |
| 2020 | 3D Point Cloud Enhancement Using Graph-Modelled Multiview Depth MeasurementsabstractA 3D point cloud is often synthesized from depth measurements collected by sensors at different viewpoints. The acquired measurements are typically both coarse in precision and corrupted by noise. To improve quality, previous works denoise a synthesized 3D point cloud a posteriori, after projecting the imperfect depth data onto the 3D space. Instead, we enhance depth measurements on the sensed images a priori, exploiting inherent 3D geometric correlation across views, before synthesizing a 3D point cloud from the improved measurements. By enhancing closer to the actual sensing process, we benefit from optimization targeting specifically the depth image formation model, before subsequent processing steps that can further obscure measurement errors. Mathematically, for each pixel row in a pair of rectified viewpoint depth images, we first construct a graph reflecting inter-pixel similarities via metric learning using data in previous enhanced rows. To optimize left and right viewpoint images simultaneously, we write a non-linear mapping function from left pixel row to the right based on 3D geometry relations. We formulate a MAP optimization problem, which, after suitable linear approximations, results in an unconstrained convex and differentiable objective, solvable using fast gradient method (FGM). Experimental results show that our method noticeably outperforms recent denoising algorithms that enhance after 3D point clouds are synthesized. Xue Zhang 0008, Gene Cheung, Jiahao Pang, Dong Tian |
ICIP | 1 |
| 2020 | IET Image Processingabstract360 video is very popular due to its 360 views of a scene. Although 360 videos are also compressed by a hybrid coding framework like 2D video, its high resolution and serious shape deformation affect coding efficiency. In equirectangular projection (ERP) format of 360 videos, if an object moves from equator regions to pole regions or vice versa, large deformation will be introduced and motion estimation cannot find the best‐matched part. To solve the above problem, the authors propose to generate a better reference frame for the current to be encoded frame. First, they project the frame prior to the current one from ERP to the sphere and rotate it at an appropriate angle depending on motion vectors. Subsequently, they insert this generated frame to the rear of the reference queue and let the encoder work as usual. The advantage is that the inserted frame has a more similar shape deformation as the current frame, which greatly helps motion estimation and makes full use of 360 video characters. Their method is simple and friendly compatible with the existing compression standard. Experiments prove that their method achieves 1.57% Bjøntegaard Delta (BD)‐gain compared with standard high efficiency video coding. Chunyu Lin, Yao Zhao 0001, Meiqin Liu 0002, Xue Zhang 0008 |
IET Image Process. | 5 |
| 2019 | Adaptive Streaming in Interactive Multiview Video SystemsabstractMultiview applications endow final users with the possibility to freely navigate within 3D scenes with minimum-delay. A real feeling of scene navigation is enabled by transmitting multiple high-quality camera views, which can be used to synthesize additional virtual views to offer a smooth navigation. However, when network resources are limited, not all camera views can be sent at high quality. It is therefore important, yet challenging, to find the right tradeoff between coding artifacts (reducing the quality of camera views) and virtual synthesis artifacts (reducing the number of camera views sent to users). To this aim, we propose an optimal transmission strategy for interactive multiview HTTP adaptive streaming. We propose a problem formulation to select the optimal set of camera views that the client requests for downloading, such that the navigation quality experienced by the user is optimized while the bandwidth constraints are satisfied. We show that our optimization problem is NP-hard, and we therefore develop an optimal solution based on the dynamic programming algorithm with polynomial time complexity. To further simplify the deployment, we present a suboptimal greedy algorithm with effective performance and lower complexity. The proposed controller is evaluated in theoretical and realistic settings characterized by realistic network statistics estimation, buffer management, and server-side representation optimization. Simulation results show significant improvement in terms of navigation quality compared with alternative baseline multiview adaptation logic solutions. Xue Zhang 0008, Laura Toni, Pascal Frossard, Yao Zhao 0001, Chunyu Lin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | A packetization strategy for interactive multiview video streaming over lossy networks
Xue Zhang 0008, Yao Zhao 0001, Tammam Tillo, Chunyu Lin |
Signal Process. | 1 |
| 2017 | Optimized receiver control in interactive multiview video streaming systemsabstractMultiview applications endow final users with the possibility to freely navigate within 3D scenes with minimum-delay. High-quality rendering of the scene is enabled by transmitting multiple high-quality camera views, which can be used to synthesize additional virtual views to offer a smooth navigation in the scene. When network resources are limited, the set of camera views needs to be properly selected by the client. The right tradeoff between coding artifacts (reducing the quality of camera views) and virtual synthesis artifacts (reducing the number of camera views sent to users) has to be optimized. Existing client adaptation logic strategies usually fail to properly consider the content characteristics and the client navigation properties in the view selection problem. We therefore propose an optimal representation selection for interactive multiview HTTP adaptive streaming (HAS), with a complete problem formulation to select the optimal set of camera views that optimize the navigation quality experienced by the user while satisfying the bandwidth constraints. We show that our optimization problem is NP-hard and develop an effective solution based on a dynamic programming algorithm with polynomial time complexity. Simulation results show significant navigation quality improvement compared to two baseline multiview adaptation logic solutions. This confirms that adaptation logics have to consider both video content and interactivity level of the user in the representation selection strategy. Xue Zhang 0008, Laura Toni, Pascal Frossard, Yao Zhao 0001, Chunyu Lin |
ICC | 1 |
| 2016 | Packetization strategies for MVD-based 3D video transmissionabstractIn multi-view video plus depth (MVD) format, virtual views are synthesized by the compressed texture videos and their associated depth through depth-image-based rendering. In this paper, we consider the setup where both the encoded texture and depth bitstreams experience packet losses during transmission. Different packetization strategies are investigated and a novel strategy is developed to improve error resilience of MVD-based video transmission, where texture data and its corresponding depth are put into the same packet. The size of texture plus associated depth data included in each packet needs to be less than the Maximum Transfer Unit (MTU). Experimental results demonstrate that our proposed packetization scheme yields a significant improvement in terms of both texture views and synthesized virtual views quality when fit in H.264/AVC. Xue Zhang 0008, Yao Zhao 0001, Tammam Tillo, Chunyu Lin, Jimin Xiao, Anhong Wang |
VCIP | 1 |