EDBT 2026 Demo / reviewers in the wild / expert
Maojun Zhang
dblp:05/2983
· DBLP profile ↗
72ranked-venue papers
9as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 19 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 7Computer networks · 6 · 5 first-author · 6 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ICWLM: A Multi-Task Wireless Large Model via In-Context Learning
Yuxuan Wen, Xiaoming Chen 0001, Maojun Zhang, Zhaohui Yang 0001, Chongwen Huang, Zhaoyang Zhang 0001 |
IEEE Trans. Commun. | 3 |
| 2026 | Semantics-Guided Diffusion for Deep Joint Source-Channel Coding in Wireless Image TransmissionabstractJoint source-channel coding (JSCC) offers a promising avenue for enhancing transmission efficiency by jointly incorporating source and channel statistics into the system design. A key advancement in this area is the deep joint source and channel coding (DeepJSCC) technique that designs a direct mapping of input signals to channel symbols parameterized by a neural network, which can be trained for arbitrary channel models and semantic quality metrics. This paper advances the DeepJSCC framework toward a semantics-aligned, high-fidelity transmission approach, called semantics-guided diffusion DeepJSCC (SGD-JSCC). Existing schemes that integrate diffusion models (DMs) with JSCC face challenges in transforming random generation into accurate reconstruction and adapting to varying channel conditions. SGD-JSCC incorporates two key innovations: (1) utilizing some inherent information that contributes to the semantics of an image, such as text description or edge map, to guide the diffusion denoising process; and (2) enabling seamless adaptability to varying channel conditions with the help of a semantics-guided DM for channel denoising. The DM is guided by diverse semantic information and integrates seamlessly with DeepJSCC. In a slow fading channel, SGD-JSCC dynamically adapts to the instantaneous channel state information (CSI) directly estimated from the channel output, thereby eliminating the need for additional pilot transmissions for channel estimation. In a fast fading channel, we introduce a training-free denoising strategy, allowing SGD-JSCC to effectively adjust to fluctuations in channel gains. Numerical results demonstrate that, guided by semantic information and leveraging the powerful DM, our method outperforms existing DeepJSCC schemes, delivering satisfactory reconstruction performance even at extremely poor channel conditions. The proposed scheme highlights the potential of incorporating diffusion models in future communication systems. The code and pretrained checkpoints will be publicly available at https://github.com/MauroZMJ/SGDJSCC, allowing integration of this scheme with existing DeepJSCC models, without the need for retraining from scratch. Maojun Zhang, Guangxu Zhu, Richeng Jin, Xiaoming Chen 0001, Deniz Gündüz |
IEEE Trans. Wirel. Commun. | 1 |
| 2025 | NTR-Gaussian: Nighttime Dynamic Thermal Reconstruction with 4D Gaussian Splatting Based on ThermodynamicsabstractThermal infrared imaging enables a non-invasive measurement of the surface temperature of objects with all-weather applicability. Leveraging such techniques for 3D reconstruction can accurately reflect the temperature distribution of a scene, thereby supporting applications such as building monitoring and energy management. However, existing approaches predominantly focus on static 3D reconstruction for a single time period, overlooking the dynamic nature of thermal radiation phenomena, and failing to predict or analyze temperature variations over time. In this paper, we introduce a novel method, termed NTR-Gaussian, grounded in thermodynamics to address the challenge of nighttime dynamic thermal reconstruction using 4D Gaussian Splatting. Specifically, We utilize neural networks to predict thermodynamic parameters, such as emissivity, convective heat transfer coefficient, and heat capacity, etc. By means of integration, we numerically solve the infrared temperature of the scene at each moment during the night, so as to predict the temperature of the nighttime scene more accurately. To further advance research in this domain, we release a comprehensive dataset of dynamic thermal reconstruction spanning four distinct regions. Extensive experiments demonstrate that NTR-Gaussian significantly outperforms comparison methods in thermal reconstruction, achieving a predicted temperature error within 1 degree Celsius. The code is available at https://github.com/NPUCVPG/NTR-Gaussian. Zeyu Cui, Yu Liu 0008, Maojun Zhang, Shen Yan 0002 |
CVPR | 5 |
| 2025 | MIMO Channel as a Neural Function: Implicit Neural Representations for Extreme CSI CompressionabstractAcquiring and utilizing accurate channel state information (CSI) is crucial for realizing the benefits of massive multiple-input multiple-output (MIMO) technology. Current CSI feedback approaches improve precision by employing advanced deep-learning methods to learn representative CSI features for a subsequent compression process. Diverging from previous works, we treat the CSI compression problem in the context of implicit neural representations. Specifically, each CSI matrix is viewed as a neural function that maps the spatial coordinates (antenna and subchannel) to the corresponding channel gains with physical significance. Rather than transmitting the parameters of the specific neural functions directly, we send low-cost modulations of the CSI matrix, derived through a meta-learning algorithm. These modulations are then applied to a shared base network at the receiver to reconstruct the CSI matrix. Numerical results show that our proposed approach achieves state-of-the-art performance and showcases flexibility in feedback strategies. Maojun Zhang, Yulin Shao, Krystian Mikolajczyk, Deniz Gündüz |
ICASSP | 2 |
| 2025 | LoD-Loc v2: Aerial Visual Localization Over Low Level-of-Detail City Models using Explicit Silhouette AlignmentabstractWe propose a novel method for aerial visual localization over low Level-of-Detail (LoD) city models. Previous wireframe-alignment-based method LoD-Loc has shown promising localization results leveraging LoD models. However, LoD-Loc mainly relies on high-LoD (LoD3 or LoD2) city models, but the majority of available models and those many countries plan to construct nationwide are low-LoD (LoD1). Consequently, enabling localization on low-LoD city models could unlock drones' potential for global urban localization. To address these issues, we introduce LoD-Loc v2, which employs a coarse-to-fine strategy using explicit silhouette alignment to achieve accurate localization over low-LoD city models in the air. Specifically, given a query image, LoD-Loc v2 first applies a building segmentation network to shape building silhouettes. Then, in the coarse pose selection stage, we construct a pose cost volume by uniformly sampling pose hypotheses around a prior pose to represent the pose probability distribution. Each cost of the volume measures the degree of alignment between the projected and predicted silhouettes. We select the pose with maximum value as the coarse pose. In the fine pose estimation stage, a particle filtering method incorporating a multi-beam tracking approach is used to efficiently explore the hypothesis space and obtain the final pose estimation. To further facilitate research in this field, we release two datasets with LoD1 city models covering 10.7 km , along with real RGB queries and ground-truth pose annotations. Experimental results show that LoD-Loc v2 improves estimation accuracy with high-LoD models and enables localization with low-LoD models for the first time. Moreover, it outperforms state-of-the-art baselines by large margins, even surpassing texture-model-based methods, and broadens the convergence basin to accommodate larger prior errors. Juelin Zhu, Shuaibang Peng, Hanlin Tan, Yu Liu 0008, Maojun Zhang, Shen Yan 0002 |
ICCV | 6 |
| 2025 | Beamforming Design for Semantic-Bit Coexisting Communication SystemabstractSemantic communication (SemCom) is emerging as a key technology for future sixth-generation (6G) systems. Unlike traditional bit-level communication (BitCom), SemCom directly optimizes performance at the semantic level, leading to superior communication efficiency. Nevertheless, the task-oriented nature of SemCom renders it challenging to completely replace BitCom. Consequently, it is desired to consider a semantic-bit coexisting communication system, where a base station (BS) serves SemCom users (sem-users) and BitCom users (bit-users) simultaneously. Such a system faces severe and heterogeneous inter-user interference. In this context, this paper provides a new semantic-bit coexisting communication framework and proposes a spatial beamforming scheme to accommodate both types of users. Specifically, we consider maximizing the semantic rate for semantic users while ensuring the quality-of-service (QoS) requirements for bit-users. Due to the intractability of obtaining the exact closed-form expression of the semantic rate, a data driven method is first applied to attain an approximated expression via data fitting. With the resulting complex transcendental function, majorization minimization (MM) is adopted to convert the original formulated problem into a multiple-ratio problem, which allows fractional programming (FP) to be used to further transform the problem into an inhomogeneous quadratically constrained quadratic programs (QCQP) problem. Solving the problem leads to a semi-closed form solution with undetermined Lagrangian factors that can be updated by a fixed point algorithm. This method is referred to as the MM-FP algorithm. Additionally, inspired by the semi-closed form solution, we also propose a low-complexity version of the MM-FP algorithm, called the low-complexity MM-FP (LP-MM-FP), which alleviates the need for iterative optimization of beamforming vectors. Extensive simulation results demonstrate that the proposed MM-FP algorithm outperforms conventional beamforming algorithms such as zero-forcing (ZF), maximum ratio transmission (MRT), and weighted minimum mean-square error (WMMSE). Moreover, the proposed LP-MMFP algorithm achieves comparable performance with the WMMSE algorithm but with lower computational complexity. Maojun Zhang, Guangxu Zhu, Richeng Jin, Xiaoming Chen 0001, Qingjiang Shi, Caijun Zhong, Kaibin Huang |
IEEE J. Sel. Areas Commun. | 1 |
| 2024 | UAVD4L: A Large-Scale Dataset for UAV 6-DoF LocalizationabstractDespite significant progress in global localization of Unmanned Aerial Vehicles (UAVs) in GPS-denied environments, existing methods remain constrained by the availability of datasets. Current datasets often focus on small-scale scenes and lack viewpoint variability, accurate ground truth (GT) pose, and UAV build-in sensor data. To address these limitations, we introduce a large-scale 6-DoF UAV dataset for localization (UAVD4L) and develop a two-stage 6-DoF localization pipeline (UAVLoc), which consists of offline synthetic data generation and online visual localization. Additionally, based on the 6DoF estimator, we design a hierarchical system for tracking ground target in $3 D$ space. Experimental results on the new dataset demonstrate the effectiveness of the proposed approach. Code and dataset are available at https://github.com/RingoWRW/UAVD4L. Rouwan Wu, Xiaoya Cheng, Juelin Zhu, Maojun Zhang, Shen Yan 0002 |
3DV | 5 |
| 2024 | LoD-Loc: Aerial Visual Localization using LoD 3D Map with Neural Wireframe AlignmentabstractWe propose a new method named LoD-Loc for visual localization in the air. Unlike existing localization algorithms, LoD-Loc does not rely on complex 3D representations and can estimate the pose of an Unmanned Aerial Vehicle (UAV) using a Level-of-Detail (LoD) 3D map. LoD-Loc mainly achieves this goal by aligning the wireframe derived from the LoD projected model with that predicted by the neural network. Specifically, given a coarse pose provided by the UAV sensor, LoD-Loc hierarchically builds a cost volume for uniformly sampled pose hypotheses to describe pose probability distribution and select a pose with maximum probability. Each cost within this volume measures the degree of line alignment between projected and predicted wireframes. LoD-Loc also devises a 6-DoF pose optimization algorithm to refine the previous result with a differentiable Gaussian-Newton method. As no public dataset exists for the studied problem, we collect two datasets with map levels of LoD3.0 and LoD2.0, along with real RGB queries and ground-truth pose annotations. We benchmark our method and demonstrate that LoD-Loc achieves excellent performance, even surpassing current state-of-the-art methods that use textured 3D models for localization. The code and dataset will be made available upon publication. Juelin Zhu, Shen Yan 0002, Shengyue Zhang, Yu Liu 0008, Maojun Zhang |
NeurIPS | 6 |
| 2024 | A survey on weakly supervised 3D point cloud semantic segmentationabstractAbstract With the popularity and advancement of 3D point cloud data acquisition technologies and sensors, research into 3D point clouds has made considerable strides based on deep learning. The semantic segmentation of point clouds, a crucial step in comprehending 3D scenes, has drawn much attention. The accuracy and effectiveness of fully supervised semantic segmentation tasks have greatly improved with the increase in the number of accessible datasets. However, these achievements rely on time‐consuming and expensive full labelling. In solve of these existential issues, research on weakly supervised learning has recently exploded. These methods train neural networks to tackle 3D semantic segmentation tasks with fewer point labels. In addition to providing a thorough overview of the history and current state of the art in weakly supervised semantic segmentation of 3D point clouds, a detailed description of the most widely used data acquisition sensors, a list of publicly accessible benchmark datasets, and a look ahead to potential future development directions is provided. Yu Liu 0008, Hanlin Tan, Maojun Zhang |
IET Comput. Vis. | 4 |
| 2024 | CLaSP: Cross-view 6-DoF localisation assisted by synthetic panoramaabstractAbstract Despite the impressive progress in visual localisation, 6‐DoF cross‐view localisation is still a challenging task in the computer vision community due to the huge appearance changes. To address this issue, the authors propose the CLaSP, a coarse‐to‐fine framework, which leverages a synthetic panorama to facilitate cross‐view 6‐DoF localisation in a large‐scale scene. The authors first leverage a segmentation map to correct the prior pose, followed by a synthetic panorama on the ground to enable coarse pose estimation combined with a template matching method. The authors finally formulate the refine localisation process as feature matching and pose refinement to obtain the final result. The authors evaluate the performance of the CLaSP and several state‐of‐the‐art baselines on the Airloc dataset, which demonstrates the effectiveness of our proposed framework. Juelin Zhu, Shen Yan 0002, Xiaoya Cheng, Rouwan Wu, Maojun Zhang |
IET Comput. Vis. | 6 |
| 2024 | Noise2Variance: Dual networks with variance constraint for self-supervised real-world image denoisingabstractAbstract Image denoising aims to restore a clean image from a noisy image. Traditional methods utilizing convolutional neural networks (CNN) for denoising are trained using pairs of noisy and clean images to comprehend the transformation from a noisy image to a clean one. However, the acquisition of such image pairs in real‐world scenarios presents a challenge. Hence, numerous self‐supervised denoising techniques have been developed that do not require clean images for training. This study demonstrates that a straightforward loss design, concentrating on variance, can effectively train a standard CNN denoiser in a self‐supervised fashion. A novel theoretical framework is introduced for training a basic CNN denoising model using three constraints: mean, variance, and augmentation. The variance constraint is crucial as it prevents the trained model from converging to trivial solutions such as identity or zero mapping. This theory provides valuable insights for the development of new self‐supervised denoising methods. Furthermore, a method that applies this theory to proposed dual networks is developed, which consist of two standard CNN models predicting both the clean image and the noise. This approach enhances model capacity during training while minimizing computational costs during inference. This method exemplifies the implementation of the variance constraint and introduces a data constraint for dual networks. Notably, the proposed method only assumes the presence of additive white noise, irrespective of the noise distribution. This minimal assumption enhances the model's robustness against noise with complex or unknown distributions in real‐world distorted images. Experimental results indicate that the proposed Noise2Variance method exhibits commendable performance on peak signal noise ratio and structural similarity metrics compared to existing self‐supervised denoising techniques. Visual comparison of results further substantiates the efficacy of the proposed method. A comparison of model complexity reveals that the method is efficient among the compared CNN‐based techniques. Hanlin Tan, Yu Liu 0008, Maojun Zhang |
IET Image Process. | 3 |
| 2024 | Spectral-Spatial Adversarial Multidomain Synthesis Network for Cross-Scene Hyperspectral Image ClassificationabstractCross-scene hyperspectral image (HSI) classification has received widespread attention due to its practicality. However, domain adaptation-based cross-scene HSI classification methods are typically tailored for a specific target scene involved in model training and require retraining for new scenes. We instead propose an novel spectral-spatial adversarial multi-domain synthetic network (S2AMSnet) that can be trained on a single source domain (SD) and generalized to unseen domains. S2AMSnet improves the robustness of the model to the unseen domain by expanding the diverse distribution of the SD. Specifically, to spatially and spectrally generate diversified generative domain (GD), the spectral-spatial domain generation network (S2DGN) is designed, and two S2DGNs with the same structure but not shared parameters are enabled to generate diversified GD through two-step min-max strategy. A Multi-domain mixing module is employed to expand the diversity of the GD further and enhance their class-domain semantic consistency information. Additionally, a multi-scale mutual information regularization network is used to constrain the S2DGN so that the intrinsic class semantic information of its generated GD does not deviate from the SD. A Semantic consistency discriminator with spectral-spatial feature extraction capability is utilized to capture class-domain semantic consistency information from diverse GD to obtain cross-domain invariant knowledge. Comparative analysis with eight state-of-the-art transfer learning methods on three real HSI datasets, along with an ablation study, validates the effectiveness of the proposed S2AMSnet in the cross-scene HSI classification task. The codes of this work will be available at https://github.com/daxichen/S2AMSnet. Xi Chen 0077, Maojun Zhang, Chen Chen 0127, Shen Yan 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Target Detection With Spectral Graph Contrast Clustering Assignment and Spectral Graph Transformer in Hyperspectral ImageryabstractHyperspectral target detection (HTD) is a method that recognizes objects of interest in a scene by a priori target spectrum. Local details and global information on the spectra are critical for accurate target identification. Detectors with excellent discrimination of spectral differences can better highlight targets while suppressing background. To this end, this article proposes an HTD method based on spectral graph contrast clustering assignment and the spectral graph transformer (SGT) to solve these problems. Specifically, for local-global feature extraction of spectra, the pixel spectra are first constructed as the spectral graph. Then, the representations of the first- or higher-order neighbors of the nodes in the spectral graph are aggregated using graph convolutional networks to extract the local detail information of the spectra. The self-attention in Transformer is utilized to learn the global information of the spectra. Second, a novel spectral graph contrast clustering assignment method is proposed to equip the model with excellent spectral discrimination ability. It maintains clustering consistency by swapping predictive clustering assignments while maximizing the similarity of semantically similar graph clusters and keeping other semantically different graph clusters away from them to better discriminate differences between spectra. Finally, comparisons with seven state-of-the-art HTD methods on four real hyperspectral datasets and ablation studies verify the effectiveness of the proposed method in HTD. Xi Chen 0077, Maojun Zhang, Yu Liu 0008 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | ATLoc: Aerial Thermal Images Localization via View SynthesisabstractWhile visual localization has made significant advances in recent years, it still lacks robustness in low-light situations. Thermal camera images, which capture temperature data, provide a potential solution for these environments. However, the scarcity of well-annotated, publicly available datasets for thermal localization, particularly those focused on absolute pose estimation, impedes further advancement in this field. In this research, we introduce a novel dataset that includes six-degree-of-freedom (6-DoF) absolute poses of query images for large-scale, realistic aerial localization of thermal images. Besides, we introduce a render-to-localization pipeline tailored for thermal image localization. This pipeline predicts the 6-DoF pose of a query using a synthetic technique based on geometric refined thermal model. Experimental results demonstrate the effectiveness of our method on this newly proposed dataset. Notably, our method achieves a median position error of less than 1.5 m and a median angle error of less than 1.5° under diverse test conditions. A comprehensive analysis of factors influencing localization accuracy is also provided. Our code and dataset will be available athttps://github.com/RingoWRW/ATLoc. Rouwan Wu, Shen Yan 0002, Xiaoya Cheng, Juelin Zhu, Yu Liu 0008, Maojun Zhang |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Joint Compression and Deadline Optimization for Wireless Federated LearningabstractFederated edge learning(FEEL) is a popular distributed learning framework for privacy-preserving at the edge, in which densely distributed edge devices periodically exchange model-updates with the server to complete the global model training. Due to limited bandwidth and uncertain wireless environment, FEEL may impose heavy burden to the current communication system. In addition, under the common FEEL framework, the server needs to wait for the slowest device to complete the update uploading before starting the aggregation process, leading to the straggler issue that causes prolonged communication time. In this paper, we propose to accelerate FEEL from two aspects: i.e., 1) performing data compression on the edge devices and 2) setting a deadline on the edge server to exclude the straggler devices. However, undesired gradient compression errors and transmission outage are introduced by the aforementioned operations respectively, affecting the convergence of FEEL as well. In view of these practical issues, we formulate a training time minimization problem, with the compression ratio and deadline to be optimized. To this end, an asymptotically unbiased aggregation scheme is first proposed to ensure zero optimality gap after convergence, and the impact of compression error and transmission outage on the overall training time are quantified through convergence analysis. Then, the formulated problem is solved in an alternating manner, based on which, the noveljoint compression and deadline optimization(JCDO) algorithm is derived. Numerical experiments for different use cases in FEEL including image classification and autonomous driving show that the proposed method is nearly 30X faster than the vanilla FedSGD algorithm, and outperforms the state-of-the-art schemes. Maojun Zhang, Yang Li 0049, Dongzhu Liu, Richeng Jin, Guangxu Zhu, Caijun Zhong, Tony Q. S. Quek |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | Long-Term Visual Localization with Mobile SensorsabstractDespite the remarkable advances in image matching and pose estimation, image-based localization of a camera in a temporally-varying outdoor environment is still a challenging problem due to huge appearance disparity between query and reference images caused by illumination, seasonal and structural changes. In this work, we propose to leverage additional sensors on a mobile phone, mainly GPS, compass, and gravity sensor, to solve this challenging problem. We show that these mobile sensors provide decent initial poses and effective constraints to reduce the searching space in image matching and final pose estimation. With the initial pose, we are also able to devise a direct 2D-3D matching network to efficiently establish 2D-3D correspondences instead of tedious 2D-2D matching in existing systems. As no public dataset exists for the studied problem, we collect a new dataset that provides a variety of mobile sensor data and significant scene appearance variations, and develop a system to acquire ground-truth poses for query images. We benchmark our method as well as several state-of-the-art baselines and demonstrate the effectiveness of the proposed approach. Our code and dataset are available on the project page: https://zju3dv.github.io/sensloc/ Shen Yan 0002, Yu Liu 0008, Zehong Shen, Haomin Liu, Maojun Zhang, Guofeng Zhang 0001, Xiaowei Zhou 0001 |
CVPR | 7 |
| 2023 | Deep Active Contours for Real-time 6-DoF Object TrackingabstractThis paper solves the problem of real-time 6-DoF object tracking from an RGB video. Prior optimization-based methods optimize the object pose by aligning the projected model to the image based on handcrafted features, which are prone to suboptimal solutions. Recent learning-based methods use neural networks to predict the pose, which suffer from limited generalizability or computational efficiency. We propose a learning-based active contour model to make the best use of both worlds. Specifically, given an initial pose, we project the object model to the image plane to obtain the initial contour and use a lightweight network to predict how the contour should move to match the true object boundary, which provides the gradients to optimize the object pose. We also devise an efficient optimization algorithm to train our model end-to-end with pose supervision. Experimental results on semi-synthetic and real-world 6-DoF object tracking datasets demonstrate that our model outperforms state-of-the-art methods by a substantial margin in pose accuracy, while achieving real-time performance on mobile devices. Code is available on our project page: https://zju3dv.github.io/deep_ac/. Shen Yan 0002, Jianan Zhen, Yu Liu 0008, Maojun Zhang, Guofeng Zhang 0001, Xiaowei Zhou 0001 |
ICCV | 5 |
| 2023 | Render-and-Compare: Cross-view 6-DoF Localization from Noisy PriorabstractDespite the significant progress in 6-DoF visual localization, researchers are mostly driven by ground-level benchmarks. Compared with aerial oblique photography capture, ground-level map collection lacks scalability and complete coverage. In this work, we propose to go beyond the traditional ground-level setting and exploit cross-view 6-DoF localization from aerial to ground. We address this problem by formulating camera pose estimation as an iterative render-and-compare pipeline and enhancing the algorithm robustness through augmenting seeds from noisy initial priors. As no public dataset exists for the studied problem, we have collected a new dataset that provides a variety of cross-view images from smartphones and low-altitude drones and developed a semi-automatic system to acquire ground-truth poses for query images. We benchmark our method as well as several state-of-the-art baselines and demonstrate that our method outperforms other approaches by a large margin. Code is available at https://github.com/Choyaa/Render2Loc. Shen Yan 0002, Xiaoya Cheng, Juelin Zhu, Rouwan Wu, Yu Liu 0008, Maojun Zhang |
ICME | 7 |
| 2022 | A noise-resilient online learning algorithm with ramp loss for ordinal regressionabstractOrdinal regression has been widely used in applications, such as credit portfolio management, recommendation systems, and ecology, where the core task is to predict the labels on ordinal scales. Due to its learning efficiency, online ordinal regression using passive aggressive (PA) algorithms has gained a much attention for solving large-scale ranking problems. However, the PA method is sensitive to noise especially in the scenario of streaming data, where the ranking of data samples may change dramatically. In this paper, we propose a noise-resilient online learning algorithm using the Ramp loss function, called PA-RAMP, to improve the performance of PA method for noisy data streams. Also, we validate the order preservation of thresholds of the proposed algorithm. Experiments on real-world data sets demonstrate that the proposed noise-resilient online ordinal regression algorithm is more robust and efficient than state-of-the-art online ordinal regression algorithms. Maojun Zhang, Cuiqing Zhang, Xijun Liang, Zhonghang Xia, Ling Jian, Jiangxia Nan |
Intell. Data Anal. | 1 |
| 2022 | View graph construction for scenes with duplicate structures via graph convolutional networkabstractAbstract View graph construction aims to effectively organise disordered image dataset through image retrieval technique before structure from motion (SfM). Existing view graph construction methods usually fail to handle scenes with duplicate structure, because these methods solely treat the construction of view graph as a process of image‐pair‐wise matching and lack in exploiting images' topological details in dataset. In this paper, we handle this problem from a novel perspective to construct view graph in a global paradigm by introducing an end‐to‐end graph convolutional network (GCN). First, a location‐aware embedding module is introduced to encode images into a feature space that takes into account the feature's location by using Vision Transformer architecture, improving the distinction between features of duplicate structure. Second, graph convolutional network that consists of topological relationship preserving module and feature metric learning module is proposed. Topological relationship preserving network is proposed to help nodes maintain their connected neighbourhood features. By merging the topological connected information into images' embedding, our method can process image matching in a global mode, thus improving the disambiguation ability for images with duplicate scenes. Then a feature metric learning network is embedded into GCN to dynamically compute the linkage prediction among nodes based on their features. Finally, our method combines these three parts to jointly optimise nodes' features and linkage prediction in an end‐to‐end paradigm. We make qualitative and quantitative comparisons based on three public benchmark datasets and demonstrate that our proposed method performs favourably against other state‐of‐the‐art methods. Shen Yan 0002, Yu Liu 0008, Maojun Zhang |
IET Comput. Vis. | 5 |
| 2022 | A Deep Learning-Based Framework for Low Complexity Multiuser MIMO Precoding DesignabstractUsing precoding to suppress multi-user interference is a well-known technique to improve spectra efficiency in multiuser multiple-input multiple-output (MU-MIMO) systems, and the pursuit of high performance and low complexity precoding method has been the focus in the last decade. The traditional algorithms including the zero-forcing (ZF) algorithm and the weighted minimum mean square error (WMMSE) algorithm failed to achieve a satisfactory trade-off between complexity and performance. In this paper, leveraging on the power of deep learning, we propose a low-complexity precoding design framework for MU-MIMO systems. The key idea is to transform the MIMO precoding problem into the multiple-input single-output precoding problem, where the optimal precoding structure can be obtained in closed-form. A customized deep neural network is designed to fit the mapping from the channels to the precoding matrix. In addition, the technique of input dimensionality reduction, network pruning, and recovery module compression are used to further improve the computational efficiency. Furthermore, the extension to the practical MIMO orthogonal frequency-division multiplexing (MIMO-OFDM) system is studied. Simulation results show that the proposed low-complexity precoding scheme achieves similar performance as the WMMSE algorithm with very low computational complexity. Maojun Zhang, Caijun Zhong |
IEEE Trans. Wirel. Commun. | 1 |
| 2022 | Communication-Efficient Federated Edge Learning via Optimal Probabilistic Device SchedulingabstractFederated edge learning (FEEL) is a popular distributed learning framework that allows privacy-preserving collaborative model training via periodic learning-updates communication between edge devices and server. Due to the constrained bandwidth, only a subset of devices can be selected to upload their updates at each training iteration. This has led to an active research area in FEEL studying the optimal device scheduling policy for communication time minimization. However, owing to the difficulty in quantifying the exact communication time, prior work in this area can only tackle the problem partially and indirectly by minimizing either the iteration rounds or per-round latency, while the total communication time is determined by both metrics. To close this research gap, we make the first attempt in this paper to formulate and solve the communication time minimization problem. We first derive a tight bound to approximate the remaining communication time through cross-disciplinary effort that combines the learning theory for convergence rate analysis and communication theory for per-round latency analysis. Building on the novel analytical result, an optimized probabilistic device scheduling policy is derived in closed-form by solving the approximate communication time minimization problem. It is found that the optimized policy gradually turns its priority from suppressing the remaining communication rounds to reducing per-round latency as the training process evolves. Extensive experiments based on real-world dataset and a use case on collaborative 3D objective detection in autonomous driving are provided to verify the superiority of the proposed policy over three benchmark policies based on the indirect solution approaches. Maojun Zhang, Guangxu Zhu, Shuai Wang 0004, Jiamo Jiang, Qing Liao 0001, Caijun Zhong, Shuguang Cui |
IEEE Trans. Wirel. Commun. | 1 |
| 2021 | PSF Estimation of Simple Lens Based on Circular Partition Strategy
Hanxiao Cai, Maojun Zhang, Zheng Zhang 0011, Wei Xu 0019 |
ICIG (3) | 3 |
| 2021 | Extremely Tiny Face Detector for Platforms with Limited Resources
Chen Chen 0127, Maojun Zhang, Hanlin Tan, Huaxin Xiao |
ICIG (2) | 2 |
| 2021 | Joint Face Detection and Landmark Localization Based on an Extremely Lightweight Network
Chen Chen 0127, Maojun Zhang, Jingbei Li |
ICIG (2) | 3 |
| 2021 | Image retrieval for Structure-from-Motion via Graph Convolutional NetworkabstractConventional image retrieval techniques for Structure-from-Motion (SfM) are limited in their ability to effectively distinguish symmetric or repetitive textured patterns and cannot guarantee an accurate generation of pairwise matches without costly redundancy. In this paper, we formulate the image retrieval task as a node binary classification problem with graph data: if a candidate node is marked as positive, it is believed to share the same scene with the query image. The key idea of our approach is that the local context in the feature space around a query image contains abundant information about the matchable relation between the image and its neighbours. By constructing a subgraph surrounding the query image as input data, we adopt a learnable Graph Convolutional Network (GCN) to determine whether nodes in the subgraph have overlapping regions with the query photograph. Experiments demonstrate that our method performs remarkably well on a challenging dataset of highly ambiguous and duplicated scenes. Furthermore, compared with state-of-the-art matchable retrieval methods , the proposed approach significantly reduces unnecessary attempted matches without sacrificing the accuracy and completeness of reconstruction. Shen Yan 0002, Maojun Zhang, Shiming Lai, Yu Liu 0008 |
Inf. Sci. | 2 |
| 2020 | DeU-Net: Deformable U-Net for 3D Cardiac MRI Video Segmentation
Shunjie Dong, Maojun Zhang, Zhengxue Shi, Jianing Deng, Yiyu Shi 0001, Cheng Zhuo |
MICCAI (4) | 3 |
| 2020 | Denoising real bursts with squeeze-and-excitation residual networkabstractThe goal of image denoising is to recover a clean image from noisy input(s). For single image denoising, utilising similarities (or priors) within and across an image dataset helps recover clean images. As the noise level increases, using multiple frames become feasible, which is defined as burst denoising. In this study, the authors propose a deep residual model with squeeze‐and‐excitation (SE) modules for the burst denoising. Unlike previous methods, the authors' model does not need an explicit aligning procedure, which is light‐weighted and fast. The network contains a noise estimation convolutional neural network, which makes it capable of blind denoising. Besides, by inverting the image processing pipeline and simulating real noise in bursts, their model can suppress real noise blindly. Since denoising performance is closely related to the noise level, frame displacement, and the number of frames (burst length), intensive experiments including ablation study are performed. Quantitative results show that the proposed method performs significantly better than previous state‐of‐the‐art methods V‐BM4D and KPN in removing Gaussian noise. Qualitative results show that the proposed method is also effective in removing real noise using bursts and the SE module is key to reduce blur in results. Hanlin Tan, Huaxin Xiao, Shiming Lai, Yu Liu 0008, Maojun Zhang |
IET Image Process. | 5 |
| 2020 | Local-adaptive and outlier-tolerant image alignment using RBF approximation
Jing Li 0014, Maojun Zhang, Zhengming Wang |
Image Vis. Comput. | 3 |
| 2020 | Online Meta Adaptation for Fast Video Object SegmentationabstractConventional deep neural networks based video object segmentation (VOS) methods are dominated by heavily fine-tuning a segmentation model on the first frame of a given video, which is time-consuming and inefficient. In this paper, we propose a novel method which rapidly adapts a base segmentation model to new video sequences with only a couple of model-update iterations, without sacrificing performance. Such attractive efficiency benefits from the meta-learning paradigm which leads to a meta-segmentation model and a novel continuous learning approach which enables online adaptation of the segmentation model. Concretely, we train a meta-learner on multiple VOS tasks such that the meta model can capture their common knowledge and gains the ability to fast adapt the segmentation model to new video sequences. Furthermore, to deal with unique challenges of VOS tasks from temporal variations in the video, e.g., object motion and appearance changes, we propose a principled online adaptation approach that continuously adapts the segmentation model across video frames by exploiting temporal context effectively, providing robustness to annoying temporal variations. Integrating the meta-learner with the online adaptation approach, the proposed VOS model achieves competitive performance against the state-of-the-arts and moreover provides faster per-frame processing speed. Huaxin Xiao, Bingyi Kang, Yu Liu 0008, Maojun Zhang, Jiashi Feng |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | A Quantum-Based Database Query Scheme for Privacy Preservation in Cloud EnvironmentabstractCloud computing is a powerful and popular information technology paradigm that enables data service outsourcing and provides higher-level services with minimal management effort. However, it is still a key challenge to protect data privacy when a user accesses the sensitive cloud data. Privacy-preserving database query allows the user to retrieve a data item from the cloud database without revealing the information of the queried data item, meanwhile limiting user’s ability to access other ones. In this study, in order to achieve the privacy preservation and reduce the communication complexity, a quantum-based database query scheme for privacy preservation in cloud environment is developed. Specifically, all the data items of the database are firstly encrypted by different keys for protecting server’s privacy, and in order to guarantee the clients’ privacy, the server is required to transmit all these encrypted data items to the client with the oblivious transfer strategy. Besides, two oracle operations, a modified Grover iteration, and a special offset encryption mechanism are combined together to ensure that the client can correctly query the desirable data item. Finally, performance evaluation is conducted to validate the correctness, privacy, and efficiency of our proposed scheme. Wenjie Liu 0001, Peipei Gao, Zhihao Liu 0001, Hanwu Chen, Maojun Zhang |
Secur. Commun. Networks | 5 |
| 2018 | Transferable Semi-Supervised Semantic SegmentationabstractThe performance of deep learning based semantic segmentation models heavily depends on sufficient data with careful annotations. However, even the largest public datasets only provide samples with pixel-level annotations for rather limited semantic categories. Such data scarcity critically limits scalability and applicability of semantic segmentation models in real applications. In this paper, we propose a novel transferable semi-supervised semantic segmentation model that can transfer the learned segmentation knowledge from a few strong categories with pixel-level annotations to unseen weak categories with only image-level annotations, significantly broadening the applicable territory of deep segmentation models. In particular, the proposed model consists of two complementary and learnable components: a Label transfer Network (L-Net) and a Prediction transfer Network (P-Net). The L-Net learns to transfer the segmentation knowledge from strong categories to the images in the weak categories and produces coarse pixel-level semantic maps, by effectively exploiting the similar appearance shared across categories. Meanwhile, the P-Net tailors the transferred knowledge through a carefully designed adversarial learning strategy and produces refined segmentation results with better details. Integrating the L-Net and P-Net achieves 96.5% and 89.4% performance of the fully-supervised baseline using 50% and 0% categories with pixel-level annotations respectively on PASCAL VOC 2012. With such a novel transfer mechanism, our proposed model is easily generalizable to a variety of new categories, only requiring image-level annotations, and offers appealing scalability in real applications. Huaxin Xiao, Yunchao Wei, Yu Liu 0008, Maojun Zhang, Jiashi Feng |
AAAI | 4 |
| 2018 | MoNet: Deep Motion Exploitation for Video Object SegmentationabstractIn this paper, we propose a novel MoNet model to deeply exploit motion cues for boosting video object segmentation performance from two aspects, i.e., frame representation learning and segmentation refinement. Concretely, MoNet exploits computed motion cue (i.e., optical flow) to reinforce the representation of the target frame by aligning and integrating representations from its neighbors. The new representation provides valuable temporal contexts for segmentation and improves robustness to various common contaminating factors, e.g., motion blur, appearance variation and deformation of video objects. Moreover, MoNet exploits motion inconsistency and transforms such motion cue into foreground/background prior to eliminate distraction from confusing instances and noisy regions. By introducing a distance transform layer, MoNet can effectively separate motion-inconstant instances/regions and thoroughly refine segmentation results. Integrating the proposed two motion exploitation components with a standard segmentation network, MoNet provides new state-of-the-art performance on three competitive benchmark datasets. Huaxin Xiao, Jiashi Feng, Guosheng Lin, Yu Liu 0008, Maojun Zhang |
CVPR | 5 |
| 2018 | FishEyeRecNet: A Multi-context Collaborative Deep Network for Fisheye Image Rectification
Xiaoqing Yin, Xinchao Wang, Jun Yu 0002, Maojun Zhang, Pascal Fua, Dacheng Tao |
ECCV (10) | 4 |
| 2018 | No-Reference Image Sharpness Assessment Using Scale and Directional ModelsabstractWe propose a new method of no-reference (NR) image sharpness assessment. Our method is based on a multiscale decomposition of non-overlapping blocks, and on analyzing the statistics of Local-Mean Magnitude (LMM) maps. With the detected high-activity blocks, a set of inter-scale and inter-direction sharpness ratios are found. These sharpness ratios are closely correlated with the level of blurring, and is capable of measuring image sharpness effectively from a multiscale and a directional view. They are linearly combined using automatic estimated weighting parameters to induce an overall effective sharpness model. To enhance accuracy, a sharpness factor based on the blur effect on DCT edges and a log-energy sharpness model are also incorporated into the method. Experiments on a large number of public images have shown that our method consistently produces predictions that are highly correlated with human perceptual of image sharpness, and outperforms the current state-of-the-art algorithms. Zheng Zhang 0011, Yu Liu 0008, Hanlin Tan, Xiaoqing Yi, Maojun Zhang |
ICME | 5 |
| 2018 | Salient object detection via robust dictionary representation
Huaxin Xiao, Weiya Ren, Wei Wang 0108, Yu Liu 0008, Maojun Zhang |
Multim. Tools Appl. | 5 |
| 2018 | Focus and Blurriness Measure Using Reorganized DCT Coefficients for an Autofocus ApplicationabstractIn this paper, two metrics for measuring image sharpness are presented and used for an autofocus (AF) application. Both measures exploit reorganized discrete cosine transform (DCT) representation. The first metric is a focus measure, which involves optimal high- and middle-frequency coefficients to evaluate relative sharpness. It is robust to noise while remaining sensitive to the best focus position. A psychometric function-based metric is introduced to quantify the focus measure. The second metric is a no-reference blurriness metric, which is used to measure absolute blurriness. It first constructs multiscale DCT edge maps using directional energy information and then determines image blurriness by combining change information in edge structures with image contrast. This metric gives predictions that are closely correlated with subjective perceived scores and shows performance comparable with that of state-of-the-art methods, especially for noisy images. For noisy situations, the two metrics are adjusted adaptively according to the estimated noise level. To prevent the introduction of extra computational load, an efficient noise-level estimation algorithm based on median absolute deviation is presented. This algorithm exploits only the available reorganized DCT coefficients. With the focus and blurriness measures, an AF method for which the two metrics play an important role was developed. Because of their high-quality performance, the realized AF function is able to locate the best focus position swiftly and reliably. Zheng Zhang 0011, Yu Liu 0008, Zhihui Xiong, Jing Li 0014, Maojun Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Parallax-Tolerant Image Stitching Based on Robust Elastic WarpingabstractImage stitching aims at generating high-quality panoramas with the lowest computational cost. In this paper, we propose a parallax-tolerant image stitching method based on robust elastic warping, which could achieve accurate alignment and efficient processing simultaneously. Given a group of point matches between images, an analytical warping function is constructed to eliminate the parallax errors. Then, the input images are warped according to the computed deformations over the meshed image plane. The seamless panorama is composed by directly reprojecting the warped images. As an important complement to the proposed method, a Bayesian model of feature refinement is proposed to adaptively remove the incorrect local matches. This ensures a more robust alignment than existing approaches. Moreover, our warp is highly compatible with different transformation types. A flexible strategy of combining it with the global similarity transformation is provided as an example. The performance of the proposed approach is demonstrated using several challenging cases. Jing Li 0014, Zhengming Wang, Shiming Lai, Yongping Zhai, Maojun Zhang |
IEEE Trans. Multim. | 5 |
| 2018 | Deep Salient Object Detection With Dense Connections and Distraction DiagnosisabstractIn this paper, we propose two novel components for improving deep salient object detection models. The first component, called saliency detection network (S-Net), introduces dense short- and long-range connections that effectively integrate multiscale features to better exploit contexts at multiple levels. Benefiting from the direct access to low- and high-level features, the S-Net can not only exploit the object context but also preserve the object boundary sharply, leading to enhanced saliency detection performance. Second, a distraction detection network (D-Net) is developed to learn to diagnose which regions of an input image are distracting and harmful for saliency prediction of the S-Net. With such distraction diagnosis, the regions that are distracting to S-Net are removed in hindsight from the input image and the resulted distraction-free image is fed to S-Net for saliency prediction. To train the D-Net, a distraction mining approach is proposed to localize the model-specific distracting regions through examining the sensitiveness of the S-Net to image regions in a principled manner. Besides, the distraction mining approach also provides a way to interpret decisions made by deep neural network (DNN) saliency detection models, which relieves the black-box issues of DNNs to some extent. Extensive experiments on seven popular benchmark datasets demonstrate the effectiveness of the combined S-Net and D-Net, which provides new state of the arts. Huaxin Xiao, Jiashi Feng, Yunchao Wei, Maojun Zhang, Shuicheng Yan |
IEEE Trans. Multim. | 4 |
| 2017 | Joint demosaicing and denoising of noisy bayer images with ADMMabstractImage demosaicing and denoising are import steps of image signal processing. Sequential executions of demosaicing and denoising have essential drawbacks that they degrade the results of each other. Joint demosaicing and denoising overcomes the difficulties by solving the two problems in one model. This paper introduces a unified object function with hidden priors and a variant of ADMM to recover a full-resolution color image with a noisy Bayer input. Experimental results demonstrate that our method performs better than state-of-the-art methods in both PSNR comparison and human vision. In addition, our method is much more robust to variations of noise level. Hanlin Tan, Xiangrong Zeng, Shiming Lai, Yu Liu 0008, Maojun Zhang |
ICIP | 5 |
| 2017 | A Simplified Low Rank and Sparse Model for Visual Tracking
Mi Wang, Huaxin Xiao, Yu Liu 0008, Wei Xu 0019, Maojun Zhang |
ICPRAM | 5 |
| 2017 | PSF Smooth Method based on Simple Lens Imaging
Dazhi Zhan, Zhihui Xiong, Mi Wang, Maojun Zhang |
ICPRAM | 5 |
| 2016 | Fast anomaly detection in traffic surveillance video based on robust sparse optical flowabstractFast abnormal events detection in video is important for intelligent analysis of video. This paper proposes a fast anomaly detection algorithm based on sparse optical flow. We improve the efficiency of optical flow computation with foreground mask and spacial sampling and increase the robustness of optical flow with good feature (TK) points selecting and forward-backward filtering. A foreground channel is also added to the feature vector to help detect static or low speed objects. The algorithm is validated on real-life traffic surveillance to prove its effectiveness. It is also evaluated on a benchmark dataset and achieve detection results comparable to state-of-art methods and outperforms them at pixel-level when the false alarm rate is low. The strength of our algorithm is that it runs real-time on the benchmark dataset which is hundreds of times faster than comparative methods. Hanlin Tan, Yongping Zhai, Yu Liu 0008, Maojun Zhang |
ICASSP | 4 |
| 2016 | Foreground Segmentation for Moving Cameras under Low Illumination ConditionsabstractA foreground segmentation method, including image enhancement, trajectory classification and object segmentation,
is proposed for moving cameras under low illumination conditions. Gradient-field-based image
enhancement is designed to enhance low-contrast images. On the basis of the dense point trajectories obtained
in long frames sequences, a simple and effective clustering algorithm is designed to classify foreground
and background trajectories. By combining trajectory points and a marker-controlled watershed algorithm,
a new type of foreground labeling algorithm is proposed to effectively reduce computing costs and improve
edge-preserving performance. Experimental results demonstrate the promising performance of the proposed
approach compared with other competing methods. Wei Wang 0068, Xiaoqing Yin, Yu Liu 0008, Maojun Zhang |
ICPRAM | 5 |
| 2016 | Data Based Color ConstancyabstractColor constancy is an important task in computer vision. By analyzing the image formation model, color gamut data under one light source can be mapped to a hyperplane whose normal vector is only determined by its light source. Thus, the canonical light source is represented through the kernel method, which trains the color data. When an image is captured under an unknown illuminant, the image-corrected matrix is obtained through optimization. After being mapped to the high-dimensional space, the corrected color data are best fit for the hyperplane of the canonical illuminant. The proposed unsupervised feature-mining kernel method only depends on the color data without any other information. The experiments on the standard test datasets show that the proposed method achieves comparable performance with other state-of-the-art methods. Wei Xu 0019, Huaxin Xiao, Yu Liu 0008, Maojun Zhang |
ICPRAM | 4 |
| 2016 | Image Denoising Using Quadtree-Based Nonlocal Means With Locally Adaptive Principal Component AnalysisabstractIn this letter, we present an efficient image denoising method combining quadtree-based nonlocal means (NLM) and locally adaptive principal component analysis. It exploits nonlocal multiscale self-similarity better, by creating sub-patches of different sizes using quadtree decomposition on each patch. To achieve spatially uniform denoising, we propose a local noise variance estimator combined with denoiser based on locally adaptive principal component analysis. Experimental results demonstrate that our proposed method achieves very competitive denoising performance compared with state-of-the-art denoising methods, even obtaining better visual perception at high noise levels. Chenglin Zuo, Ljubomir Jovanov, Bart Goossens, Hiêp Quang Luong, Wilfried Philips, Yu Liu 0008, Maojun Zhang |
IEEE Signal Process. Lett. | 7 |
| 2015 | Data Separation of L1-minimization for Real-time Motion Detection
Yu Liu 0008, Huaxin Xiao, Zheng Zhang 0011, Wei Xu 0019, Maojun Zhang, Jianguo Zhang 0001 |
BMVC | 5 |
| 2015 | A robust motion detection algorithm on noisy videosabstractThe applicability and performance of motion detection methods dramatically degrade with the increasing noise. In this paper, we propose a robust dictionary-based background subtraction approach, which formulates background modeling as a linear and sparse combination of atoms in a pre-learned dictionary. Motion detection is then implemented to compare the difference between sparse representations of the current frame and the background model. The projection of noise over the dictionary being irregular and random guarantees the adaptability of our approach. Experimental results on synthetic and real noisy videos demonstrate the robustness of the proposed approach compared to other methods. Yu Liu 0008, Huaxin Xiao, Wei Wang 0068, Maojun Zhang |
ICASSP | 4 |
| 2015 | Rotation invariant similarity measure for non-local self-similarity based image denoisingabstractNon-local self-similarity based image denoising depends strictly on similarity measure. The denoising performance is determined based on the ability to reliably find sufficient number of similar patches. In this paper, we propose a rotation invariant similarity measure to fully exploit the image non-local self-similarity. Instead of using image patches, we employ local frequency descriptors, that are rotation invariant and robust to noise, to measure the similarity. Thus, both translational and rotational similarity can be handled even at high noise level. The comparative experimental results show that the proposed method is effective as a rotation invariant similarity measure, and it can consistently improve the performance of non-local means algorithm to achieve better denoising results. Chenglin Zuo, Ljubomir Jovanov, Hiêp Quang Luong, Bart Goossens, Wilfried Philips, Yu Liu 0008, Maojun Zhang |
ICIP | 7 |
| 2015 | Efficient Video Stitching Based on Fast Structure DeformationabstractIn computer vision, video stitching is a very challenging problem. In this paper, we proposed an efficient and effective wide-view video stitching method based on fast structure deformation that is capable of simultaneously achieving quality stitching and computational efficiency. For a group of synchronized frames, firstly, an effective double-seam selection scheme is designed to search two distinct but structurally corresponding seams in the two original images. The seam location of the previous frame is further considered to preserve the interframe consistency. Secondly, along the double seams, 1-D feature detection and matching is performed to capture the structural relationship between the two adjacent views. Thirdly, after feature matching, we propose an efficient algorithm to linearly propagate the deformation vectors to eliminate structure misalignment. At last, image intensity misalignment is corrected by rapid gradient fusion based on the successive over relaxation iteration (SORI) solver. A principled solution to the initialization of the SORI significantly reduced the number of iterations required. We have compared favorably our method with seven state-of-the-art image and video stitching algorithms as well as traditional ones. Experimental results show that our method outperforms the existing ones compared in terms of overall stitching quality and computational efficiency. Jing Li 0014, Wei Xu 0019, Jianguo Zhang 0001, Maojun Zhang, Zhengming Wang, Xuelong Li 0001 |
IEEE Trans. Cybern. | 4 |
| 2014 | Robust Sharpness Metrics Using Reorganized DCT Coefficients for Auto-Focus Application
Zheng Zhang 0011, Yu Liu 0008, Maojun Zhang |
ACCV (4) | 4 |
| 2014 | A Low Illumination Environment Motion Detection Method based on Dictionary LearningabstractThis paper proposes a dictionary-based motion detection method on video images captured under low light with serious noise. The proposed approach trains a dictionary by background images without foreground. It then reconstructs the test image according to the theory of sparse coding, and introduces the Structural Similarity Index Measurement (SSIM) as the detection standard to identify the detection caused by the brightness and contrast ratio changes. Experimental results show that compared to the mixture of Gaussian model and frame difference method, the proposed method can reach a better result under extreme low illumination circumstance. Huaxin Xiao, Yu Liu 0008, Bin Wang 0043, Shuren Tan, Maojun Zhang |
ICPRAM | 5 |
| 2014 | Omni-gradient-based total variation minimisation for sparse reconstruction of omni-directional imageabstractTotal variation (TV) minimisation algorithms have been successfully applied in compressive sensing (CS) recovery for natural images owing to its advantage of preserving edges. However, traditional TV is no longer appropriate for omni‐directional image processing because of the distortions in catadioptric imaging systems. The omni‐gradient computing method combined with the characteristics of omni‐directional imaging is proposed in this study. To reconstruct the image from its compressive samples, the omni‐total variation (omni‐TV) regularisation based on omni‐gradient is utilised instead of traditional TV during the image restoration. The experimental results show that the omni‐directional images can be reconstructed effectively and accurately. Compared with the classical TV minimisation model, the images recovered based on omni‐TV model can provide higher quality both in subjective evaluation and objective evaluation. Jingtao Lou, Yongle Li, Yu Liu 0008, Shuren Tan, Maojun Zhang |
IET Image Process. | 5 |
| 2013 | Human action recognition with Optimized Video Densely SamplingabstractDense sample video patches have been used for video representation in action recognition and achieve better performance than sparse spatiotemporal local features. However, two problems of this method must be considered. First one, many video patches are from background other than human body. Second one, the descriptor is not reliable, since it is neither shift nor scale invariant. To solve these two problems, we proposed an Optimized Video Dense Sampling (OVDS) method combing with dense sampling and spatiotemporal interest points detector. OVDS densely sampled video patches with optimizing the position and scale parameters to guarantee the features are shift and scale invariant. To omit the action unrelated features, we extracted video patches only from human body regions instead of the whole videos. Experimental results on KTH, Weizmann, UCF, Hoollywood2 datasets showed that the features detected by OVDS are informative and reliable for action recognition, and achieve better performance over the existing spatiotemporal local features. Bin Wang 0043, Yu Liu 0008, Wenhua Xiao, Zhihui Xiong, Wei Wang 0068, Maojun Zhang |
ICME | 6 |
| 2013 | Action recognition using Feature Position Constrained Linear CodingabstractRecently space-time interest points (STIPs) using bag-of-feature (BOF) in action recognition has been highly successful. Despite its popularity, The quantization error and the lost of semantic meaning among STIPs are the main weaknesses that severely limit the effectiveness of this method. To overcome these limitations, this paper incorporated the feature position information into coding procedure and proposed a novel Feature Position Constrained Linear Coding (FPLC) method by extending the Locality Constrained Linear Coding (LLC) approach. It first project the features into the human ROI, then codes the features locally using FPLC with the consideration of feature position. Owning to that the local area of human ROI often aggregate features extracted from the same part of human body and those features should exhibit similar values, this local coding strategy helps to alleviate the quantization error and enhance correlation between features at the same time, which helps to improve the recognition accuracy. Compared with the state-of-the-art action recognition method, experiment results demonstrated the effectiveness of the proposed method. Wenhua Xiao, Bin Wang 0043, Yu Liu 0008, Wei Xu 0019, Wei Wang 0068, Weidong Bao 0001, Maojun Zhang |
ICME | 7 |
| 2013 | Angle consistency for registration between catadioptric omni-images and orthorectified aerial imagesabstractRegistration between catadioptric omni‐images and orthorectified aerial images is the key step to integrate them to achieve three‐dimensional urban construction. This problem becomes very challenging because of the non‐linearity of the imaging model of catadioptric omni‐cameras. In this study, the authors attempt to address this problem. The authors first study the properties of horizontal line structure under catadioptric omni‐cameras to prove and extend the theorem of catadioptric distance, and then present angle consistency of horizontal lines between a catadioptric omni‐image and an orthorectified aerial image. The authors further employ them to achieve registration between catadioptric omni‐images and orthorectified aerial images. To the best of authors’ knowledge, this study has not been done before. Experimental results on both simulated data and real scene images confirm the effectiveness of this approach. Wei Xu 0019, Jianguo Zhang 0001, Maojun Zhang |
IET Image Process. | 4 |
| 2012 | Cross-selection kernel regression for super-resolution fusion of complementary panoramic imagesabstractComplementary catadioptric imaging technique was proposed to solve the problem of low and non-uniform resolution in omnidirectional imaging. To enhance this research, our paper focuses on how to generate a high-resolution panoramic image from the captured omnidirectional image. To avoid the interference between the inner and outer images while fusing the two complementary views, a cross-selection kernel regression method is proposed. First, in view of the complementarity of sampling resolution in the tangential and radial directions between the inner and the outer images respectively, the horizontal gradients in the expected panoramic image are estimated based on the scattered neighboring pixels mapped from the outer, while the vertical gradients are estimated using the inner image. Then, the size and shape of the regression kernel are adaptively steered based on the local gradients. Furthermore, the neighboring pixels in the next interpolation step of kernel regression are also selected based on the comparison between the horizontal and vertical gradients. In simulation and real-image experiments, the proposed method outperforms existing kernel regression methods and our previous wavelet-based fusion method in terms of both visual quality and objective evaluation. Lidong Chen, Anup Basu, Maojun Zhang, Wei Wang 0068 |
SMC | 3 |
| 2012 | Fish-eye distortion correction based on midpoint circle algorithmabstractThis paper presents a novel embedded real-time fisheye image distortion correction algorithm with application in IP network camera. A fast and simple distortion correction method is introduced based on Midpoint Circle Algorithm (MCA) which aims to determine the pixel positions along a circle circumference based on incremental calculation of decision parameters. Although only the vertical is rectilinearised, experimental results show that our correction method based on MCA is efficient and effective. In particular, our method can be applied without considering planar checkerboard, iterative fitting of model parameters, complex computation, or traditional lookup tables. Therefore, our algorithm is suitable for embedded camera platform without any extra hardware resources. Yongle Li, Maojun Zhang, Yu Liu 0008, Zhihui Xiong |
SMC | 2 |
| 2012 | Coded aperture techniques for catadioptric omni-directional image defocus deblurringabstractThe defocus blur in catadioptric omni-directional imaging is caused by large apertures and mirror curvatures. This problem becomes more obviously when introducing high resolution sensors. In order to overcome this drawback, this work proposes a simple modification to a conventional catadioptric system that allows for the recovery of an all-focus omni-directional image. The modification is designed by inserting a patterned occluder within the aperture of the camera lens, creating a coded aperture. Then this work introduces a specific deconvolution method, which can recover an all-focus omni-directional image from photograph(s) taken by the camera with coded aperture. Comparing to the conventional aperture, the coded aperture techniques identify the blur scale easier and more accurate. The recovered sharp image eliminates the defocus blur and shows the efficiency of the algorithm. The obtained sharp image can be combined for various catadioptric applications, including omni-directional monitoring systems, intelligent omni-directional systems and robotics, etc. Yu Liu 0008, Yongle Li, Maojun Zhang |
SMC | 4 |
| 2012 | Framework of wargame CGF system based on multi-agentabstractCompared with other large-scale warfare simulation systems, wargame, as a traditional type of warfare simulations, has the advantages of low cost, convenience, practicability, etc. In order to realize Human-Computer Confrontation, Computer Generated Forces (CGF) technology has been added to wargame, which would enhance wargame's training and decision support ability. Based on computerization of “Future: Korea War” which is a strategic-level wargame, this paper presents a framework of wargame CGF system based on multi-agent and defines its Atomic Action Library. The framework and the Atomic Action Library provide the basic platform for farther research on the intelligent behaviors of each agent in wargame CGF system. Shiming Lai, Wei Wang 0068, Maojun Zhang |
SMC | 4 |
| 2012 | HW/SW co-design of an embedded omni-imaging systemabstractOmni-imaging can be used in many practical applications that need a wide field of view, therefore a real-time and high-definition embedded system design and implementation of omni-imaging is desired. In this study, we propose a hardware/software co-design method for the design and implementation of embedded omni-imaging systems. In order to achieve real-time and high-definition goals, we perform hardware/software partitioning based on the analysis of functional modules in a basic embedded omni-imaging system. In the experiments, the proposed hardware/software co-design omni-imaging system is implemented in a FPGA (Field Programmable Gate Array) plus DSP (Digital Signal Processor) system architecture. Results indicate that the omni-imaging speed achieved is 39fps with the imaging resolution set at 1024×768 for the original omni-image and 1280×288 for the unwarped image. Zhihui Xiong, Irene Cheng 0001, Maojun Zhang, Anup Basu |
SMC | 3 |
| 2012 | An application oriented and shape feature based multi-touch gesture description and recognition method
De-xin Wang, Zhihui Xiong, Maojun Zhang |
Multim. Tools Appl. | 3 |
| 2012 | Efficient omni-image unwarping using geometric symmetry
Zhihui Xiong, Irene Cheng 0001, Anup Basu, Wei Wang 0068, Wei Xu 0019, Maojun Zhang |
Mach. Vis. Appl. | 6 |
| 2011 | A multi-touch platform based on four corner cameras and methods for accurately locating contact points
De-xin Wang, Qing-bao Liu, Maojun Zhang |
Multim. Tools Appl. | 3 |
| 2011 | A 2-point algorithm for 3D reconstruction of horizontal lines from a single omni-directional image
Irene Cheng 0001, Zhihui Xiong, Anup Basu, Maojun Zhang |
Pattern Recognit. Lett. | 5 |
| 2010 | Application Oriented Semantic Multi-touch Gesture Description Method
De-xin Wang, Maojun Zhang |
ICIC (2) | 2 |
| 2009 | Object Tracking and Primitive Event Detection by Spatio-Temporal Tracklet AssociationabstractAccurate object tracking is a challenging problem in visual surveillance due to noise segmentation, partial and full object occlusions. In this paper, we present a method for object tracking and primitive event detection by associating tracklet caused by these problems. The aim is to keep track identity across tracking gaps and detect object's motion changes (identify primitive event) that cause tracklet gaps. We first detect moving objects and generate tracklet, then grow these tracklets by finding the best spatial and temporal association of observations to track object across tracklet gaps and indentify the video event they involved. We successfully track multiple moving vehicles and persons under occlusion, noisy detections and split-merge situations and can identify the event that cause tracking gaps. Jiangfeng Wang, Maojun Zhang, Anthony G. Cohn 0001 |
ICIG | 2 |
| 2009 | Adaptive Method for Early Detecting Zero Quantized DCT Coefficients in H.264/AVC Video EncodingabstractIn H.264/AVC video encoding, zero quantized discrete cosine transform (DCT) coefficients are quite common in low bit-rate video application. The computations of DCT and quantization can be remarkably reduced if zero quantized DCT coefficients are detected prior to DCT and quantization. Many methods have been applied to early detect most zero quantized DCT coefficients. But the detection ratios of zero quantized DCT coefficients need to be improved when quantization parameter (QP) is small. We present an adaptive method to detect zero quantized DCT coefficients. When QP exceeds a certain value, a new threshold is derived to efficiently detect the all-zero blocks (AZBs) without any video degradation. Otherwise, a concept of fourteen-zeros block (FZB), which means only two coefficients are non-zero in 16 DCT coefficients of a 4 times 4 block, is proposed. An innovative method to process FZBs is proposed to reduce the redundant computations. Experimental results show that the proposed adaptive method achieves approximately 8%-20% computational savings, compared with that of the existing methods. Maojun Zhang, Wei Wang 0068 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2008 | View Synthesis for Virtual Walk through in Real Scene Based on Catadioptric Omnidirectional ImagesabstractThis paper proposes an approach to more conveniently and efficiently create virtual walk through in real scene from catadioptric omni directional images via view synthesis. Acquisition and unwarping process of omnidirectional image are discussed firstly, then, according to the speciality of cylindrical panoramic image, we utilize the advantage of epiline-sampling method for cylindrical image, which samples reference images along epilines as much as possible grounding on epipolar geometry. As to novel view generation after image rectification, a corresponding interpolation algorithm based on rectified images is developed. Results of proposed approach applied to both synthetic and real scene are given at the end of this paper. Experiments show that our method can reduce distortion and pixel information losing from image rectification obviously meanwhile maintain scene information well. Wei Xu 0019, Zhihui Xiong, Maojun Zhang |
CW | 4 |
| 2007 | Fast Panorama Unrolling of Catadioptric Omni-Directional Images for Cooperative Robot Vision SystemabstractFor omni-directional imaging based cooperative robot vision system, panorama unrolling is an important problem. We present an eight direction symmetry reuse algorithm for this problem, principles of this algorithm are: 1) Treat the original image to be unrolled as a series of co-centric circles, and uniformly partition them into eight parts of symmetrical sector regions. 2) According to symmetry principle in spatial geometry, we construct the symmetrical transform relations pixel coordinates among these regions. So, we need compute only one region of pixel coordinate for panorama unrolling, by use of complex ray-trace coordinate mapping, while the other seven parts of regions can be determined by symmetrically reusing results from the first region, which reduces seven-eighths computational burden for ray-tracing method. Experiments indicate that, under the same precision, our algorithm improves 3.28 times of unrolling speed compared with ray-trace unrolling method. In omnidirectional video with look-up table method, the idea of eight direction symmetry reuse can also be used to reduce the look-up table size to only one-eighth of the original table. Zhihui Xiong, Maojun Zhang, Yunli Wang, Sikun Li |
CSCWD | 2 |
| 2004 | An Orientation Update Message Filtering Algorithm in Collaborative Virtual Environments
Maojun Zhang, Nicolas D. Georganas |
J. Comput. Sci. Technol. | 1 |
| 1999 | VCS: a virtual environment support for awareness and collaborationabstractArticle Free Access Share on VCS: a virtual environment support for awareness and collaboration Authors: Maojun Zhang Department of System Engineering and Mathematics, National University of Defense Technology, Changsha, Hunan, P. R. China Department of System Engineering and Mathematics, National University of Defense Technology, Changsha, Hunan, P. R. ChinaView Profile , Linda Wu Department of System Engineering and Mathematics, National University of Defense Technology, Changsha, Hunan, P. R. China Department of System Engineering and Mathematics, National University of Defense Technology, Changsha, Hunan, P. R. ChinaView Profile , Lifeng Sun Department of System Engineering and Mathematics, National University of Defense Technology, Changsha, Hunan, P. R. China Department of System Engineering and Mathematics, National University of Defense Technology, Changsha, Hunan, P. R. ChinaView Profile , Yunhao Li Department of System Engineering and Mathematics, National University of Defense Technology, Changsha, Hunan, P. R. China Department of System Engineering and Mathematics, National University of Defense Technology, Changsha, Hunan, P. R. ChinaView Profile , Bing Yang Department of System Engineering and Mathematics, National University of Defense Technology, Changsha, Hunan, P. R. China Department of System Engineering and Mathematics, National University of Defense Technology, Changsha, Hunan, P. R. ChinaView Profile Authors Info & Claims MULTIMEDIA '99: Proceedings of the seventh ACM international conference on Multimedia (Part 2)October 1999 Pages 163–165https://doi.org/10.1145/319878.319922Published:01 October 1999Publication History 2citation221DownloadsMetricsTotal Citations2Total Downloads221Last 12 Months2Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Maojun Zhang, Linda Wu, Lifeng Sun |
ACM Multimedia (2) | 1 |