EDBT 2026 Demo / reviewers in the wild / expert
Mingjin Zhang
dblp:136/8003
· DBLP profile ↗
65ranked-venue papers
43as first author
47since 2021 · last 2026
0000-0002-1473-9784ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 26 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 23 first-author · 20 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | S-DAG: A Subject-Based Directed Acyclic Graph for Multi-Agent Heterogeneous ReasoningabstractLarge Language Models (LLMs) have achieved impressive performance in complex reasoning problems. Their effectiveness highly depends on the specific nature of the task, especially the required domain knowledge. Existing approaches, such as mixture-of-experts, typically operate at the task level; they are too coarse to effectively solve the heterogeneous problems involving multiple subjects. This work proposes a novel framework that performs fine-grained analysis at subject level equipped with a designated multi-agent collaboration strategy for addressing heterogeneous problem reasoning. Specifically, given an input query, we first employ a Graph Neural Network to identify the relevant subjects and infer their interdependencies to generate an Subject-based Directed Acyclic Graph (S-DAG), where nodes represent subjects and edges encode information flow. Then we profile the LLM models by assigning each model a subject-specific expertise score, and select the top-performing one for matching corresponding subject of the S-DAG. Such subject-model matching enables graph-structured multi-agent collaboration where information flows from the starting model to the ending model over S-DAG. We curate and release multi-subject subsets of standard benchmarks (MMLU-Pro, GPQA, MedMCQA) to better reflect complex, real-world reasoning tasks. Extensive experiments show that our approach significantly outperforms existing task-level model selection and multi-agent collaboration baselines in accuracy and efficiency. These results highlight the effectiveness of subject-aware reasoning and structured collaboration in addressing complex and multi-subject problems. Jiangwen Dong, Wanyu Lin, Mingjin Zhang |
AAAI | 4 |
| 2026 | CO²IF: Language-Bridging Hyperspectral-Multispectral Image Fusion with Coordinated and Cross-modal Optimal TransportabstractDue to the difficulties of directly obtaining high-resolution hyperspectral images (HR-HSI), the fusion of low-resolution hyperspectral images (LR-HSI) and high-resolution multispectral images (HR-MSI) has emerged as an effective approach. While existing methods leverage image-level priors from HR-MSI, they often lack explicit semantic guidance for precise detail reconstruction. Recognizing that textual scene descriptions encapsulate valuable object attributes and contextual information, we introduce the first Language-Bridging framework for Hyperspectral and Multispectral image fusion (CO²IF). CO²IF leverages language semantics as prior knowledge to explicitly guide the reconstruction process. To bridge the modality gap between textual descriptions and high-dimensional hyperspectral data, we design a Cross-modal Optimal Transport (COT) module. COT establishes precise semantic correspondences between language features and the visual cues of individual spectral bands. Building upon this semantic alignment, we develop a Multimodal Coordinated State Space Model (CoMamba). CoMamba effectively integrates the language-derived priors with spatial information from HR-MSI and spectral information from LR-HSI. This language-guided reconstruction significantly enhances the extraction of crucial spatial-spectral details, leading to superior fidelity in the generated HR-HSI. In addition, this paper adds text descriptions for three widely used datasets. Both qualitative and quantitative experimental results on the public datasets confirm the superiority of the proposed method compared to the SOTA methods. Mingjin Zhang, Zhongkai Yang, Fei Gao 0006 |
AAAI | 1 |
| 2026 | Efficient KV Cache Migration for Geo-Distributed LLM Inference in Collaborative Edge Computing
Mingjin Zhang, Jiannong Cao 0001, Xiangchun Chen, Ne Wang |
ICDCS | 1 |
| 2026 | Edge large language models: a comprehensive survey
Shan Jiang 0005, Xuecheng Zhou, Mingjin Zhang, Changfu Xu, Guocheng Liao, Jianguo Chen 0001, Jiannong Cao 0001 |
CCF Trans. Pervasive Comput. Interact. | 3 |
| 2026 | Research on multi-step prediction of wind speed under mixed intense wind climate in mountainous areas: focusing on cooling windstorms
Yiyan Dai, Mingjin Zhang, Fanying Jiang, Jinxiang Zhang, Haoxiang Zheng |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | TF-SNN: Temporal focus-based dynamic neuron regulation framework for spiking neural networks
Jie Guo 0009, Junxiang Wu, Mingjin Zhang, Yunsong Li 0001 |
Expert Syst. Appl. | 5 |
| 2026 | ViCGA: High-fidelity human Gaussian avatar extraction from monocular video without camera intrinsics
Jingyuan Gao, Boqian Zhang, Yumeng Hu, Mingjin Zhang |
Neurocomputing | 6 |
| 2026 | SPHSR: A super-resolution reconstruction method for PG-SPECT inspired by smoothed particle hydrodynamics
Mingjin Zhang, Haojuan Yuan |
Neurocomputing | 2 |
| 2026 | Snapshot Compressive Imaging via Degradation Cue and Spectral Latent DiffusionabstractThe goal of snapshot spectral compressive imaging reconstruction is to recover the 3D hyperspectral image from a 2D measurement. However, current reconstruction methods still face significant challenges in fully leveraging degradation and image prior. Many methods estimate degradation solely from a single measurement rather than learning from the real imaging process, resulting in inaccurate prior modeling. Moreover, the high compression of the CASSI measurement leads to the loss of spectral-spatial context, and the existing priors fail to fully capture it - for instance, in complex scenarios (such as S5, S9 in Table I), the performance gap can be as high as 3 dB. To address these issues, this paper introduces a novel reconstruction method with Degradation Cue Learning and Spectral Latent Diffusion (DCL-SLD), which comprises two key components: the Degradation Cue Learning (DCL) module and the Spectral Latent Diffusion (SLD) module. In the spatial domain, the DCL module employs a pre-trained image encoder and a feature distribution transmission strategy to extract degraded information and integrate it into the feature, enabling reconstruction through learned visual context. In the spectral domain, the SLD module leverages a latent diffusion model based on spectral correlations to generate a low-rank vector representation, effectively preserving contextual relationships within the high-dimensional structure. By enhancing priors in both dimensions, the model significantly improves its ability to exploit contextual information for more accurate recovery. Extensive experimental results on both simulation and real datasets demonstrate the superior performance of DCL-SLD over state-of-the-art methods. Mingjin Zhang, Longyi Li, Jie Guo 0009, Yunsong Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2026 | Degradation-Adaptive Denoising: Aligning Diffusion Models With Physics of Video Snapshot Compressive ImagingabstractVideo Snapshot Compressive Imaging (SCI) captures multiple video frames in a single exposure, enabling efficient reconstruction of high-speed scenes for motion analysis and event detection. Existing SCI in coded aperture compressive temporal imaging (CACTI) methods predominantly rely on feedforward deep networks with fixed denoising strategies. However, they lack alignment with the SCI physical inverse model and struggle to balance motion detail recovery and static background smoothing. In this paper, we propose PCD-Diffusion for Video SCI, the first diffusion-based reconstruction framework for Video SCI, which reformulates the inverse problem as a progressive denoising process. Specifically, we design a Physically-Constrained Dynamic Diffusion (PCD-Diffusion) model, introducing a region-adaptive diffusion schedule and spatiotemporal residual estimation. This method explicitly aligns the denoising process with SCI's spatially non-uniform and temporally evolving residual distribution. Additionally, a motion prior-guided diffusion schedule and a Gauss-guided spatiotemporal adaptive residual estimation dynamically steer the denoising trajectory, ensuring accurate motion detail restoration and physically consistent reconstructions. Extensive results on simulated and real datasets verify the superior reconstruction fidelity and temporal coherence of the proposed PCD-Diffusion framework over existing approaches. Code will be released upon publication. Mingjin Zhang, Jie Guo 0009, Yunsong Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2026 | Decentralized Task Offloading in Collaborative Edge Computing: A Digital Twin Assisted Multi-Agent Reinforcement Learning ApproachabstractDecentralized Edge Computing (DEC) has emerged as a computing paradigm leveraging computational resources of edge nodes for complex, data-intensive applications. Decentralized task offloading decides when and at which edge node each task is executed without a central coordinator. However, ensuring reliability for decentralized task offloading is crucial, especially in critical applications like video analytics. Existing centralized approaches often face single points of failure and high communication overhead. Current decentralized methods often ignore task dependencies and bandwidth allocation, leading to suboptimal resource utilization and low reliability. We address the Reliability-aware Dependent Task Offloading (RDTO) problem in DEC, jointly optimizing bandwidth allocation, to maximize task success rate. The challenge of RDTO lies in optimizing dynamic task offloading and bandwidth allocation with task dependencies. We propose a Digital Twin assisted Multi-agent Reinforcement Learning (DT-MARL) algorithm. Our approach integrates a novel digital twin model that provides real-time estimation of task completion time and edge node failure rates. By integrating digital twin with multi-agent reinforcement learning, we enable each edge node to make informed decisions for offloading strategies, effectively improving the task success rate. Extensive experiments using real-world and synthetic datasets demonstrate that DT-MARL outperforms state-of-the-art baselines on task success rate up to 32.00% and 32.43%, respectively. Xiangchun Chen, Jiannong Cao 0001, Yuvraj Sahni, Mingjin Zhang, Yusheng Ji |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | IRMamba: Pixel Difference Mamba with Layer Restoration for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) focuses on identifying small targets in infrared images. Despite advancements with deep learning, challenges persist due to the IR long-range imaging mechanism, where targets are small, dim, and easily lost in noise and background clutter. Current deep learning methods struggle to suppress noise and background interference while preserving fine details, leading to missed detections and false alarms. To address these issues, we propose IRMamba, an encoder-decoder architecture featuring Pixel Difference Mamba (PDMamba) and a Layer Restoration Module (LRM). Specifically, PDMamba integrates the intensity and directional information of pixel differences between scanning positions and their central neighborhoods into the state equation of the state space model (SSM). This enhances target detail representation and suppresses background interference by capturing local 2D dependencies from a global perspective. In addition, LRM incorporates the double-depth image prior into the iterative convergence algorithm, and utilizes the inter-layer interrelationships to gradually reverse the separation of the target layer, achieving noise suppression and refined reconstruction of the image mask. Experiments conducted on multiple public datasets, including NUAA-SIRST, NUDT-SIRST, and IRSTD-1K, demonstrate the significant advantages of IRMamba over SOTA methods. Mingjin Zhang, Fei Gao 0006, Jie Guo 0009 |
AAAI | 1 |
| 2025 | MOCID: Motion Context and Displacement Information Learning for Moving Infrared Small Target DetectionabstractIn the field of Moving Infrared Small Target Detection (MIRSTD), current methods typically use sequential modeling with two individual modules for spatial and temporal processing. However, such a modeling strategy lacks clear guidance on the motion and displacement difference between moving targets and background noise, thereby limiting the feature discriminability and resulting in error-prone target localization. This paper addresses this issue from clip and frame levels and proposes a novel architecture MOCID for MIRSTD. For clip-level feature fusion, we design a spatio-temporal backbone consisting of several proposed Fourier-inspired Spatio-temporal Attention (FISTA) layers. Each FISTA layer sequentially processes the features from spatial and temporal views to capture clip-level temporal motion context, where Fourier Transformation and Inverse Fourier Transformation are employed for each view. This context is then embedded into dynamic convolutional kernels for subsequent spatial feature extraction, thereby enabling clear motion difference guidance and generating comprehensive features. For frame-level feature fusion, we design a Displacement-aware Mamba Module (DAM) to capture detailed frame-to-frame displacement information. DAM utilizes an innovative Temporal Interpolation and Displacement-aware Scan technique to perform spatio-temporal difference-aware displacement modeling, introducing elaborate temporal indicators into feature extraction. Combining the above improvements, our model captures comprehensive motion and displacement contexts, significantly improving the detection of the small target. Extensive experiments demonstrate that MOCID achieves state-of-the-art detection accuracy on popular IRDST and DAUB datasets. Furthermore, MOCID offers a superior balance between throughput and performance compared to other methods. The code for this work will be made publicly available. Mingjin Zhang, Yuanjun Ouyang, Fei Gao 0006, Jie Guo 0009, Qiming Zhang 0001, Jing Zhang 0037 |
AAAI | 1 |
| 2025 | Semi-supervised Infrared Small Target Detection with Thermodynamic-Inspired Uneven Perturbation and Confidence AdaptationabstractSingle-frame Infrared Small Target (SIRST) detection has made significant advancements, but it still faces challenges due to limited labeled data and the foreground-background class imbalance. To address these issues, we introduce a novel Semi-Supervised SIRST Detection (S^3D) pipeline in this paper. First, drawing inspiration from thermodynamics, we propose augmenting infrared images using both chromatically and spatially uneven perturbations. This dual-stream perturbation enhances the diversity and balance of infrared samples, contributing to the robustness of detection models. Additionally, we develop a confidence-adaptive matching method to maintain weighted consistency among perturbed unlabeled samples. Second, to tackle class imbalance in labeled data, we compel the model to generate discriminative predictions for challenging, misclassified examples while down-weighting well-classified examples. We achieve this by modifying the standard cross-entropy loss to squeeze the detector and truncating the loss on well-classified examples. Our innovative Truncated Squeeze (TS) loss focuses on learning discriminative representations for difficult cases and prevents over-optimization for simpler ones. To assess the effectiveness of the perturbation techniques and loss functions, we apply them to various SIRST detectors and conduct comprehensive experiments on two benchmark datasets. Notably, our proposed methods consistently and significantly improve accuracy. Remarkably, our approach achieves over 98% performance of the state-of-the-art fully-supervised method using only 1/8 of the labeled samples. Mingjin Zhang, Wenteng Shang, Fei Gao 0006, Qiming Zhang 0001, Fengqin Lu, Jing Zhang 0037 |
AAAI | 1 |
| 2025 | SAIST: Segment Any Infrared Small Target Model Guided by Contrastive Language-Image PretrainingabstractInfrared Small Target Detection (IRSTD) aims to identify low signal-to-noise ratio small targets in infrared images with complex backgrounds, which is crucial for various applications. However, existing IRSTD methods typically rely solely on image modalities for processing, which fail to fully capture contextual information, leading to limited detection accuracy and adaptability in complex environments. Inspired by vision-language models, this paper proposes a novel framework, SAIST, which integrates textual information with image modalities to enhance IRSTD performance. The framework consists of two main components: Scene Recognition Contrastive Language-Image Pretraining (SR-CLIP) and CLIP-guided Segment Anything Model (CG-SAM). SR-CLIP generates a set of visual descriptions through object-object similarity and object-scene relevance, embedding them into learnable prompts to refine the textual description set. This reduces the domain gap between vision and language, generating precise textual and visual prompts. CG-SAM utilizes the prompts generated by SR-CLIP to accurately guide the Mask Decoder in learning prior knowledge of background features, while incorporating infrared imaging equations to improve small target recognition in complex backgrounds and significantly reduce the false alarm rate. Additionally, this paper introduces the first multimodal IRSTD dataset, MIRSTD, which contains abundant image-text pairs. Experimental results demonstrate that the proposed SAIST method outperforms existing state-of-the-art approaches. Mingjin Zhang, Fei Gao 0006, Jie Guo 0009, Xinbo Gao 0001, Jing Zhang 0037 |
CVPR | 1 |
| 2025 | Joint UAV Deployment and Model Partition for Efficient Collaborative InferenceabstractDeep learning-based intelligent perception has become pivotal in enhancing the effectiveness of UAV monitoring systems. However, deploying complex models on resource-constrained UAV swarms presents significant shortcomings: existing approaches either compromise model accuracy to enable lightweight deployment or introduce communication delays through cloud offloading. More critically, they generally overlook the fundamental prerequisite of monitoring tasks: maintaining stable coverage of the target area. To address this issue, we proposes AirInfer, an innovative collaborative UAV inference framework with critical zones coverage. We formulate a joint optimization problem of deep learning model partitioning and UAV swarm deployment, to minimize end-to-end inference latency. This complex problem can be decomposed into two sub-problems: model partitioning and UAV deployment, which can be solved efficiently using dynamic programming and successive convex approximation, respectively. On this basis, an iterative algorithm is devised to provide guarantees of$\epsilon$-local convergence. Theoretical analysis and experimental results demonstrate that AirInfer not only guarantees blind spot-free monitoring but also reduces inference latency by at least 37 % compared to existing solutions, achieving a balance between perception performance and mission reliability. Wenjing Xia, Tao Wu 0011, Hongjun Wang 0010, Ruhao Jiang, Mingjin Zhang, Yuben Qu |
ICPADS | 5 |
| 2025 | Multimodal Prior Learning with Double Constraint Alignment for Snapshot Spectral Compressive ImagingabstractThe objective of snapshot spectral compressive imaging reconstruction is to recover the 3D hyperspectral image (HSI) from a 2D measurement. Existing methods either focus on network architecture design or simply introduce image-level prior to the model. However, these methods lack guiding information for accurate reconstruction. Recognizing that textual description contain rich semantic information that can significantly enhance details, this paper introduces a novel framework, CAMM, which integrates text information into the model to improve the performance. The framework comprises two key components: Fine-grained Alignment Module (FAM) and Multimodal Fusion Mamba (MFM). Specifically, FAM is used to reduce the knowledge gap between the RGB domain obtained by the pre-trained vision-language model and the HSI domain. Through the double constraints of distribution similarity and entropy, the adaptive alignment of different complexity features is realized, which makes the encoded features more accurate. MFM aims to identify the guiding effect of RGB features and text features on HSI in space and channel dimensions. Instead of fusing features directly, it integrates prior at image-level and text-level prior into Mamba's state-space equation, so that each scanning step can be accurately guided. This kind of positive feedback adjustment ensures the authenticity of the guiding information. To our knowledge, this is the first text-guided model for compressive spectral imaging. Extensive experimental results the public datasets demonstrate the superior performance of CAMM, validating the effectiveness of our proposed method. Mingjin Zhang, Longyi Li, Fei Gao 0006, Qiming Zhang 0001, Jie Guo 0009 |
IJCAI | 1 |
| 2025 | IIRNet: Infinite impulse response inspired network for compressed video quality enhancement
Mingjin Zhang, Lingping Zheng, Yunsong Li 0001, Xinbo Gao 0001 |
Neurocomputing | 1 |
| 2025 | EdgeShard: Efficient LLM Inference via Collaborative Edge ComputingabstractLarge language models (LLMs) have shown great success in content generation and intelligent intelligent decision making for IoT systems. Traditionally, LLMs are deployed on the cloud, incurring prolonged latency, high bandwidth costs, and privacy concerns. More recently, edge computing has been considered promising in addressing such concerns because the edge devices are closer to data sources. However, edge devices are cursed by their limited resources and can hardly afford LLMs. Existing studies address such a limitation by offloading heavy workloads from edge to cloud or compressing LLMs via model quantization. These methods either still rely heavily on the remote cloud or suffer substantial accuracy loss. This work is the first to deploy LLMs on a collaborative edge computing environment, in which edge devices and cloud servers share resources and collaborate to infer LLMs with high efficiency and no accuracy loss. We design EdgeShard, a novel approach to partition a computation-intensive LLM into affordable shards and deploy them on distributed devices. The partition and distribution are nontrivial, considering device heterogeneity, bandwidth limitations, and model complexity. To this end, we formulate an adaptive joint device selection and model partition problem and design an efficient dynamic programming algorithm to optimize the inference latency and throughput. Extensive experiments of the popular Llama2 serial models on a real-world testbed reveal that EdgeShard achieves up to 50% latency reduction and$2 \times $throughput improvement over the state-of-the-art. Mingjin Zhang, Xiaoming Shen, Jiannong Cao 0001, Zeyang Cui, Shan Jiang 0005 |
IEEE Internet Things J. | 1 |
| 2025 | WMRNet: Wavelet Mamba With Reversible Structure for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) is of great practical significance in many real-world applications, such as maritime rescue and early warning systems, benefiting from the unique and excellent infrared imaging ability in adverse weather and low-light conditions. Nevertheless, segmenting small targets from the background remains a challenge. When the subsampling frequency during image processing does not satisfy the Nyquist criterion, the aliasing effect occurs, which makes it extremely difficult to identify small targets. To address this challenge, we propose a novel Wavelet Mamba with Reversible Structure Network (WMRNet) for infrared small target detection in this paper. Specifically, WMRNet consists of a Discrete Wavelet Mamba (DW-Mamba) module and a Third-order Difference Equation guided Reversible (TDE-Rev) structure. DW-Mamba employs the Discrete Wavelet Transform to decompose images into multiple subbands, integrating this information into the state equations of a state space model. This method minimizes frequency interference while preserving a global perspective, thereby effectively reducing background aliasing. The TDE-Rev aims to suppress edge aliasing effects by refining the target edges, which first processes features with an explicit neural structure derived from the second-order difference equations and then promotes feature interactions through a reversible structure. Extensive experiments on the public IRSTD-1k and SIRST datasets demonstrate that the proposed WMRNet outperforms the state-of-the-art methods. Mingjin Zhang, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Mobility-Aware Dependent Task Offloading in Edge Computing: A Digital Twin-Assisted Reinforcement Learning ApproachabstractCollaborative edge computing (CEC) has emerged as a promising paradigm, enabling edge nodes to collaborate and execute tasks from end devices. Task offloading is a fundamental problem in CEC that decides when and where tasks are executed upon the arrival of tasks. However, the mobility of users often results in unstable connections, leading to network failures and resource underutilization. Existing works have not adequately addressed joint mobility-aware dependent task offloading and network flow scheduling, resulting in network congestion and suboptimal performance. To address this, we formulate an online joint mobility-aware dependent task offloading and bandwidth allocation problem, to improve the quality of service by reducing task completion time and energy consumption. We introduce a Mobility-aware Digital Twin-assisted Deep Reinforcement Learning (MDT-DRL) algorithm. Our digital twin model equips the reinforcement learning process by providing future states of mobile users, enabling efficient offloading plans for adapting to the mobile CEC system. Experimental results on real-world and synthetic datasets show that MDT-DRL surpasses state-of-the-art baselines on average task completion time and energy consumption. Xiangchun Chen, Jiannong Cao 0001, Yuvraj Sahni, Mingjin Zhang, Zhixuan Liang, Lei Yang 0024 |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | MDEformer: Mixed Difference Equation Inspired Transformer for Compressed Video Quality EnhancementabstractDeep learning methods have achieved impressive performance in compressed video quality enhancement tasks. However, these methods rely excessively on practical experience by manually designing the network structure and do not fully exploit the potential of the feature information contained in the video sequences, i.e., not taking full advantage of the multiscale similarity of the compressed artifact information and not seriously considering the impact of the partition boundaries in the compressed video on the overall video quality. In this article, we propose a novel Mixed Difference Equation inspired Transformer (MDEformer) for compressed video quality enhancement, which provides a relatively reliable principle to guide the network design and yields a new insight into the interpretable transformer. Specifically, drawing on the graphical concept of the mixed difference equation (MDE), we utilize multiple cross-layer cross-attention aggregation (CCA) modules to establish long-range dependencies between encoders and decoders of the transformer, where partition boundary smoothing (PBS) modules are inserted as feedforward networks. The CCA module can make full use of the multiscale similarity of compression artifacts to effectively remove compression artifacts, and recover the texture and detail information of the frame. The PBS module leverages the sensitivity of smoothing convolution to partition boundaries to eliminate the impact of partition boundaries on the quality of compressed video and improve its overall quality, while not having too much impacts on non-boundary pixels. Extensive experiments on the MFQE 2.0 dataset demonstrate that the proposed MDEformer can eliminate compression artifacts for improving the quality of the compressed video, and surpasses the state-of-the-arts (SOTAs) in terms of both objective metrics and visual quality. Mingjin Zhang, Haichen Bai, Wenteng Shang, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | IRPruneDeXt: Efficient Infrared Small Target Detection via Musical Wavelet-Regularized Channel PruningabstractInfrared small target detection (IRSTD) refers to detecting faint targets in infrared (IR) images, which has achieved notable progress with the advent of deep learning. However, the drive for improved detection accuracy has led to larger, intricate models with redundant parameters, causing storage and computation inefficiencies. In this pioneering study, we introduce the concept of utilizing network pruning to enhance the efficiency of IRSTD. Due to the challenge posed by low signal-to-noise ratios (SNRs) and the absence of detailed semantic information in IR images, directly applying existing pruning techniques yields suboptimal performance. To address this, we propose a novel wavelet structure-regularized multidimensional musical scale soft channel pruning (SCP) method, giving rise to the efficient IRPruneDeXt model. Our approach involves representing the weight matrix in the wavelet domain and formulating a wavelet channel pruning (WCP) strategy. We incorporate wavelet regularization to induce structural sparsity without incurring extra memory usage. Additionally, we design a multidimensional musical scale soft channel reconstruction (MMSCR) method that adapts the strategy across temporal and spatial dimensions to preserve key target information and prevent premature pruning. By leveraging interactions between criteria, it balances pruning and reconstruction through a musical scale feedback effect, achieving an optimal sparse structure while maintaining overall sparsity. Through extensive experiments on many widely used benchmarks, our IRPruneDeXt method surpasses established techniques in both model complexity and accuracy. Specifically, when employing U-net as the baseline network, IRPruneDeXt achieves a 65.68% reduction in parameters and a 51.77% decrease in floating-point operations (FLOPs) while improving intersection over union (IoU) from 73.31% to 76.17% and normalized IoU (nIoU) from 70.92% to 75.08%. The code is available at github.com/hd0013/IRPruneDet. Mingjin Zhang, Jin Feng, Handi Yang, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Computational Fluid Dynamic Network for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) aims to identify and locate small targets amidst background noise. It is highly valuable in various practical application domains, such as maritime rescue and early warning systems deployed in challenging conditions such as harsh weather, low illumination, and long imaging distances. Different from existing works that either adopt well-designed backbone networks or devise specific modules to improve them from different aspects, in this article, we formulate the learning process of IRSTD from a novel perspective, i.e., the mechanism of pixel movement. Considering that the movement of pixels passing through the layers of the network for IRSTD can be analogized to the flow of particles in a fluid dynamic system, we propose a computational fluid dynamic network (CFD-Net) derived from computational fluid dynamics. Technically, we leverage the superiority of the unilateral difference equation with third-order accuracy and devise a unilateral differential residual structure as the backbone of CFD-Net. This design ensures that the pixel stream only flows in the forward direction. In addition, a switch-controlled multidirectional treatment tank (SMTT) is introduced to CFD-Net to dynamically guide the pixel stream to the appropriate path for different targets with varying shapes and orientations, facilitating learning robust target representation and improving detection performance. The proposed CFD-Net is evaluated on the IRSTD-1k and SIRST datasets and is found to outperform existing state-of-the-art (SOTA) methods. Mingjin Zhang, Ke Yue, Jie Guo 0009, Qiming Zhang 0001, Jing Zhang 0037, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | IRPruneDet: Efficient Infrared Small Target Detection via Wavelet Structure-Regularized Soft Channel PruningabstractInfrared Small Target Detection (IRSTD) refers to detecting faint targets in infrared images, which has achieved notable progress with the advent of deep learning. However, the drive for improved detection accuracy has led to larger, intricate models with redundant parameters, causing storage and computation inefficiencies. In this pioneering study, we introduce the concept of utilizing network pruning to enhance the efficiency of IRSTD. Due to the challenge posed by low signal-to-noise ratios and the absence of detailed semantic information in infrared images, directly applying existing pruning techniques yields suboptimal performance. To address this, we propose a novel wavelet structure-regularized soft channel pruning method, giving rise to the efficient IRPruneDet model. Our approach involves representing the weight matrix in the wavelet domain and formulating a wavelet channel pruning strategy. We incorporate wavelet regularization to induce structural sparsity without incurring extra memory usage. Moreover, we design a soft channel reconstruction method that preserves important target information against premature pruning, thereby ensuring an optimal sparse structure while maintaining overall sparsity. Through extensive experiments on two widely-used benchmarks, our IRPruneDet method surpasses established techniques in both model complexity and accuracy. Specifically, when employing U-net as the baseline network, IRPruneDet achieves a 64.13% reduction in parameters and a 51.19% decrease in FLOPS, while improving IoU from 73.31% to 75.12% and nIoU from 70.92% to 74.30%. The code is available at https://github.com/hd0013/IRPruneDet. Mingjin Zhang, Handi Yang, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001, Jing Zhang 0037 |
AAAI | 1 |
| 2024 | IRSAM: Advancing Segment Anything Model for Infrared Small Target Detection
Mingjin Zhang, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001, Jing Zhang 0037 |
ECCV (67) | 1 |
| 2024 | Explore Hybrid Modeling for Moving Infrared Small Target Detection
Mingjin Zhang, Shilong Liu 0005, Yuanjun Ouyang, Jie Guo 0009, Zhihong Tang, Yunsong Li 0001 |
ACM Multimedia | 1 |
| 2024 | VmambaSCI: Dynamic Deep Unfolding Network with Mamba for Compressive Spectral ImagingabstractSnapshot spectral compressive imaging can capture spectral information across multiple wavelengths in one imaging. The coded aperture snapshot spectral imaging (CASSI) method, aims to recover 3D spectral cubes from 2D measurements. Most existing approaches employ a deep unfolding framework based on Transformer, which alternately address a data subproblem and a prior subproblem. However, these frameworks lack flexibility regarding the sensing matrix and inter-stage interactions. In addition, the quadratic computational complexity of global Transformer and the restricted receptive field of local Transformer impact reconstruction efficiency and accuracy. In this paper, we propose a dynamic deep unfolding network with mamba for compressive spectral imaging, called VmambaSCI. We integrate spatial-spectral information from the sensing matrix into the data module and utilizes spatial adaptive operations in the stage interaction of the prior module. Furthermore, recognizing that the imaging process causes aliasing of spatial and spectral information, we develop a dual-domain scanning mamba (DSMamba), featuring a novel spatial-channel scanning method for enhanced efficiency and accuracy. To our knowledge, VmambaSCI is the first Mamba-based model for compressive spectral imaging. Experimental results on the public databases, CAVE and KAIST, demonstrate the superiority of the proposed VmambaSCI over the state-of-the-art approaches. Mingjin Zhang, Longyi Li, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001 |
ACM Multimedia | 1 |
| 2024 | Unleashing the Power of Generic Segmentation Model: A Simple Baseline for Infrared Small Target DetectionabstractRecent advancements in deep learning have greatly advanced the field of infrared small object detection (IRSTD). Despite their remarkable success, a notable gap persists between these IRSTD methods and generic segmentation approaches in natural image domains. This gap primarily arises from the significant modality differences and the limited availability of infrared data. In this study, we aim to bridge this divergence by investigating the adaptation of generic segmentation models, such as the Segment Anything Model (SAM), to IRSTD tasks. Our investigation reveals that many generic segmentation models can achieve comparable performance to state-of-the-art IRSTD methods. However, their full potential in IRSTD remains untapped. To address this, we propose a simple, lightweight, yet effective baseline model for segmenting small infrared objects. Through appropriate distillation strategies, we empower smaller student models to outperform state-of-the-art methods, even surpassing fine-tuned teacher results. Furthermore, we enhance the model's performance by introducing a novel query design comprising dense and sparse queries to effectively encode multi-scale features. Through extensive experimentation across four popular IRSTD datasets, our model demonstrates significantly improved performance in both accuracy and throughput compared to existing approaches, surpassing SAM and Semantic-SAM by over 14 IoU on NUDT and 4 IoU on IRSTD1k. The source code and models will be released at https://github.com/O937-blip/SimIR. Mingjin Zhang, Chi Zhang 0080, Qiming Zhang 0001, Yunsong Li 0001, Xinbo Gao 0001, Jing Zhang 0037 |
ACM Multimedia | 1 |
| 2024 | Wind speed multi-step prediction based on the comparison of wind characteristics and error correction: Focusing on periodic thermally-developed winds
Yiyan Dai, Mingjin Zhang, Fanying Jiang, Jinxiang Zhang, Maoyi Liu, Weicheng Hu |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Novas: Tackling Online Dynamic Video Analytics With Service Adaptation at Mobile Edge ServersabstractVideo analytics at mobile edge servers offers significant benefits like reduced response time and enhanced privacy. However, guaranteeing various quality-of-service (QoS) requirements of dynamic video analysis requests on heterogeneous edge devices remains challenging. In this paper, we propose a scalable online video analytics scheme, called Novas, which automatically makes precise service configuration adjustments upon constant video content changes. Specifically, Novas leverages the filtered confidence sum and a two-window t-test to online detect accuracy fluctuations without ground truth information. In such cases, Novas efficiently estimates the performance of all potential service configurations through a singular value decomposition (SVD)-based collaborative filtering method. Finally, given the NP-hardness of the optimal scheduling problem, a heuristic scheduling strategy that maximizes the minimum remaining resources is devised to schedule the most suitable configurations to servers for execution. We evaluate the effectiveness of Novas through extensive hybrid experiments conducted on a dedicated testbed. Results show that Novas can achieve a substantial over 27$\times$improvement in satisfying the accuracy requirements compared with existing methods adopting fixed configurations, while ensuring latency requirements. Moreover, Novas improves the goodput of the system by an average of 37.86% compared to existing state-of-the-art scheduling solutions. Liang Zhang 0027, Hongzi Zhu, Wen Fei, Yunzhe Li 0001, Mingjin Zhang, Jiannong Cao 0001, Minyi Guo |
IEEE Trans. Computers | 5 |
| 2024 | SPH-Net: Hyperspectral Image Super-Resolution via Smoothed Particle Hydrodynamics ModelingabstractReconstructing a high-resolution hyperspectral image (HSI) from a low-resolution HSI is significant for many applications, such as remote sensing and aerospace. Most deep learning-based HSI super-resolution methods pay more attention to developing novel network structures but rarely study the HSI super-resolution problem from the perspective of image dynamic evolution. In this article, we propose that the HSI pixel motion during the super-resolution reconstruction process can be analogized to the particle movement in the smoothed particle hydrodynamics (SPH) field. To this end, we design an SPH network (SPH-Net) for HSI super-resolution in light of the SPH theory. Specifically, we construct a smooth function based on SPH and design a smooth convolution in multiscales to exploit spectral correlation and preserve the spectral information in the super-resolved image. In addition, we apply the SPH approximation method to discretize the Navier-Stokes motion equation into SPH equation form, which can guide the HSI pixel motion in the desired direction during super-resolution reconstruction, thereby producing clear edges in the spatial domain. Experiments on three public hyperspectral datasets demonstrate that the proposed SPH-Net outperforms the state-of-the-art methods in terms of objective metrics and visual quality. Mingjin Zhang, Jiamin Xu, Jing Zhang 0037, Haimei Zhao, Wenteng Shang, Xinbo Gao 0001 |
IEEE Trans. Cybern. | 1 |
| 2024 | Single-Frame Infrared Small Target Detection via Gaussian Curvature Inspired NetworkabstractSingle-frame infrared small target detection (SIRSTD) is in urgent demand for many practical tasks, such as fire rescue and urban management systems, benefiting from the excellent performance of infrared (IR) imaging in harsh climates and low-light environments. SIRSTD strives to segment small targets from the background as accurately as possible. However, in a real-world application, complex background environments with high brightness and strong edges have similar physical characteristics to small IR targets, which makes it extremely difficult to separate small targets. To address this challenge, we propose a novel Gaussian Curvature Inspired Network (GCI-Net). Inspired by the well-known Gaussian curvature, we develop a Gaussian curvature-based branch (GCB) to eliminate the smoothing noise and preserve the target structure texture information. In addition, we design a complementary patch-group attention (PGA) module that relies on the complementary relationship between low-level and high-level features to provide accurate guidance for GCB. The curvature information generated by the GCB is continuously optimized under the constraint of the curvature information of the ground truth. The proposed GCI-Net provides a reliable guarantee for accurate separation of small targets from the background. We conduct extensive experiments on the public IRSTD-1k and SIRST datasets. The experimental results demonstrate that the proposed GCI-Net outperforms the state-of-the-art (SOTA) methods. Mingjin Zhang, Ke Yue, Boyang Li 0007, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Heat Transfer-Inspired Network for Image Super-Resolution ReconstructionabstractImage super-resolution (SR) is a critical image preprocessing task for many applications. How to recover features as accurately as possible is the focus of SR algorithms. Most existing SR methods tend to guide the image reconstruction process with gradient maps, frequency perception modules, etc. and improve the quality of recovered images from the perspective of enhancing edges, but rarely optimize the neural network structure from the system level. In this article, we conduct an in- depth exploration for the inner nature of the SR network structure. In light of the consistency between thermal particles in the thermal field and pixels in the image domain, we propose a novel heat-transfer-inspired network (HTI-Net) for image SR reconstruction based on the theoretical basis of heat transfer. With the finite difference theory, we use a second-order mixed-difference equation to redesign the residual network (ResNet), which can fully integrate multiple information to achieve better feature reuse. In addition, according to the thermal conduction differential equation (TCDE) in the thermal field, the pixel value flow equation (PVFE) in the image domain is derived to mine deep potential feature information. The experimental results on multiple standard databases demonstrate that the proposed HTI-Net has superior edge detail reconstruction effect and parameter performance compared with the existing SR methods. The experimental results on the microscope chip image (MCI) database consisting of realistic low-resolution (LR) and high-resolution (HR) images show that the proposed HTI-Net for image SR reconstruction can improve the effectiveness of the hardware Trojan detection system. Mingjin Zhang, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | ESSAformer: Efficient Transformer for Hyperspectral Image Super-resolutionabstractSingle hyperspectral image super-resolution (single-HSI-SR) aims to restore a high-resolution hyperspectral image from a low-resolution observation. However, the prevailing CNN-based approaches have shown limitations in building long-range dependencies and capturing interaction information between spectral features. This results in inadequate utilization of spectral information and artifacts after upsampling. To address this issue, we propose ES-SAformer, an ESSA attention-embedded Transformer network for single-HSI-SR with an iterative refining structure. Specifically, we first introduce a robust and spectral-friendly similarity metric, i.e., the spectral correlation coefficient of the spectrum (SCC), to replace the original attention matrix and incorporates inductive biases into the model to facilitate training. Built upon it, we further utilize the kernelizable attention technique with theoretical support to form a novel efficient SCC-kernel-based self-attention (ESSA) and reduce attention computation to linear complexity. ESSA enlarges the receptive field for features after upsampling without bringing much computation and allows the model to effectively utilize spatial-spectral information from different scales, resulting in the generation of more natural high-resolution images. Without the need for pretraining on large-scale datasets, our experiments demonstrate ESSA’s effectiveness in both visual quality and quantitative results. The code will be released at ESSAformer. Mingjin Zhang, Chi Zhang 0080, Qiming Zhang 0001, Jie Guo 0009, Xinbo Gao 0001, Jing Zhang 0037 |
ICCV | 1 |
| 2023 | Fluid Micelle Network for Image Super-Resolution ReconstructionabstractMost existing convolutional neural-network-based super-resolution (SR) methods focus on designing effective neural blocks but rarely describe the image SR mechanism from the perspective of image evolution in the SR process. In this study, we explore a new research routine by abstracting the movement of pixels in the reconstruction process as the flow of fluid in the field of fluid dynamics (FD), where explicit motion laws of particles have been discovered. Specifically, a novel fluid micelle network is devised for image SR based on the theory of FD that follows the residual learning scheme but learns the residual structure by solving the finite difference equation in FD. The pixel motion equation in the SR process is derived from the Navier-Stokes (N-S) FD equation, establishing a guided branch that is aware of edge information. Thus, the second-order residual drives the network for feature extraction, and the guided branch corrects the direction of the pixel stream to supplement the details. Experiments on popular benchmarks and a real-world microscope chip image dataset demonstrate that the proposed method outperforms other modern methods in terms of both objective metrics and visual quality. The proposed method can also reconstruct clear geometric structures, offering the potential for real-world applications. Mingjin Zhang, Jing Zhang 0037, Xinbo Gao 0001, Jie Guo 0009, Dacheng Tao |
IEEE Trans. Cybern. | 1 |
| 2023 | Dim2Clear Network for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) is important for many practical applications such as hazardous aircraft warning, especially when the target is not visible in visible light image due to atmospheric conditions such as fog and cloud. However, IRSTD is challenging due to noises, small and dim targets. To address this challenge, we propose a novel Dim2Clear Network (Dim2Clear) for IRSTD in this paper. Specifically, the Dim2Clear consists of a U-Net backbone encoder, a context mixer decoder (CMD) based on spatial and frequency attention (SFA), and an eyeball-shaped enhancement module (EEM). The CMD is composed of cascaded regular residual blocks where two SFA modules are inserted. Each SFA module receives features from different residual blocks and generates spatial attention map from them to modulate the low-level features, which are then decomposed into low and high frequencies using the discrete cosine transformation. Accordingly, features are further modulated according to the generated frequency attention maps. In this way, SFA can extract both spatial context and frequency context to improve the feature representation capacity. In addition, we design an EEM to suppress the noise and enhance the signal-to-noise ratio in the segmentation results from the perspective of image super-resolution. Experiments on the SIRST dataset and our newly constructed IRSTD-1k dataset show that the proposed Dim2Clear outperforms state-of-the-art methods. Mingjin Zhang, Rui Zhang 0124, Jing Zhang 0037, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Blockchain-based Collaborative Edge Intelligence for Trustworthy and Real-Time Video SurveillanceabstractTrustworthy and real-time video surveillance aims to analyze the live camera streams in a privacy-preserving manner for the decision-making of various advanced services, such as pedestrian reidentification and traffic monitoring. In recent years, edge computing has been identified as a promising technology for trustworthy and real-time video surveillance because it keeps confidential video data locally and reduces the latency caused by massive data transmission. Generally, a single edge device can hardly afford the computation-intensive video analytics tasks. Most existing solutions incorporate cloud servers to handle the overloaded tasks. However, such an edge-cloud collaboration approach still suffers from unpredictable latency and privacy concerns because the remote cloud is centralized and distant from the cameras. In this work, we designed a blockchain-based collaborative edge intelligence (BCEI) approach for trustworthy and real-time video surveillance. In BCEI, geo-distributed edge devices form a peer-to-peer network to maintain a permissioned blockchain and share data and computation resources to perform computation-intensive video analytics tasks. The video analytics results are written on the blockchain in an immutable manner to guarantee trustworthiness. To reduce task execution time, we formulate and solve a joint stream mapping and task scheduling problem to schedule video streams and machine learning models among edge devices. A pedestrian reidentification prototype is implemented and deployed based on BCEI with the extensive performance evaluation, indicating the superiority of BCEI in latency reduction and system throughput improvement by leveraging collaboration among edge devices. Mingjin Zhang, Jiannong Cao 0001, Yuvraj Sahni, Qianyi Chen, Shan Jiang 0005, Lei Yang 0024 |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Curvature Consistent Network for Microscope Chip Image Super-ResolutionabstractDetecting hardware Trojan (HT) from a microscope chip image (MCI) is crucial for many applications, such as financial infrastructure and transport security. It takes an inordinate cost in scanning high-resolution (HR) microscope images for HT detection. It is useful when the chip image is in low-resolution (LR), which can be acquired faster and at a lower cost than its HR counterpart. However, the lost details and noises due to the electric charge effect in LR MCIs will affect the detection performance, making the problem more challenging. In this article, we address this issue by first discussing why recovering curvature information matters for HT detection and then proposing a novel MCI super-resolution (SR) method via a curvature consistent network (CCN). It consists of a homogeneous workflow and a heterogeneous workflow, where the former learns a mapping between homogeneous images, i.e., LR and HR MCIs, and the latter learns a mapping between heterogeneous images, i.e., MCIs and curvature images. Besides, a collaborative fusion strategy is used to leverage features learned from both workflows level-by-level by recovering the HR image eventually. To mitigate the issue of lacking an MCI dataset, we construct a new benchmark consisting of realistic MCIs at different resolutions, called MCI. Experiments on MCI demonstrate that the proposed CCN outperforms representative SR methods by recovering more delicate circuit lines and yields higher HT detection performance. The dataset is available at github.com/RuiZhang97/CCN. Mingjin Zhang, Jingwei Xin, Jing Zhang 0037, Dacheng Tao, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | ISNet: Shape Matters for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) refers to extracting small and dim targets from blurred backgrounds, which has a wide range of applications such as traffic management and marine rescue. Due to the low signal-to-noise ratio and low contrast, infrared targets are easily submerged in the background of heavy noise and clutter. How to detect the precise shape information of infrared targets remains challenging. In this paper, we propose a novel infrared shape network (ISNet), where Taylor finite difference (TFD) -inspired edge block and two-orientation attention aggregation (TOAA) block are devised to address this problem. Specifically, TFD-inspired edge block aggregates and enhances the comprehensive edge information from different levels, in order to improve the contrast between target and background and also lay a foundation for extracting shape information with mathematical interpretation. TOAA block calculates the lowlevel information with attention mechanism in both row and column directions and fuses it with the high-level information to capture the shape characteristic of targets and suppress noises. In addition, we construct a new benchmark consisting of 1, 000 realistic images in various target shapes, different target sizes, and rich clutter backgrounds with accurate pixel-level annotations, called IRSTD-1k. Experiments on public datasets and IRSTD-1 k demonstrate the superiority of our approach over representative state-of-the-art IRSTD methods. The dataset and code are available at github.com/RuiZhang97/ISNet. Mingjin Zhang, Rui Zhang 0124, Yuxiang Yang 0001, Haichen Bai, Jing Zhang 0037, Jie Guo 0009 |
CVPR | 1 |
| 2022 | ENTS: An Edge-native Task Scheduling System for Collaborative Edge ComputingabstractCollaborative edge computing (CEC) is an emerging paradigm enabling sharing of the coupled data, computation, and networking resources among heterogeneous geo-distributed edge nodes. Recently, there has been a trend to orchestrate and schedule containerized application workloads in CEC, while Kubernetes has become the de-facto standard broadly adopted by the industry and academia. However, Kubernetes is not preferable for CEC because its design is not dedicated to edge computing and neglects the unique features of edge nativeness. More specifically, Kubernetes primarily ensures resource provision of workloads while neglecting the performance requirements of edge-native applications, such as throughput and latency. Furthermore, Kubernetes neglects the inner dependencies of edge-native applications and fails to consider data locality and networking resources, leading to inferior performance. In this work, we design and develop ENTS, the first edge-native task scheduling system, to manage the distributed edge resources and facilitate efficient task scheduling to optimize the performance of edge-native applications. ENTS extends Kubernetes with the unique ability to collaboratively schedule computation and networking resources by comprehensively considering job profile and resource status. We showcase the superior efficacy of ENTS with a case study on data streaming applications. We mathematically formulate a joint task allocation and flow scheduling problem that maximizes the job throughput. We design two novel online scheduling algorithms to optimally decide the task allocation, bandwidth allocation, and flow routing policies. The extensive experiments on a real-world edge video analytics application show that ENTS achieves 43% -220% higher average job throughput compared with the state-of-the-art. Mingjin Zhang, Jiannong Cao 0001, Lei Yang 0024, Liang Zhang 0027, Yuvraj Sahni, Shan Jiang 0005 |
SEC | 1 |
| 2022 | SAR-to-Optical Image Translation via Neural Partial Differential EquationsabstractSynthetic Aperture Radar (SAR) becomes prevailing in remote sensing while SAR images are challenging to interpret by human visual perception due to the active imaging mechanism and speckle noise. Recent researches on SAR-to-optical image translation provide a promising solution and have attracted increasing attentions, though still suffering from low optical image quality with geometric distortion due to the large domain gap. In this paper, we mitigate this issue from a novel perspective, i.e., neural partial differential equations (PDE). First, based on the efficient numerical scheme for solving PDE, i.e., Taylor Central Difference (TCD), we devise a basic TCD residual block to build the backbone network, which promotes the extraction of useful information in SAR images by aggregating and enhancing features from different levels. Furthermore, inspired by the Perona-Malik Diffusion (PMD), we devise a PMD neural module to implement feature diffusion through layers, aiming at removing the noises in smooth regions while preserving the geometric structures. Assembling them together, we propose a novel SAR-to-Optical image translation network named S2O-NPDE, which delivers optical images with finer structures and less noise while enjoying an explainability advantage from explicit mathematical derivation. Experiments on the popular SEN1-2 dataset show that our model outperforms state-of-the-art methods in terms of both objective metrics and visual quality. Mingjin Zhang, Chengyu He, Jing Zhang 0037, Yuxiang Yang 0001, Xiaoqi Peng, Jie Guo 0009 |
IJCAI | 1 |
| 2022 | RKformer: Runge-Kutta Transformer with Random-Connection Attention for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) refers to segmenting the small targets from infrared images, which is of great significance in practical applications. However, due to the small scale of targets as well as noise and clutter in the background, current deep neural network-based methods struggle in extracting features with discriminative semantics while preserving fine details. In this paper, we address this problem by proposing a novel RKformer model with an encoder-decoder structure, where four specifically designed Runge-Kutta transformer (RKT) blocks are stacked sequentially in the encoder. Technically, it has three key designs. First, we adopt a parallel encoder block (PEB) of the transformer and convolution to take their advantages in long-range dependency modeling and locality modeling for extracting semantics and preserving details. Second, we propose a novel random-connection attention (RCA) block, which has a reservoir structure to learn sparse attention via random connections during training. RCA encourages the target to attend to sparse relevant positions instead of all the large-area background pixels, resulting in more informative attention scores. It has fewer parameters and computations than the original self-attention in the transformer while performing better. Third, inspired by neural ordinary differential equations (ODE), we stack two PEBs with several residual connections as the basic encoder block to implement the Runge-Kutta method for solving ODE, which can effectively enhance the feature and suppress noise. Experiments on the public NUAA-SIRST dataset and IRSTD-1k dataset demonstrate the superiority of the RKformer over state-of-the-art methods. Mingjin Zhang, Haichen Bai, Jing Zhang 0037, Rui Zhang 0124, Jie Guo 0009, Xinbo Gao 0001 |
ACM Multimedia | 1 |
| 2022 | Exploring Feature Compensation and Cross-level Correlation for Infrared Small Target DetectionabstractSingle frame infrared small target (SIRST) detection is useful for many practical applications, such as maritime rescue. However, SIRST detection is challenging due to the low-contrast between small targets and noisy background in infrared images. To address this challenge, we propose a novel FC3-Net by exploring feature compensation and cross-level correlation for SIRST detection. Specifically, FC3-Net consists of a Fine-detail guided Multi-level Feature Compensation (F-MFC) module, and a Cross-level Feature Correlation (CFC) module. The F-MFC module aims to compensate the information loss of details caused by the downsampling layers in convolutional neural networks (CNN) via aggregating features from multiple adjacent levels, so that the detail features of small targets can be propagated to the deeper layers of the network. Besides, to suppress the side impact of background noise, the CFC module constructs an energy filtering kernel based on the higher-level features with less background noise to filter out the noise in the middle-level features, and fuse them with the low-level ones to learn a strong target representation. Putting them together into the encoder-decoder structure, our FC3-Net could produce an accurate target mask with fine shape and details. Experiment results on the public NUAA-SIRST and IRSTD-1k datasets demonstrate that the proposed FC3-Net outperforms state-of-the-art methods in terms of both pixel-level and object-level metrics. The code will be released at https://github.com/IPIC-Lab/SIRST-Detection-FC3-Net. Mingjin Zhang, Ke Yue, Jing Zhang 0037, Yunsong Li 0001, Xinbo Gao 0001 |
ACM Multimedia | 1 |
| 2022 | COCO-Net: A Dual-Supervised Network With Unified ROI-Loss for Low-Resolution Ship Detection From Optical Satellite Image SequencesabstractLow-resolution ship detection from optical satellite image sequences is critical in high-orbit remote sensing satellite applications. However, it is still a difficult problem due to the following challenges: 1) the size of the ship is tiny in the low-resolution image; 2) the ship target is dim and the contrast with the background is low; 3) the interference of cloud and fog covering is complex and changeable. For these reasons, the targets are easily lost during the detection. In fact, the Clearer the Objects against to the background, the more Confident the Observers can detect it. In light of these considerations, we propose a COCO-Net to detect the small dynamic objects on low-resolution images in this paper. First, the multi-frame images are associated by introducing motion information as an effective compensation for small object features. Second, an integrated dual-supervised network that processes single-level tasks hierarchically is presented to adaptively enhance the input data quality of object detection without being limited by diverse scene disturbances. Third, a unified ROI-loss scheme that modulates the loss function of the first component by introducing ROI-masks from the second component is utilized to make the first component also work for object detection. In addition, we construct a new dataset for the small dynamic object detection based on the GaoFen-4 satellite imagery. Comprehensive experiments on a self-assembled dataset from the GaoFen-4 satellite show the superior performance of the proposed method compared to state-of-the-art object detectors. Qizhi Xu, Yuan Li 0037, Mingjin Zhang, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | E-Tree Learning: A Novel Decentralized Model Learning Framework for Edge AIabstractTraditionally, Artificial Intelligence (AI) models are trained on the central cloud with data collected from end devices. This leads to high communication cost, long response time, and privacy concerns. Recently Edge-empowered AI, namely, Edge AI, has been proposed to support AI model learning and deployment at the network edge closer to the data sources. Existing research, including federated learning adopts a centralized architecture for model learning, where a central server aggregates the model updates from the clients/workers. The centralized architecture has drawbacks, such as performance bottleneck, poor scalability, and single point of failure. In this article, we propose a novel decentralized model learning approach, namely, E-Tree, which makes use of a well-designed tree structure imposed on the edge devices. The tree structure and the locations and orders of the aggregation on the tree are optimally designed to improve the training convergency and model accuracy. In particular, we design an efficient device clustering algorithm, named by K-Means and average accuracy, for E-Tree by taking into account the data distribution on the devices as well as the network distance. Evaluation results show that E-Tree significantly outperforms the benchmark approaches, such as federated learning and gossip learning under nonindependently and identically distributed (Non-i.i.d.) data in terms of model accuracy and convergency. Lei Yang 0024, Jiannong Cao 0001, Mingjin Zhang |
IEEE Internet Things J. | 5 |
| 2021 | Make complex CAPTCHAs simple: A fast text captcha solver based on a small number of samples
Yao Wang 0027, Yuliang Wei, Mingjin Zhang, Bailing Wang |
Inf. Sci. | 3 |
| 2020 | A Contribution Algorithm from LDRI to HDRIabstractHigh dynamic range image (HDRI) which is combined with low dynamic range image (LDRI) needs to be mapped to a low dynamic area to display. In the process of mapping, it is impossible to determine the contribution of low dynamic image sequences in the display images, so that it results in a problem that the low dynamic images cannot be accurately selected. In this paper, for the first time, a contribution algorithm from LDRI to HDRI according to the corresponding response curve of the camera is proposed. Junsong Luo, Shi Qiu 0002, Yizhang Jiang, Keyang Cheng, Huping Ye, Mingjin Zhang |
Int. J. Pattern Recognit. Artif. Intell. | 6 |
| 2020 | Bionic Face Sketch GeneratorabstractFace sketch synthesis is a crucial technique in digital entertainment. However, the existing face sketch synthesis approaches usually generate face sketches with coarse structures. The fine details on some facial components fail to be generated. In this paper, inspired by the artists during drawing face sketches, we propose a bionic face sketch generator. It includes three parts: 1) a coarse part; 2) a fine part; and 3) a finer part. The coarse part builds the facial structure of a sketch by a generative adversarial network in the U-Net. In the middle part, the noise produced by the coarse part is erased and the fine details on the important face components are generated via a probabilistic graphic model. To compensate for the fine sketch with distinctive edge and area of shadows and lights, we learn a mapping relationship at the high-frequency band by a convolutional neural network in the finer part. The experimental results show that the proposed bionic face sketch generator can synthesize the face sketch with more delicate and striking details, satisfy the requirement of users in the digital entertainment, and provide the students with the coarse, fine, and finer face sketch copies when learning sketches. Compared with the state-of-the-art methods, the proposed approach achieves better results in both visual effects and quantitative metrics. Mingjin Zhang, Nannan Wang 0001, Yunsong Li 0001, Xinbo Gao 0001 |
IEEE Trans. Cybern. | 1 |
| 2020 | Cascaded Face Sketch Synthesis Under Various IlluminationsabstractFace sketch synthesis from a photo is of significant importance in digital entertainment. An intelligent face sketch synthesis system requires a strong robustness to lighting variations. Under uncontrolled lighting conditions in real-world settings, such a system will perform consistently well and have little restriction on the lighting conditions. However, previous face sketch synthesis methods tend to synthesize sketches under well-controlled lighting conditions. These methods are sensitive to lighting variations and produce unsatisfactory results when the lighting condition varies. In this paper, we propose a novel cascaded face sketch synthesis framework composed of a multiple feature generator and a cascaded low-rank representation. The multiple feature generator not only produces a generated sketch feature consistent with an artist's drawing style but also extracts a photo feature that is robust to various illuminations. Both features ensure that given a photo patch, the optimal sketch candidates can be selected from the database. The cascaded low-rank representation enables a gradual reduction in the gap between the synthesized face sketch and the corresponding artistdrawn sketch. Experimental results illustrate that the proposed cascaded framework generates realistic sketches on par with the current methods on the Chinese University of Hong Kong face sketch database under well-controlled illuminations. Moreover, this framework exhibits greatly improved performance compared to these methods on the extended Chinese University of Hong Kong face sketch database and Chinese celebrity face photos from the web under different illuminations. We argue that this framework paves a novel way for the implementation of computer-aided optical systems that are of essential importance in both face sketch synthesis and optical imaging. Mingjin Zhang, Yunsong Li 0001, Nannan Wang 0001, Yuan Chi, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Neural Probabilistic Graphical Model for Face Sketch SynthesisabstractNeural network learning for face sketch synthesis from photos has attracted substantial attention due to its favorable synthesis performance. However, most existing deep-learning-based face sketch synthesis models stacked only by multiple convolutional layers without structured regression often lose the common facial structures, limiting their flexibility in a wide range of practical applications, including intelligent security and digital entertainment. In this article, we introduce a neural network to a probabilistic graphical model and propose a novel face sketch synthesis framework based on the neural probabilistic graphical model (NPGM) composed of a specific structure and a common structure. In the specific structure, we investigate a neural network for mapping the direct relationship between training photos and sketches, yielding the specific information and characteristic features of a test photo. In the common structure, the fidelity between the sketch pixels generated by the specific structure and their candidates selected from the training data are considered, ensuring the preservation of the common facial structure. Experimental results on the Chinese University of Hong Kong face sketch database demonstrate, both qualitatively and quantitatively, that the proposed NPGM-based face sketch synthesis approach can more effectively capture specific features and recover common structures compared with the state-of-the-art methods. Extensive experiments in practical applications further illustrate that the proposed method achieves superior performance. Mingjin Zhang, Nannan Wang 0001, Yunsong Li 0001, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Dual-Transfer Face Sketch-Photo SynthesisabstractRecognizing the identity of a sketched face from a face photograph dataset is a critical yet challenging task in many applications, not least law enforcement and criminal investigations. An intelligent sketched face identification system would rely on automatic face sketch synthesis from photographs, thereby avoiding the cost of artists manually drawing sketches. However, conventional face sketch-photo synthesis methods tend to generate sketches that are consistent with the artists'drawing styles. Identity-specific information is often overlooked, leading to unsatisfactory identity verification and recognition performance. In this paper, we discuss the reasons why conventional methods fail to recover identity-specific information. Then, we propose a novel dual-transfer face sketch-photo synthesis framework composed of an inter-domain transfer process and an intra-domain transfer process. In the inter-domain transfer, a regressor of the test photograph with respect to the training photographs is learned and transferred to the sketch domain, ensuring the recovery of common facial structures during synthesis. In the intra-domain transfer, a mapping characterizing the relationship between photographs and sketches is learned and transferred across different identities, such that the loss of identity-specific information is suppressed during synthesis. The fusion of information recovered by the two processes is straightforward by virtue of an ad hoc information splitting strategy. We employ both linear and nonlinear formulations to instantiate the proposed framework. Experiments on The Chinese University of Hong Kong face sketch database demonstrate that compared to the current state-of-the-art the proposed framework produces more identifiable facial structures and yields higher face recognition performance in both the photo and sketch domains. Mingjin Zhang, Ruxin Wang 0002, Xinbo Gao 0001, Jie Li 0001, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2019 | Deep Latent Low-Rank Representation for Face Sketch SynthesisabstractFace sketch synthesis is useful and profitable in digital entertainment. Most existing face sketch synthesis methods rely on the assumption that facial photographs/sketches form a low-dimensional manifold. Once the training data are insufficient, the manifold could not characterize the identity-specific information that is included in a test photograph but excluded in the training data. Thus, the synthesized sketch would lose this information, such as glasses, earrings, hairstyles, and hairpins. To provide the sufficient data and satisfy the assumption on manifold, we propose a novel face sketch synthesis framework based on deep latent low-rank representation (DLLRR) in this paper. The DLLRR induces the hidden training sketches with the identity-specific information as the hidden data to the insufficient original training sketches as the observed data. And it searches the lowest rank representation on the candidates of a test photograph from the both hidden and observed data. For the strong representational capability of the coupled autoencoder, we leverage it to reveal the hidden data. Experiment results on face photograph-sketch database illustrate that the proposed method can successfully provide the sufficient training data with the identity-specific information. And compared to the state of the arts, the proposed method synthesizes more clean and vivid face sketches. Mingjin Zhang, Nannan Wang 0001, Yunsong Li 0001, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Face Sketch Synthesis From Coarse to FineabstractSynthesizing fine face sketches from photos is a valuable yet challenging problem in digital entertainment. Face sketches synthesized by conventional methods usually exhibit coarse structures of faces, whereas fine details are lost especially on some critical facial components. In this paper, by imitating the coarse-to-fine drawing process of artists, we propose a novel face sketch synthesis framework consisting of a coarse stage and a fine stage. In the coarse stage, a mapping relationship between face photos and sketches is learned via the convolutional neural network. It ensures that the synthesized sketches keep coarse structures of faces. Given the test photo and the coarse synthesized sketch, a probabilistic graphic model is designed to synthesize the delicate face sketch which has fine and critical details. Experimental results on public face sketch databases illustrate that our proposed framework outperforms the state-of-the-art methods in both quantitive and visual comparisons. Mingjin Zhang, Nannan Wang 0001, Yunsong Li 0001, Ruxin Wang 0002, Xinbo Gao 0001 |
AAAI | 1 |
| 2018 | Markov Random Neural Fields for Face Sketch SynthesisabstractSynthesizing face sketches with both common and specific information from photos has been recently attracting considerable attentions in digital entertainment. However, the existing approaches either make the strict similarity assumption on face sketches and photos, leading to lose some identity-specific information, or learn the direct mapping relationship from face photos to sketches by the simple neural network, resulting in the lack of some common information. In this paper, we propose a novel face sketch synthesis based on the Markov random neural fields including two structures. In the first structure, we utilize the neural network to learn the non-linear photo-sketch relationship and obtain the identity-specific information of the test photo, such as glasses, hairpins and hairstyles. In the second structure, we choose the nearest neighbors of the test photo patch and the sketch pixel synthesized in the first structure from the training data which ensure the common information of Miss or Mr Average. Experimental results on the Chinese University of Hong Kong face sketch database illustrate that our proposed framework can preserve the common structure and capture the characteristic features. Compared with the state-of-the-art methods, our method achieves better results in terms of both quantitative and qualitative experimental evaluations. Mingjin Zhang, Nannan Wang 0001, Xinbo Gao 0001, Yunsong Li 0001 |
IJCAI | 1 |
| 2018 | Compositional Model-Based Sketch Generator in Facial EntertainmentabstractFace sketch synthesis (FSS) plays an important role in facial entertainment, which includes face sketch morphing among two styles, multiview FSS and face sketch expression manipulation. For facial entertainment, most existing FSS methods generate sketches with over-smoothing effects, i.e., fine details are suppressed more or less. In this paper, we propose a face sketch generator based on the compositional model to handle this issue. It decomposes a face into different components instead of patches as before, and each component has several candidate templates. Multilevel B-spline approximation is utilized to delicately polish the chosen templates of all components. To fuse these components, Poisson blending is employed instead of the weighted average operator. The proposed compositional method crucially reduces the high frequency loss and improves the synthesis performance in comparison to the state-of-the-art methods. Experiments on face sketch morphing, expression manipulation, and multiview FSS, make further efforts to demonstrate the effectiveness of the proposed method. Mingjin Zhang, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Cybern. | 1 |
| 2015 | Moving Video Mapper and City Recorder with Geo-Referenced Videos
Guangqiang Zhao, Mingjin Zhang, Tao Li 0001, Shu-Ching Chen, Ouri Wolfson, Naphtali Rishe |
WISE (2) | 2 |
| 2015 | Recognition of facial sketch styles
Mingjin Zhang, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001 |
Neurocomputing | 1 |
| 2015 | Face Sketch Synthesis via Sparse Representation-Based Greedy SearchabstractFace sketch synthesis has wide applications in digital entertainment and law enforcement. Although there is much research on face sketch synthesis, most existing algorithms cannot handle some nonfacial factors, such as hair style, hairpins, and glasses if these factors are excluded in the training set. In addition, previous methods only work on well controlled conditions and fail on images with different backgrounds and sizes as the training set. To this end, this paper presents a novel method that combines both the similarity between different image patches and prior knowledge to synthesize face sketches. Given training photo-sketch pairs, the proposed method learns a photo patch feature dictionary from the training photo patches and replaces the photo patches with their sparse coefficients during the searching process. For a test photo patch, we first obtain its sparse coefficient via the learnt dictionary and then search its nearest neighbors (candidate patches) in the whole training photo patches with sparse coefficients. After purifying the nearest neighbors with prior knowledge, the final sketch corresponding to the test photo can be obtained by Bayesian inference. The contributions of this paper are as follows: 1) we relax the nearest neighbor search area from local region to the whole image without too much time consuming and 2) our method can produce nonfacial factors that are not contained in the training set and is robust against image backgrounds and can even ignore the alignment and image size aspects of test photos. Our experimental results show that the proposed method outperforms several state-of-the-arts in terms of perceptual and objective metrics. Shengchuan Zhang, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001, Mingjin Zhang |
IEEE Trans. Image Process. | 5 |
| 2015 | TerraFly GeoCloud: An Online Spatial Data Analysis and Visualization SystemabstractWith the exponential growth of the usage of web map services, geo-data analysis has become more and more popular. This article develops an online spatial data analysis and visualization system, TerraFly GeoCloud, which helps end-users visualize and analyze spatial data and share the analysis results. Built on the TerraFly Geo spatial database, TerraFly GeoCloud is an extra layer running upon the TerraFly map and can efficiently support many different visualization functions and spatial data analysis models. Furthermore, users can create unique URLs to visualize and share the analysis results. TerraFly GeoCloud also enables the MapQL technology to customize map visualization using SQL-like statements. The system is available at http://terrafly.fiu.edu/GeoCloud/. Mingjin Zhang, Huibo Wang, Yun Lu 0001, Tao Li 0001, Yudong Guang, Erik Edrosa, Hongtai Li, Naphtali Rishe |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2013 | TerraFly GeoCloud: online spatial data analysis systemabstractWith the exponential growth of the usage of web map services, the geo data analysis has become more and more popular. This paper develops an online Spatial Data Analysis System, TerraFly GeoCloud, which facilitates the end user to visualize and analyze spatial data, and to share the analysis results. Built on the TerraFly Geo spatial database, TerraFly GeoCloud is an extra layer running upon TerraFly map supporting many different visualization functions and spatial data analysis models. TerraFly GeoCloud also enables the MapQL technology to create maps using SQL-like statements. The TerraFly GeoCloud system is available at http://terrafly.fiu.edu/GeoCloud/. Yun Lu 0001, Mingjin Zhang, Tao Li 0001, Erik Edrosa, Naphtali Rishe |
CIKM | 2 |
| 2013 | SksOpen: Efficient Indexing, Querying, and Visualization of Geo-spatial Big DataabstractWith the fast growing use of web-based map services, the performance of indexing and querying of location-based data is becoming a critical quality of service aspect. Spatial indexing is typically time-consuming and is not available to end-users. To address this challenge, we have developed and open-sourced an Online Indexing and Querying System for Big Geospatial Data, sksOpen. Integrated with the TerraFly Geospatial database [1], TerraFly sksOpen is an efficient indexing and query engine for processing Top-k Spatial Boolean Queries. Further, we provide ergonomic visualization of query results on interactive maps to facilitate the user's data analysis. Yun Lu 0001, Mingjin Zhang, Shonda Witherspoon, Yelena Yesha, Yaacov Yesha, Naphtali Rishe |
ICMLA (2) | 2 |
| 2013 | Epidemiological Data Analysis in TerraFly Geo-spatial CloudabstractGIS systems and online services are growing at a very fast pace, however, there are few online services for the analysis of geospatial epidemiology and their functionality is limited. We present a geospatial epidemiology analysis system on the TerraFly Geo-spatial Cloud platform. The system provides comprehensive spatial analysis methods and visualization. In this system, the user is not required to program in order to employ the functionality. All the datasets are stored in the Geo-spatial Cloud. This system is accessible at http://terrafly.fiu.edu/GeoCloud/. The system API algorithms adapted to geospatial epidemiology. The application utilizes the GeoCloud distributed storage system for the Big Data to be analyzed, it utilizes an interactive mapping API to display results. Huibo Wang, Yun Lu 0001, Yudong Guang, Erik Edrosa, Mingjin Zhang, Raul Camarca, Yelena Yesha, Tajana Lucic, Naphtali Rishe |
ICMLA (2) | 5 |
| 2007 | Neuro-Adaptive Formation Control of Multi-Mobile Vehicles: Virtual Leader Based Path Planning and Tracking
Mingjin Zhang, Xiaohong Liao, Wenchuan Cai, Yongduan Song 0001 |
ISNN (1) | 2 |
| 2007 | Neural-Memory Based Control of Micro Air Vehicles (MAVs) with Flapping Wings
Liguo Weng, Wenchuan Cai, Mingjin Zhang, Xiaohong Liao, David Y. Song |
ISNN (1) | 3 |