EDBT 2026 Demo / reviewers in the wild / expert
Qiang Ling 0001
dblp:99/4581-1
· DBLP profile ↗
72ranked-venue papers
5as first author
46since 2021 · last 2026
0000-0001-5688-4130ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 4 first-author · 24 since 2021Artificial intelligence and machine learning · 24 · 1 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Computer networks · 5 · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DenoiseGS: Delta-Based 3D Gaussian Splatting with B-spline Trajectory Optimization for Dynamic Driving Scene Reconstructionabstract3D street scene reconstruction is a challenging yet crucial task for autonomous driving. Many reconstruction methods often overlook two key limitations for high-quality driving scene reconstruction, sensitivity to camera parameter noise from high-speed vehicles and heavy reliance on precise dynamic object annotations of datasets. To resolve these issues, we propose DenoiseGS, a simple yet effective approach based on explicit 3D Gaussian splatting. Specifically, we propose a novel learnable Delta attribute per Gaussian primitive that operates on the image plane during rasterization to mitigate the impact of noisy camera parameters through modulating the inputs of the alpha-blending process. To enhance the representation of this Delta attribute, we propose a DeltaEstimator that encodes viewing direction and contextual cues to facilitate view dependence. We also extend additional CUDA operations to enable efficient gradient update for the delta attribute. Furthermore, to overcome the limitation of inaccurate annotations for dynamic objects, we propose a learnable B-spline trajectory optimization with a few control points to model the trajectory of a moving object. Comprehensive experiments conducted on nuScenes and Waymo Open Dataset demonstrate that our DenoiseGS outperforms some state-of-the-art methods across all metrics of both reconstruction quality and novel view synthesis. Junjie Linghu, Qiang Ling 0001 |
AAAI | 2 |
| 2026 | DGKD: Depth-Guided knowledge distillation network for monocular 3D object detection
Qiang Ling 0001 |
Appl. Intell. | 2 |
| 2026 | DR-Net: Dual-Representation Network With Motion-Aware Augmentation for 3D Object Detection Based on 4D RadarsabstractNowadays 4D millimeter-wave radars can generate LiDAR-equivalent 3D point clouds with superior cost efficiency and enhanced stability in harsh environments, and have found wide applications in 3D object detection. However, existing 4D radar-based object detection methods may overlook the detrimental effects of information loss during feature processing and fail to effectively leverage the velocity information of 4D radars. These limitations hinder the full exploitation of 4D radar’s potential. To resolve this issue, we propose a dual-representation network with motion-aware augmentation, named as DR-Net. Specifically, to compensate the information loss, we propose a dual-representation encoder (DRE) and a sampling fusion backbone (SFB). The encoder extracts radar features from both pillars’ and points’ perspectives, well exploiting the complementarity between pillar-level structural context and point-level fine-grained details. The backbone fuses features of the pillar representation and the point representation, effectively mitigating the feature-level information loss. Instead of simply taking the velocity information as an additional input, we design a motion-aware augmentation (MAA) module to augment 4D radar point cloud data from the perspectives of both raw point cloud features and training instances. Finally, we extend the original 3D detection head by incorporating an additional velocity supervision branch to enhance the capability of perceiving both static and dynamic objects. We conduct comprehensive experiments on the View-of-Delft (VoD) and TJ4DRadSet datasets. Experimental results reveal that compared with some state-of-the-art 4D radar-based approaches, our DR-Net achieves significant performance advancement. Jinrong Cao, Qiang Ling 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | OCC-Exoskeleton: A Plug-and-Play Module to Enhance CNN-Based Occupancy Prediction NetworksabstractIn 3D semantic occupancy prediction, both the task-specific characteristics and input data critically influence network perception performance. It encounters many challenges, such as data-label misalignment, spatial-significance variance, and long-tailed distribution in semantic occupancy prediction labels. To deal with these challenges, we propose a generalized auxiliary enhancement module, termed OCC-Exoskeleton, for semantic occupancy prediction. The proposed module demonstrates remarkable adaptability, enabling seamless integration with diverse occupancy prediction models while maintaining architectural compatibility. Our module is made up of three parts, each of which is specifically designed to address one of the three mentioned challenges: 1) Virtual point cloud distillation. We generate the virtual point cloud, teaching realistic modalities to concentrate on the data-label misalignment positions. 2) Dual-expert occupancy head. We allocate the occupancy prediction task to two expert heads according to spatial significance to obtain more targeted outcomes. 3) Scene-level frames mixture augmentation. We propose a frames mixture augmentation method that introduces additional foreground objects to create more complex driving scenes, alleviating the long-tailed distribution and enhancing the model’s robustness. Furthermore, the proposed module functions as an efficient plug-and-play module, capable of enhancing the performance of existing network architectures while maintaining minimal computational overhead. Extensive experiments demonstrate that our module achieves significant performance improvement in a range of methods with different input modalities. Siyuan Wang 0032, Yuejie Lu, Qiang Ling 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Learned Image Compression with Frequency Feature Interaction and Non-local Cross-similarity PriorabstractRecently, learned image compression models have achieved better compression performance than traditional non-learning image compression standards. Those learned models usually utilize spatial self-attention and Convolutional Neural Network (CNN) to extract non-local and local features and generate the latent representation. However, previous methods adopt a linear layer to fuse non-local and local features and lack the flexibility to adaptively adjust feature weights and capture complex non-linear interactions between distinct feature representations. Additionally, how to more effectively compress the latent representation based on its channel similarity characteristics remains unexplored. To solve the above issues, we propose a novel image compression method with frequency feature interaction and non-local cross-similarity prior. More specifically, we extend the previous spatial self-attention module and alternately use spatial and channel self-attention modules to extract non-local spatial and channel features, respectively, and depth-wise convolution is utilized to extract local features. As local features focus on high-frequency detail information and non-local features concentrate on low-frequency structural information, we propose a frequency interaction module (FIM) that generates two weight maps to dynamically fuse non-local and local features. Moreover, we observe the non-local cross-similarity in different channels of the latent representation, which indicates that different channels share similar non-local semantic and structural information, but have distinct local detail information. So, we design a dual transformer entropy model to emphasize non-local features and remove local features. Experiment results validate our method achieves promising compression performance on the Kodak, CLIC, and Tecnick datasets. Jian Wang 0132, Qiang Ling 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Data-Free Knowledge Distillation with Diffusion ModelsabstractRecently Data-Free Knowledge Distillation (DFKD) has garnered attention and can transfer knowledge from a teacher neural network to a student neural network without requiring any access to training data. Although diffusion models are adept at synthesizing high-fidelity photorealistic images across various domains, existing methods cannot be easiliy implemented to DFKD. To bridge that gap, this paper proposes a novel approach based on diffusion models, DiffDFKD. Specifically, DiffDFKD involves targeted optimizations in two key areas. Firstly, DiffDFKD utilizes valuable information from teacher models to guide the pre-trained diffusion models’ data synthesis, generating datasets that mirror the training data distribution and effectively bridge domain gaps. Secondly, to reduce computational burdens, DiffDFKD introduces Latent CutMix Augmentation, an efficient technique, to enhance the diversity of diffusion model-generated images for DFKD while preserving key attributes for effective knowledge transfer. Extensive experiments validate the efficacy of DiffDFKD, yielding state-of-the-art results exceeding existing DFKD approaches.We release our code at https://github.com/xhqi0109/DiffDFKD. Xiaohua Qi, Renda Li, Qiang Ling 0001, Jun Yu 0001, Ziyi Chen 0005, Peng Chang 0002, Jing Xiao 0006 |
ICME | 4 |
| 2025 | LVLM-MIR: Large Vision-Language Model with Parameter-Efficient Fine-Tuning for Multimodal Interleaved ReasoningabstractMultimodal interleaved reasoning, which requires models to understand interleaved image-text sequences and multiple images, is a critical challenge in contemporary AI. This paper proposes a parameter-efficient fine-tuning framework based on Large Vision-Language Models, with Qwen2.5-VL as the backbone and Low-Rank Adaptation for task-specific adaptation. The framework integrates four stages: multimodal input preprocessing to align with pre-training distributions, visual feature extraction via a modified Vision Transformer, cross-modal fusion via attention mechanisms, and response generation via an autoregressive decoder. By freezing pre-trained weights and fine-tuning low-rank adapters in both visual and language modules, it balances preserving general multimodal knowledge with optimizing target tasks, achieving high performance with low computational overhead. On the MIRAGE Challenge Track A Dataset, it performs strongly across subtasks, achieving an aggregate score of 0.7857 and securing second place in the challenge. Ablation studies confirm that joint LoRA fine-tuning of visual and language modules yields optimal results; limitations in fine-grained visual difference tasks indicate future directions in enhancing subtle feature capture and adaptive cross-modal alignment. Jun Yu 0001, Xilong Lu, Cong Wang 0039, Qiang Ling 0001 |
ACM Multimedia | 4 |
| 2025 | CMA-VC: Large Vision-Language Model for Cross-Modal Alignment in Intention-Oriented Video CaptioningabstractTraditional video captioning methods often produce generic descriptions that fail to align with specific user intentions, limiting their applicability in scenarios requiring customized information extraction. This paper proposes a novel intention-oriented controllable video captioning approach, which leverages large-scale vision-language models (InternViT and InternLM) and achieves parameter-efficient fine-tuning through Low-Rank Adaptation (LoRA). The proposed framework processes both video content and user-specified intentions via a unified cross-modal pipeline, dynamically aligning visual features with intent semantics to generate focused and contextually accurate captions. Experiments on the IntentVC dataset validate the effectiveness of the proposed method in generating intention-aligned captions, with the following performance metrics: BLEU@4 scores 44.38 on the public test set and 40.21 on the private test set; METEOR scores 63.79 and 60.07 respectively; CIDEr scores 230.33 and 208.15 respectively; ROUGE-L scores 61.75 and 57.14 respectively. Ablation studies confirm the significant effectiveness of joint LoRA adaptation on vision and language modules, as well as the sensitivity of performance to LoRA parameters. This work advances the field by enabling precise control over caption generation, enhancing the practical utility of video understanding systems in applications such as accessibility services and targeted video retrieval. Jun Yu 0001, Xilong Lu, Qiang Ling 0001 |
ACM Multimedia | 4 |
| 2025 | LVLM-HBA: Large Vision-Language Model with Cross-Modal Alignment for Human Behavior AnalysisabstractBodily Behaviour Recognition (BBR) and Eye Contact Detection (ECD) in multi-person group conversations are critical for understanding social dynamics, but traditional methods often rely solely on visual cues, lacking integration with semantic context. To address this, we propose a novel framework based on Large Vision-Language Models (LVLMs), leveraging their cross-modal alignment capability to fuse visual features (e.g., body postures, gaze directions) and linguistic semantics (e.g., behavioral category descriptions). A parameter-efficient tuning strategy using Low-Rank Adaptation (LoRA) is adopted, adapting only a subset of parameters in both the Language Model (LM) and Vision Transformer (ViT) modules, thus retaining pre-trained knowledge while reducing computational costs. The framework incorporates multi-task output heads to simultaneously predict BBR and ECD results. Experiments on the MPIIGroupInteraction dataset demonstrate superior performance: our method achieves 0.65 accuracy on BBR and 0.82 accuracy on ECD, outperforming state-of-the-art approaches by 0.02-0.03 in absolute terms. Ablation studies validate that applying LoRA to both LM and ViT with optimal hyperparameters yields the best results, confirming the importance of cross-modal synergy. This work highlights the potential of LVLMs in social behavior analysis, providing a lightweight and effective solution for understanding complex group interactions. Jun Yu 0001, Xilong Lu, Lingsi Zhu, Qiang Ling 0001 |
ACM Multimedia | 4 |
| 2025 | DRGNet: Dual-Relation Graph Network for point cloud analysis
Ce Zhou, Qiang Ling 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | Pyramidal structure-correlated refinement for robust face alignment
Qiyuan Dai 0002, Qiang Ling 0001 |
Knowl. Based Syst. | 2 |
| 2025 | GLALLM: Adapting LLMs for spatio-temporal wind speed forecasting via global-local aware modeling
Tangjie Wu, Qiang Ling 0001 |
Knowl. Based Syst. | 2 |
| 2025 | Adaptive depth position encoding for sparse query-based 3D object detection from multi-camera images
Junjie Linghu, Qiang Ling 0001 |
Pattern Recognit. | 2 |
| 2025 | Hybrid Representation Learning for End-to-End Multi-Person Pose EstimationabstractRecent multi-person pose estimation methods design end-to-end pipelines under the DETR framework. However, these methods involve complex keypoint decoding processes because the DETR framework cannot be directly used for pose estimation, which results in constrained performance and ineffective information interaction between human instances. To tackle this issue, we propose a hybrid representation learning method for end-to-end multi-person pose estimation. Our method represents instance-level and keypoint-level information as hybrid queries based on point set prediction and can facilitate parallel interaction between instance-level and keypoint-level representations in a unified decoder. We also employ the instance segmentation task for auxiliary training to enrich the spatial context of hybrid representations. Furthermore, we introduce a pose-unified query selection (PUQS) strategy and an instance-gated module (IGM) to improve the keypoint decoding process. PUQS predicts local pose proposals to produce scale-aware instance initializations and can avoid the scale assignment mistake of one-to-one matching. IGM refines instance contents and filters out invalid information with the message of cross-instance interaction and can enhance the decoder’s capability to handle queries of instances. Compared with current end-to-end multi-person pose estimation methods, our method can detect human instances and body keypoints simultaneously through a concise decoding process. Extensive experiments on COCO Keypoint and CrowdPose benchmarks demonstrate that our method outperforms some state-of-the-art methods. Qiyuan Dai 0002, Qiang Ling 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Mask-Guided Siamese Tracking With a Frequency-Spatial Hybrid NetworkabstractCurrent tracking methods often adopt a compact template to emphasize target-specific features, alongside an expansive search region to encapsulate surrounding environmental information. However, the employment of a small template size may result in the loss of critical contextual information, which can be particularly harmful in challenging scenarios. Moreover, current tracking methods predominantly focus on spatial or channel operations, neglecting the potential of the frequency domain. To resolve those issues, we propose a novel Mask-Guided Siamese Tracking (MGTrack) framework to enhance tracking efficacy from two perspectives. Firstly, we propose an innovative Template Mask Encoder (TME) that employs a large template to produce a learnable mask embedding, thus preserving more surrounding contextual cues while focusing on target-oriented discriminative features. Secondly, we propose a frequency-spatial hybrid network, which is composed of a Frequency-Spatial Fusion (FSF) module and a Frequency-Spatial Attention (FSA) module. Particularly, the FSF module integrates frequency blocks with local and global fusion blocks, effectively aggregating deep semantic features from the backbone network with shallow texture features. Additionally, the FSA module enables bidirectional information exchange between spatial and frequency attention during the feature interaction process. Experiments across short-term and long-term tracking benchmarks demonstrate that our MGTrack can achieve better tracking performance with fewer parameters and FLOPs than some state-of-the-art tracking frameworks. The code of our MGTrack is available athttps://github.com/jiabingxiing/MGTrack. Jiabing Xiong, Qiang Ling 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Dual Geometry Learning and Adaptive Sparse Attention for Point Cloud AnalysisabstractPoint cloud analysis is essential in accurately perceiving and analyzing real-world scenarios. Recently, transformer-based models have demonstrated great performance superiority in diverse domains. Nonetheless, directly applying transformers to point clouds is still challenging, primarily due to the computational intensity of transformers, which may significantly compromise their efficacy. Moreover, most methods typically rely on the relative 3D coordinates of point pairs to generate geometric information without fully exploiting the inherent local geometric properties. To tackle these challenges, we propose DGAS-Net, a novel architecture to enhance point cloud analysis. Specifically, we propose a Dual Geometry Learning (DGL) module to generate explicit geometric descriptors from triangular representations. These descriptors capture the local shape and geometric details of each point, serving as the foundation for deriving informative geometric features. Subsequently, we introduce a Dual Geometry Context Aggregation (DGCA) module to efficiently merge local geometric and semantic information. Furthermore, we design an Adaptive Sparse Attention (ASA) module to capture long-range information and expand the effective receptive field. ASA adaptively selects globally representative points and employs a novel vector attention mechanism for efficient global information fusion, thereby significantly reducing the computational complexity. Extensive experiments on four datasets demonstrate the superiority of DGAS-Net for various point cloud analysis tasks. The codes of DGAS-Net are available athttps://github.com/zcustc-10/DGAS-Net Ce Zhou, Qiang Ling 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Weakly Supervised Large-Scale Point Cloud Semantic Segmentation Based on Dual Consistency Constraints and Uncertainty-Aware FusionabstractWeakly supervised point cloud semantic segmentation has attracted more and more attention due to its capability to circumvent the time-consuming and expensive full labeling annotation process, which is required by fully supervised learning. However, existing weakly supervised methods typically rely solely on the sparsely labeled points for network training and cannot fully exploit the vast amount of unlabeled data. Moreover, recent weakly supervised segmentation methods are not efficient in enhancing network generalization and extracting discriminative features. To resolve these issues, we propose a novel framework (DCUF-Net) for weakly supervised point cloud semantic segmentation based on dual consistency constraints and uncertainty-aware fusion. Specifically, we first design dual perturbations in the data and feature domains to improve the network generalization and employ prediction consistency constraints between the two perturbed branches and the original branch. Additionally, we propose to jointly represent the prediction reliability of unlabeled points according to distribution confidence and uncertainty, which enables us to introduce an uncertainty-aware loss for all unlabeled points, providing additional supervised signals to optimize the network. To further extract more discriminative features, we propose an efficient regional context adaptive aggregation (RCAA) module to enhance information interaction between points. Extensive experiments on three large-scale datasets, including S3DIS, Toronto3D, and STPLS3D, demonstrate the superior performance of our method against some state-of-the-art methods. Ce Zhou, Qiang Ling 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Sparse LaneFormer: End-to-End Lane Detection With Sparse Proposals and InteractionsabstractLane detection is a challenging task due to the inherently sparse nature of lanes, which are characterized by their thin and long structures. However, existing lane detection methods often rely on dense anchors or interactions and may not well handle the sparse characteristics of lanes. Dense anchor-based methods are not fully end-to-end, necessitating non-maximum suppression to eliminate near-duplicate detection results. Recently, DETR-like methods have emerged, aiming to achieve end-to-end lane detection with sparse queries. Nevertheless, these methods establish dense interactions among all image tokens, which not only leads to substantial computational redundancy, but also complicates the extraction of sparse lane features from global feature maps. To tackle these issues, we propose a novel end-to-end lane detection framework called Sparse LaneFormer, which leverages both sparse proposals and interactions. Specifically, we introduce a Lane Aware Query (LAQ) as sparse lane proposals, which can predict learnable and explicit lane positions. Additionally, we introduce a Lane Token Selector (LTS) and a Linear Deformable Sampler (LDS) to establish sparse interactions among lane tokens. LTS selectively updates sparse lane tokens in the encoder, while LDS, in the decoder, extracts discriminative lane representations from global feature maps by sampling lane features along LAQ. By concentrating on the most informative tokens, Sparse LaneFormer can achieves state-of-the-art performance with only 10% of encoder tokens, resulting in an 80% reduction in overall computational cost compared to its dense counterpart. Our Sparse LaneFormer represents a pioneering effort in exploring a fully sparse design for lane detection, setting a new benchmark for efficiency and accuracy in this field. Qiang Ling 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Contrastive Learning With Multiple Prototypes for Unsupervised Domain Adaptive Semantic SegmentationabstractUnsupervised domain adaptive semantic segmentation aims to transfer knowledge from the annotated source domain to the unlabeled target domain. Recently, self-training methods have gained substantial attention, which leverage high-confidence predictions in the target domain as pseudo labels for supervision. However, limited exploration of intra-class variations across domains, including significant visual differences within each category, has led to misalignment between feature distribution across domains. In this article, we present a unified non-parametric distance-based online clustering method to efficiently maintain multiple centroid-based prototypes within each category subspace instead of one prototype for each category subspace, which enables prototypes to possess the capacity for richer feature representation. Then, considering the variance across different dimensions of a feature representation, we then extend the prototypes from centroid-based ones to distribution-based ones. Specifically, each subspace is modeled using a Gaussian mixture model which includes several anisotropic Gaussian distributions, aimed at prioritizing discriminative dimensions and obtaining a finer measurement of the pixel-to-prototype similarity. Meanwhile, a category-aware feature space is achieved through pixel-to-prototype contrastive learning to ensure the compactness of pixel features in the same subcategory and drive the separation between pixel features of different subcategories. What's more, multi-resolution features are utilized to promote diversity and robustness among intra-class prototypes. Experiments validate the competitiveness of our two prototype-based methods against existing state-of-the-art methods, with a mIoU of 76.8% on GTA$\rightarrow$Cityscapes, 68.4% on Synthia$\rightarrow$Cityscapes, 54.5% on Cityscapes$\rightarrow$DarkZurich and 56.4% on Cityscapes$\rightarrow$ACDC. Notably, our method is able to seamlessly integrate with existing UDA methods. Jun Yu 0001, Guochen Xie, Quansheng Liu, Zhen Kan, Lei Wang 0203, Qiang Ling 0001, Fang Gao 0001 |
IEEE Trans. Multim. | 7 |
| 2025 | Online Adaptive Optimal Control Algorithm Based on Weighted Policy IterationabstractIn this article, we propose a novel online learning algorithm based on weighted policy iteration (WPI) for addressing optimal control problems of nonlinear systems. WPI is proposed to deal with the influence of the neural network (NN) approximation error on the admissibility of the improved control policy. It is shown that the new iterative method can converge to the optimal solution uniformly. Utilizing NN approximation and experience replaying techniques, a WPI-based online learning algorithm is proposed. The new online algorithm distinguishes from previously known ones in that there can be fewer neurons in the hidden layer, giving rise to significant computational improvement. The assumption that the number of neurons needs to be sufficiently large can be dropped. Besides, instead of a standard persistently excited (PE) condition, only a relaxed PE condition is needed, which is also easier to check. Finally, numerical experiments are conducted to verify the effectiveness of our method. Wanlin Tan, Rui Luo 0003, Zhinan Peng, Qiang Ling 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Part-level Reconstruction for Self-Supervised Category-level 6D Object Pose Estimation with Coarse-to-Fine Correspondence OptimizationabstractSelf-supervised category-level 6D pose estimation stands as a fundamental task in computer vision. However, current self-supervised methods face two major challenges. Firstly, existing networks struggle to reconstruct precise object models due to significant part-level shape variations among specific categories. Secondly, they are impacted by the many-to-one ambiguity in the correspondences between pixels and point clouds. To address these challenges, we propose a novel approach that includes a Part-level Shape Reconstruction (PSR) module and a Coarse-to-Fine Correspondence Optimization (CFCO) module. In the (PSR) module, we introduce a part-level discrete shape memory to capture more fine-grained shape variations of different objects and use it to perform precise reconstruction. In the (CFCO) module, we utilize Hungarian matching to generate one-to-one pseudo labels at both region and pixel levels, which provides explicit supervision for the corresponding similarity matrices. We evaluate our method on the REAL275 and WILD6D datasets. Our extensive experiments show that our self-supervised approach outperforms existing methods and achieves new state-of-the-art results within the self-supervised framework. Jun Yu 0001, Liangxian Cui, Qiang Ling 0001 |
ACM Multimedia | 4 |
| 2024 | Enhanced encoder-decoder architecture for visual perception multitasking of autonomous driving
Muhammad Usman 0013, Zaka-Ud-Din Muhammad, Qiang Ling 0001 |
Expert Syst. Appl. | 3 |
| 2024 | Semantic segmentation for large-scale point clouds based on hybrid attention and dynamic fusion
Ce Zhou, Zhaokun Shu, Qiang Ling 0001 |
Pattern Recognit. | 4 |
| 2024 | A new metaheuristic honey badger-based maximum energy harvesting algorithm for thermoelectric generation system under dynamic operating conditions
Maryam Ejaz, Qiang Ling 0001 |
Soft Comput. | 2 |
| 2024 | Pixel-Wise Gamma Correction Mapping for Low-Light Image EnhancementabstractLow-light image enhancement aims to improve the visual quality of images captured under poor illumination and has caught much attention these years. However, existing low-light enhancement methods encounter many problems, e.g., they may not be robust to diverse low-light conditions or have to sacrifice computational efficiency for enhancement performance, which hinder their practical applications. To solve these problems, this paper proposes a novel enhancement method, called Pixel-Wise Gamma Correction Mapping (PWGCM), which combines our innovative pixel-wise Gamma Correction (GC) and deep learning. Compared with conventional GC, our pixel-wise GC is characterized by a set of gamma correction maps, which have the same size as the input image and are taken to replace the single global GC parameter of conventional GC. These gamma correction maps are generated from the low-light image input by a lightweight convolutional neural network at low computational cost. New no-reference loss functions are provided to train the network, ensuring reliable unsupervised learning. Furthermore, our PWGCM is enhanced by an iterative strategy, under which the low-light input image is iteratively enhanced based on the generated gamma correction maps and can yield visually pleasant results. Extensive experiments are done to compare our PWGCM with several state-of-the-art methods in terms of visual quality, efficiency, and auxiliary effects on high-level tasks. The comparison results confirm the superiority of our PWGCM. Xiangsheng Li, Manlu Liu 0001, Qiang Ling 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | FDNet: Frequency Decomposition Network for Learned Image CompressionabstractRecently learned image compression methods have achieved better rate-distortion performance than traditional non-learning image compression standards. Some previous image compression methods combine the local modeling capability of CNN with the long-range attention of Transformer to generate the latent representation. However, previous methods ignored the fact that Transformer pays attention to low-frequency feature learning while CNN focuses on high-frequency feature learning, resulting in insufficient fusion of these two structures. In this paper, we propose a novel image compression method with Frequency Decomposition Network (FDNet), which processes low-frequency and high-frequency components in different ways. More specifically, FDNet initially implements a dynamic frequency filter to adaptively decompose the features into low-frequency and high-frequency components. As invertible neural networks do not lose any information during the feature transformation and can be implemented by CNN residual networks, the invertible neural network block (INNB) is used to extract high-frequency local information. Then FDNet takes a hybrid attention block (HAB), which is composed of window-based multi-head self-attention (W-MSA) and channel attention, to extract window-based and global spatial low-frequency information. Besides, previous channel entropy models adopt CNN networks to remove high-frequency redundancy of the latent representation. However, there exists low-frequency redundancy between different channels of the latent representation. To solve this issue, FDNet further introduces the hybrid attention block to the channel entropy model. W-MSA and channel attention of the hybrid attention block can remove the window-based and global low-frequency redundancy, respectively. Extensive experiments demonstrate that FDNet achieves promising rate-distortion performance on the Kodak, CLIC and Tecnick datasets. Jian Wang 0132, Qiang Ling 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Hyper-Anchor Based Lane DetectionabstractAs a critical task in autonomous driving, lane detection has caught increasing attention. Due to the inherently thin and long structure of lanes and complex external environments, current lane detection methods may not perform well, particularly in challenging driving scenarios, where lanes might be hardly visible because of extreme illumination, occlusion, absence of lane marks, and so on. To tackle the lane detection problem, we propose a novel hyper-anchor, which can take flexible shapes and provide coarse estimates of lane points. Based on hyper-anchors, multi-level lane-aware feature aggregations are proposed to integrate local lane details with global context. Specifically, lane descriptors of an adaptive receptive field are initially constructed by extracting local lane features from neighboring rows and columns. Intra-lane aggregation establishes an information flow across rows and columns within each lane descriptor. By aggregating local lane features, accurate offsets between hyper-anchors and lane points can be obtained to fine-tune the lane locations. Additionally, inter-lane feature aggregation is conducted to capture the global structural relationship of lanes. Subsequently, lane confidence is predicted to measure the overall reliability of a lane prediction and reduce false positive predictions. Extensive experiments demonstrate that our hyper-anchor based method outperforms some state-of-the-art lane detection methods, especially in challenging driving scenarios. Meanwhile, our method also achieves desirable speed due to its efficient hyper-anchor generation and lane-aware feature aggregations. Qiang Ling 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Fusing hybrid attentive network with self-supervised dual-channel heterogeneous graph for knowledge tracing
Tangjie Wu, Qiang Ling 0001 |
Expert Syst. Appl. | 2 |
| 2023 | Attention-based global and local spatial-temporal graph convolutional network for vehicle emission prediction
Xihong Fei, Qiang Ling 0001 |
Neurocomputing | 2 |
| 2023 | A Bayesian Semisupervised Approach to Keyword Extraction with Only Positive and Unlabeled DataabstractIn the era of big data, people benefit from the existence of tremendous amounts of information. However, availability of said information may pose great challenges. For instance, one big challenge is how to extract useful yet succinct information in an automated fashion. As one of the first few efforts, keyword extraction methods summarize an article by identifying a list of keywords. Many existing keyword extraction methods focus on the unsupervised setting, with all keywords assumed unknown. In reality, a (small) subset of the keywords may be available for a particular article. To use such information, we propose a rigorous probabilistic model based on a semisupervised setup. Our method incorporates the graph-based information of an article into a Bayesian framework via an informative prior so that our model facilitates formal statistical inference, which is often absent from existing methods. To overcome the difficulty arising from high-dimensional posterior sampling, we develop two Markov chain Monte Carlo algorithms based on Gibbs samplers and compare their performance using benchmark data. We use a false discovery rate (FDR)-based approach for selecting the number of keywords, whereas the existing methods use ad hoc threshold values. Our numerical results show that the proposed method compared favorably with state-of-the-art methods for keyword extraction. History: Accepted by Ramaswamy Ramesh, Area Editor for Data Science and Machine Learning. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2023.1283 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2021.0234 ) at ( http://dx.doi.org/10.5281/zenodo.7348935 ). Guanshen Wang, Yichen Cheng, Yusen Xia, Qiang Ling 0001, Xinlei Wang 0001 |
INFORMS J. Comput. | 4 |
| 2023 | A viable framework for semi-supervised learning on realistic dataset
Guochen Xie, Jun Yu 0001, Qiang Ling 0001, Fang Gao 0001 |
Mach. Learn. | 4 |
| 2023 | GAF-Net: Geometric Contextual Feature Aggregation and Adaptive Fusion for Large-Scale Point Cloud Semantic SegmentationabstractLarge-scale point cloud semantic segmentation is a challenging task due to the complexity and diversity of real-world 3D scenes. Most existing methods primarily rely on spatial coordinates to learn geometric representations without fully exploring local structural relationships. Additionally, the semantic gap between the encoder and decoder in segmentation networks is an important factor that constrains model performance. To address these challenges, we propose a novel network architecture called GAF-Net, which comprises a Geometric Contextual Feature Aggregation (GCFA) module and a Multi-scale Feature Adaptive Fusion (MFAF) module. The GCFA module consists of three primary blocks: (1) a Geometric Edge Representation block, designed to leverage spatial relative position and orientation information between the center point and its neighbors to capture detailed local geometric structural relations; (2) a Point Geometry Prior block, aimed at extracting explicit geometric priors for each point from raw point clouds. This block is lightweight and parameter-free; (3) a Geometry-Aware Attentive Pooling block, which combines semantic features with learned geometric representations, enabling the learning and aggregation of informative local contextual features. Our proposed MFAF module integrates multi-scale features by introducing an adaptive fusion approach. It effectively bridges the semantic gap between the encoder and decoder and mitigates the information loss caused by random sampling. Extensive experimental results on three large-scale benchmark datasets including S3DIS, Toronto3D, and SemanticKITTI demonstrate the superior performance of our proposed GAF-Net. Ce Zhou, Qiang Ling 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Robust Video Stabilization based on Motion DecompositionabstractVideo stabilization aims to eliminate camera jitter and improve the visual experience of shaky videos. Video stabilization methods often ignore the active movement of the foreground objects and the camera, and may result in distortion and over-smoothing problems. To resolve these issues, this paper proposes a novel video stabilization method based on motion decomposition. Since the inter-frame movement of foreground objects is different from that of the background, we separate foreground feature points from background feature points by modifying the classic density based spatial clustering method of applications with noise (DBSCAN). The movement of background feature points is consistent with the movement of the camera, which can be decomposed into the camera jitter and the active movement of the camera. And the movement of foreground feature points can be decomposed into the movement of the camera and the active movement of foreground objects. Based on motion decomposition, we design first-order and second-order trajectory smoothing constraints to eliminate the high-frequency and low-frequency components of the camera jitter. To reduce content distortion, shape-preserving constraints, and regularization constraints are taken to generate stabilized views of all feature points. Experimental results demonstrate the effectiveness and robustness of the proposed video stabilization method on a variety of challenging videos. Jian Wang 0132, Qiang Ling 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Ranking-Based Siamese Visual TrackingabstractCurrent Siamese-based trackers mainly formulate the visual tracking into two independent subtasks, including classification and localization. They learn the classification subnetwork by processing each sample separately and neglect the relationship among positive and negative samples. Moreover, such tracking paradigm takes only the classification confidence of proposals for the final prediction, which may yield the misalignment between classification and localization. To resolve these issues, this paper proposes a ranking-based optimization algorithm to explore the relationship among different proposals. To this end, we introduce two ranking losses, including the classification one and the IoU-guided one, as optimization constraints. The classification ranking loss can ensure that positive samples rank higher than hard negative ones, i.e., distractors, so that the trackers can select the foreground samples successfully without being fooled by the distractors. The IoU-guided ranking loss aims to align classification confidence scores with the Intersection over Union(IoU) of the corresponding localization prediction for positive samples, enabling the well-localized prediction to be represented by high classification confidence. Specifically, the proposed two ranking losses are compatible with most Siamese trackers and incur no additional computation for inference. Extensive experiments on seven tracking benchmarks, including OTB100, UAV123, TC128, VOT2016, NFS30, GOT-10k and LaSOT, demonstrate the effectiveness of the proposed ranking-based optimization algorithm. The code and raw results are available at https://github.com/sansanfree/RBO. Qiang Ling 0001 |
CVPR | 2 |
| 2022 | Micro Expression Generation with Thin-plate Spline Motion Model and Face ParsingabstractMicro-expression generation aims at transfering the expression from the driving videos to the source images, which can be viewed as a motion transfer task. Recently, several works have been proposed to tackle this problem and achieve great performance. However, due to the intrinsic complexity of the face motion and different attributes of face regions, the task still remains challenging. In this paper, we propose an end-to-end unsupervised motion transfer network to tackle this challenge. As the motion of the face is non-rigid, we adopt an effective and flexible thin-plate spline motion estimation method to estimate the optical flow of the face motion. What's more, we find that several faces with eyeglasses show weird deformation in motion transfering. Thus, we introduce face parsing method to pay specific attention to the eyeglasses regions to ensure the reasonability of the deformation. We conduct several experiments on the provided datasets of the ACM MM 2022 micro-expression grand challenge (MEGC2022) and compare our method with several other typical methods. In comparison, our method shows the best performance. We (Team: USTC-IAT-United) also compare our method with other competitors' in MEGC2022, and the expert evaluation results show that our method performs best, which verifies the effectiveness of our method. Our code is available at https://github.com/HowToNameMe/micro-expression Jun Yu 0001, Guochen Xie, Zhongpeng Cai, Peng He 0004, Fang Gao 0001, Qiang Ling 0001 |
ACM Multimedia | 6 |
| 2022 | Event-Triggered Stabilization of a Linear System With Model Uncertainty and i.i.d. Feedback DropoutsabstractThis article investigates the input-to-state stability (ISS) of continuous-time networked control systems with model uncertainty and bounded noise based on event triggering. The feedback loop is closed over an unreliable digital communication network. Feedback packets suffer from network delay and may be dropped in an independent and identically distributed (i.i.d.) way, which may harm to the concerned stability. This article focuses on a Lyapunov-based event-triggered control design scheme with the consideration of i.i.d. packet dropouts. By designing a state-dependent event-triggering threshold and updating methods, it can still ensure ISS for the concerned multidimensional system in the presence of i.i.d. packet dropouts and model uncertainty without exhibiting the Zeno behavior. Simulations are done to verify the effectiveness of the achieved results. Xianghua Jiang, Qiang Ling 0001 |
IEEE Trans. Cybern. | 2 |
| 2022 | Bit-Rate Conditions for the Consensus of Quantized Multiagent Systems Based on Event TriggeringabstractThis article investigates the asymptotic consensus problem for multiple discrete-time agents with general dynamics under event-triggered sampling. In such multiagent systems (MASs), the information exchanged between neighboring agents is quantized and transmitted through a digital communication network. Due to the limited bandwidth, it is impossible for agents to capture the precise states of neighbors and only their quantized version can be received, which may degrade or even break the consensus of these MASs. In order to ensure the desired consensus and save the occupied network bandwidth, a model-based event-triggering strategy is well co-designed with a dynamic quantization method. Under the proposed strategy, the error between the real state and its estimated version at neighbors is always bounded by an event-triggering function, which is also taken to compute an upper bound on the control inputs of all agents. We derive a sufficient bit-rate condition required for the desired asymptotic consensus. That condition is mainly determined by the dynamics and communication network topology of agents. Since our event-triggering strategy takes advantage of the extra information extracted from the event-triggering time instants of all feedback packets, the derived bit rate can be lower than that of the conventional periodic sampling strategies, while still ensuring the consensus of the concerned MAS. Simulations are done to illustrate the advantages of the derived bit-rate condition. Na Lin 0004, Qiang Ling 0001 |
IEEE Trans. Cybern. | 2 |
| 2022 | Event-Triggered Stabilizing Bit Rate Conditions for an n-Dimensional Linear System With i.i.d. Feedback DropoutsabstractThis article aims to stabilize an n -dimensional linear time-invariant (LTI) system, whose feedback packets are transmitted through a digital communication network. The digital network suffers from network delay and independent and identically distributed (i.i.d.) feedback dropouts, which may destabilize the system. The coupling among multiple state variables may further harm the stability of the system. In order to deal with these issues and save the occupied bandwidth of the feedback network, we propose a periodic event-triggering strategy. In our strategy, the state is measured periodically, but only quantized and transmitted when a certain condition is triggered. By well balancing the state coupling and making full use of both the information inside transmitted feedback packets and the one carried by sampling time instants, our strategy can maintain the desired mean square stability at a lower bit rate than conventional periodic sampling policies. The obtained stabilizing bit rate conditions are determined by the processing and network delays, the dropout rate, and the unstable eigenvalues of the system matrix, but independent of the process noise. Moreover, the lack of the direct state access does not incur any additional stabilizing bit rate. Simulations are done to confirm the effectiveness of the obtained stabilizing bit rate conditions. Yuan Liu 0031, Qiang Ling 0001 |
IEEE Trans. Cybern. | 2 |
| 2021 | Recurrent knowledge tracing machine based on the knowledge state of studentsabstractAbstract Knowledge tracing (KT) aims to closely trace the knowledge level of students during their learning. KT is often implemented by intelligent tutoring systems (ITS) to predict the student performance, and schedule the individual learning plan for each student. However, most existing KT models cannot predict the knowledge forgetting well due to the lack of temporal information and some KT models for knowledge forgetting prediction are not accurate enough. To resolve these issues, this paper proposes a recurrent knowledge tracing machine (RKTM), which temporally enriches the encoding of knowledge tracing machine (KTM) and difficulty, student ability, skill, and student skill practice history (DAS3H) with the knowledge state of students. RKTM consists of two major components, including the tracing component and the predicting component. The tracing component temporally traces the knowledge state of a student, while the predicting component captures the interaction between the current learning scenario and the current knowledge state of that student and provides accurate prediction of student performance. Experiments show that the proposed RKTM well outperforms KTM, DAS3H, and some other models. Zefeng Lai, Lei Wang 0203, Qiang Ling 0001 |
Expert Syst. J. Knowl. Eng. | 3 |
| 2021 | Multimodal Inputs Driven Talking Face Generation With Spatial-Temporal DependencyabstractGiven an arbitrary speech clip or text information as input, the proposed work aims to generate a talking face video with accurate lip synchronization. Existing works mainly have three limitations. (1) A single-modal learning is adopted with either audio or text as input, hence it lacks the complementarity ofmultimodal inputs. (2) Each frame is generated independently, hence it ignores thetemporal dependencybetween consecutive frames. (3) Each face image is generated by the traditional convolution neural network (CNN) with a local receptive field, hence it cannot effectively capture thespatial dependencywithin internal representations of face images. To overcome these problems above, we decompose the talking face generation task into two steps: mouth landmarks prediction and video synthesis. First, a multimodal learning method is proposed to generate accurate mouth landmarks with multimedia inputs (both text and audio). Second, a network named Face2Vid is proposed to generate video frames conditioned on the predicted mouth landmarks. In Face2Vid, the optical flow is employed to model the temporal dependency between frames, meanwhile, a self-attention mechanism is introduced to model the spatial dependency across image regions. Extensive experiments demonstrate that our approach can generate photo-realistic video frames with the background, and exhibit the superiorities on accurate synchronization of lip movements and smooth transition of facial movements. Lingyun Yu 0002, Jun Yu 0001, Qiang Ling 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Adaptively Meshed Video StabilizationabstractVideo stabilization is essential for improving the visual quality of shaky videos. Current video stabilization methods usually take feature trajectories in the background to estimate one global transformation matrix or several transformation matrices based on a fixed mesh, and warp shaky frames into their stabilized views. However, these methods may not model the shaky camera motion well in complicated scenes, such as scenes containing large foreground objects or strong parallax, and may result in notable visual artifacts in the stabilized videos. To resolve the above issues, this paper proposes an adaptively meshed method to stabilize a shaky video based on all of its feature trajectories and an adaptive blocking strategy. More specifically, we first extract the feature trajectories of the shaky video and then generate a triangle mesh according to the distribution of the feature trajectories in each frame. Then, the transformations between shaky frames and their stabilized views over all triangular grids of the mesh are calculated to stabilize the shaky video. Since more feature trajectories can usually be extracted from all of the regions, including both the background and foreground regions, a finer mesh will be obtained and provided for camera motion estimation and frame warping. We estimate the mesh-based transformations of each frame by solving a two-stage optimization problem. Moreover, foreground and background feature trajectories are no longer distinguished and both contribute to the estimation of the camera motion in the proposed optimization problem, yielding better estimation performance than previous works, particularly in challenging videos with large foreground objects or strong parallax. To further enhance the robustness of our method, we propose two adaptive weighting mechanisms to improve its spatial and temporal adaptability. Experimental results demonstrate the effectiveness of our method in producing visually pleasing stabilization effects in various challenging videos. Minda Zhao, Qiang Ling 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Bit-Rate Conditions for the Consensus of Quantized Multiagent Systems With Network-Induced Delays Based on Event TriggeringabstractThis paper is mainly concerned with bit-rate conditions to guarantee the consensus of unstable scalar multiagent systems by event-triggering strategies. In such systems, state information is quantized and transmitted among agents through digital communication networks which suffer from network-induced delays. In order to save communication resources, we implement a periodic event-triggering scheme, under which agents quantize and transmit their own states when a predefined event is triggered. By introducing a well-designed internal saturation function, the proposed strategy can place a priori bounds on the control inputs of agents and guarantee the asymptotic consensus of agents at a finite bit rate, even under network-induced delays. By extracting extra information from the receive time instants of data packets, our strategy can require less information carried by data packets and a lower bit rate to maintain the desired asymptotic consensus than some state-of-the-art time-triggered consensus strategies. Under our strategy, the obtained bit-rate conditions depend on the network-induced delays, the unstable eigenvalue of agent dynamics, and the network topology. The simulation results are provided to illustrate the effectiveness of these bit-rate conditions. Jiayu Chen 0004, Qiang Ling 0001 |
IEEE Trans. Cybern. | 2 |
| 2021 | Learning to Rank Proposals for Siamese Visual TrackingabstractRecently, Siamese network based trackers with region proposal networks(RPN) decompose the visual tracking task into classification and regression, and have drawn much attention. However, previous Siamese trackers process all the training samples equally to learn the desired network, and only take the classification scores of proposals to locate the tracked target at the inference stage. To address the above issues, we propose a simple, yet effective strategy to rank the importance of training samples, and pay more attention to the important samples, which can facilitate the classification optimization. Moreover, we propose a lightweight ranking network to generate the ranking scores for proposals. Higher scores are assigned to proposals whose Intersection over Union(IoU) with the ground-truth are larger. The combination of classification and ranking scores serves as a new proposal selection criterion for online tracking, and can boost the tracking performance significantly. Our proposed method could be easily integrated into existing RPN-based Siamese networks in an end-to-end fashion. Extensive experiments are conducted on 10 tracking benchmarks, including NFS, UAV123, OTB2015, Temple-Color, VOT2016, VOT2017, VOT2019, TrackingNet, GOT-10K and LaSOT. The proposed method achieves a state-of-the-art tracking accuracy with a real-time speed. Qiang Ling 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Bilateral Upsampling Network for Single Image Super-Resolution With Arbitrary Scaling FactorsabstractSingle Image Super-Resolution (SISR) is essential for many computer vision tasks. In some real-world applications, such as object recognition and image classification, the captured image size can be arbitrary while the required image size is fixed, which necessitates SISR with arbitrary scaling factors. It is a challenging problem to take a single model to accomplish the SISR task under arbitrary scaling factors. To solve that problem, this paper proposes a bilateral upsampling network which consists of a bilateral upsampling filter and a depthwise feature upsampling convolutional layer. The bilateral upsampling filter is made up of two upsampling filters, including a spatial upsampling filter and a range upsampling filter. With the introduction of the range upsampling filter, the weights of the bilateral upsampling filter can be adaptively learned under different scaling factors and different pixel values. The output of the bilateral upsampling filter is then provided to the depthwise feature upsampling convolutional layer, which upsamples the low-resolution (LR) feature map to the high-resolution (HR) feature space depthwisely and well recovers the structural information of the HR feature map. The depthwise feature upsampling convolutional layer can not only efficiently reduce the computational cost of the weight prediction of the bilateral upsampling filter, but also accurately recover the textual details of the HR feature map. Experiments on benchmark datasets demonstrate that the proposed bilateral upsampling network can achieve better performance than some state-of-the-art SISR methods. Menglei Zhang, Qiang Ling 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Supervised Pixel-Wise GAN for Face Super-ResolutionabstractFor many face-related multimedia applications, low-resolution face images may greatly degrade the face recognition performance and necessitate face super-resolution (SR). Among the current SR methods, MSE-oriented SR methods often produce over-smoothed outputs and could miss some texture details while GAN-oriented SR methods may generate artifacts which are harmful to face recognition. To resolve the above issues, this paper presents a supervised pixel-wise Generative Adversarial Network (SPGAN) that can resolve a very low-resolution face image of$16\times 16$or smaller pixel-size to its larger version of multiple scaling factors ($2\times$,$4\times$,$8\times$and even$16\times$) in a unified framework. Being different from traditional unsupervised discriminators which generate a single number to represent the likelihood whether the input image is real or fake, the proposed supervised pixel-wise discriminator mainly focus on whether each pixel of the generated SR face image is as photo-realistic as its corresponding pixel in the ground-truth HR (high-resolution) face image. To further improve the face recognition performance of SPGAN, we take advantage of the face identity prior by sending two inputs to the discriminator, including an input face image (either a real HR face image or its corresponding SR face image) and its face features which are extracted from a pre-trained face recognition model. Due to the introduced face identity prior, the identity-based discriminator can pay more attention to texture details which are closely related to face recognition. Extensive experiments demonstrate that the proposed SPGAN can achieve more photo-realistic SR images and higher face recognition accuracy than some state-of-the-art methods. Menglei Zhang, Qiang Ling 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | Model-Based Periodic Event-Triggered Control Strategy to Stabilize a Scalar Nonlinear SystemabstractThis article focuses on the problem of stabilizing a scalar continuous-time nonlinear system under bounded network delay and process noise. In order to save the feedback network's bandwidth, a model-based periodic event-triggered control policy is utilized to maintain longer intersampling intervals, which are at least as long as the sampling period of periodic policies. Furthermore, the event-triggering condition is only checked intermittently at fixed time instants, i.e., the sampling time instants. Without acknowledgment (ACK), the updating of nominal models at the sensor and at the controller are asynchronous. The two cases, where the network delay is either less than the sampling period or larger than the sampling period, are investigated. In comparison with periodic sampling methods, our scheme can make full use of the received data packets, particularly their sampling time instant information, which yields a lower occupied bit rate while guaranteeing the desired input-to-state stability. Note that the obtained bit rate conditions are only related to the Lipschitz parameter, the bound of network delay, the number of quantization bits, and the sampling period. The bounded process noise will not incur any increase of the stabilizing bit rate under the proposed strategy. Rundong Dou, Qiang Ling 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2020 | Contour-Aware Long-Term Tracking With Reliable Re-DetectionabstractRecently discriminative correlation filter (DCF) based methods have gained much popularity for their excellent performance and high efficiency. However, most of them perform poorly in long-term tracking as they are not equipped with an effective mechanism to evaluate the quality of tracking results and correct tracking errors. To resolve such issue, this paper proposes a long-term tracking method, which consists of two components, including tracking-by-detection and re-detection. The tracking-by-detection part is built upon the DCF framework by incorporating a contour constraint map, which could identify non-target samples and refine the tracking results efficiently in presence of some challenging situations, such as deformation and occlusion. Benefited by our proposed re-detection strategy, the missing/occluded target could be captured immediately after it reappears. Moreover, a re-detected result is allowed to replace the original tracking result only when it owns a performance gap against the original one, which can reduce the risk of wrong substitution and well enhance the long-term tracking robustness. Extensive experiments on OTB2015, Temple-Color, UAV20L and VOT-LT2018 show that the proposed long-term tracking method outperforms state-of-the-art hand-crafted based methods, and even some deep learning based methods. Qiang Ling 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | PWStableNet: Learning Pixel-Wise Warping Maps for Video StabilizationabstractAs the videos captured by hand-held cameras are often perturbed by high-frequency jitters, stabilization of these videos is an essential task. Many video stabilization methods have been proposed to stabilize shaky videos. However, most methods estimate one global homography or several homographies based on fixed meshes to warp the shaky frames into their stabilized views. Due to the existence of parallax, such single or a few homographies can not well handle the depth variation. In contrast to these traditional methods, we propose a novel video stabilization network, called PWStableNet, which comes up pixel-wise warping maps, i.e., potentially different warping for different pixels, and stabilizes each pixel to its stabilized view. To our best knowledge, this is the first deep learning based pixel-wise video stabilization. The proposed method is built upon a multi-stage cascade encoder-decoder architecture and learns pixel-wise warping maps from consecutive unstable frames. Inter-stage connections are also introduced to add feature maps of a former stage to the corresponding feature maps at a latter stage, which enables the latter stage to learn the residual from the feature maps of former stages. This cascade architecture can produce more precise warping maps at latter stages. To ensure the correct learning of pixel-wise warping maps, we use a well-designed loss function to guide the training procedure of the proposed PWStableNet. The proposed stabilization method achieves comparable performance with traditional methods, but stronger robustness and much faster processing speed. Moreover, the proposed stabilization method outperforms some typical CNN-based stabilization methods, especially in videos with strong parallax. Codes will be provided at https://github.com/mindazhao/pix-pix-warping-video-stabilization. Minda Zhao, Qiang Ling 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Mining Audio, Text and Visual Information for Talking Face GenerationabstractProviding methods to support audio-visual interaction with growing volumes of video data is an increasingly important challenge for data mining. To this end, there has been some success in speech-driven lip motion generation or talking face generation. Among them, talking face generation aims to generate realistic talking heads synchronized with the audio or text input. This task requires mining the relationship between audio signal/text and lip-sync video frames and ensures the temporal continuity between frames. Due to the issues such as polysemy, ambiguity, and fuzziness of sentences, creating visual images with lip synchronization is still challenging. To overcome the problems above, we present a data-mining framework to learn the synchronous pattern between different channels from large recorded audio/text dataset and visual dataset, and apply it to generate realistic talking face animations. Specifically, we decompose this task into two steps: mouth landmarks prediction and video synthesis. First, a multimodal learning method is proposed to generate accurate mouth landmarks with multimedia inputs (both text and audio). Second, a network named Face2Vid is proposed to generate video frames conditioned on the predicted mouth landmarks. In Face2Vid, optical flow is employed to model the temporal dependency between frames, meanwhile, a self-attention mechanism is introduced to model the spatial dependency across image regions. Extensive experiments demonstrate that our method can generate realistic videos with background, and exhibit the superiorities on accurate synchronization of lip movements and smooth transition of facial movements. Lingyun Yu 0002, Jun Yu 0001, Qiang Ling 0001 |
ICDM | 3 |
| 2019 | Deep Neural Network Based 3D Articulatory Movement Prediction Using Both Text and Audio Inputs
Lingyun Yu 0002, Jun Yu 0001, Qiang Ling 0001 |
MMM (1) | 3 |
| 2019 | Spatial-aware correlation filters with adaptive weight maps for visual tracking
Qiang Ling 0001 |
Neurocomputing | 2 |
| 2019 | An Iterative Feedback-Based Change Detection Algorithm for Flood Mapping in SAR ImagesabstractThis letter proposes a novel algorithm for the unsupervised detection of flood mapping in synthetic aperture radar (SAR) images. In the literature, unsupervised change detection of SAR images mainly consists of two steps, i.e., first generating a difference image from two given images and then binarizing the difference image to produce the desired change map. Conventional change detection algorithms usually execute these two steps sequentially and separately. In contrast, our algorithm introduces the feedback of the obtained intermediate change maps into both generation and binarization of the difference image. More specifically, we adjust the weights of neighboring pixels in generating the difference image according to the intermediate change maps. With the fed-back intermediate change maps, we also extend the conventional single binarizing threshold for all pixels of the difference image to threshold maps, i.e., two individual binarizing thresholds are defined for each pixel of the difference image and the threshold maps are adjusted accordingly. Due to such feedback of the intermediate change maps, we may obtain a better difference image and generate more precise change maps, which can surely be fed back again. This iterative execution of the above-mentioned generation and binarization of difference images is terminated when a predefined Markov energy function stops decreasing, i.e., it reaches a local minimum. Experiments with several SAR image data sets with floods show that our algorithm consistently outperforms several state-of-the-art algorithms. Minda Zhao, Qiang Ling 0001, Feng Li 0042 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2019 | Improving hierarchical mobile video caching through distributed cross-layer coordination
Feng Li 0042, Lixiang Xu, Shihui Duan, Wenfu Wu, Qiang Ling 0001 |
Multim. Tools Appl. | 6 |
| 2019 | Synthesizing 3D Trump: Predicting and Visualizing the Relationship Between Text, Speech, and Articulatory MovementsabstractThe movements of articulators, such as lips, tongue and teeth, play an important role in increasing the language expression capability by unmasking the information hid in text or speech. Hence, it is necessary to deeply mine and visualize the relationship between text, speech and articulatory movements for understanding language in multi-modality and multi-level. As a case study, given text and audio of President Donald John Trump, this paper synthesizes a high quality 3D animation of him speaking with accurate synchronicity between speech and articulators. First, visual co-articulation is modeled by predicting the mapping from text/speech to articulatory movements. Then, based on a reconstructed 3D head model, physiological characteristics and statistical learning are combined to visualize each phoneme. Finally, the visualization results of consecutive phonemes are fused by visual co-articulation model to generate synchronized articulatory animations. Experiments show that the system can not only produce photo-realistic results in front but also distinguish the visual differences among phonemes from unconstrained views. Jun Yu 0001, Qiang Ling 0001, Changwei Luo, Chang Wen Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Stabilization of Traffic Videos Based on Both Foreground and Background Feature TrajectoriesabstractThis paper considers stabilizing traffic videos, which are recorded by cameras mounted on moving vehicles. Compared with videos captured by hand-held cameras, traffic videos are more difficult to stabilize due to dynamic scenes, higher frequency camera jitter, more moving foreground objects and more serious parallax. The conventional video stabilization methods usually estimate the camera jitter from the background feature trajectories which are mainly determined by the camera motion, and then stabilize videos with the estimated jitter. These methods have to correctly distinguish background and foreground feature trajectories, which is not trivial, and may suffer from performance degradation when large foreground objects exist and no enough number of background feature trajectories can be obtained. To resolve these issues, this paper proposes a novel stabilization method, under which background and foreground feature trajectories are no longer distinguished and work together to yield stabilized trajectories. More specifically, the movement of all feature trajectories is modeled as the summation of the camera motion and the object motion. By solving an optimization problem, we can remove the high frequency components of the camera motion, i.e., the camera jitter, and stabilize videos. Parallax is also treated as object motion and contributes feature trajectories for video stabilization. As our method makes use of both foreground and background feature trajectories, it can outperform the conventional stabilization methods which use only background feature trajectories, especially when there are large foreground objects and the number of extracted background feature trajectories is small. Furthermore, some refinements are proposed to speed up our method and enhance its robustness. Experiments are done and confirm its performance superiority. Qiang Ling 0001, Minda Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Adaptive Convolution for Object DetectionabstractIt is quite challenging to detect objects, especially, small objects, in complex scenes. To solve this problem, we propose a novel module named as adaptive convolution block (ACB), which adaptively adjusts the parameters of convolutional filters according to the current feature maps, and then, filter these feature maps with the obtained adaptive convolutional filters to generate enhanced features. Due to such adaptive convolution, the enhanced features can pay more attention to the concerned objects, suppress the interference information caused by irrelevant surroundings, and efficiently improve the detection accuracy. The proposed ACB is light weight and fast. By directly embedding the ACB into the single shot detection framework, we construct a novel real-time adaptive convolutional detector (ACD). Experiments on PASCAL VOC and MS COCO benchmarks confirm that our ACD outperforms the existing state-of-the-art single-stage detection models, and achieves a better tradeoff between accuracy and speed. Chunlin Chen 0002, Qiang Ling 0001 |
IEEE Trans. Multim. | 2 |
| 2019 | BLTRCNN-Based 3-D Articulatory Movement Prediction: Learning Articulatory Synchronicity From Both Text and Audio InputsabstractPredicting articulatory movements from audio or text has diverse applications, such as speech visualization. Various approaches have been proposed to solve the acoustic-articulatory mapping problem. However, their precision is not high enough with only acoustic features available. Recently, deep neural network (DNN) has brought tremendous success in various fields, like speech recognition and image processing. To increase the accuracy, we propose a new network architecture for articulatory movement prediction with both text and audio inputs, called a bottleneck long-term recurrent convolutional neural network (BLTRCNN). To the best of our knowledge, it is the first time to predict articulatory movements based on DNN by fusing text and audio inputs. Our BLTRCNN consists of two networks. The first is the bottleneck network, generating a compact bottleneck features of text information for each frame independently. The second, including convolutional neural network, long short-term memory and skip connection, is called the long-term recurrent convolutional neural network (LTRCNN). LTRCNN is used for articulatory movement prediction when bottleneck features, acoustic features, and text features are integrated as inputs together. Experiments show that the proposed BLTRCNN achieves the state-of-the-art root-mean-square error (RMSE) 0.528 mm and the correlation coefficient 0.961. Moreover, we also demonstrate how text information complements acoustic features in this prediction task. Lingyun Yu 0002, Jun Yu 0001, Qiang Ling 0001 |
IEEE Trans. Multim. | 3 |
| 2018 | Synthesizing 3D Acoustic-Articulatory Mapping Trajectories: Predicting Articulatory Movements by Long-Term Recurrent Convolutional Neural NetworkabstractRobust and accurate predicting of articulatory movements has various important applications, such as 3D articulatory animations and visual communication. Various approaches have been proposed to solve the acoustic-articulatory mapping problem. However, their precision is not high enough. Recently, deep neural network (DNN), especially convolutional neural network (CNN) and recurrent neural network (RNN), has brought tremendous success in speech recognition and synthesis. To increase the accuracy, we propose a new network architecture for acoustic-articulatory mapping, called long-term recurrent convolutional neural network (LTRCNN). The network consists of CNN, RNN and a skip connection. CNN can model the spectral correlation among acoustic features efficiently. RNN, like long short-term memory (LSTM), can learn the temporal context information from sequential data powerfully. Besides, skip connections can increase the input representation from different levels to preserve the feature information. Experiments show that LTRCNN achieves the state-of-the-art root-mean-squared error (RMSE) with 0.690 mm and the correlation coefficient with 0.949 in this prediction task. Lingyun Yu 0002, Jun Yu 0001, Qiang Ling 0001 |
VCIP | 3 |
| 2018 | Boosting image sentiment analysis with visual attention
Kaikai Song, Ting Yao 0003, Qiang Ling 0001, Tao Mei 0001 |
Neurocomputing | 3 |
| 2018 | A Feedback-Based Robust Video Stabilization Method for Traffic VideosabstractTraffic videos are often recorded by vehicle-mounted cameras. Compared with videos recorded by handheld cameras, traffic videos suffer from more challenges, such as higher frequency and more violent jitters, dynamic scenes, large moving objects, and parallax, which can result in significant visual quality degradation. To address these challenges for traffic videos, we propose a special stabilization method. The key aspect of our method is a feedback strategy that divides the extracted feature trajectories into background trajectories and foreground trajectories by feeding back the previous trajectory classification results. The method can perform robustly, even in the case of large moving objects and parallax. Furthermore, our method maintains the number of available background trajectories within a reasonable range via two refinement strategies. One strategy attempts to reliably recover background trajectories from misjudged foreground trajectories when there are an insufficient number of background trajectories. The other strategy can adaptively adjust the number of feature points in each frame to efficiently avoid too many or too few background trajectories. With the obtained background trajectories, a homography matrix between each frame and its stabilized view is computed and implemented to warp the frame image to produce smooth videos. Experimental results confirm that our method is both effective in stabilizing traffic videos and quite robust against large moving objects and parallax. Qiang Ling 0001, Sibin Deng, Feng Li 0042, Qinghua Huang, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Detecting shot boundary with sparse coding for video summarization
Ting Yao 0003, Qiang Ling 0001, Tao Mei 0001 |
Neurocomputing | 3 |
| 2016 | Multi-Scale Triplet CNN for Person Re-IdentificationabstractPerson re-identification aims at identifying a certain person across non-overlapping multi-camera networks. It is a fundamental and challenging task in automated video surveillance. Most existing researches mainly rely on hand-crafted features, resulting in unsatisfactory performance. In this paper, we propose a multi-scale triplet convolutional neural network which captures visual appearance of a person at various scales. We propose to optimize the network parameters by a comparative similarity loss on massive sample triplets, addressing the problem of small training set in person re-identification. In particular, we design a unified multi-scale network architecture consisting of both deep and shallow neural networks, towards learning robust and effective features for person re-identification under complex conditions. Extensive evaluation on the real-world Market-1501 dataset have demonstrated the effectiveness of the proposed approach. Jiawei Liu 0001, Zhengjun Zha, Q. I. Tian, Dong Liu 0002, Ting Yao 0003, Qiang Ling 0001, Tao Mei 0001 |
ACM Multimedia | 6 |
| 2016 | A feedback-based adaptive data migration method for hybrid storage VOD caching systems
Qiang Ling 0001, Lixiang Xu, Jinfeng Yan, Yicheng Zhang 0001, Feng Li 0042 |
Multim. Tools Appl. | 1 |
| 2015 | To go or not to go green: An economic analysisabstractGreen, i.e., energy-efficient, communications technologies have been a trend for next-generation cellular communications systems. An imperative question for network operators to address is whether and to what extent one should move to embrace such newly proposed green communications technologies. On one hand, the introduction of green communications technologies saves the operating cost, and on the other hand, it may also lead to some extent of degradation of the quality of service, which would drive users away towards other operators. In this paper, a preliminary economic analysis is developed to address such a tension. An operator chooses to upgrade a proportion of its infrastructure from the legacy technology to some green communications technology, and the goal of analysis is to figure out how large the proportion should be, depending upon system parameters including quality of service, price, and initial user distribution. The analysis reveals that a variety of possibilities exist. Wenyi Zhang 0001, Qiang Ling 0001 |
WiOpt | 3 |
| 2015 | An adaptive caching algorithm suitable for time-varying user accesses in VOD systems
Qiang Ling 0001, Lixiang Xu, Jinfeng Yan, Yicheng Zhang 0001 |
Multim. Tools Appl. | 1 |
| 2014 | A background modeling and foreground segmentation approach based on the feedback of moving objects in traffic surveillance systems
Qiang Ling 0001, Jinfeng Yan, Feng Li 0042, Yicheng Zhang 0001 |
Neurocomputing | 1 |
| 2012 | Capacity offload game over unlicensed spectrumabstractWith the blasting increase of wireless data traffic, incumbent wireless service providers (WSPs) face critical challenges in provisioning spectrum resource. Given the permission of unlicensed access to TV white spaces, WSPs can alleviate their burden by exploiting the concept of “capacity offload” to transfer part of their traffic load to unlicensed spectrum. For such use cases, a central problem is for WSPs to coexist with others, since all of them may access the unlicensed spectrum without coordination thus interfering each other. Game theory provides tools for predicting the behavior of WSPs, and we formulate the coexistence problem under the framework of non-cooperative games as a capacity offload game (COG). We show that a COG always possesses at least one pure-strategy Nash equilibrium (NE). The analysis provides a full characterization of the structure of the NEs in two-player COGs. When the game is played many times and each WSP individually updates its strategy based on its best-response function, the resulting process forms a best-response dynamic. We establish that if the network configuration satisfies certain conditions so that the resulting best-response dynamics become linear, both simultaneous-move and alternating-move best-response dynamics are guaranteed to converge to the unique NE. Wenyi Zhang 0001, Qiang Ling 0001 |
ICC | 3 |
| 2012 | Non-Cooperative Game for Capacity OffloadabstractWith the dramatic increase of wireless data traffic, incumbent wireless service providers (WSPs) face critical challenges in provisioning spectrum resource. Given the permission of unlicensed access to TV white spaces, WSPs can alleviate their burden by exploiting the concept of "capacity offload" to transfer part of their traffic load to unlicensed spectrum. For such use cases, a central problem is for WSPs to coexist with others, since all of them may access the unlicensed spectrum without coordination thus interfering with each other. Game theory provides tools for predicting the behavior of WSPs, and we formulate the coexistence problem under the framework of non-cooperative games as a capacity offload game (COG). We show that a COG always possesses at least one pure-strategy Nash equilibrium (NE), and does not have any non-degenerate mixed-strategy NE. The analysis provides a detailed characterization of the structure of the NEs in two-player COGs. When the game is played repeatedly and each WSP individually updates its strategy based on its best-response function, the resulting process forms a best-response dynamic. We establish that, for two-player COGs, alternating-move best-response dynamics always converge to an NE, while simultaneous-move best-response dynamics do not always converge to an NE when multiple NEs exist. When there are more than two players in a COG, if the network configuration satisfies certain conditions so that the resulting best-response dynamics become linear, both simultaneous-move and alternating-move best-response dynamics are guaranteed to converge to the unique NE. Wenyi Zhang 0001, Qiang Ling 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2006 | Firm Real-Time System Scheduling Based on a Novel QoS ConstraintabstractMany real-time systems have firm real-time requirements which allow occasional deadline violations but discard any jobs that are not finished by their deadlines. To measure the performance of such a system, a quality of service (QoS) metric is needed. Examples of often used QoS metrics for firm real-time systems are average deadline miss rates and (m, k)-firm constraints. However, for certain applications, these metrics may not be adequate measures of system performance. This paper introduces a novel QoS constraint for firm real-time systems. The new QoS constraint generalizes existing firm real-time constraints. Furthermore, using networked control system as an example, we show that this constraint can be directly related to the control system's performance. We then present three different scheduling approaches with respect to this QoS constraint. Experimental results are provided to show the effectiveness of these approaches. Xiaobo Sharon Hu, Michael Lemmon 0001, Qiang Ling 0001 |
IEEE Trans. Computers | 4 |
| 2005 | Scheduling Tasks with Markov-Chain Based ConstraintsabstractMarkov-chain (MC) based constraints have been shown to be an effective QoS measure for a class of real-time systems, particularly those arising from control applications. Scheduling tasks with MC constraints introduces new challenges because these constraints require not only specific task finishing patterns but also certain task completion probability. Multiple tasks with different MC constraints competing for the same resource further complicates the problem. In this paper, we study the problem of scheduling multiple tasks with different MC constraints. We present two scheduling approaches which (i) lead to improvements in "overall" system performance, and (ii) allow the system to achieve graceful degradation as system load increases. The two scheduling approaches differ in their complexities and performances. We have implemented our scheduling algorithms in the QNX real-time operating system environment and used the setup for several realistic control tasks. Data collected from the experiments as well as simulation all show that our new scheduling algorithms outperform algorithms designed for window-based constraints as well as previous algorithms designed for handling MC constraints. Xiaobo Sharon Hu, Michael Lemmon 0001, Qiang Ling 0001 |
ECRTS | 4 |
| 2003 | Firm Real-Time System Scheduling Based on a Novel QoS ConstraintabstractMany real-time systems have firm real-time requirements which allow occasional deadline violations but discard tasks that are not finished by their deadlines. To measure the goodness of such a system, a quality of service (QoS) metric is needed. Examples of often used QoS metrics for firm real-time systems are average deadline miss rates and (m,k)-firm constraint. However, for certain applications, these metrics may not be adequate measures of system performance. This paper introduces a novel QoS constraint for networked feedback control systems. We show that this constraint can be directly related to the control system's performance. We then present three different scheduling approaches with respect to this QoS constraint. Experimental results are provided to compare these approaches. Xiaobo Sharon Hu, Michael Lemmon 0001, Qiang Ling 0001 |
RTSS | 4 |
| 2003 | Overload management in sensor-actuator networks used for spatially-distributed control systemsabstractOverload management policies avoid network congestion by actively dropping packets. This paper studies the effect that such data dropouts have on the performance of spatially distributed control systems. We formally relate the spatially-distributed system's performance (as measured by the average output signal power) to the data dropout rate. This relationship is used to pose an optimization problem whose solution is a Markov chain characterizing a dropout process that maximizes control system performance subject to a specified lower bound on the dropout rate. We then use this Markov chain to formulate an overload management policy that enables nodes to enforce the "optimal" dropout process identified in our optimization problem. Simulation experiments are used to verify the paper's claims. Michael Lemmon 0001, Qiang Ling 0001, Yashan Sun |
SenSys | 2 |