VLDB 2026 Research / reviewers in the wild / expert
Peng Qiao
dblp:45/10050
· DBLP profile ↗
62ranked-venue papers
6as first author
45since 2021 · last 2026
0000-0001-6752-7892ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 3 first-author · 27 since 2021Artificial intelligence and machine learning · 19 · 1 first-author · 13 since 2021Systems, architecture and hardware · 5 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Transolver Is a Linear Transformer: Revisiting Physics-Attention Through the Lens of Linear AttentionabstractRecent advances in Transformer-based Neural Operators have enabled significant progress in data-driven solvers for Partial Differential Equations (PDEs). Most current research has focused on reducing the quadratic complexity of attention to address the resulting low training and inference efficiency. Among these works, Transolver stands out as a representative method that introduces Physics-Attention to reduce computational costs. Physics-Attention projects grid points into slices for slice attention, then maps them back through deslicing. However, we observe that Physics-Attention can be reformulated as a special case of linear attention, and that the slice attention may even hurt the model performance. Based on these observations, we argue that its effectiveness primarily arises from the slice and deslice operations rather than interactions between slices. Building on this insight, we propose a two-step transformation to redesign Physics-Attention into a canonical linear attention, which we call Linear Attention Neural Operator (LinearNO). Our method achieves state-of-the-art performance on six standard PDE benchmarks, while reducing the number of parameters by an average of 40.0% and computational cost by 36.2%. Additionally, it delivers superior performance on two challenging, industrial-level datasets: AirfRANS and Shape-Net Car. Sidun Liu, Peng Qiao, Zhenglun Sun, Yong Dou |
AAAI | 3 |
| 2026 | VTalker: Text-Driven Synthesis of Talking Head with Vision Diffusion TransformerabstractText-driven Talking Head Generation (THG) marks a significant advancement in the video production industry by enabling the creation of realistic talking head videos with minimal data input. While previous research has explored few-shot talking head synthesis, these methods often fall short in terms of lip-sync consistency and expression diversity, both of which are essential for practical applications. In this article, we present a novel multi-modal framework, VTalker, for synthesizing talking heads with specific vocal tones and speaking styles. First, a Text-to-Speech (T2S) model is developed for generating speech with a given voice tone from text. Second, we design a novel multi-modal fusion module to effectively integrate speech and text features. Third, we propose the V-DiT (Video Diffusion Transformer) as the backbone to frame the generation of talking heads as a temporally iterative denoising task. To further enhance performance, the appearance and temporal conditions are incorporated into the backbone as tokens, along with patches, to improve image fidelity and ensure smooth spatio-temporal motion. This innovative design enables our model to generate high-fidelity, text-synchronized talking head videos that generalize smoothly across various identities. Extensive experiments demonstrate the effectiveness of our approach in generating high-quality text-driven talking head videos. Yali Cai, Peng Qiao, Dongsheng Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Highly Parallelized Reinforcement Learning Training with Relaxed Assignment DependenciesabstractAs the demands for superior agents grow, the training complexity of Deep Reinforcement Learning (DRL) becomes higher. Thus, accelerating training of DRL has become a major research focus. Dividing the DRL training process into sub-tasks and using parallel computation can effectively reduce training costs. However, current DRL training systems lack sufficient parallelization due to data assignment between sub-task components. This assignment issue has been ignored, but addressing it can further boost training efficiency. Therefore, we propose a high-throughput distributed RL training system called TianJi. It relaxes assignment dependencies between sub-task components and enables event-driven asynchronous communication. Meanwhile, TianJi maintains clear boundaries between sub-task components. To address convergence uncertainty from relaxed assignment dependencies, TianJi proposes a distributed strategy based on the balance of sample production and consumption. The strategy controls the staleness of samples to correct their quality, ensuring convergence. We conducted extensive experiments. TianJi achieves a convergence time acceleration ratio of up to 4.37 compared to related comparison frameworks. When scaled to eight computational nodes, TianJi shows a convergence time speedup of 1.6 and a throughput speedup of 7.13 relative to XingTian, emonstrating its capability to accelerate training and scalability. In data transmission efficiency experiments, TianJi significantly outperforms other frameworks, approaching hardware limits. TianJi also shows effectiveness in on-policy algorithms, achieving convergence time acceleration ratios of 4.36 and 2.95 compared to RLlib and XingTian. Zhouyu He, Peng Qiao, Rongchun Li, Yong Dou, Yusong Tan |
AAAI | 2 |
| 2025 | Rethinking Incision Segmentation with Geometry-Structure Aligned Polygon PromptabstractThe incision plays a critical role in clinical surgery and has recently emerged as a complex segmentation objective in medical image analysis. The Segment Anything Model (SAM) demonstrates strong generalization and performs well across various medical imaging tasks. With task-specific finetuning, SAM can generate rough incision segmentations based on prompts. However, the irregular morphology of incisions presents challenges to existing prompting strategies. First, coarse prompts such as points and bounding boxes fail to capture the shape characteristics of incisions. Second, while existing convex polygon prompts show potential in guiding incision segmentation, their structural constraints limit the representation of concave regions. In contrast, unconstrained naive polygons often suffer from improper vertex sampling, resulting in ineffective guidance. Moreover, the prompt encoding schemes with fixed vertex order tend to introduce ambiguity when applied to flexible polygon structures. To address these challenges, we explore how to deliver more effective prompts under limited interactions and propose the Geometry-Structure Aligned Polygon Prompt (GSAPP) method, which maintains both geometric and structural consistency with the segmentation target. GSAPP comprises two core modules: Structure-Aligned Polygon Prompt (SAPP) and Geometry-Aligned Encoding (GAE) module. SAPP introduces a polygon prompt adapted to incision morphology, mitigating structural mismatches with the target shape. Additionally, we develop a heuristic generator to automatically produce high-quality polygon prompts. The GAE module encodes prompts using extreme vertex matching, effectively reducing geometric ambiguity. Extensive experiments demonstrate that GSAPP improves prompt effectiveness and achieves more than a 3% Dice gain over the previous state-of-the-art in incision segmentation. Keran Ding, Peng Qiao, Zhenglun Sun, Tongrui Hu, Yong Dou |
BIBM | 2 |
| 2025 | Beyond Synthetic Data: Leveraging Natural Image Pretraining and Finetuning for Mixed Exposure Correction in Capsule EndoscopyabstractWireless capsule endoscopy (WCE) has become standard in gastrointestinal diagnostics, but its frames often exhibit mixed exposure problem with both over- and underexposed regions. Existing methods mostly use pixel-level supervision in RGB space and introduce unnatural color shifts because exposure, color, and texture are tightly entangled. Another challenge is the lack of reliable paired WCE training data. Available synthetic datasets are small, low in quality, and lack diversity, while real paired examples are hard to obtain. Fully supervised models trained on such data tend to overfit and generalize poorly. We first extract exposure priors from a model pretrained on large natural-image datasets and then fine-tune it on WCE data. Leveraging hue stability and pixel dispersion nature in HSV space, our separated exposure and color correction method comprises three significant components: the training-free Single-Image Exposure Fusion (SIEF) module for precise exposure partitioning, the Sequential Interaction Correction (SIC) module for efficient information exchange between the exposure and color correction branches, and the Phase-Shifting Coder (PSC) to resolve discontinuities in the circular H channel and ensure smooth, stable hue prediction. Extensive experiments demonstrate that our method achieves SOTA results on natural-image exposure correction and transfers effectively to WCE mixed-exposure cases. It outperforms prior supervised approaches and shows robust recovery across multiple WCE datasets. Tongrui Hu, Peng Qiao, Keran Ding, Yong Dou, Rongchun Li |
BIBM | 2 |
| 2025 | Partial Order-centered Hyperbolic Representation Learning for Few-shot Relation ExtractionabstractPrototype network-based methods have made substantial progress in few-shot relation extraction (FSRE) by enhancing relation prototypes with relation descriptions. However, the distribution of relations and instances in distinct representation spaces isolates the constraints of relations on instances, making relation prototypes biased. In this paper, we propose an end-to-end partial order-centered hyperbolic representation learning (PO-HRL) framework, which imposes the constraints of relations on instances by modeling partial order in hyperbolic space, so as to effectively learn the distribution of instance representations. Specifically, we develop the hyperbolic supervised contrastive learning based on Lorentzian cosine similarity to align representations of relations and instances, and model the partial order by constraining instances to reside within the Lorentzian entailment cone of their respective relation. Experiments on three benchmark datasets show that PO-HRL outperforms the strong baselines, especially in 1-shot settings lacking relation descriptions. Zhen Huang 0006, Minghao Hu 0001, Pinglv Yang, Peng Qiao, Yong Dou, Zhilin Wang |
COLING | 5 |
| 2025 | DA-NeRF: High-Fidelity Talking Face Generation From Speech With Neural Radiance Fields
Yali Cai, Peng Qiao, Dongsheng Li 0001 |
ICANN (4) | 2 |
| 2025 | GPIS: Geometric Informed Polygon Prompt for Incision Segmentation
Keran Ding, Peng Qiao, Zhenglun Sun, Yong Dou |
ICANN (2) | 2 |
| 2025 | A Counterfactual Ultrasound Anti-Interference Self-Supervised Network for B-mode Ultrasound Tongue ExtractionabstractB-mode ultrasound tongue imaging is a non-invasive and real-time method for visualizing vocal tract deformation. However, accurately extracting the tongue’s surface contour remains a significant challenge due to the low signal-to-noise ratio (SNR) and prevalent speckle noise in ultrasound images. Traditional supervised learning models often require large labeled datasets, which are labor-intensive to produce and susceptible to noise interference. To address these limitations, we present a novel Counterfactual Ultrasound Anti-Interference Self-Supervised Network (CUAI-SSN), which integrates self-supervised learning (SSL) with counterfactual data augmentation, progressively disentangles confounding factors, ensuring that the model generalizes well across varied ultrasound conditions. Our approach leverages causal reasoning to decouple noise from relevant features, enabling the model to learn robust representations that focus on essential tongue structures. By generating counterfactual image-label pairs, our method introduces alternative, noise-independent scenarios that enhance model training. Furthermore, we introduce attention mechanisms to enhance the network’s ability to capture fine-grained details even in noisy conditions. Extensive experiments on real ultrasound tongue images demonstrate that CUAI-SSN outperforms existing methods, setting a new benchmark for automated contour extraction in ultrasound tongue imaging. Our code is publicly available at https://github.com/inexhaustible419/CounterfactualultrasoundAI. Yan Jia 0001, Yuqing Cheng, Kele Xu, Yong Dou, Peng Qiao, Zhouyu He |
ICASSP | 5 |
| 2025 | Scaling Bioacoustic Signal Pre-training with Million Samples Via Mask-ModelingabstractDeep learning-based bioacoustic audio analysis holds immense potential across various applications. However, existing studies in bioacoustics often focus on a limited number of species, potentially hindering the transferability of models across different species. Furthermore, the manual annotation of bioacoustic data is both costly and labor-intensive. To address these challenges, self-supervised learning on large-scale bioacoustic audio data presents a promising solution. In this paper, we introduce GPM-BT (General Pre-training Model for Bioacoustic Tasks), a self-supervised, Transformer-based model pre-trained on approximately 1.2 million unannotated bioacoustic audio samples. We evaluate the scalability and effectiveness of this pre-training approach through comprehensive experiments across a broad range of classification and detection tasks. Our results demonstrate that pre-training on large-scale bioacoustic data significantly enhances model performance, improving both generalization and robustness. Notably, GPM-BT achieves state-of-the-art performance on the BEANS benchmark and secures first place in the Few-shot Bioacoustic Event Detection task at the IEEE DCASE 2024 Challenge1. To further advance research in bioacoustics, we have open-sourced our models and code2. Xuyao Deng, Tianjiao Wan, Kele Xu, Peng Qiao, Yong Dou |
ICASSP | 5 |
| 2025 | MonoIR: Inpainting and Reconstruction for Monocular Endoscope Deformation ScenesabstractMonocular endoscopic scene reconstruction is challenging due to limited viewpoints and interference from surgical instruments. While 3D Gaussian-based methods are popular for their strong reconstruction capabilities and efficiency, they often rely on sensors or stereo depth, resulting in blurred tissue areas when instruments obstruct the view. To overcome these issues, we propose MonoIR, an inpainting and reconstruction method that produces spatiotemporally consistent videos and high-quality reconstructions. Our approach employs a propagation-based method to inpaint video holes, enhanced by optical flow constraints for robustness and an additional transformer-based module to address detail loss. We also introduce the normal and depth regularization with confidence to improve reconstruction quality. Extensive experiments demonstrate that MonoIR efficiently reconstructs deformed tissues and outperforms both monocular and binocular methods in key metrics. Ziteng Zhang, Sidun Liu, Peng Qiao, Yong Dou |
ICASSP | 4 |
| 2025 | End-To-End Casual Video Reconstruction: Geometry, Pose and MotionabstractFrom casual videos in daily life, humans can effortlessly recognize object shapes, perceive variations in viewpoints, decompose dynamic objects and infer their motions. This suggests that reconstruction algorithms should also, like humans, simultaneously achieve these capabilities. However, existing methods tackle the aforementioned tasks in multiple stages. The phased processing approaches mean that each stage’s performance heavily depends on the preceding one, and the entire reconstruction process cannot be optimized in an end-to-end manner, which limits its overall potential. In response, we present an algorithm, designed to integrate these capabilities for casual videos in an end-to-end manner. Specifically, we represent the 4D scene in a video as the combination of local multi-view depth maps and a shared canonical space, where a continuous bijective mapping is used to model motions between local and canonical space. We also extend the depth estimation network to decompose scenes into static and dynamic parts, which helps to avoid degenerate cases and leads to accurate tracking. Experiments demonstrate that our method performs well on casual videos, and achieves performance comparable to state-of-the-art methods across all tasks. Our project page is https://fullre.github.io/FullRe. Peng Qiao, Sidun Liu, Zongxin Ye, Ziteng Zhang, Zhenglun Sun, Yong Dou |
ICME | 2 |
| 2025 | Only One Stage: A Chemical-Aware Model for Accurate Combustion Chemical Kinetics PredictionabstractThe combustion chemical kinetics simulation, which focuses on the change in species mass fractions during reactions, is vital for clean energy development. Due to the sample complexities introduced by chemical kinetics, current multi-stage methods employ data preprocessing stage to separate it into subspaces, aiming to ease training. However, the current approaches to sample space separation are not effective, which not only affects the training accuracy of the network but also introduces a complex preprocessing procedure. To solve this, we propose a one-stage, end-to-end model with chemical-aware capabilities, using an auto dividing mechanism and spatio-temporal convolution for feature extraction. In hydrogen simulation experiments, our model achieved an L1 error of 10−6level, which is nearly identical to the standard results of numerical calculations, outperforming the usually used Multilayer Perceptron (MLP) method by 41.9 times. Zhenglun Sun, Peng Qiao, Yong Dou, Rongchun Li, Sidun Liu |
ICME | 2 |
| 2025 | Knowledge Bridges the Intent Gap: Contextual Fusion in Medical Fine-Grained Segmentation
Peng Qiao, Yan Jia 0001, Yong Dou |
MICCAI (1) | 2 |
| 2025 | AnchorTalk: High-Fidelity Upper-Body Talking Human Generation From SpeechabstractWhile most existing speech-driven talking head generation methods provide effective solutions, they primarily focus on the facial area. However, producing upper-body talking videos from speech remains challenging. Addressing how to use speech to simultaneously drive subtle facial motion and large-scale body motion while generating naturally synchronized upper body video frames is urgent. In this study, we propose AnchorTalk, a novel system based on tri-plane hash NeRF, capable of producing high-quality anchor-style talking videos. Firstly, to integrate both rigid and non-rigid motion within a unified system, we introduce a coarse-to-fine framework that consists of coarse pose generation and facial details optimization. A speech disentanglement encoder decouples speech features into pose-related and head-related features to drive the motion of the body and head. Secondly, during the coarse pose generation phase, we propose a geometry correction module to obtain precise body parameters to guide the body motion. Thirdly, less detailed head parameters can lead to facial distortion and disjointed motion during facial optimization. To mitigate this issue, we propose a head controller to capture facial expressions accurately. By fine-tuning the model on a one-minute video, the system can generalize to novel identities. Experimental results validate the effectiveness and feasibility of our method in generating high-quality, coherent upper-body talking human videos from speech. Yali Cai, Peng Qiao, Dongsheng Li 0001 |
ICMR | 2 |
| 2025 | Mono3R: Exploiting Monocular Cues for Geometric 3D ReconstructionabstractRecent advances in data-driven geometric multi-view 3D reconstruction foundation models (e.g., DUSt3R) have shown remarkable performance across various 3D vision tasks, facilitated by the release of large-scale, high-quality 3D datasets. However, as we observed, constrained by their matching-based principles, the reconstruction quality of existing models suffers significant degradation in challenging regions with limited matching cues, particularly in weakly textured areas and low-light conditions. To mitigate these limitations, we propose to harness the inherent robustness of monocular geometry estimation to compensate for the shortcomings. Specifically, we introduce a monocular-guided refinement module that integrates monocular geometric priors into multi-view reconstruction frameworks. This integration substantially enhances the robustness of multi-view reconstruction systems, leading to high-quality feed-forward reconstructions. Comprehensive experiments across multiple benchmarks demonstrate that our method achieves substantial improvements in both multi-view camera pose estimation and point cloud accuracy. Sidun Liu, Peng Qiao, Yong Dou |
ACM Multimedia | 3 |
| 2025 | Regist3R: Incremental Registration with Stereo Foundation ModelabstractMulti-view 3D reconstruction has remained an essential yet challenging problem in the field of computer vision. While DUSt3R and its successors have achieved breakthroughs in 3D reconstruction from unposed images, these methods exhibit significant limitations when scaling to multi-view scenarios, including high computational cost and cumulative error induced by global alignment. To address these challenges, we propose Regist3R, a novel stereo foundation model tailored for efficient and scalable incremental reconstruction. Regist3R leverages an incremental reconstruction paradigm, enabling large-scale 3D reconstructions from unordered and many-view image collections. We evaluate Regist3R on public datasets for camera pose estimation and 3D reconstruction. Our experiments demonstrate that Regist3R achieves comparable performance with optimization-based methods while significantly improving computational efficiency, and outperforms existing multi-view reconstruction models. Furthermore, to assess its performance in real-world applications, we introduce a challenging oblique aerial dataset which has long spatial spans and hundreds of views. The results highlight the effectiveness of Regist3R. We also demonstrate the first attempt to reconstruct large-scale scenes encompassing over thousands of views through pointmap-based foundation models, showcasing its potential for practical applications in large-scale 3D reconstruction tasks, including urban modeling, aerial mapping, and beyond. Sidun Liu, Peng Qiao, Yong Dou |
ACM Multimedia | 3 |
| 2025 | Segment Anything for Visual Bird Sound DenoisingabstractCurrent audio denoising methods perform well with synthetic noise but struggle with complex natural noise, especially for bird sounds, which contain natural environmental sounds such as wind and rain, making it challenging to extract clean bird sounds. This issue becomes more pronounced with short and faint bird sounds, where existing methods are less effective. In this paper, we introduceBudSAM, a novel audio denoising model that incorporates theSegment Anything Model (SAM), originally designed for image segmentation task, into the field of visual bird sound denoising. By treating audio denoising as a segmentation task, BudSAM utilizes SAM's powerful segmentation capabilities and we incorporates BCE and Dice losses to enhance the model's ability to segment weak signals, effectively isolating the clean bird sounds that are often masked by background noise. Our method is evaluated on the BirdSoundsDenoising dataset, achieving a 4.0% improvement in IoU and a 0.77 dB increase in SDR compared to state-of-the-art methods. To the best knowledge of the authors, BudSAM marks the first attempt which employs SAM in audio denoising task, offering a promising direction for future research and real-world bird sound processing tasks. Tianjiao Wan, Kele Xu, Peng Qiao, Yong Dou |
IEEE Signal Process. Lett. | 4 |
| 2024 | A Connectivity-Enhanced Multi-Task Learning based on Anatomical Priors for 3D Class-Balanced Pulmonary Airway SegmentationabstractAccurate and efficient airway segmentation is essential for evaluating pulmonary diseases, aiding diagnosis, reducing the preoperative burden of airway identification, and minimizing patient discomfort during prolonged surgeries. However, current pulmonary airway reconstruction techniques are hindered by two major challenges: difficulty in accurately reconstructing fine airway branches due to the tendency to overlook small targets, and insufficient structural connectivity leading to frequent branch discontinuities within the airway tree. These limitations directly affect the clinical applicability of reconstructed airways. To overcome these challenges, a novel 3D pulmonary airway segmentation multi-task framework is proposed, designed to enhance the performance of existing backbone models. This approach integrates Anatomical Prior-Based Multi-Task Learning (AP-MTL) through the use of Gaussian-constructed connectivity-enhanced isosurfaces, significantly improving the network’s ability to maintain airway continuity. Additionally, a Class-Balanced CT Density Distribution Reconstruction mechanism (DDR-CB) is introduced, further refining the model’s capability to detect and segment fine airway branches. As a result of these enhancements, the model demonstrates a 11.5% average improvement in segmentation accuracy and connectivity compared to the baseline. The source code is publicly accessible at https://github.com/inexhaustible419/APMTLAirwaySegment. Yan Jia 0001, Yong Dou, Peng Qiao, Yuqing Cheng, Kele Xu, Zhouyu He |
BIBM | 3 |
| 2024 | SAM-NeRF: NeRF-Based 3D Instance Segmentation with Segment Anything Model
Linglin Xie, Peng Qiao, Yong Dou, Sidun Liu, Kaijun Yang |
ICANN (2) | 3 |
| 2024 | Dual Dreamer: Extending Single-View Dreamer with Few Shot of Complementary Views
Ziteng Zhang, Peng Qiao, Dou Yong, Sidun Liu |
ICANN (3) | 2 |
| 2024 | KnowMIM: a Self-supervised Pre-training Framework Based on Knowledge-Guided Masked Image Modeling for Retinal Vessel Segmentation
Jiuyuan Zhu, Tianci Xun, Chunjiao Tan, Yingqi Xu, Peng Qiao |
ICANN (8) | 8 |
| 2024 | Adapter-Based Incremental Learning for Face Forgery DetectionabstractMany existing face forgery detection methods primarily revolve around learning general representations on predefined datasets and subsequently crossing these static representations to other datasets. However, these approaches could lead to catastrophic forgetting in real-world scenarios, especially when new forgery methods continually emerge. In this paper, we proposed a novel incremental learning framework for face forgery detection, where we design an adapter-based incremental learning scheme combined with a confidence-based ensemble prediction mechanism. When confronted with new forgery methods, we incorporate small trainable adapter modules, which are retrained along with their corresponding classification layers, yielding a series of task-specific modules. Then we incorporate a confidence-based ensemble prediction mechanism to aggregate all predictions. Through comprehensive evaluations on multiple benchmark datasets (FF++, DFD, and Celeb-DF), our method successfully mitigates the catastrophic forgetting problem in a cost-effective manner and attains state-of-the-art performance in cross-dataset scenario. Caili Gao, Qisheng Xu, Peng Qiao, Kele Xu, Xifu Qian, Yong Dou |
ICASSP | 3 |
| 2024 | Improving Motion Deblur By Multi-Output LearningabstractImage deblurring is an ill-posed task, where exists infinite feasible solutions for blurry images. Modem deep learning approaches usually discard the learning of blur kernels and directly employ end-to-end supervised learning. However, supervised learning can’t handle ill-posed tasks appropriately. It regresses the average thus losing sharp details. Therefore, we propose an extension to the network to learn from stochastic supervisions, where a novel multi-output architecture and Min-Out loss function are designed. Our approach enables the model to output multiple feasible solutions to fit various non-uniform motions. We then propose a novel parameter multiplexing method that reduces computations while improving performance with fewer parameters. After training, the best-performed head is fine-tuned to be used for inference where the sharp label is absent. The proposed approach is evaluated with multiple image deblur models on the GoPro motion deblur dataset. On average, the multi-output extension improves the PSNR by 0.08 dB. When applied to the popular attention-based model Restormer, the multi-output helps it achieve 33.05 dB (+0.13 dB) PSNR without modification on network architecture. Sidun Liu, Peng Qiao, Yong Dou |
ICASSP | 2 |
| 2024 | ParaSurRe: Parallel Surface Reconstruction with No Pose PriorabstractSurface reconstruction from multi-view images without pose prior is challenging. Recent advances integrate incremental Structure from Motion (SfM) pipeline into surface optimization, enabling simultaneous surface reconstruction and pose estimation. However, due to the inherent incremental registration scheme of SfM, the efficiency of these methods is far from satisfactory, e.g., reconstruction of an object captured by 49 images costs over 9 hours using a high-end GPU. Inspired by divide-and-conquer strategy, we present a Parallel Surface Reconstruction method, coined as ParaSurRe, where image collections are divided into non-overlapped clusters and the incremental reconstructions are performed individually in each cluster. Owing to image partitioning, each cluster only accurately reconstructs a part of the surface. The major challenge is to merge multiple partial neural implicit surfaces into a complete one. We propose a confidence-aware surface fusion strategy and a geometry-guided refinement to tackle this issue. Experiments on real-world datasets demonstrate that ParaSurRe reconstructs delicate surfaces from unposed images, and achieves competitive pose estimation performance compared with state-of-the-art methods, with up to 6.5× speedup on a scene captured by 81 images. Zongxin Ye, Sidun Liu, Ziteng Zhang, Peng Qiao, Yong Dou |
ICME | 6 |
| 2024 | FedEKT: Ensemble Knowledge Transfer for Model-Heterogeneous Federated LearningabstractFederated Learning (FL) enables multiple clients to collaboratively train a shared server model while preserving data privacy. Most existing FL systems rely on the assumption that the server model and client models have homogeneous architecture. However, intensive resource requirements during the training process prevent low-end devices from contributing to the server model with their own data. On the other hand, the resource constraints on participating clients can significantly limit the size of the server model in the model-homogeneous setting, thereby restricting the application scope of FL. In this work, we propose FedEKT, a novel model-heterogeneous FL system designed to obtain a high-performance large server model while benefiting heterogeneous small client models. Specifically, a new aggregation approach is designed to enable the integration of knowledge from heterogeneous client models to a large server model while mitigating the adverse effects of biases stemming from data heterogeneity. Subsequently, to enhance the performance of client models by benefiting from the high-performance server model, FedEKT distills this large server model into multiple heterogeneous client models, facilitating the transfer of integrated knowledge back to the client models. In addition, we design specialized modules within the model and communication strategy to accomplish aggregation and transfer of knowledge in a data-free manner. The evaluation results demonstrate that FedEKT enhances the accuracy of the server model and client models by up to 53.96% and 12.35%, respectively, compared with the state-of-the-art FL approach on CIFAR-100. Meihan Wu, Li Li 0064, Tao Chang, Peng Qiao, Cui Miao, Jie Zhou 0032, Jingnan Wang, Xiaodong Wang 0002 |
IWQoS | 4 |
| 2024 | AbsGS: Recovering Fine Details in 3D Gaussian Splatting
Zongxin Ye, Sidun Liu, Peng Qiao, Yong Dou |
ACM Multimedia | 4 |
| 2024 | End-To-End High-Quality Transformer Object Detection Model Applied to Human Head Detection
Rongchun Li, Peng Qiao, Jingfei Jiang |
PRCV (12) | 3 |
| 2024 | ER-SFM: Efficient and Robust Cluster-Based Structure from Motion
Zongxin Ye, Sidun Liu, Peng Qiao, Yong Dou |
PRCV (6) | 4 |
| 2024 | Introducing diminutive causal structure into graph representation learningabstractWhen engaging in end-to-end graph representation learning with Graph Neural Networks (GNNs), the intricate causal relationships and rules inherent in graph data pose a formidable challenge for the model in accurately capturing authentic data relationships. A proposed mitigating strategy involves the direct integration of rules or relationships corresponding to the graph data into the model. However, within the domain of graph representation learning, the inherent complexity of graph data obstructs the derivation of a comprehensive causal structure that encapsulates universal rules or relationships governing the entire dataset. Instead, only specialized diminutive causal structures, delineating specific causal relationships within constrained subsets of graph data, emerge as discernible. Motivated by empirical insights, it is observed that GNN models exhibit a tendency to converge towards such specialized causal structures during the training process. Consequently, we posit that the introduction of these specific causal structures is advantageous for the training of GNN models. Building upon this proposition, we introduce a novel method that enables GNN models to glean insights from these specialized diminutive causal structures, thereby enhancing overall performance. Our method specifically extracts causal knowledge from the model representation of these diminutive causal structures and incorporates interchange intervention to optimize the learning process. Theoretical analysis serves to corroborate the efficacy of our proposed method. Furthermore, empirical experiments consistently demonstrate significant performance improvements across diverse datasets. Hang Gao 0004, Peng Qiao, Fengge Wu, Jiangmeng Li, Changwen Zheng |
Knowl. Based Syst. | 2 |
| 2023 | MANet: An Architecture Adaptive Method for Sparse Matrix Format Selection
Zhenglun Sun, Peng Qiao, Yong Dou |
ICA3PP (2) | 2 |
| 2023 | VPPT: Visual Pre-Trained Prompt Tuning Framework for Few-Shot Image ClassificationabstractLarge-scale pre-trained transformers have recently achieved remarkable success in several computer vision tasks. However, it remains highly challenging to fully fine-tune models for downstream tasks, due to the expensive computational and storage cost. Recently, Parameter-Efficient Tuning (PETuning) techniques, e.g., Visual Prompt Tuning (VPT), have significantly reduced the computation cost by inserting lightweight prompt modules including prompt tokens or adapter layers, into the pre-trained models and tuning these prompt modules with a small number of trainable parameters, while keeping the transformer backbone freeze. Although encouraging results were achieved, existing PETuning methods cannot perform well under the few-shot learning settings (i.e., extremely limited training data, with only 1 or 2 shots per class), due to the scarce supervision signal. To this end, we first empirically identify the poor performance is mainly due to the inappropriate way of initializing prompt modules, which has also been verified in the pre-trained language models. Next, we propose a Visual Pre-trained Prompt Tuning framework (VPPT), which pre-trains the prompt modules first and then leverages the pre-trained modules along with the pre-trained transformer backbone to perform prompt tuning on downstream tasks. Extensive experiments show that our VPPT framework achieves 16.08% average accuracy absolute improvement under 1 shot setting on five fine-grained visual classification datasets, compared with the previous PETuning techniques, e.g., VPT, in few-shot image classification. Zhao Song 0011, Ke Yang 0004, Naiyang Guan, Peng Qiao, Qingyong Hu |
ICASSP | 5 |
| 2023 | Spatial and Frequency Domains Inconsistency Learning for Face Forgery Detection
Caili Gao, Peng Qiao, Yong Dou, Qisheng Xu, Xifu Qian |
ICONIP (12) | 2 |
| 2023 | HAAN: Human Action Aware Network for Multi-label Temporal Action DetectionabstractThe task of multi-label temporal action detection aims to accurately detect dense action instances in untrimmed videos. Previous methods focused on modeling the appearance features of RGB images have struggled to capture the fine details and subtle variations in human actions, resulting in three critical issues: overlapping action confusion, intra-class appearance diversity, and background interferences. These issues have significantly undermined the accuracy and generalization of detection models. To tackle these issues, we propose incorporating the human skeleton into the feature design of the detection model. By utilizing multi-person skeletons, our proposed method can accurately represent various human actions in the scene, balance the salience of overlapping actions, and reduce the impact of changes in human appearance and background interferences on action features. Overall, we propose a novel two-stream human action aware network~(HAAN) for multi-label temporal action detection based on the original RGB frames and the estimated skeleton frames. To leverage the complementary advantages of RGB features and skeleton features, we design a cross-modality fusion module that allows the two features to guide each other and enhance their representation of human actions. On the popular benchmarks MultiTHUMOS and Charades, our HAAN achieves state-of-the-art performance with 56.9% (+5.4%) and 32.1% (+3.3%) mean average precision (mAP) compared to the best available methods. Importantly, HAAN shows superior improvements of +6.83%, +22.35%, and +2.56% on the challenging sample subsets of the three critical issues. Zikai Gao, Peng Qiao, Yong Dou |
ACM Multimedia | 2 |
| 2022 | Weight-Aware Graph Contrastive Learning
Hang Gao 0004, Jiangmeng Li, Peng Qiao, Changwen Zheng |
ICANN (2) | 3 |
| 2022 | Qrelation: an Agent Relation-Based Approach for Multi-Agent Reinforcement Learning Value Function FactorizationabstractThe Centralized Training with Decentralized Execution paradigm (CTDE), which trains policies centrally with additional information, is important for Multi-Agent Reinforcement Learning (MARL). For CTDE, value function factorization methods make use of state during training and factorize the value function into multiple local value functions for decentralized execution. These approaches do not fully consider the relational information among agents, resulting in sub-optimal models for complex tasks. To remedy this issue, we propose QRelation which is a graph neural network approach for value function factorization. It considers both the static relations (e.g., agent types) and dynamic relations (e.g., close-by). We show that QRelation can obtain better results than state-of-the-art methods on challenging StarCraft II benchmarks. Mengwei Qiu, Weiquan Liu, Cheng Wang 0003, Yongquan Fu, Peng Qiao |
ICASSP | 8 |
| 2022 | Searching Latent Sub-Goals in Hierarchical Reinforcement Learning as Riemannian Manifold OptimizationabstractHierarchical Reinforcement Learning (HRL) is promising to tackle the long-term sparse reward problem. However, goal conditioned HRL, which decomposes the goal into a series of sub-goals, suffers from sub-goal search inefficiency problems when the observation space is too large. This problem is more severe in a visual observation space, since its high latent dimensions, where the complete dynamics information is preserved, exponentially increase the difficulty of sub-goal search. In view of this, we propose to treat the latent space as a manifold, i.e., a Riemannian manifold. Assisted by the Riemannian manifold optimization, sub-goals can be efficiently searched in the higher-dimensional latent space, with the help of preserving the dynamics information efficiently. Experiments on a series of MuJoCo tasks with visual observation show that the proposed Riemannian manifold optimization, compared with the baseline that directly searches for sub-goals in bounded latent space, improves the success rate by 1.5 times on average. In much higher dimensions where the baseline no longer converges, the success rate of the proposed method is maintained. Sidun Liu, Peng Qiao, Yong Dou, Ruochun Jin |
ICME | 2 |
| 2022 | MLPs: Efficient Training of MiniGo on Large-scale Heterogeneous Computing SystemabstractDeep Reinforcement Learning has been successfully applied in various applications and achieved impressive performance compared with previous traditional methods but suffers from high computation cost and long training time. MLPerf takes deep reinforcement learning as one of the benchmark tracks and provides a single node training version of MiniGo as a reference. A key challenge is to achieve efficient MiniGo training on a large-scale computing system. According to the training computation pattern in MiniGo and the characteristics of our large-scale heterogeneous computing system, we propose a MultiLevel Parallel strategy, MLPs, including task-level parallelism between nodes, CPU-DSP heterogeneous parallelism, and DSP multi-core parallelism. The proposed method reduces the overall execution time from 43 hours to 16 hours while scaling the node size from 1067 to 4139. The scaling efficiency is 69.1%. According to our fitting method, the scaling efficiency is 46.5% when scaling to 8235 nodes. The experimental results show that the proposed method achieves the efficient training of MiniGo on the largescale heterogeneous computing system. Peng Qiao, Zhouyu He, Rongchun Li, Jingfei Jiang, Yong Dou, Dongsheng Li 0001 |
ICPADS | 1 |
| 2022 | An automatic learning rate decay strategy for stochastic gradient descent optimization methods in neural networksabstractStochastic Gradient Descent (SGD) series optimization methods play the vital role in training neural networks, attracting growing attention in science and engineering fields of the intelligent system. The choice of learning rates affects the convergence rate of SGD series optimization methods. Currently, learning rate adjustment strategies mainly face the following problems: (1) The traditional learning rate decay method mainly adopts manual manner during training iterations, the small learning rate produced from which causes slow convergence in training neural networks. (2) Adaptive method (e.g., Adam) has poor generalization performance. To alleviate the above issues, we propose a novel automatic learning rate decay strategy for SGD optimization methods in neural networks. On the basis of the observation that the convergence rate's upper bound enjoys minimization in a specific iteration concerning the current learning rate, we first present the expression of the current learning rate determined by historical learning rates. And merely one extra parameter is initialized to generate automatic decreasing learning rates during the training process. Our proposed approach is applied to SGD and Momentum SGD optimization algorithms, and concrete theoretical proof explains its convergence. Numerical simulations are conducted on the MNIST and Cifar-10 data sets with different neural networks. Experimental results show that our algorithm outperforms existing classical ones, achieving faster convergence rate, better stability, and generalization performance in neural network training. It also lays a foundation for large-scale parallel search of initial parameters in intelligent systems. Yong Dou, Tao Sun 0005, Peng Qiao, Dong Wen 0004 |
Int. J. Intell. Syst. | 4 |
| 2022 | Fixed-Size Objects Encoding for Visual Relationship Detection
Hengyue Pan, Xin Niu 0002, Yixin Chen 0004, Peng Qiao, Zhen Huang 0006, Dongsheng Li 0001 |
Neural Process. Lett. | 5 |
| 2022 | Hierarchical learning with backtracking algorithm based on the Visual Confusion Label Tree for large-scale image classification
Yuntao Liu 0004, Yong Dou, Ruochun Jin, Rongchun Li, Peng Qiao |
Vis. Comput. | 5 |
| 2021 | Global-Localized Agent Graph Convolution for Multi-Agent Reinforcement LearningabstractA lot of efforts have been devoted to solving the problem about complex relationship and localized cooperation among a large number of agents in large-scale multi-agent systems. However, global cooperation among all agents is also important while interactions between agents often happen locally. It is a challenging problem to enable agent to learn global and localized cooperate information simultaneously in multi-agent systems. In this paper, we model the global and localized cooperation among agents by global and localized agent graphs and propose a novel graph convolutional reinforcement learning mechanism based on these two graphs which allows each agent to communicate with neighbors and all a-gents to cooperate at the high level. Experiments on the large-scale multi-agent scenarios in StarCraft II show that our pro-posed method gets better performance compared with state-of-the-art algorithms and allows agents learning to cooperate efficiently. Yuntao Liu 0004, Yong Dou, Peng Qiao |
ICASSP | 4 |
| 2021 | Graphcomm: A Graph Neural Network Based Method for Multi-Agent Reinforcement LearningabstractThe communication among agents is important for Multi-Agent Reinforcement Learning (MARL). In this work, we propose GraphComm, a method makes use of the relation-ships among agents for MARL communication. GraphComm takes the explicit relations (e.g., agent types), which can be provided through some knowledge background, into account to better model the relationships among agents. Besides explicit relations, GraphComm considers implicit relations, which are formed by agent interactions. GraphComm use Graph Neural Networks (GNNs) to model the relational information, and use GNNs to assist the learning of agent communication. We show that GraphComm can obtain better results than state-of-the-art methods on the challenging StarCraft II unit micromanagement tasks through extensive experimental evaluation. Yongquan Fu, Huayou Su, Hengyue Pan, Peng Qiao, Yong Dou, Cheng Wang 0003 |
ICASSP | 5 |
| 2021 | Ddper: Decentralized Distributed Prioritized Experience ReplayabstractIn off-policy reinforcement learning, prioritized experience replay plays an important role. However, the centralized prioritized experience replay becomes the bottleneck for efficient training. We propose to approximate the centralized prioritized experience replay in a distributed and decentralized way under certain mild assumptions. To be specific, each actor stores samples in its local replay in the same way as prioritized experience replay, the learner fetches a batch of samples from these replays following a certain strategy. We implement a Deep Q-Learning off-policy algorithm upon the proposed framework. The comparison experiments are performed on a commonly used subset of the Atari-57 learning environment. The experimental results show that the proposed framework speeds up training as the number of actors increases. With the same algorithm and hyper-parameter settings, the proposed framework with 16 actors achieves superior performance that Ape-X with 32 and even more actors does. Sidun Liu, Peng Qiao, Yong Dou, Rongchun Li |
ICME | 2 |
| 2021 | An Edge Traffic Flow Detection Scheme Based on Deep Learning in an Intelligent Transportation SystemabstractAn intelligent transportation system (ITS) plays an important role in public transport management, security and other issues. Traffic flow detection is an important part of the ITS. Based on the real-time acquisition of urban road traffic flow information, an ITS provides intelligent guidance for relieving traffic jams and reducing environmental pollution. The traffic flow detection in an ITS usually adopts the cloud computing mode. The edge of the network will transmit all the captured video to the cloud computing center. However, the increasing traffic monitoring has brought great challenges to the storage, communication and processing of traditional transportation systems based on cloud computing. To address this issue, a traffic flow detection scheme based on deep learning on the edge node is proposed in this article. First, we propose a vehicle detection algorithm based on the YOLOv3 (You Only Look Once) model trained with a great volume of traffic data. We pruned the model to ensure its efficiency on the edge equipment. After that, the DeepSORT (Deep Simple Online and Realtime Tracking) algorithm is optimized by retraining the feature extractor for multiobject vehicle tracking. Then, we propose a real-time vehicle tracking counter for vehicles that combines the vehicle detection and vehicle tracking algorithms to realize the detection of traffic flow. Finally, the vehicle detection network and multiple-object tracking network are migrated and deployed on the edge device Jetson TX2 platform, and we verify the correctness and efficiency of our framework. The test results indicate that our model can efficiently detect the traffic flow with an average processing speed of 37.9 FPS (frames per second) and an average accuracy of 92.0% on the edge device. Chen Chen 0006, Bin Liu 0070, Shaohua Wan 0001, Peng Qiao, Qingqi Pei |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2020 | NeuronFlow: a neuromorphic processor architecture for Live AI applicationsabstractNeuronflow is a neuromorphic, many core, data flow architecture that exploits brain-inspired concepts to deliver a scalable event-based processing engine for neuron networks in Live AI applications. Its design is inspired by brain biology, but not necessarily biologically plausible. The main design goal is the exploitation of sparsity to dramatically reduce latency and power consumption as required by sensor processing at the Edge. Orlando Moreira, Amirreza Yousefzadeh, Fabian Chersi, Gokturk Cinserin, Rik-Jan Zwartenkot, Ajay Kapoor, Peng Qiao, Peter Kievits, Mina A. Khoei, Louis Rouillard, Aimee Ferouge, Jonathan Tapson, Ashoka Visweswara |
DATE | 7 |
| 2020 | Attentional Fused Temporal Transformation Network for Video Action RecognitionabstractEffective spatiotemporal feature representation is crucial to the video-based action recognition task. Focusing on discriminate spatiotemporal feature learning, we propose Attentional Fused Temporal Transformation Network (AttnTTN) for action recognition on top of popular Temporal Segment Network (TSN) framework. In the network, Attentional Fusion Module (AttnFM) is designed to fuse the appearance and motion features at multiple ConvNet levels for each video snippet, forming a short-term video descriptor. With fused features as inputs, Temporal Transformation Networks (TTN) are employed to model middle-term temporal transformation between the neighboring temporal snippets following a sequential order. AttnTTN achieves the state-of-the-art results on two most popular action recognition datasets: UCF101 and HMDB51. Ke Yang 0004, Huadong Dai, Tianlong Shen, Peng Qiao, Xin Niu 0002, Dongsheng Li 0001, Yong Dou |
ICASSP | 5 |
| 2020 | Objectness Consistent Representation for Weakly Supervised Object DetectionabstractWeakly supervised object detection aims at learning object detectors with only image-level category labels. Most existing methods tend to solve this problem by using a multiple instance learning detector which is usually trapped to discriminate object parts. In order to select high-quality proposals, recent works leverage objectness scores derived from weakly-supervised segmentation maps to rank the object proposals. Base on our observation, this kind of segmentation guided method always fails due to neglect of the fact that the objectness of all proposals inside the ground-truth box should be consistent. In this paper, we propose a novel object representation named Objectness Consistent Representation (OCRepr) to meet the consistency criterion of objectness. Specifically, we project the segmentation confidence scores into two orthogonal directions, namely vertical and horizontal, to get the OCRepr. With the novel object representation, more high-quality proposals can be mined for learning a much stronger object detector. We obtain 54.6% and 51.1% mAP scores on VOC 2007 and 2012 datasets, significantly outperforming the state-of-the-art and demonstrating the superiority of OCRepr for weakly supervised object detection. Ke Yang 0004, Peng Zhang 0035, Peng Qiao, Dongsheng Li 0001, Yong Dou |
ACM Multimedia | 3 |
| 2020 | Beyond top-N accuracy indicator: a comprehensive evaluation indicator of CNN models in image classificationabstractNowadays, a large number of deep convolutional neural network (CNN) models are applied to image classification tasks. However, the authors find that the most widely used evaluation indicator, the Top‐ N Accuracy indicator, cannot discriminate these models effectively. In this study, they propose a new indicator called Maximum‐Spanning‐Confusion‐Tree indicator to solve this problem. The Maximum‐Spanning‐Confusion‐Tree indicator is computed based on the hierarchical structure of the Maximum Spanning Confusion Tree of the deep CNN model on the dataset and reflect the ability of deep CNN models to discriminate confused categories in the dataset. The hierarchical structure of the Maximum Spanning Confusion Tree can reveal the confused category set of one selected category in the dataset efficiently and flexibly. Experiments show that they can discriminate ten different deep CNN models more accurately with the Maximum Spanning Confusion Tree indicator than the Top‐ N Accuracy indicator and the Maximum Spanning Confusion Tree intuitively shows the distribution of confused category sets in the dataset so they can find out the weakness of deep CNN models effectively. Yuntao Liu 0004, Yong Dou, Peng Qiao |
IET Comput. Vis. | 3 |
| 2019 | Spatial Attention Network for Few-Shot Learning
Xianhao He, Peng Qiao, Yong Dou, Xin Niu 0002 |
ICANN (2) | 2 |
| 2019 | Exploring frame segmentation networks for temporal action localization
Ke Yang 0004, Xiaolong Shen, Peng Qiao, Shijie Li 0002, Dongsheng Li 0001, Yong Dou |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Exploring Temporal Preservation Networks for Precise Temporal Action LocalizationabstractTemporal action localization is an important task of computer vision. Though a variety of methods have been proposed, it still remains an open question how to predict the temporal boundaries of action segments precisely. Most works use segment-level classifiers to select video segments pre-determined by action proposal or dense sliding windows. However, in order to achieve more precise action boundaries, a temporal localization system should make dense predictions at a fine granularity. A newly proposed work exploits Convolutional-Deconvolutional-Convolutional (CDC) filters to upsample the predictions of 3D ConvNets, making it possible to perform per-frame action predictions and achieving promising performance in terms of temporal action localization. However, CDC network loses temporal information partially due to the temporal downsampling operation. In this paper, we propose an elegant and powerful Temporal Preservation Convolutional (TPC) Network that equips 3D ConvNets with TPC filters. TPC network can fully preserve temporal resolution and downsample the spatial resolution simultaneously, enabling frame-level granularity action localization with minimal loss of time information. TPC network can be trained in an end-to-end manner. Experiment results on public datasets show that TPC network achieves significant improvement in both per-frame action prediction and segment-level temporal action localization. Ke Yang 0004, Peng Qiao, Dongsheng Li 0001, Shaohe Lv, Yong Dou |
AAAI | 2 |
| 2018 | Learning Generic Diffusion Processes for Image Restoration
Peng Qiao, Yong Dou, Yunjin Chen, WenSen Feng |
BMVC | 1 |
| 2018 | Deep Image Clustering Using Convolutional Autoencoder Embedding with Inception-Like BlockabstractImage clustering is one of the challenging tasks in machine learning, and has been extensively used in various applications. Recently, various deep clustering methods has been proposed. These methods take a two-stage approach, feature learning and clustering, sequentially or jointly. We observe that these works usually focus on the combination of reconstruction loss and clustering loss, relatively little work has focused on improving the learning representation of the neural network for clustering. In this paper, we propose a deep convolutional embedded clustering algorithm with inception-like block (DCECI). Specifically, an inception-like block with different type of convolution filters are introduced in the symmetric deep convolutional network to preserve the local structure of convolution layers. We simultaneously minimize the reconstruction loss of the convolutional autoencoders with inception-like block and the clustering loss. Experimental results on multiple image datasets exhibit the promising performance of our proposed algorithm compared with other competitive methods. Qiang Wang 0006, Rongchun Li, Peng Qiao, Ke Yang 0004, Shijie Li 0002, Yong Dou |
ICIP | 4 |
| 2018 | Temporal Pyramid Relation Network for Video-Based Gesture RecognitionabstractGesture recognition in video is an important application of computer vision. However, there are few works talked about the temporal order or relation of the frames in video, which is important for model gestures. In this paper, we propose Temporal Pyramid Relation Network (TPRN) which can model the temporal relation of video frames effectively and efficiently. First, we use Temporal Pyramid Pooling (TPP) layer to get temporal feature sequences of multiple scale pyramids. Then, a Temporal Relation Network (TRN) is stacked on the feature sequence of each scale respectively to model the temporal relations of video frames at multiple scales. At last, representations of all scales are aggregated to get the final prediction. TPRN can take video clips of various length as input and is scalable for video length. We evaluate TPRN on a recently released very large video-based gesture recognition dataset - 20BN-Jester dataset v1, and TPRN achieves competitive performance. Ke Yang 0004, Rongchun Li, Peng Qiao, Qiang Wang 0006, Dongsheng Li 0001, Yong Dou |
ICIP | 3 |
| 2018 | Visual Tree Convolutional Neural Network in Image ClassificationabstractIn image classification, Convolutional Neural Net-work(CNN) models have achieved high performance with the rapid development in deep learning. However, some categories in the image datasets are more difficult to distinguished than others. Improving the classification accuracy on these confused categories is benefit to the overall performance. In this paper, we build a Confusion Visual Tree(CVT) based on the confused semantic level information to identify the confused categories. With the information provided by the CVT, we can lead the CNN training procedure to pay more attention on these confused categories. Therefore, we propose Visual Tree Convolutional Neural Networks(VT-CNN) based on the original deep CNN embedded with our CVT. We evaluate our VT-CNN model on the benchmark datasets CIFAR-10 and CIFAR-100. In our experiments, we build up 3 different VT-CNN models and they obtain improvement over their based CNN models by 1.36%, 0.89% and 0.64%, respectively. Yuntao Liu 0004, Yong Dou, Ruochun Jin, Peng Qiao |
ICPR | 4 |
| 2018 | Fast and Accurate Poisson Denoising With Trainable Nonlinear DiffusionabstractThe degradation of the acquired signal by Poisson noise is a common problem for various imaging applications, such as medical imaging, night vision, and microscopy. Up to now, many state-of-the-art Poisson denoising techniques mainly concentrate on achieving utmost performance, with little consideration for the computation efficiency. Therefore, in this paper we aim to propose an efficient Poisson denoising model with both high computational efficiency and recovery quality. To this end, we exploit the newly developed trainable nonlinear reaction diffusion (TNRD) model which has proven an extremely fast image restoration approach with performance surpassing recent state-of-the-arts. However, the straightforward direct gradient descent employed in the original TNRD-based denoising task is not applicable in this paper. To solve this problem, we resort to the proximal gradient descent method. We retrain the model parameters, including the linear filters and influence functions by taking into account the Poisson noise statistics, and end up with a well-trained nonlinear diffusion model specialized for Poisson denoising. The trained model provides strongly competitive results against state-of-the-art approaches, meanwhile bearing the properties of simple structure and high efficiency. Furthermore, our proposed model comes along with an additional advantage, that the diffusion process is well-suited for parallel computation on graphics processing units (GPUs). For images of size , our GPU implementation takes less than 0.1 s to produce state-of-the-art Poisson denoising performance. WenSen Feng, Peng Qiao, Yunjin Chen |
IEEE Trans. Cybern. | 2 |
| 2017 | Platform-Adaptive High-Throughput Surveillance Video Condensation on Heterogeneous Processor Clusters
Peng Qiao, Teng Li 0010, Yong Dou, Yuanwu Lei, Hongbing Luo |
APPT | 1 |
| 2017 | Learning Non-local Image Diffusion for Image DenoisingabstractImage diffusion plays a fundamental role for the task of image denoising. The recently proposed trainable nonlinear reaction diffusion (TNRD) model defines a simple but very effective framework for image denoising. However, as the TNRD model is a local model, whose diffusion behavior is purely controlled by information of local patches, it is prone to create artifacts in the homogenous regions and over-smooth highly textured regions, especially in the case of strong noise levels. Meanwhile, it is widely known that the non-local self-similarity (NSS) prior stands as an effective image prior for image denoising, which has been widely exploited in many non-local methods. In this work, we are highly motivated to embed the NSS prior into the TNRD model to tackle its weaknesses. In order to preserve the expected property that end-to-end training remains available, we exploit the NSS prior by defining a set of non-local filters, and derive our proposed trainable non-local reaction diffusion (TNLRD) model for image denoising. Together with the local filters and influence functions, the non-local filters are learned by employing loss-specific training. The experimental results show that the trained TNLRD model produces visually plausible recovered images with more textures and less artifacts, compared to its local versions. Moreover, the trained TNLRD model can achieve strongly competitive performance to recent state-of-the-art image denoising methods in terms of peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM). Peng Qiao, Yong Dou, WenSen Feng, Rongchun Li, Yunjin Chen |
ACM Multimedia | 1 |
| 2017 | Image Denoising via Multiscale Nonlinear Diffusion ModelsabstractImage denoising is a fundamental operation in image processing and holds considerable practical importance for various real-world applications. Arguably several thousands of papers are dedicated to image denoising. In the past decade, state-of-the-art denoising algorithms have been clearly dominated by nonlocal patch-based methods, which explicitly exploit patch self-similarity within the targeted image. However, in the past two years, discriminatively trained local approaches have started to outperform previous nonlocal models and have been attracting increasing attention due to the additional advantage of computational efficiency. Successful approaches include cascade of shrinkage fields (CSF) and trainable nonlinear reaction diffusion (TNRD). These two methods are built on the filter response of linear filters of small size using feed forward architectures. Due to the locality inherent in local approaches, the CSF and TNRD models become less effective when the noise level is high and consequently introduce some noise artifacts. In order to overcome this problem, in this paper we introduce a multiscale strategy. To be specific, we build on our newly developed TNRD model, adopting the multiscale pyramid image representation to devise a multiscale nonlinear diffusion process. As expected, all the parameters in the proposed multiscale diffusion model, including the filters and the influence functions across scales, are learned from training data through a loss-based approach. Numerical results on Gaussian and Poisson denoising substantiate that the exploited multiscale strategy can successfully boost the performance of the original TNRD model with a single scale. As a consequence, the resulting multiscale diffusion models can significantly suppress the typical incorrect features for those noisy images with heavy noise. It turns out that multiscale TNRD variants achieve better performance than state-of-the-art denoising methods. WenSen Feng, Peng Qiao, Xuanyang Xi, Yunjin Chen |
SIAM J. Imaging Sci. | 2 |
| 2017 | Variational single image interpolation with time-varying regularization
Peng Qiao, Yunjin Chen, Yong Dou |
Signal Process. Image Commun. | 1 |
| 2011 | A 0.964mW digital hearing aid systemabstractThis paper concerns the design and optimization of a digital hearing aid application. It aims to show that a suitably adapted ASIP can be constructed to create a highly optimized solution for the wide variety of complex algorithms that play a role in this domain. These algorithms are configurable to fit the various hearing impairments of different users. They pose significant challenges to digital hearing aids, having strict area and power consumption constraints. First, a typical digital hearing aid application is proposed and implemented, comprising all critical parts of today's products. Then a small area and ultra low-power 16-bit processor is designed for the application domain. The resulting hearing aid system achieves a power reduction of >; 56 × over the RISC implementation and can operate for >; 300 hours on a typical battery. Peng Qiao, Henk Corporaal, Menno Lindwer |
DATE | 1 |