VLDB 2026 Research / reviewers in the wild / expert
Yong Dou
dblp:76/305
· DBLP profile ↗
196ranked-venue papers
9as first author
83since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 69 · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 65 · 35 since 2021Systems, architecture and hardware · 52 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 9 · 9 since 2021Computer networks · 3Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Transolver Is a Linear Transformer: Revisiting Physics-Attention Through the Lens of Linear AttentionabstractRecent advances in Transformer-based Neural Operators have enabled significant progress in data-driven solvers for Partial Differential Equations (PDEs). Most current research has focused on reducing the quadratic complexity of attention to address the resulting low training and inference efficiency. Among these works, Transolver stands out as a representative method that introduces Physics-Attention to reduce computational costs. Physics-Attention projects grid points into slices for slice attention, then maps them back through deslicing. However, we observe that Physics-Attention can be reformulated as a special case of linear attention, and that the slice attention may even hurt the model performance. Based on these observations, we argue that its effectiveness primarily arises from the slice and deslice operations rather than interactions between slices. Building on this insight, we propose a two-step transformation to redesign Physics-Attention into a canonical linear attention, which we call Linear Attention Neural Operator (LinearNO). Our method achieves state-of-the-art performance on six standard PDE benchmarks, while reducing the number of parameters by an average of 40.0% and computational cost by 36.2%. Additionally, it delivers superior performance on two challenging, industrial-level datasets: AirfRANS and Shape-Net Car. Sidun Liu, Peng Qiao, Zhenglun Sun, Yong Dou |
AAAI | 5 |
| 2026 | Quantitative Analysis and Performance Optimization of Graph Neural Networks on Multi-core CPUsabstractGraph Neural Networks (GNNs) are becoming increasingly popular in graph data processing due to their excellent performance in feature extraction on graph datasets. Compared to GPUs, CPUs are more widely accessible and serve as a practical platform for GNN inference. However, achieving efficient GNN execution on CPUs remains a challenge. We first comprehensively evaluate and quantitatively analyze the performance of GNN inference on multi-core CPUs using the state-of-the-art frameworks, identifying four key performance bottlenecks: inefficient sparse computation, poor data locality, workload imbalance, and inefficient General Matrix Multiplication (GEMM). To tackle these issues, we introduce a set of joint optimizations. Specifically, for the aggregation phase, we propose three optimizations: a register padding and tiling Graph Sparse-dense Matrix Multiplication (GSpMM) algorithm that leverages the computation capability of long vector processing units on modern multi-core CPUs, a destination node-oriented indexes reorganization to enhance data locality, and a boundary buffer-based method to balance the workloads. Additionally, for the update phase, we develop an efficient bias fusion GEMM algorithm, tailored for the irregular matrices. We evaluate the proposed optimizations extensively with three popular GNN models on three typical multi-core CPU platforms. Experimental results on Intel, AMD, and ARM platforms show that our optimizations outperform the state-of-the-art GNN framework DGL by an average factor of 2.41×, 1.58×, and 2.04× (up to 4.75×, 2.70×, and 3.55×), respectively. Compared to PyG, our implementations achieve an average speedup of 1.70×, 1.86×, and 2.44×, respectively. Kangkang Chen, Huayou Su, Xi Yang 0020, Zitong An, Yong Dou, Dongsheng Li 0001 |
ACM Trans. Archit. Code Optim. | 5 |
| 2025 | Maintaining Fairness in Logit-based Knowledge Distillation for Class-Incremental LearningabstractLogit-based knowledge distillation (KD) is commonly used to mitigate catastrophic forgetting in class-incremental learning (CIL) caused by data distribution shifts. However, the strict match of logit values between student and teacher models conflicts with the cross-entropy (CE) loss objective of learning new classes, leading to significant recency bias (i.e. unfairness). To address this issue, we rethink the overlooked limitations of KD-based methods through empirical analysis. Inspired by our findings, we introduce a plug-and-play pre-process method that normalizes the logits of both the student and teacher across all classes, rather than just the old classes, before distillation. This approach allows the student to focus on both old and new classes, capturing intrinsic inter-class relations from the teacher. By doing so, our method avoids the inherent conflict between KD and CE, maintaining fairness between old and new classes. Additionally, recognizing that overconfident teacher predictions can hinder the transfer of inter-class relations (i.e., dark knowledge), we extend our method to capture intra-class relations among different instances, ensuring fairness within old classes. Our method integrates seamlessly with existing logit-based KD approaches, consistently enhancing their performance across multiple CIL benchmarks without incurring additional training costs. Zijian Gao, Shanhao Han, Xingxing Zhang 0001, Kele Xu, Dulan Zhou, Xinjun Mao, Yong Dou, Huaimin Wang 0001 |
AAAI | 7 |
| 2025 | Highly Parallelized Reinforcement Learning Training with Relaxed Assignment DependenciesabstractAs the demands for superior agents grow, the training complexity of Deep Reinforcement Learning (DRL) becomes higher. Thus, accelerating training of DRL has become a major research focus. Dividing the DRL training process into sub-tasks and using parallel computation can effectively reduce training costs. However, current DRL training systems lack sufficient parallelization due to data assignment between sub-task components. This assignment issue has been ignored, but addressing it can further boost training efficiency. Therefore, we propose a high-throughput distributed RL training system called TianJi. It relaxes assignment dependencies between sub-task components and enables event-driven asynchronous communication. Meanwhile, TianJi maintains clear boundaries between sub-task components. To address convergence uncertainty from relaxed assignment dependencies, TianJi proposes a distributed strategy based on the balance of sample production and consumption. The strategy controls the staleness of samples to correct their quality, ensuring convergence. We conducted extensive experiments. TianJi achieves a convergence time acceleration ratio of up to 4.37 compared to related comparison frameworks. When scaled to eight computational nodes, TianJi shows a convergence time speedup of 1.6 and a throughput speedup of 7.13 relative to XingTian, emonstrating its capability to accelerate training and scalability. In data transmission efficiency experiments, TianJi significantly outperforms other frameworks, approaching hardware limits. TianJi also shows effectiveness in on-policy algorithms, achieving convergence time acceleration ratios of 4.36 and 2.95 compared to RLlib and XingTian. Zhouyu He, Peng Qiao, Rongchun Li, Yong Dou, Yusong Tan |
AAAI | 4 |
| 2025 | Rethinking Incision Segmentation with Geometry-Structure Aligned Polygon PromptabstractThe incision plays a critical role in clinical surgery and has recently emerged as a complex segmentation objective in medical image analysis. The Segment Anything Model (SAM) demonstrates strong generalization and performs well across various medical imaging tasks. With task-specific finetuning, SAM can generate rough incision segmentations based on prompts. However, the irregular morphology of incisions presents challenges to existing prompting strategies. First, coarse prompts such as points and bounding boxes fail to capture the shape characteristics of incisions. Second, while existing convex polygon prompts show potential in guiding incision segmentation, their structural constraints limit the representation of concave regions. In contrast, unconstrained naive polygons often suffer from improper vertex sampling, resulting in ineffective guidance. Moreover, the prompt encoding schemes with fixed vertex order tend to introduce ambiguity when applied to flexible polygon structures. To address these challenges, we explore how to deliver more effective prompts under limited interactions and propose the Geometry-Structure Aligned Polygon Prompt (GSAPP) method, which maintains both geometric and structural consistency with the segmentation target. GSAPP comprises two core modules: Structure-Aligned Polygon Prompt (SAPP) and Geometry-Aligned Encoding (GAE) module. SAPP introduces a polygon prompt adapted to incision morphology, mitigating structural mismatches with the target shape. Additionally, we develop a heuristic generator to automatically produce high-quality polygon prompts. The GAE module encodes prompts using extreme vertex matching, effectively reducing geometric ambiguity. Extensive experiments demonstrate that GSAPP improves prompt effectiveness and achieves more than a 3% Dice gain over the previous state-of-the-art in incision segmentation. Keran Ding, Peng Qiao, Zhenglun Sun, Tongrui Hu, Yong Dou |
BIBM | 5 |
| 2025 | Beyond Synthetic Data: Leveraging Natural Image Pretraining and Finetuning for Mixed Exposure Correction in Capsule EndoscopyabstractWireless capsule endoscopy (WCE) has become standard in gastrointestinal diagnostics, but its frames often exhibit mixed exposure problem with both over- and underexposed regions. Existing methods mostly use pixel-level supervision in RGB space and introduce unnatural color shifts because exposure, color, and texture are tightly entangled. Another challenge is the lack of reliable paired WCE training data. Available synthetic datasets are small, low in quality, and lack diversity, while real paired examples are hard to obtain. Fully supervised models trained on such data tend to overfit and generalize poorly. We first extract exposure priors from a model pretrained on large natural-image datasets and then fine-tune it on WCE data. Leveraging hue stability and pixel dispersion nature in HSV space, our separated exposure and color correction method comprises three significant components: the training-free Single-Image Exposure Fusion (SIEF) module for precise exposure partitioning, the Sequential Interaction Correction (SIC) module for efficient information exchange between the exposure and color correction branches, and the Phase-Shifting Coder (PSC) to resolve discontinuities in the circular H channel and ensure smooth, stable hue prediction. Extensive experiments demonstrate that our method achieves SOTA results on natural-image exposure correction and transfers effectively to WCE mixed-exposure cases. It outperforms prior supervised approaches and shows robust recovery across multiple WCE datasets. Tongrui Hu, Peng Qiao, Keran Ding, Yong Dou, Rongchun Li |
BIBM | 4 |
| 2025 | Partial Order-centered Hyperbolic Representation Learning for Few-shot Relation ExtractionabstractPrototype network-based methods have made substantial progress in few-shot relation extraction (FSRE) by enhancing relation prototypes with relation descriptions. However, the distribution of relations and instances in distinct representation spaces isolates the constraints of relations on instances, making relation prototypes biased. In this paper, we propose an end-to-end partial order-centered hyperbolic representation learning (PO-HRL) framework, which imposes the constraints of relations on instances by modeling partial order in hyperbolic space, so as to effectively learn the distribution of instance representations. Specifically, we develop the hyperbolic supervised contrastive learning based on Lorentzian cosine similarity to align representations of relations and instances, and model the partial order by constraining instances to reside within the Lorentzian entailment cone of their respective relation. Experiments on three benchmark datasets show that PO-HRL outperforms the strong baselines, especially in 1-shot settings lacking relation descriptions. Zhen Huang 0006, Minghao Hu 0001, Pinglv Yang, Peng Qiao, Yong Dou, Zhilin Wang |
COLING | 6 |
| 2025 | Knowledge Memorization and Rumination for Pre-trained Model-based Class-Incremental LearningabstractClass-Incremental Learning (CIL) enables models to continuously learn new classes while mitigating catastrophic forgetting. Recently, Pre-Trained Models (PTMs) have greatly enhanced CIL performance, even when fine-tuning is limited to the first task. This advantage is particularly beneficial for CIL methods that freeze the feature extractor after first-task fine-tuning, such as analytic learning-based approaches using a least squares solution-based classification head to acquire knowledge recursively. In this work, we revisit the analytical learning approach combined with PTMs and identify its limitations in adapting to new classes, leading to sub-optimal performance. To address this, we propose the Momentum-based Analytical Learning (MoAL) approach. MoAL achieves robust knowledge memorization via an analytical classification head and improves adaptivity to new classes through momentum-based adapter weight interpolation, leading to forgetting outdated knowledge. Importantly, we introduce a knowledge rumination mechanism that leverages refined adaptivity, allowing the model to revisit and reinforce old knowledge, thereby improving performance on old classes. MoAL facilitates the acquisition of new knowledge and consolidates old knowledge, achieving a win-win outcome between plasticity and stability. Extensive experiments on various incremental settings show MoAL’s state-of-the-art performance1. Zijian Gao, Wangwang Jia, Xingxing Zhang 0001, Dulan Zhou, Kele Xu, Yong Dou, Xinjun Mao, Huaimin Wang 0001 |
CVPR | 7 |
| 2025 | GPIS: Geometric Informed Polygon Prompt for Incision Segmentation
Keran Ding, Peng Qiao, Zhenglun Sun, Yong Dou |
ICANN (2) | 6 |
| 2025 | A Counterfactual Ultrasound Anti-Interference Self-Supervised Network for B-mode Ultrasound Tongue ExtractionabstractB-mode ultrasound tongue imaging is a non-invasive and real-time method for visualizing vocal tract deformation. However, accurately extracting the tongue’s surface contour remains a significant challenge due to the low signal-to-noise ratio (SNR) and prevalent speckle noise in ultrasound images. Traditional supervised learning models often require large labeled datasets, which are labor-intensive to produce and susceptible to noise interference. To address these limitations, we present a novel Counterfactual Ultrasound Anti-Interference Self-Supervised Network (CUAI-SSN), which integrates self-supervised learning (SSL) with counterfactual data augmentation, progressively disentangles confounding factors, ensuring that the model generalizes well across varied ultrasound conditions. Our approach leverages causal reasoning to decouple noise from relevant features, enabling the model to learn robust representations that focus on essential tongue structures. By generating counterfactual image-label pairs, our method introduces alternative, noise-independent scenarios that enhance model training. Furthermore, we introduce attention mechanisms to enhance the network’s ability to capture fine-grained details even in noisy conditions. Extensive experiments on real ultrasound tongue images demonstrate that CUAI-SSN outperforms existing methods, setting a new benchmark for automated contour extraction in ultrasound tongue imaging. Our code is publicly available at https://github.com/inexhaustible419/CounterfactualultrasoundAI. Yan Jia 0001, Yuqing Cheng, Kele Xu, Yong Dou, Peng Qiao, Zhouyu He |
ICASSP | 4 |
| 2025 | Scaling Bioacoustic Signal Pre-training with Million Samples Via Mask-ModelingabstractDeep learning-based bioacoustic audio analysis holds immense potential across various applications. However, existing studies in bioacoustics often focus on a limited number of species, potentially hindering the transferability of models across different species. Furthermore, the manual annotation of bioacoustic data is both costly and labor-intensive. To address these challenges, self-supervised learning on large-scale bioacoustic audio data presents a promising solution. In this paper, we introduce GPM-BT (General Pre-training Model for Bioacoustic Tasks), a self-supervised, Transformer-based model pre-trained on approximately 1.2 million unannotated bioacoustic audio samples. We evaluate the scalability and effectiveness of this pre-training approach through comprehensive experiments across a broad range of classification and detection tasks. Our results demonstrate that pre-training on large-scale bioacoustic data significantly enhances model performance, improving both generalization and robustness. Notably, GPM-BT achieves state-of-the-art performance on the BEANS benchmark and secures first place in the Few-shot Bioacoustic Event Detection task at the IEEE DCASE 2024 Challenge1. To further advance research in bioacoustics, we have open-sourced our models and code2. Xuyao Deng, Tianjiao Wan, Kele Xu, Peng Qiao, Yong Dou |
ICASSP | 7 |
| 2025 | MonoIR: Inpainting and Reconstruction for Monocular Endoscope Deformation ScenesabstractMonocular endoscopic scene reconstruction is challenging due to limited viewpoints and interference from surgical instruments. While 3D Gaussian-based methods are popular for their strong reconstruction capabilities and efficiency, they often rely on sensors or stereo depth, resulting in blurred tissue areas when instruments obstruct the view. To overcome these issues, we propose MonoIR, an inpainting and reconstruction method that produces spatiotemporally consistent videos and high-quality reconstructions. Our approach employs a propagation-based method to inpaint video holes, enhanced by optical flow constraints for robustness and an additional transformer-based module to address detail loss. We also introduce the normal and depth regularization with confidence to improve reconstruction quality. Extensive experiments demonstrate that MonoIR efficiently reconstructs deformed tissues and outperforms both monocular and binocular methods in key metrics. Ziteng Zhang, Sidun Liu, Peng Qiao, Yong Dou |
ICASSP | 5 |
| 2025 | End-To-End Casual Video Reconstruction: Geometry, Pose and MotionabstractFrom casual videos in daily life, humans can effortlessly recognize object shapes, perceive variations in viewpoints, decompose dynamic objects and infer their motions. This suggests that reconstruction algorithms should also, like humans, simultaneously achieve these capabilities. However, existing methods tackle the aforementioned tasks in multiple stages. The phased processing approaches mean that each stage’s performance heavily depends on the preceding one, and the entire reconstruction process cannot be optimized in an end-to-end manner, which limits its overall potential. In response, we present an algorithm, designed to integrate these capabilities for casual videos in an end-to-end manner. Specifically, we represent the 4D scene in a video as the combination of local multi-view depth maps and a shared canonical space, where a continuous bijective mapping is used to model motions between local and canonical space. We also extend the depth estimation network to decompose scenes into static and dynamic parts, which helps to avoid degenerate cases and leads to accurate tracking. Experiments demonstrate that our method performs well on casual videos, and achieves performance comparable to state-of-the-art methods across all tasks. Our project page is https://fullre.github.io/FullRe. Peng Qiao, Sidun Liu, Zongxin Ye, Ziteng Zhang, Zhenglun Sun, Yong Dou |
ICME | 7 |
| 2025 | Only One Stage: A Chemical-Aware Model for Accurate Combustion Chemical Kinetics PredictionabstractThe combustion chemical kinetics simulation, which focuses on the change in species mass fractions during reactions, is vital for clean energy development. Due to the sample complexities introduced by chemical kinetics, current multi-stage methods employ data preprocessing stage to separate it into subspaces, aiming to ease training. However, the current approaches to sample space separation are not effective, which not only affects the training accuracy of the network but also introduces a complex preprocessing procedure. To solve this, we propose a one-stage, end-to-end model with chemical-aware capabilities, using an auto dividing mechanism and spatio-temporal convolution for feature extraction. In hydrogen simulation experiments, our model achieved an L1 error of 10−6level, which is nearly identical to the standard results of numerical calculations, outperforming the usually used Multilayer Perceptron (MLP) method by 41.9 times. Zhenglun Sun, Peng Qiao, Yong Dou, Rongchun Li, Sidun Liu |
ICME | 3 |
| 2025 | Improving the Continuity of Goal-Achievement Ability via Policy Self-Regularization for Goal-Conditioned Reinforcement LearningabstractThis paper addresses the challenge of discontinuity in goal-achievement capabilities observed in Goal-conditioned Reinforcement Learning (GCRL) algorithms. Through a theoretical analysis, we identify that the reuse of successful trajectories or policies during training can aid in achieving adjacent goals of achievable goals. However, the policy discrepancy between achievable and adjacent goals must be carefully managed to avoid both overly trivial and excessively large differences, which can respectively hinder policy performance. To tackle this issue, we propose a margin-based policy self-regularization approach that optimizes the policy discrepancies between adjacent desired goals to a minimal acceptable threshold. This method can be integrated into popular GCRL algorithms, such as GC-SAC, HER, and GC-PPO. Systematic evaluations across two robotic arm control tasks and a complex fixed-wing aircraft control task demonstrate that our approach significantly improves the continuity of goal-achievement abilities of GCRL algorithms, thereby enhancing their overall performance. Xudong Gong, Sen Yang 0003, Kele Xu, Bo Ding 0001, Huaimin Wang 0001, Yong Dou |
ICML | 7 |
| 2025 | Multi-view Fusion and Parameter Perturbation for Few-Shot Class-Incremental Audio Classification
Yulu Fang, Mingyue He, Qisheng Xu, Jianqiao Zhao, Cheng Yang 0004, Kele Xu, Yong Dou |
INTERSPEECH | 7 |
| 2025 | Knowledge Bridges the Intent Gap: Contextual Fusion in Medical Fine-Grained Segmentation
Peng Qiao, Yan Jia 0001, Yong Dou |
MICCAI (1) | 5 |
| 2025 | Mono3R: Exploiting Monocular Cues for Geometric 3D ReconstructionabstractRecent advances in data-driven geometric multi-view 3D reconstruction foundation models (e.g., DUSt3R) have shown remarkable performance across various 3D vision tasks, facilitated by the release of large-scale, high-quality 3D datasets. However, as we observed, constrained by their matching-based principles, the reconstruction quality of existing models suffers significant degradation in challenging regions with limited matching cues, particularly in weakly textured areas and low-light conditions. To mitigate these limitations, we propose to harness the inherent robustness of monocular geometry estimation to compensate for the shortcomings. Specifically, we introduce a monocular-guided refinement module that integrates monocular geometric priors into multi-view reconstruction frameworks. This integration substantially enhances the robustness of multi-view reconstruction systems, leading to high-quality feed-forward reconstructions. Comprehensive experiments across multiple benchmarks demonstrate that our method achieves substantial improvements in both multi-view camera pose estimation and point cloud accuracy. Sidun Liu, Peng Qiao, Yong Dou |
ACM Multimedia | 4 |
| 2025 | Regist3R: Incremental Registration with Stereo Foundation ModelabstractMulti-view 3D reconstruction has remained an essential yet challenging problem in the field of computer vision. While DUSt3R and its successors have achieved breakthroughs in 3D reconstruction from unposed images, these methods exhibit significant limitations when scaling to multi-view scenarios, including high computational cost and cumulative error induced by global alignment. To address these challenges, we propose Regist3R, a novel stereo foundation model tailored for efficient and scalable incremental reconstruction. Regist3R leverages an incremental reconstruction paradigm, enabling large-scale 3D reconstructions from unordered and many-view image collections. We evaluate Regist3R on public datasets for camera pose estimation and 3D reconstruction. Our experiments demonstrate that Regist3R achieves comparable performance with optimization-based methods while significantly improving computational efficiency, and outperforms existing multi-view reconstruction models. Furthermore, to assess its performance in real-world applications, we introduce a challenging oblique aerial dataset which has long spatial spans and hundreds of views. The results highlight the effectiveness of Regist3R. We also demonstrate the first attempt to reconstruct large-scale scenes encompassing over thousands of views through pointmap-based foundation models, showcasing its potential for practical applications in large-scale 3D reconstruction tasks, including urban modeling, aerial mapping, and beyond. Sidun Liu, Peng Qiao, Yong Dou |
ACM Multimedia | 4 |
| 2025 | UniEmotion: A Unified Framework for Multimodal Emotion Recognition with Iterative Consensus-based TrainingabstractTraditional emotion recognition methods struggle with complex emotional dynamics including multi-emotion states, transitions, and contextual reasoning. While multimodal large language models demonstrate great potential for understanding such complex scene dynamics, they still face challenges in adapting to emotion recognition tasks. We propose UniEmotion, a unified framework that simultaneously addresses conventional categorical emotion recognition, open-vocabulary fine-grained emotion recognition, and descriptive emotion understanding. Our approach leverages an iterative consensus-based training pipeline where pseudo-labels and model parameters co-evolve, maximizing large models' utility while mitigating downstream limitations. The framework integrates a selector module that identifies high-quality samples through prediction variance analysis, coupled with a pseudo-labeling module employing consistency regularization and class-wise adaptive mapping. This dual mechanism reduces error accumulation during self-training while aligning open-vocabulary output with task-specific labels. Experimental results demonstrate the effectiveness of our framework, achieving state-of-the-art performance across all three tracks, including 1st place on the MER-SEMI track with a significant improvement of 11.97% over the best baseline, and 2nd place on the MER-DES track. Yanjie Sun, Wuyang Chen 0002, Yong Dou |
ACM Multimedia | 3 |
| 2025 | AudioSet-R: A Refined AudioSet with Multi-Stage LLM Label ReannotationabstractAudioSet is a widely used benchmark in the audio research community and has significantly advanced various audio-related tasks. However, persistent issues with label accuracy and completeness remain critical bottlenecks that limit performance in downstream applications. To address the aforementioned challenges, we propose a three-stage reannotation framework that harnesses general-purpose audio-language foundation models to systematically improve the label quality of AudioSet. The framework employs a cross-modal prompting strategy, inspired by the concept of prompt chaining, wherein prompts are sequentially composed to execute subtasks (audio comprehension, label synthesis, and semantic alignment). Leveraging this framework, we construct a high-quality, structured relabeled version of AudioSet-R. Extensive experiments conducted on representative audio classification models-including AST, PANNs, SSAST, and AudioMAE-consistently demonstrate substantial performance improvements, thereby validating the generalizability and effectiveness of the proposed approach in enhancing label reliability.The code is publicly available at: https://github.com/colaudiolab/AudioSet-R. Qisheng Xu, Yi Su 0010, Yong Dou, Xinwang Liu 0002, Kele Xu |
ACM Multimedia | 5 |
| 2025 | A cause fusion framework with information bottleneck for conversational causal emotion entailment
Xinxin Su, Zhen Huang 0006, Menglong Lu, Sisi Dai, Yong Dou |
Neural Networks | 6 |
| 2025 | Segment Anything for Visual Bird Sound DenoisingabstractCurrent audio denoising methods perform well with synthetic noise but struggle with complex natural noise, especially for bird sounds, which contain natural environmental sounds such as wind and rain, making it challenging to extract clean bird sounds. This issue becomes more pronounced with short and faint bird sounds, where existing methods are less effective. In this paper, we introduceBudSAM, a novel audio denoising model that incorporates theSegment Anything Model (SAM), originally designed for image segmentation task, into the field of visual bird sound denoising. By treating audio denoising as a segmentation task, BudSAM utilizes SAM's powerful segmentation capabilities and we incorporates BCE and Dice losses to enhance the model's ability to segment weak signals, effectively isolating the clean bird sounds that are often masked by background noise. Our method is evaluated on the BirdSoundsDenoising dataset, achieving a 4.0% improvement in IoU and a 0.77 dB increase in SDR compared to state-of-the-art methods. To the best knowledge of the authors, BudSAM marks the first attempt which employs SAM in audio denoising task, offering a promising direction for future research and real-world bird sound processing tasks. Tianjiao Wan, Kele Xu, Peng Qiao, Yong Dou |
IEEE Signal Process. Lett. | 5 |
| 2025 | SPDFA: A Novel Dataflow Fusion Sparse Deep Neural Network AcceleratorabstractUnstructured sparse pruning significantly reduces the computational and parametric complexities of deep neural network models. Nevertheless, the highly irregular nature of sparse models limits their performance and efficiency on traditional computing platforms, thereby prompting the development of specialized hardware solutions. To improve computational efficiency, we introduce the Sparse Dataflow Fusion Accelerator (SPDFA), a specialized architecture meticulously designed for sparse deep neural networks. Firstly, we present a non-blocking data distribution-computing engine that integrates inner product and column product. This engine boosts computational efficiency by decomposing matrix multiplication and convolution into rectangular matrix-vector multiplications. Secondly, we implement a computation array to further exploit the parallelism, and design an on-chip buffer structure that supports multi-line memory access mode. Lastly, to bolster the adaptability of our accelerator, we propose an innovative macroinstruction set coupled with a micro-kernel scheme. Furthermore, we refine the macroinstruction issue strategy, thereby further enhancing computational efficiency. Our evaluation results demonstrate that SPDFA achieves an average 1.29 \(\times\) –2.38 \(\times\) improvement in computational efficiency compared to the state-of-the-art SpMM accelerators when applied to unstructured sparse deep neural network models. Furthermore, its performance outperforms existing sparse neural network accelerators by a factor of 1.03 \(\times\) –1.83 \(\times\) . Additionally, SPDFA exhibits excellent scalability with a scaling efficiency exceeding 80%. Jinwei Xu, Jingfei Jiang, Xifu Qian, Yong Dou |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2024 | A Connectivity-Enhanced Multi-Task Learning based on Anatomical Priors for 3D Class-Balanced Pulmonary Airway SegmentationabstractAccurate and efficient airway segmentation is essential for evaluating pulmonary diseases, aiding diagnosis, reducing the preoperative burden of airway identification, and minimizing patient discomfort during prolonged surgeries. However, current pulmonary airway reconstruction techniques are hindered by two major challenges: difficulty in accurately reconstructing fine airway branches due to the tendency to overlook small targets, and insufficient structural connectivity leading to frequent branch discontinuities within the airway tree. These limitations directly affect the clinical applicability of reconstructed airways. To overcome these challenges, a novel 3D pulmonary airway segmentation multi-task framework is proposed, designed to enhance the performance of existing backbone models. This approach integrates Anatomical Prior-Based Multi-Task Learning (AP-MTL) through the use of Gaussian-constructed connectivity-enhanced isosurfaces, significantly improving the network’s ability to maintain airway continuity. Additionally, a Class-Balanced CT Density Distribution Reconstruction mechanism (DDR-CB) is introduced, further refining the model’s capability to detect and segment fine airway branches. As a result of these enhancements, the model demonstrates a 11.5% average improvement in segmentation accuracy and connectivity compared to the baseline. The source code is publicly accessible at https://github.com/inexhaustible419/APMTLAirwaySegment. Yan Jia 0001, Yong Dou, Peng Qiao, Yuqing Cheng, Kele Xu, Zhouyu He |
BIBM | 2 |
| 2024 | A New Pipeline for Knowledge Graph Reasoning Enhanced by Large Language Models Without Fine-TuningabstractConventional Knowledge Graph Reasoning (KGR) models learn the embeddings of KG components over the structure of KGs, but their performances are limited when the KGs are severely incomplete.Recent LLM-enhanced KGR models input KG structural information into LLMs.However, they require fine-tuning on open-source LLMs and are not applicable to closed-source LLMs.Therefore, in this paper, to leverage the knowledge in LLMs without fine-tuning to assist and enhance conventional KGR models, we propose a new three-stage pipeline, including knowledge alignment, KG reasoning and entity reranking.Specifically, in the alignment stage, we propose three strategies to align the knowledge in LLMs to the KG schema by explicitly associating unconnected nodes with semantic relations.Based on the enriched KGs, we train structure-aware KGR models to integrate aligned knowledge to original knowledge existing in KGs.In the reranking stage, after obtaining the results of KGR models, we rerank the top-scored entities with LLMs to recall correct answers further.Experiments show our pipeline can enhance the KGR performance in both incomplete and general situations. Zhongwu Chen, Long Bai 0002, Zixuan Li 0001, Zhen Huang 0002, Xiaolong Jin 0001, Yong Dou |
EMNLP | 6 |
| 2024 | SAM-NeRF: NeRF-Based 3D Instance Segmentation with Segment Anything Model
Linglin Xie, Peng Qiao, Yong Dou, Sidun Liu, Kaijun Yang |
ICANN (2) | 4 |
| 2024 | Adapter-Based Incremental Learning for Face Forgery DetectionabstractMany existing face forgery detection methods primarily revolve around learning general representations on predefined datasets and subsequently crossing these static representations to other datasets. However, these approaches could lead to catastrophic forgetting in real-world scenarios, especially when new forgery methods continually emerge. In this paper, we proposed a novel incremental learning framework for face forgery detection, where we design an adapter-based incremental learning scheme combined with a confidence-based ensemble prediction mechanism. When confronted with new forgery methods, we incorporate small trainable adapter modules, which are retrained along with their corresponding classification layers, yielding a series of task-specific modules. Then we incorporate a confidence-based ensemble prediction mechanism to aggregate all predictions. Through comprehensive evaluations on multiple benchmark datasets (FF++, DFD, and Celeb-DF), our method successfully mitigates the catastrophic forgetting problem in a cost-effective manner and attains state-of-the-art performance in cross-dataset scenario. Caili Gao, Qisheng Xu, Peng Qiao, Kele Xu, Xifu Qian, Yong Dou |
ICASSP | 6 |
| 2024 | Improving Motion Deblur By Multi-Output LearningabstractImage deblurring is an ill-posed task, where exists infinite feasible solutions for blurry images. Modem deep learning approaches usually discard the learning of blur kernels and directly employ end-to-end supervised learning. However, supervised learning can’t handle ill-posed tasks appropriately. It regresses the average thus losing sharp details. Therefore, we propose an extension to the network to learn from stochastic supervisions, where a novel multi-output architecture and Min-Out loss function are designed. Our approach enables the model to output multiple feasible solutions to fit various non-uniform motions. We then propose a novel parameter multiplexing method that reduces computations while improving performance with fewer parameters. After training, the best-performed head is fine-tuned to be used for inference where the sharp label is absent. The proposed approach is evaluated with multiple image deblur models on the GoPro motion deblur dataset. On average, the multi-output extension improves the PSNR by 0.08 dB. When applied to the popular attention-based model Restormer, the multi-output helps it achieve 33.05 dB (+0.13 dB) PSNR without modification on network architecture. Sidun Liu, Peng Qiao, Yong Dou |
ICASSP | 3 |
| 2024 | Voice-to-Face Generation: Couple of Self-Supervised Representation Learning with Diffusion ModelabstractIn this study, we explore the challenging task of generating facial images from unheard voices, aiming to synthesize similar faces that correspond to the voice identity. We design a novel framework that encompasses voice-face self-supervised representation learning and extends to voice-based face generation. The key idea behind the feasibility of cross-modal generation is that we not only enrich the voice representations by modeling the locally inherent correlations within voice data but also establish the cross-modal connections through aligning voice with paired face data. To enhance the association between voice and face, we further promote a false negative mitigation method. The learned voice representations are then fed into the diffusion model using cross-attention to produce an image. Experiments show that our framework outperforms previous state-of-the-art methods on various voice-face association evaluation tasks and yields substantially better images than prior approaches. Wuyang Chen 0002, Kele Xu, Yong Dou |
ICME | 3 |
| 2024 | ParaSurRe: Parallel Surface Reconstruction with No Pose PriorabstractSurface reconstruction from multi-view images without pose prior is challenging. Recent advances integrate incremental Structure from Motion (SfM) pipeline into surface optimization, enabling simultaneous surface reconstruction and pose estimation. However, due to the inherent incremental registration scheme of SfM, the efficiency of these methods is far from satisfactory, e.g., reconstruction of an object captured by 49 images costs over 9 hours using a high-end GPU. Inspired by divide-and-conquer strategy, we present a Parallel Surface Reconstruction method, coined as ParaSurRe, where image collections are divided into non-overlapped clusters and the incremental reconstructions are performed individually in each cluster. Owing to image partitioning, each cluster only accurately reconstructs a part of the surface. The major challenge is to merge multiple partial neural implicit surfaces into a complete one. We propose a confidence-aware surface fusion strategy and a geometry-guided refinement to tackle this issue. Experiments on real-world datasets demonstrate that ParaSurRe reconstructs delicate surfaces from unposed images, and achieves competitive pose estimation performance compared with state-of-the-art methods, with up to 6.5× speedup on a scene captured by 81 images. Zongxin Ye, Sidun Liu, Ziteng Zhang, Peng Qiao, Yong Dou |
ICME | 7 |
| 2024 | Self-Supervised Learning-Based General Fine-tuning Framework For Audio Classification and Event DetectionabstractRecently, self-supervised learning (SSL) has made remarkable progress in signal representation and has become a de facto solution for different audio processing tasks. Generally, the SSL consists of the foundation pre-training and downstream fine-tuning phases. However, fine-tuning frameworks may lack universality due to the distinct learning paradigms and model designs employed in audio signal processing tasks. Furthermore, the varying degrees of dataset labeling across different tasks challenge unifying a fine-tuning framework. To address these issues, we propose vec2task, a cross-task general fine-tuning framework based on the SSL pre-trained model. It employs a semantic-aware module and an alternating training strategy, enabling the framework to generalize across various audio signal processing tasks. Additionally, the framework employs automatic audio augmentation strategies, eliminating the requirement for individually tailored algorithms to improve task performance. Experimental validations of the vec2task framework outperformed previous methods in audio classification and event detection tasks, showcasing its generalization ability across tasks. Yanjie Sun, Kele Xu, Yong Dou |
ICME | 3 |
| 2024 | CRNet: Cross-Reconstruction Network for Inconsistent Point Cloud RegistrationabstractDeep learning methods have made significant advancements in point cloud registration, achieving excellent performance on consistent point clouds. However, these methods face challenges when dealing with point clouds exhibiting inconsistent spatial distributions. To address this issue, we propose the Cross-Reconstruction Network (CRNet), a novel approach designed to register two inconsistent point clouds by reconstructing them into a consistent shape. CRNet utilizes a cross-learning framework that facilitates feature interaction between input point clouds at both global and point-wise levels. This interaction network enables the bidirectional generation of corresponding points to reconstruct consistent point clouds for transformation estimation. Furthermore, the transformation parameters can be refined by a regression network to achieve more accurate registration. The experimental results valuated on benchmark datasets demonstrate that CRNet outperforms state-of-the-art methods in inconsistent scenarios. Yunzhe Xiao, Xueqiong Li, Shaowu Yang, Wenjing Yang 0002, Yong Dou |
ICME | 5 |
| 2024 | Contrastive Learning-based Chaining-Cluster for Multilingual Voice-Face Association
Wuyang Chen 0002, Yanjie Sun, Kele Xu, Yong Dou |
ACM Multimedia | 4 |
| 2024 | AbsGS: Recovering Fine Details in 3D Gaussian Splatting
Zongxin Ye, Sidun Liu, Peng Qiao, Yong Dou |
ACM Multimedia | 5 |
| 2024 | ER-SFM: Efficient and Robust Cluster-Based Structure from Motion
Zongxin Ye, Sidun Liu, Peng Qiao, Yong Dou |
PRCV (6) | 5 |
| 2024 | Sensing the diversity of rumors: Rumor detection with hierarchical prototype contrastive learningabstractThe proliferation of rumors on social networks poses a serious threat to cybersecurity, justice and public trust, increasing the urgent need for rumor detection. Existing detection methods typically treat all rumors as a single homogeneous category, neglecting the diverse semantic hierarchies within rumors. Rumors pervade various domains, each with its distinct characteristics. These methods tend to lag in expressiveness when confronted with real-world scenarios involving multiple semantic levels . Furthermore, the diversity of rumors also complicates the collection of datasets, and inevitably introduces noisy data, which hinders the correctness of the learned representations. To address these challenges, we propose a rumor detection framework with Hierarchical Prototype Contrastive Learning (HPCL). In this framework, we construct a set of dynamically updated hierarchical prototypes through contrastive learning to encourage capturing the hierarchical semantic structure within rumors. Additionally, we design a difficulty metric function based on the distance between instances and prototypes, and introduce curriculum learning to mitigate the adverse effects of noisy data. Experiments on four public datasets demonstrate that our approach achieves state-of-the-art performance. Our code is publicly released at https://github.com/Coder-HenryZa/HPCL . Peng Zheng 0003, Yong Dou, Yeqing Yan |
Inf. Process. Manag. | 2 |
| 2024 | FedGKD: Federated Graph Knowledge Distillation for privacy-preserving rumor detection
Peng Zheng 0003, Yong Dou, Yeqing Yan |
Knowl. Based Syst. | 2 |
| 2024 | Hierarchical Shared Encoder With Task-Specific Transformer Layer Selection for Emotion-Cause Pair ExtractionabstractEmotion Cause Pair Extraction (ECPE) aims to extract emotions and their causes from a document. Powerful emotion and cause extraction abilities have proven essential in achieving accurate ECPE. However, most existing methods employ shared feature learning of emotion extraction and cause extraction, which can harm the abilities of both tasks as they focus on different information (i.e., task-specific features). Moreover, shared feature learning of the two tasks also leads to the label imbalance problem. To address these issues, this paper proposes a multi-task learning framework named Hierarchical Shared Encoder with Task-specific Transformer Layer Selection (HSE-TTLS). The model achieves ECPE via two subtasks: Emotion Extraction (EE) and Emotion Cause Extraction (ECE). The design of two subtasks for ECPE corresponds to the fact that cause clauses are emotion-dependent and significantly alleviates the label imbalance problem. To effectively extract task-specific features for EE and ECE, we employ BERT as the token-level encoder and select task-specific optimal layers for the two subtasks. Focal loss is used as the objective function for EE to further alleviate the label imbalance problem. Extensive experiments on benchmark ECPE corpus demonstrate the effectiveness of HSE-TTLS, which outperforms state-of-the-art baseline methods by at least 1.56% on the F1 score. Xinxin Su, Zhen Huang 0006, Yixin Su 0001, Bayu Distiawan Trisedya, Yong Dou |
IEEE Trans. Affect. Comput. | 5 |
| 2024 | Automated Data Augmentation for Audio ClassificationabstractAudio classification is a challenging task that requires categorizing audio data based on its content or characteristics. Existing approaches for audio classification rely either on supervised learning or fine-tuning based on self-supervised learning, both of which require manually labeled data. However, manually labeling audio datasets is a time-consuming and expensive process that limits the dataset's size. Moreover, the diversity of sound categories and class imbalances can further impede classification performance. To overcome these challenges, researchers have proposed various audio data augmentation methods. However, most of these methods focus less on augmentations combination and design and rely solely on waveform-based or spectrogram-based approaches. This paper presents an Automated Audio Augmentation (AAA) method for audio classification, which generates learnable and composable augmentation policies suitable for the audio classification task and can be employed in a plug-and-play manner. This method leverages both waveform-level and spectrogram-level augmentation, and a Bayesian optimization algorithm is proposed to search for composed augmentation policies. To the best of our knowledge, this is the first attempt to propose an automatic data augmentation method for audio classification tasks. Through large-scale empirical studies, we demonstrate that the proposed method outperforms previous competitive methods by a significant margin. We improve the average performance of multiple datasets by 6.421% and by 7.330% on few-shot scenarios, respectively. Yanjie Sun, Kele Xu, Chaorun Liu, Yong Dou, Huaimin Wang 0001, Bo Ding 0001, Qinghua Pan |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | MANet: An Architecture Adaptive Method for Sparse Matrix Format Selection
Zhenglun Sun, Peng Qiao, Yong Dou |
ICA3PP (2) | 3 |
| 2023 | Rumor Detection Via Assessing the Spreading Propensity of UsersabstractThe explosion of rumors on social media has adversely affected cyber security and our lives, increasing the urgent demand for rumor detection. Existing detection methods focus on exploring signals of deception from textual contents and rumor propagation structures, without fully considering user’s long-term spreading propensity. Sociology and psychology have demonstrated that when a rumor satisfies a user’s inner requirements, he/she has more propensity to spread it. User context, such as historical posts, offers extensive details about propensities for spreading rumors, which has great potential to promote rumor detection. Therefore, we explore a new feature space by extracting the spreading propensity from user context, and combine it with social interaction information to construct a creative detection algorithm. Experiments on three Twitter datasets show that our approach achieves significant improvements compared to strong baselines and displays a superior capacity for detecting rumors at early stages. Our code is publicly released at https://github.com/Coder-HenryZa/RDSPU. Peng Zheng 0003, Zhen Huang 0006, Yong Dou, Yeqing Yan |
ICASSP | 3 |
| 2023 | SNCSE: Contrastive Learning for Unsupervised Sentence Embedding with Soft Negative Samples
Yong Dou |
ICIC (4) | 2 |
| 2023 | Spatial and Frequency Domains Inconsistency Learning for Face Forgery Detection
Caili Gao, Peng Qiao, Yong Dou, Qisheng Xu, Xifu Qian |
ICONIP (12) | 3 |
| 2023 | HAAN: Human Action Aware Network for Multi-label Temporal Action DetectionabstractThe task of multi-label temporal action detection aims to accurately detect dense action instances in untrimmed videos. Previous methods focused on modeling the appearance features of RGB images have struggled to capture the fine details and subtle variations in human actions, resulting in three critical issues: overlapping action confusion, intra-class appearance diversity, and background interferences. These issues have significantly undermined the accuracy and generalization of detection models. To tackle these issues, we propose incorporating the human skeleton into the feature design of the detection model. By utilizing multi-person skeletons, our proposed method can accurately represent various human actions in the scene, balance the salience of overlapping actions, and reduce the impact of changes in human appearance and background interferences on action features. Overall, we propose a novel two-stream human action aware network~(HAAN) for multi-label temporal action detection based on the original RGB frames and the estimated skeleton frames. To leverage the complementary advantages of RGB features and skeleton features, we design a cross-modality fusion module that allows the two features to guide each other and enhance their representation of human actions. On the popular benchmarks MultiTHUMOS and Charades, our HAAN achieves state-of-the-art performance with 56.9% (+5.4%) and 32.1% (+3.3%) mean average precision (mAP) compared to the best available methods. Importantly, HAAN shows superior improvements of +6.83%, +22.35%, and +2.56% on the challenging sample subsets of the three critical issues. Zikai Gao, Peng Qiao, Yong Dou |
ACM Multimedia | 3 |
| 2023 | Automatic Audio Augmentation for Requests Sub-ChallengeabstractThis paper presents our solution for the Requests Sub-challenge of the ACM Multimedia 2023 Computational Paralinguistics Challenge. Drawing upon the framework of self-supervised learning, we put forth an automated data augmentation technique for audio classification, accompanied by a multi-channel fusion strategy aimed at enhancing overall performance. Specifically, to tackle the issue of imbalanced classes in complaint classification, we propose an audio data augmentation method that generates appropriate augmentation strategies for the challenge dataset. Furthermore, recognizing the distinctive characteristics of the dual-channel HC-C dataset, we individually evaluate the classification performance of the left channel, right channel, channel difference, and channel sum, subsequently selecting the optimal integration approach. Our approach yields a significant improvement in performance when compared to the competitive baselines, particularly in the context of the complaint task. Moreover, our method demonstrates noteworthy cross-task transferability. Yanjie Sun, Kele Xu, Chaorun Liu, Yong Dou, Kun Qian 0003 |
ACM Multimedia | 4 |
| 2023 | Incorporating Structured Sentences with Time-enhanced BERT for Fully-inductive Temporal Relation PredictionabstractTemporal relation prediction in incomplete temporal knowledge graphs (TKGs) is a popular temporal knowledge graph completion (TKGC) problem in both transductive and inductive settings. Traditional embedding-based TKGC models (TKGE) rely on structured connections and can only handle a fixed set of entities, i.e., the transductive setting. In the inductive setting where test TKGs contain emerging entities, the latest methods are based on symbolic rules or pre-trained language models (PLMs). However, they suffer from being inflexible and not time-specific, respectively. In this work, we extend the fully-inductive setting, where entities in the training and test sets are totally disjoint, into TKGs and take a further step towards a more flexible and time-sensitive temporal relation prediction approach SST-BERT,incorporating Structured Sentences with Time-enhanced BERT. Our model can obtain the entity history and implicitly learn rules in the semantic space by encoding structured sentences, solving the problem of inflexibility. We propose to use a time masking MLM task to pre-train BERT in a corpus rich in temporal tokens specially generated for TKGs, enhancing the time sensitivity of SST-BERT. To compute the probability of occurrence of a target quadruple, we aggregate all its structured sentences from both temporal and semantic perspectives into a score. Experiments on the transductive datasets and newly generated fully-inductive benchmarks show that SST-BERT successfully improves over state-of-the-art baselines. Zhongwu Chen, Chengjin Xu, Fenglong Su, Zhen Huang 0006, Yong Dou |
SIGIR | 5 |
| 2023 | Meta-Learning Based Knowledge Extrapolation for Temporal Knowledge GraphabstractIn the last few years, the solution to Knowledge Graph (KG) completion via learning embeddings of entities and relations has attracted a surge of interest. Temporal KGs(TKGs) extend traditional Knowledge Graphs (KGs) by associating static triples with timestamps forming quadruples. Different from KGs and TKGs in the transductive setting, constantly emerging entities and relations in incomplete TKGs create demand to predict missing facts with unseen components, which is the extrapolation setting. Traditional temporal knowledge graph embedding (TKGE) methods are limited in the extrapolation setting since they are trained within a fixed set of components. In this paper, we propose a Meta-Learning based Temporal Knowledge Graph Extrapolation (MTKGE) model, which is trained on link prediction tasks sampled from the existing TKGs and tested in the emerging TKGs with unseen entities and relations. Specifically, we meta-train a GNN framework that captures relative position patterns and temporal sequence patterns between relations. The learned embeddings of patterns can be transferred to embed unseen components. Experimental results on two different TKG extrapolation datasets show that MTKGE consistently outperforms both the existing state-of-the-art models for knowledge graph extrapolation and specifically adapted KGE and TKGE baselines. Zhongwu Chen, Chengjin Xu, Fenglong Su, Zhen Huang 0006, Yong Dou |
WWW | 5 |
| 2023 | Rumor detection on social media through mining the social circles with high homogeneity
Peng Zheng 0003, Zhen Huang 0006, Yong Dou, Yeqing Yan |
Inf. Sci. | 3 |
| 2023 | Isolate Sets Based Parallel Louvain Method for Community Detection
Hang Qie, Yong Dou, Zhen Huang 0006, Yunsheng Xiong |
J. Comput. Sci. Technol. | 2 |
| 2023 | Recent Trends in Deep Learning Based Textual Emotion Cause ExtractionabstractEmotion Cause Extraction Field (ECEF) focuses on the cause that triggers an emotion in a document and mainly includes Emotion Cause Extraction (ECE) and Emotion Cause Pair Extraction (ECPE). Traditional ECE aims to extract the cause based on a given emotion while ECPE aims to extract both the emotion and its corresponding cause. Recently, ECEF has attracted a lot of attention and most of the advances have benefited from significant developments in deep learning techniques, especially machine reading comprehension and neural-network-based information retrieval. The large pre-trained language model of BERT has also shown effectiveness in this field. Following the proposal of ECPE, the development of ECEF has accelerated. However, a comprehensive review of existing approaches and recent trends in the field is lacking. To address this issue, this survey presents a thorough review to summarise existing methods and recent key advances, illustrate the general technical architecture of traditional ECE, introduce several important variants, in particular ECPE, and provide a detailed comparison of several public datasets. Finally, the limitations of existing work and the prospects for further technological advances in ECEF are discussed. Xinxin Su, Zhen Huang 0006, Yong Dou, Hengyue Pan |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | ArtVerse: A Paradigm for Parallel Human-Machine Collaborative Painting Creation in MetaversesabstractCurrently, the development of the foundation model, metaverse, nonfungible token (NFT), and other emerging technologies has brought profound effects on the whole art field, including art creation, dissemination, transaction, etc. However, there is no research focusing on the framework, methodologies, and applications of the human–machine collaborative creation in the metaverse era. Based on parallel theory, this article proposes a novel human–machine collaborative creation paradigm called ArtVerse, in which machines take on the roles of humans to perform creation exploration and evolution and build decentralized art organizations. Besides, the operational processes involving several key technologies are designed to achieve the proposed ArtVerse. Then, a prototype system of ArtVerse, our long-term efforts toward the human–machine collaborative painting, is presented. Finally, a new ecology of artistic creation in the metaverse era is demonstrated through the applications of the ArtVerse. Chao Guo 0006, Yong Dou, Tianxiang Bai, Xingyuan Dai, Chunfa Wang |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2022 | Temporal Dynamic Weighted Graph Convolution for Multi-agent Reinforcement Learning
Yuntao Liu 0004, Yong Dou, Yuan Li 0011, Xinhai Xu, Donghong Liu |
CogSci | 2 |
| 2022 | Adaptive Threshold Selective Self-Attention for Chinese NERabstractRecently, Transformer has achieved great success in Chinese named entity recognition (NER) owing to its good parallelism and ability to model long-range dependencies, which utilizes self-attention to encode context. However, the fully connected way of self-attention may scatter the attention distribution and allow some irrelevant character information to be integrated, leading to entity boundaries being misidentified. In this paper, we propose a data-driven Adaptive Threshold Selective Self-Attention (ATSSA) mechanism that aims to dynamically select the most relevant characters to enhance the Transformer architecture for Chinese NER. In ATSSA, the attention score threshold of each query is automatically generated, and characters with attention score higher than the threshold are selected by the query while others are discarded, so as to address irrelevant attention integration. Experiments on four benchmark Chinese NER datasets show that the proposed ATSSA brings 1.68 average F1 score improvements to the baseline model and achieves state-of-the-art performance. Zhen Huang 0006, Minghao Hu 0001, Yong Dou |
COLING | 5 |
| 2022 | IMCI: Integrate Multi-view Contextual Information for Fact Extraction and VerificationabstractWith the rapid development of automatic fake news detection technology, fact extraction and verification (FEVER) has been attracting more attention. The task aims to extract the most related fact evidences from millions of open-domain Wikipedia documents and then verify the credibility of corresponding claims. Although several strong models have been proposed for the task and they have made great process, we argue that they fail to utilize multi-view contextual information and thus cannot obtain better performance. In this paper, we propose to integrate multi-view contextual information (IMCI) for fact extraction and verification. For each evidence sentence, we define two kinds of context, i.e. intra-document context and inter-document context. Intra-document context consists of the document title and all the other sentences from the same document. Inter-document context consists of all other evidences which may come from different documents. Then we integrate the multi-view contextual information to encode the evidence sentences to handle the task. Our experimental results on FEVER 1.0 shared task show that our IMCI framework makes great progress on both fact extraction and verification, and achieves state-of-the-art performance with a winning FEVER score of 73.96% and label accuracy of 77.25% on the online blind test set. We also conduct ablation study to detect the impact of multi-view contextual information. Yangguang Li 0001, Zhen Huang 0006, Yong Dou |
COLING | 4 |
| 2022 | RSGT: Relational Structure Guided Temporal Relation ExtractionabstractTemporal relation extraction aims to extract temporal relations between event pairs, which is crucial for natural language understanding. Few efforts have been devoted to capturing the global features. In this paper, we propose RSGT: Relational Structure Guided Temporal Relation Extraction to extract the relational structure features that can fit for both inter-sentence and intra-sentence relations. Specifically, we construct a syntactic-and-semantic-based graph to extract relational structures. Then we present a graph neural network based model to learn the representation of this graph. After that, an auxiliary temporal neighbor prediction task is used to fine-tune the encoder to get more comprehensive node representations. Finally, we apply a conflict detection and correction algorithm to adjust the wrongly predicted labels. Experiments on two well-known datasets, MATRES and TB-Dense, demonstrate the superiority of our method (2.3% F1 improvement on MATRES, 3.5% F1 improvement on TB-Dense). Jie Zhou 0032, Shenpo Dong, Hongkui Tu, Xiaodong Wang 0002, Yong Dou |
COLING | 5 |
| 2022 | Improving the Classification of Phonetic Segments from Raw Ultrasound Using Self-Supervised Learning and Hard Example MiningabstractUltrasound tongue imaging is an attractive way for speech production study as it provides an effective visualization for the vocal tract. Automatic classification of phonetic segments (tongue shapes) from raw ultrasound data is vital for further interpretation. Recently, deep learning-based approaches have been adopted in this task, which required a large-scale annotated dataset for the training, and it is not easy to be obtained in practical settings. Moreover, the data may contain many hard examples for the classification task, due to contamination of speckle noise. In this paper, we aim to address these issues: firstly, self-supervised learning is adopted to utilize the unlabeled datasets and extract the features without any human annotations; secondly, hard example mining is applied to imitate the learning path of the clinical linguists. To empirically demonstrate the proposed method’s effectiveness, we evaluate the method on the Ultrax Typically Developing dataset (UXTD) under different scenarios. The results show that the proposed method outperforms the other methods and achieves superior performance. To better promote the study in this field, we release our code publicly at1. Yunsheng Xiong, Kele Xu, Yong Dou, Jinjia Wang |
ICASSP | 5 |
| 2022 | ROGC: Role-Oriented Graph Convolution Based Multi-Agent Reinforcement LearningabstractThe role-oriented learning approach could improve the performance of multi-agent reinforcement learning by decomposing complex multi-agent tasks into different roles. However, due to the dynamic environment and interactions among agents, the role undertaken by an agent changes rapidly with time going on. Therefore, the roles of agents should be adapted to the varying situation during the learning process. In this paper, we propose a role-oriented graph convolution based multi-agent reinforcement learning framework (ROGC). Firstly, we design a role assigner based on samples generated from the environment to learn roles for classifying agents into different groups. To further enhance cooperation among agents in the same group for higher performance, we design a graph convolutional module to achieve intra-role communications based on discovered roles. With roles and extracted role features, we design a role-oriented policy learning module that embeds the role information into the algorithm and generates effective policies for individuals. Further, we introduce an auto-encoder to learn the intra-role cooperation knowledge in the graph convolutional module, which ensures our framework executes in a decentralized way. Extensive experiments show that our framework can learn dynamic roles and make full use of learned roles, which makes it outperform popular MARL methods. Yuntao Liu 0004, Yuan Li 0011, Xinhai Xu, Donghong Liu, Yong Dou |
ICME | 5 |
| 2022 | Searching Latent Sub-Goals in Hierarchical Reinforcement Learning as Riemannian Manifold OptimizationabstractHierarchical Reinforcement Learning (HRL) is promising to tackle the long-term sparse reward problem. However, goal conditioned HRL, which decomposes the goal into a series of sub-goals, suffers from sub-goal search inefficiency problems when the observation space is too large. This problem is more severe in a visual observation space, since its high latent dimensions, where the complete dynamics information is preserved, exponentially increase the difficulty of sub-goal search. In view of this, we propose to treat the latent space as a manifold, i.e., a Riemannian manifold. Assisted by the Riemannian manifold optimization, sub-goals can be efficiently searched in the higher-dimensional latent space, with the help of preserving the dynamics information efficiently. Experiments on a series of MuJoCo tasks with visual observation show that the proposed Riemannian manifold optimization, compared with the baseline that directly searches for sub-goals in bounded latent space, improves the success rate by 1.5 times on average. In much higher dimensions where the baseline no longer converges, the success rate of the proposed method is maintained. Sidun Liu, Peng Qiao, Yong Dou, Ruochun Jin |
ICME | 3 |
| 2022 | MLPs: Efficient Training of MiniGo on Large-scale Heterogeneous Computing SystemabstractDeep Reinforcement Learning has been successfully applied in various applications and achieved impressive performance compared with previous traditional methods but suffers from high computation cost and long training time. MLPerf takes deep reinforcement learning as one of the benchmark tracks and provides a single node training version of MiniGo as a reference. A key challenge is to achieve efficient MiniGo training on a large-scale computing system. According to the training computation pattern in MiniGo and the characteristics of our large-scale heterogeneous computing system, we propose a MultiLevel Parallel strategy, MLPs, including task-level parallelism between nodes, CPU-DSP heterogeneous parallelism, and DSP multi-core parallelism. The proposed method reduces the overall execution time from 43 hours to 16 hours while scaling the node size from 1067 to 4139. The scaling efficiency is 69.1%. According to our fitting method, the scaling efficiency is 46.5% when scaling to 8235 nodes. The experimental results show that the proposed method achieves the efficient training of MiniGo on the largescale heterogeneous computing system. Peng Qiao, Zhouyu He, Rongchun Li, Jingfei Jiang, Yong Dou, Dongsheng Li 0001 |
ICPADS | 5 |
| 2022 | CNA: A Dataset for Parsing Discourse Structure on Chinese News ArticlesabstractDiscourse structure analysis has shown to be useful for many artificial intelligence (AI) tasks such as text sum-marization and text categorization. However, for the Chinese news domain, the discourse structure analysis system is still immature due to the limitation of the lack of expert-annotated datasets. In this paper, we present CNA, a Chinese news corpus containing 1155 news articles annotated by human experts, which covers four domains and four news media sources. Next, we implement several text classification methods as baselines. Experimental results demonstrate that document-level method can achieve a better performance, and we further propose a document-level neural network model with multiple sentence features which achieves the state-of-the-art performance. In the end, we analyze the content type distribution of each sentence in CNA and the prediction errors of our model that occurred on the test set. The codes and dataset will be open-sourced at https://github.com/gzl98/Chinese_Discourse_Profiling. Zhenliang Guo, Zhen Huang 0006, Yong Dou, Xiubin Yu, Zhongwu Chen, Xinxin Su |
ICTAI | 3 |
| 2022 | Discourse Component Recognition via Graph Neural Network in Chinese Student Argumentative Essays
Yong Dou, Zhen Huang 0006 |
KSEM (1) | 3 |
| 2022 | Topic and Reference Guided Keyphrase Generation from Social Media
Xiubin Yu, Xingjun Chen, Zhen Huang 0006, Yong Dou |
KSEM (2) | 4 |
| 2022 | Heterogeneous Skill Learning for Multi-agent TasksabstractHeterogeneous behaviours are widespread in many multi-agent tasks, which have not been paid much attention in the community of multi-agent reinforcement learning. It would be a key factor for improving the learning performance to efficiently characterize and automatically find heterogeneous behaviours. In this paper, we introduce the concept of the skill to explore the ability of heterogeneous behaviours. We propose a novel skill-based multi-agent reinforcement learning framework to enable agents to master diverse skills. Specifically, our framework consists of the skill representation mechanism, the skill selector and the skill-based policy learning mechanism. We design an auto-encoder model to generate the latent variable as the skill representation by incorporating the environment information, which ensures the distinguishable of agents for skill selection and the discriminability for the skill learning. With the representation, a skill selection mechanism is invented to realize the assignment from agents to skills. Meanwhile, diverse skill-based policies are generated through a novel skill-based policy learning method. To promote efficient skill discovery, a mutual information based intrinsic reward function is constructed. Empirical results show that our framework obtains the best performance on three challenging benchmarks, i.e., StarCraft II micromanagement tasks, Google Research Football and GoBigger, over state-of-the-art MARL methods. Yuntao Liu 0004, Yuan Li 0011, Xinhai Xu, Yong Dou, Donghong Liu |
NeurIPS | 4 |
| 2022 | A pipelining strategy for accelerating convolution neural networks on ARM CPUsabstractAbstract Convolution is a primary operation in convolution neural networks. The speed of inference is mainly decided by the speed of the convolutional layer. Improving the performance of embedded processors makes it possible to process the inference on embedded devices. In this article, a pipelining strategy of single instruction and multiple data (SIMD) instructions is proposed to finely optimize the process of the 3 × 3 convolution on ARM‐based CPUs. We implement the SIMD group to improve the efficiency of the SIMD pipeline. A tiling method is exploited to increase data reuse during the process. An evaluation model is proposed to guide the design of the tiling method and register allocation. The speed of our implementation is 5.18 times of the GNU compiler collection compiled unoptimized version on RK3288. The effect of our optimizing method is measured by a performance profiling tool, the performance information suggests that the pipelining strategy has a significant effect for both normal and depthwise separable convolution. By implementing multithread processing, the speedup achieves 18.3 compared with the single thread unoptimized version. Yong Dou, Rongchun Li, Peng Zhang 0035, Yuntao Liu 0004 |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | An automatic learning rate decay strategy for stochastic gradient descent optimization methods in neural networksabstractStochastic Gradient Descent (SGD) series optimization methods play the vital role in training neural networks, attracting growing attention in science and engineering fields of the intelligent system. The choice of learning rates affects the convergence rate of SGD series optimization methods. Currently, learning rate adjustment strategies mainly face the following problems: (1) The traditional learning rate decay method mainly adopts manual manner during training iterations, the small learning rate produced from which causes slow convergence in training neural networks. (2) Adaptive method (e.g., Adam) has poor generalization performance. To alleviate the above issues, we propose a novel automatic learning rate decay strategy for SGD optimization methods in neural networks. On the basis of the observation that the convergence rate's upper bound enjoys minimization in a specific iteration concerning the current learning rate, we first present the expression of the current learning rate determined by historical learning rates. And merely one extra parameter is initialized to generate automatic decreasing learning rates during the training process. Our proposed approach is applied to SGD and Momentum SGD optimization algorithms, and concrete theoretical proof explains its convergence. Numerical simulations are conducted on the MNIST and Cifar-10 data sets with different neural networks. Experimental results show that our algorithm outperforms existing classical ones, achieving faster convergence rate, better stability, and generalization performance in neural network training. It also lays a foundation for large-scale parallel search of initial parameters in intelligent systems. Yong Dou, Tao Sun 0005, Peng Qiao, Dong Wen 0004 |
Int. J. Intell. Syst. | 2 |
| 2022 | Focus on Hard Categories and Hard Examples: Remote Sensing Image Scene Classification via Expert Model and Hard Example MiningabstractDeep learning has seen dramatic improvements in remote-sensing image scene classification. However, hard categories and hard examples widely exist in the data sets, due to the intraclass diversity and interclass similarity. In this letter, we propose a novel framework to address these issues. Specifically, our method first trains a general model to obtain the confusion matrix and select the hard categories. Then a sampling strategy is proposed to restructure the training set and an expert model is trained to focus on the hard categories. Finally, the knowledge of the expert model is distilled into the student model through a novel loss function, which encourages the student model to predict the hard label provided by manual annotation. Thus, the model can match the soft label provided by the expert model and pay more attention to the hard examples simultaneously. With this method, the student model cannot only deal with hard categories but also hard examples existing in other easy categories. To empirically demonstrate the effectiveness of the proposed method, we comprehensively evaluate the method on three publicly available benchmark data sets, the obtained results show that the proposed method outperforms the existing baseline methods and achieves superior results on all three data sets. Yunsheng Xiong, Peng Zhang 0035, Yong Dou, Kele Xu, Xin Niu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | An Adaptive Learning Rate Schedule for SIGNSGD Optimizer in Neural Networks
Tao Sun 0005, Yong Dou |
Neural Process. Lett. | 3 |
| 2022 | WRMatch: Improving FixMatch With Weighted Nuclear-Norm Regularization for Few-Shot Remote Sensing Scene ClassificationabstractSemisupervised learning (SSL), such as FixMatch, has been successfully applied to remote sensing scene classification to relieve the burden of data annotation. However, in some extreme settings, only very few samples available, e.g., one to ten labels per remote sensing scene, can be used. When meeting this “few-shot” scenario, the deep model may be overfitting and prone to generate confusing predictions due to the lack of labels and strong augmentation-based perturbations. Thus, the prediction’s diversity may collapse, and the discriminability exceeds the reasonable interval. How to improve the performance of few-shot learning is underexplored for the remote sensing scene classification in previous studies. In this article, we present a novel framework for the task by utilizing the improved FixMatch and the weighted nuclear-norm regularization (WNNR). Specifically, we regularize the prediction matrix by exploiting the nuclear-norm, which is an approximation of the matrix rank and a relaxed boundary for the Shannon entropy. We further provide two weighting schemes to improve the nuclear-norm-based regularization. First, the random-weighting scheme for nuclear-norm (RWNNR) is proposed based on the Dirichlet distribution to improve the model’s generalization. Second, we present the self-weighting scheme (SWNNR) to weight the singular values according to singular values themselves and adjust the relaxed degree for the boundary between the nuclear-norm and the Shannon entropy. Maximizing the weighted nuclear-norm can improve the prediction diversity and optimize the prediction discriminability simultaneously. Combining the advantages of SSL and the aforementioned improvements, we can reliably classify the remote sensing scene image with very limited annotated datasets. To empirically demonstrate the proposed method’s effectiveness, we comprehensively evaluate the method on three publicly available benchmark datasets. The results show that the proposed method outperforms the baseline methods by a large margin and achieves superior performance on all three datasets. Our method can be an effective alternative to metalearning in few-shot scene classification, with the advantage lying in the competitive performance and the absence of metatraining stage associated with a large number of labels. Yunsheng Xiong, Kele Xu, Yong Dou, Yang Zhao 0003, Zikai Gao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Hierarchical learning with backtracking algorithm based on the Visual Confusion Label Tree for large-scale image classification
Yuntao Liu 0004, Yong Dou, Ruochun Jin, Rongchun Li, Peng Qiao |
Vis. Comput. | 2 |
| 2021 | RFC-HyPGCN: A Runtime Sparse Feature Compress Accelerator for Skeleton-Based GCNs Action Recognition Model with Hybrid PruningabstractSkeleton-based Graph Convolutional Networks (GCNs) models for action recognition have achieved excellent prediction accuracy in the field. However, limited by large model and computation complexity, GCNs for action recognition like 2s-AGCN have insufficient power-efficiency and throughput on GPU. Thus, the demand of model reduction and hardware acceleration for low-power GCNs action recognition application becomes continuously higher.To address challenges above, this paper proposes a runtime sparse feature compress accelerator with hybrid pruning method: RFC-HyPGCN. First, this method skips both graph and spatial convolution workloads by reorganizing the multiplication order. Following spatial convolutions channel-pruning dataflow, a coarse-grained pruning method on temporal filters is designed, together with sampling-like fine-grained pruning on time dimension. Later, we come up with an architecture where all convolutional layers are mapped on chip to pursue high throughput. To further reduce storage resource utilization, online sparse feature compress format is put forward. Features are divided and encoded into several banks according to presented format, then bank storage is split into depth-variable mini-banks. Furthermore, this work applies quantization, input-skipping and intra-PE dynamic data scheduling to accelerate the model. In experiments, proposed pruning method is conducted on 2s-AGCN, acquiring 3.0x-8.4x model compression ratio and 73.20% graph-skipping efficiency with balancing weight pruning. Implemented on Xilinx XCKU-115 FPGA, the proposed architecture has the peak performance of 1142 GOP/s and achieves up to 9.19x and 3.91x speedup over high-end GPU NVIDIA 2080Ti and NVIDIA V100, respectively. Compared with latest accelerator for action recognition GCNs models, our design reaches 22.9x speedup and 28.93% improvement on DSP efficiency. Dong Wen 0004, Jingfei Jiang, Jinwei Xu, Yang Zhao 0003, Yong Dou |
ASAP | 7 |
| 2021 | LADstackING: Stacking Ensemble Learning-based Computational Model for Predicting Potential LncRNA-disease AssociationsabstractIn recent years, accumulation of researches have proved many diseases that seriously endanger human health originate from mutations or dysfunctions in LncRNA (Long non-encoding RNA). Therefore, it is important to discover the intrinsic associations between the LncRNAs and the diseases. Meantime, accurately identifying potential associations between diseases and LncRNAs remains a highly challenging task. In this paper, we proposed a model based on the stacking ensemble learning framework called LADstackING to predict the potential LncRNA associated disease. LADstackING effectively integrates different types of strong predictive performance models rather than the same type models with weak predictive performance.LADstackING is able to exploit the respective advantages of different base models in its framework and significantly improve the overall predictive performance. The multi-perspective features bring by different base models allow LADstackING remain stable in facing of sparse data sources. Moreover, the overall predictive performance of LADstackING is greatly improved compare to the stat-of-art models. Experimental results and case study result demonstrate that LADstackING performs promising in predicting the potential LncRNA-disease associations. Jiechen Li, Xiangxiang Zeng, Yong Dou, Fei Xia 0003, Shaoliang Peng |
BIBM | 3 |
| 2021 | Local and Non-local Context Graph Convolutional Networks for Skeleton-Based Action Recognition
Zikai Gao, Yang Zhao 0003, Yong Dou |
ICANN (3) | 5 |
| 2021 | Global-Localized Agent Graph Convolution for Multi-Agent Reinforcement LearningabstractA lot of efforts have been devoted to solving the problem about complex relationship and localized cooperation among a large number of agents in large-scale multi-agent systems. However, global cooperation among all agents is also important while interactions between agents often happen locally. It is a challenging problem to enable agent to learn global and localized cooperate information simultaneously in multi-agent systems. In this paper, we model the global and localized cooperation among agents by global and localized agent graphs and propose a novel graph convolutional reinforcement learning mechanism based on these two graphs which allows each agent to communicate with neighbors and all a-gents to cooperate at the high level. Experiments on the large-scale multi-agent scenarios in StarCraft II show that our pro-posed method gets better performance compared with state-of-the-art algorithms and allows agents learning to cooperate efficiently. Yuntao Liu 0004, Yong Dou, Peng Qiao |
ICASSP | 2 |
| 2021 | Graphcomm: A Graph Neural Network Based Method for Multi-Agent Reinforcement LearningabstractThe communication among agents is important for Multi-Agent Reinforcement Learning (MARL). In this work, we propose GraphComm, a method makes use of the relation-ships among agents for MARL communication. GraphComm takes the explicit relations (e.g., agent types), which can be provided through some knowledge background, into account to better model the relationships among agents. Besides explicit relations, GraphComm considers implicit relations, which are formed by agent interactions. GraphComm use Graph Neural Networks (GNNs) to model the relational information, and use GNNs to assist the learning of agent communication. We show that GraphComm can obtain better results than state-of-the-art methods on the challenging StarCraft II unit micromanagement tasks through extensive experimental evaluation. Yongquan Fu, Huayou Su, Hengyue Pan, Peng Qiao, Yong Dou, Cheng Wang 0003 |
ICASSP | 6 |
| 2021 | Ddper: Decentralized Distributed Prioritized Experience ReplayabstractIn off-policy reinforcement learning, prioritized experience replay plays an important role. However, the centralized prioritized experience replay becomes the bottleneck for efficient training. We propose to approximate the centralized prioritized experience replay in a distributed and decentralized way under certain mild assumptions. To be specific, each actor stores samples in its local replay in the same way as prioritized experience replay, the learner fetches a batch of samples from these replays following a certain strategy. We implement a Deep Q-Learning off-policy algorithm upon the proposed framework. The comparison experiments are performed on a commonly used subset of the Atari-57 learning environment. The experimental results show that the proposed framework speeds up training as the number of actors increases. With the same algorithm and hyper-parameter settings, the proposed framework with 16 actors achieves superior performance that Ape-X with 32 and even more actors does. Sidun Liu, Peng Qiao, Yong Dou, Rongchun Li |
ICME | 3 |
| 2021 | A Framework of Data Augmentation While Active Learning for Chinese Named Entity Recognition
Zhen Huang 0006, Yong Dou |
KSEM | 3 |
| 2021 | COVID Edge-Net: Automated COVID-19 Lung Lesion Edge Detection in Chest CT Images
Yang Zhao 0003, Yong Dou, Dong Wen 0004, Zikai Gao |
ECML/PKDD (4) | 3 |
| 2021 | Beyond AP: a new evaluation index for multiclass classification task accuracy
Kaifang Zhang, Huayou Su, Yong Dou |
Appl. Intell. | 3 |
| 2021 | A high-throughput scalable BNN accelerator with fully pipelined architecture
Jingfei Jiang, Jinwei Xu, Peng Zhang 0035, Dong Wen 0004, Yong Dou |
CCF Trans. High Perform. Comput. | 7 |
| 2021 | An energy-efficient convolutional neural network accelerator for speech classification based on FPGA and quantization
Dong Wen 0004, Jingfei Jiang, Yong Dou, Jinwei Xu |
CCF Trans. High Perform. Comput. | 3 |
| 2021 | Representation learning on textual network with personalized PageRank
Teng Li 0010, Yong Dou |
Sci. China Inf. Sci. | 2 |
| 2021 | Multilevel parallelism optimization of stencil computations on SIMDlized NUMA architectures
Kaifang Zhang, Huayou Su, Yong Dou |
J. Supercomput. | 3 |
| 2020 | Argumentation Mining on Essays at Multi ScalesabstractArgumentation mining on essays is a new challenging task in natural language processing, which aims to identify the types and locations of argumentation components. Recent research mainly models the task as a sequence tagging problem and deal with all the argumentation components at word level. However, this task is not scale-independent. Some types of argumentation components which serve as core opinions on essays or paragraphs, are at essay level or paragraph level. Sequence tagging method conducts reasoning by local context words, and fails to effectively mine these components. To this end, we propose a multi-scale argumentation mining model, where we respectively mine different types of argumentation components at corresponding levels. Besides, an effective coarse-to-fine argumentation fusion mechanism is proposed to further improve the performance. We conduct a serial of experiments on the Persuasive Essay dataset (PE2.0). Experimental results indicate that our model outperforms existing models on mining all types of argumentation components. Zhen Huang 0006, Yong Dou |
COLING | 3 |
| 2020 | End-to-end Spatial Attention Network with Feature Mimicking for Head DetectionabstractHuman head detection is a widely used task and suitable for identifying persons in practical applications. Although existing methods have achieved significant progress, the problems of false alarm and miss detection are still challenging, which arise from weak classification power of detector in the face of variability in occlusion, illumination, etc. In this paper, we present an effective end-to-end head detector called Spatial Attention Network with feature Mimicking(SANM) that can obtain better feature and enhanced classification power, through attention mechanism and a feature mimic method. The spatial-wise attention is extracted from several levels of feature and supervised by the bounding-box annotated heat map. The attention improves the quality of the features in the head and opposite area. To further improve the classification ability, we utilize the feature mimicking method to drive network learning the feature refined by a deep cascading classifier. Compared with the baseline model, our method achieves better performance and produces leading results on head detection benchmarks. Yuntao Liu 0004, Rongchun Li, Yong Dou |
FG | 4 |
| 2020 | Learning Network Representation Through Reinforcement LearningabstractNetwork Representation Learning embeds each node in a network into a low-dimensional real-value vector which can be used for downstream tasks such as link prediction and recommendation. Many existing approaches use unsupervised or (semi-)supervised methods to explore the network topology and learn representations from it. In contrast, we propose, reinforcement learning network representations (RLNet), which explores the idea of using reinforcement learning to learn to explore the network and to obtain network representations. Based on reward signals, RLNet learns an actor which uses a policy to determine the network navigation actions. RLNet uses node representations to parameterize its policy, and the representations are learned together with the policy. Through experiments based on multiple datasets, we show that RLNet can obtain promising results in link prediction tasks. Yongquan Fu, Adele Lu Jia, Huayou Su, Chengsong Wang, Yong Dou |
ICASSP | 7 |
| 2020 | Attentional Fused Temporal Transformation Network for Video Action RecognitionabstractEffective spatiotemporal feature representation is crucial to the video-based action recognition task. Focusing on discriminate spatiotemporal feature learning, we propose Attentional Fused Temporal Transformation Network (AttnTTN) for action recognition on top of popular Temporal Segment Network (TSN) framework. In the network, Attentional Fusion Module (AttnFM) is designed to fuse the appearance and motion features at multiple ConvNet levels for each video snippet, forming a short-term video descriptor. With fused features as inputs, Temporal Transformation Networks (TTN) are employed to model middle-term temporal transformation between the neighboring temporal snippets following a sequential order. AttnTTN achieves the state-of-the-art results on two most popular action recognition datasets: UCF101 and HMDB51. Ke Yang 0004, Huadong Dai, Tianlong Shen, Peng Qiao, Xin Niu 0002, Dongsheng Li 0001, Yong Dou |
ICASSP | 9 |
| 2020 | Towards Precise End-to-end Semi-Supervised Human Head Detection NetworkabstractHead detection, as a fundamental task in practice for many head-related problems, requires an enormous number of annotated boxes to maintain the performance. To alleviate the time and cost of labeling each image in the dataset, we propose an end-to-end semi-supervised head detection frame-work, which shows competitive results with only a small set of data. Specifically, under the setting of semi-supervised, we introduce a weak boxes generate branch and a weak boxes refine branch to produce pseudo ground truth label for unlabeled images with the guidance of annotated images. The weak boxes generate branch is embedded in the detection framework taking the proposals as input and outputting the initial weak boxes that coarsely locate the place of the head. Then, the weak boxes refine branch adjusts the weak boxes more accurate gradually by training a transferred sub-network with the established relation between proposals, weak boxes and labeled boxes. In the training process, we jointly train the two branches in an end-to-end manner, which can generate better pseudo bounding boxes with a small dataset online to avoid over-fitting and obtain a more precise head detector. The results on the public head detection benchmark Brainwash and SCUT-HEAD show the effectiveness of our method. Rongchun Li, Yuntao Liu 0004, Yong Dou |
IJCNN | 4 |
| 2020 | A High-Throughput LDPC Decoder Based on GPUs for 5G New RadioabstractIn this paper, we propose a GPU-based QC-LDPC decoder for 5G New Radio(NR). Different from existing LDPC decoders based on GPUs, our decoder achieves high throughput when decoding LDPC codes with high code rates. Moreover, we implement the shortening and puncturing techniques which are exploited by 5G NR. The decoding algorithm Min-Sum approximation algorithm(MSA) is optimized to implement efficient parallel decoding on the GPU. In order to save the on-chip and the off-chip bandwidth, we propose the two-level quantization scheme and implement data packing on the GPU. We also analyse the optimum thread assignment for different code rates based on our implementation. By using the optimum settings on the GPU, the decoding throughput achieves 1.38 Gbps in the case of (2080, 1760), r=5/6 on Nvidia RTX 2080Ti. Rongchun Li, Hengyue Pan, Huayou Su, Yong Dou |
ISCC | 5 |
| 2020 | Objectness Consistent Representation for Weakly Supervised Object DetectionabstractWeakly supervised object detection aims at learning object detectors with only image-level category labels. Most existing methods tend to solve this problem by using a multiple instance learning detector which is usually trapped to discriminate object parts. In order to select high-quality proposals, recent works leverage objectness scores derived from weakly-supervised segmentation maps to rank the object proposals. Base on our observation, this kind of segmentation guided method always fails due to neglect of the fact that the objectness of all proposals inside the ground-truth box should be consistent. In this paper, we propose a novel object representation named Objectness Consistent Representation (OCRepr) to meet the consistency criterion of objectness. Specifically, we project the segmentation confidence scores into two orthogonal directions, namely vertical and horizontal, to get the OCRepr. With the novel object representation, more high-quality proposals can be mined for learning a much stronger object detector. We obtain 54.6% and 51.1% mAP scores on VOC 2007 and 2012 datasets, significantly outperforming the state-of-the-art and demonstrating the superiority of OCRepr for weakly supervised object detection. Ke Yang 0004, Peng Zhang 0035, Peng Qiao, Dongsheng Li 0001, Yong Dou |
ACM Multimedia | 6 |
| 2020 | Beyond top-N accuracy indicator: a comprehensive evaluation indicator of CNN models in image classificationabstractNowadays, a large number of deep convolutional neural network (CNN) models are applied to image classification tasks. However, the authors find that the most widely used evaluation indicator, the Top‐ N Accuracy indicator, cannot discriminate these models effectively. In this study, they propose a new indicator called Maximum‐Spanning‐Confusion‐Tree indicator to solve this problem. The Maximum‐Spanning‐Confusion‐Tree indicator is computed based on the hierarchical structure of the Maximum Spanning Confusion Tree of the deep CNN model on the dataset and reflect the ability of deep CNN models to discriminate confused categories in the dataset. The hierarchical structure of the Maximum Spanning Confusion Tree can reveal the confused category set of one selected category in the dataset efficiently and flexibly. Experiments show that they can discriminate ten different deep CNN models more accurately with the Maximum Spanning Confusion Tree indicator than the Top‐ N Accuracy indicator and the Maximum Spanning Confusion Tree intuitively shows the distribution of confused category sets in the dataset so they can find out the weakness of deep CNN models effectively. Yuntao Liu 0004, Yong Dou, Peng Qiao |
IET Comput. Vis. | 2 |
| 2020 | Annealed gradient descent for deep learning
Hengyue Pan, Xin Niu 0002, Rongchun Li, Yong Dou |
Neurocomputing | 4 |
| 2020 | DropFilterR: A Novel Regularization Method for Learning Convolutional Neural Networks
Hengyue Pan, Xin Niu 0002, Rongchun Li, Yong Dou |
Neural Process. Lett. | 5 |
| 2020 | Absent Multiple Kernel Learning AlgorithmsabstractMultiple kernel learning (MKL) has been intensively studied during the past decade. It optimally combines the multiple channels of each sample to improve classification performance. However, existing MKL algorithms cannot effectively handle the situation where some channels of the samples are missing, which is not uncommon in practical applications. This paper proposes three absent MKL (AMKL) algorithms to address this issue. Different from existing approaches where missing channels are first imputed and then a standard MKL algorithm is deployed on the imputed data, our algorithms directly classify each sample based on its observed channels, without performing imputation. Specifically, we define a margin for each sample in its own relevant space, a space corresponding to the observed channels of that sample. The proposed AMKL algorithms then maximize the minimum of all sample-based margins, and this leads to a difficult optimization problem. We first provide two two-step iterative algorithms to approximately solve this problem. After that, we show that this problem can be reformulated as a convex one by applying the representer theorem. This makes it readily be solved via existing convex optimization packages. In addition, we provide a generalization error bound to justify the proposed AMKL algorithms from a theoretical perspective. Extensive experiments are conducted on nine UCI and six MKL benchmark datasets to compare the proposed algorithms with existing imputation-based methods. As demonstrated, our algorithms achieve superior performance and the improvement is more significant with the increase of missing ratio. Xinwang Liu 0002, Lei Wang 0001, Xinzhong Zhu, Miaomiao Li 0001, En Zhu, Tongliang Liu, Li Liu 0002, Yong Dou, Jianping Yin |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2020 | Temporally Refined Graph U-Nets for Human Shape and Pose Estimation From Monocular VideosabstractThis work addresses a challenging problem of estimating the full 3D human shape and pose from monocular videos. Since real-world 3D mesh-labeled datasets are limited, most current methods in 3D human shape reconstruction only focus on single RGB images, losing all the temporal information. In contrast, we propose temporally refined Graph U-Nets, including an image-level module and a video-level module, to solve this problem. The image-level module is Graph U-Nets for human shape and pose estimation from images, where the Graph Convolutional Neural Network (Graph CNN) helps the information communication of neighboring vertices, and the U-Nets architecture enlarges the receptive field of each vertex and fuses high-level and low-level features. The video-level module is a small Residual Temporal Graph CNN (Residual TG-CNN), which learns temporal dynamics from both structural and temporal neighbors. The temporal dynamics of each vertex are continuous in the temporal dimension and highly relevant to the structural neighbors, so it is helpful to diminish the ambiguity of the body in single images by fusing temporal dynamics. Our algorithm makes full use of labels from image-level datasets and refines the image-level results through video-level module. Evaluated on Human3.6 M and 3DPW datasets, our model produces accurate 3D human meshes and achieves superior 3D human pose estimation accuracy when compared with state-of-the-art methods. Yang Zhao 0003, Yong Dou, Jiashi Feng |
IEEE Signal Process. Lett. | 2 |
| 2019 | Spatial Attention Network for Few-Shot Learning
Xianhao He, Peng Qiao, Yong Dou, Xin Niu 0002 |
ICANN (2) | 3 |
| 2019 | Towards Precise End-to-End Weakly Supervised Object Detection NetworkabstractIt is challenging for weakly supervised object detection network to precisely predict the positions of the objects, since there are no instance-level category annotations. Most existing methods tend to solve this problem by using a two-phase learning procedure, i.e., multiple instance learning detector followed by a fully supervised learning detector with bounding-box regression. Based on our observation, this procedure may lead to local minima for some object categories. In this paper, we propose to jointly train the two phases in an end-to-end manner to tackle this problem. Specifically, we design a single network with both multiple instance learning and bounding-box regression branches that share the same backbone. Meanwhile, a guided attention module using classification loss is added to the backbone for effectively extracting the implicit location information in the features. Experimental results on public datasets show that our method achieves state-of-the-art performance. Ke Yang 0004, Dongsheng Li 0001, Yong Dou |
ICCV | 3 |
| 2019 | Heavy-ball Algorithms Always Escape Saddle PointsabstractNonconvex optimization algorithms with random initialization have attracted increasing attention recently. It has been showed that many first-order methods always avoid saddle points with random starting points. In this paper, we answer a question: can the nonconvex heavy-ball algorithms with random initialization avoid saddle points? The answer is yes! Direct using the existing proof technique for the heavy-ball algorithms is hard due to that each iteration of the heavy-ball algorithm consists of current and last points. It is impossible to formulate the algorithms as iteration like xk+1= g(xk) under some mapping g. To this end, we design a new mapping on a new space. With some transfers, the heavy-ball algorithm can be interpreted as iterations after this mapping. Theoretically, we prove that heavy-ball gradient descent enjoys larger stepsize than the gradient descent to escape saddle points to escape the saddle point. And the heavy-ball proximal point algorithm is also considered; we also proved that the algorithm can always escape the saddle point. Tao Sun 0005, Dongsheng Li 0001, Zhe Quan, Hao Jiang 0001, Shengguo Li, Yong Dou |
IJCAI | 6 |
| 2019 | Accelerated Inference Framework of Sparse Neural Network Based on Nested Bitmask StructureabstractIn order to satisfy the ever-growing demand for high-performance processors for neural networks, the state-of-the-art processing units tend to use application-oriented circuits to replace Processing Engine (PE) on the GPU under circumstances where low-power solutions are required. The application-oriented PE is fully optimized in terms of the circuit architecture and eliminates incorrect data dependency and instructional redundancy. In this paper, we propose a novel encoding approach on a sparse neural network after pruning. We partition the weight matrix into numerous blocks and use a low-rank binary map to represent the validation of these blocks. Furthermore, the elements in each nonzero block are also encoded into two submatrices: one is the binary stream discriminating the zero/nonzero position, while the other is the pure nonzero elements stored in the FIFO. In the experimental part, we implement a well pre-trained sparse neural network on the Xilinx FPGA VC707. Experimental results show that our algorithm outperforms the other benchmarks. Our approach has successfully optimized the throughput and the energy efficiency to deal with a single frame. Accordingly, we contend that Nested Bitmask Neural Network (NBNN), is an efficient neural network structure with only minor accuracy loss on the SoC system. Yipeng Zhang 0001, Bo Du 0001, Lefei Zhang, Rongchun Li, Yong Dou |
IJCAI | 5 |
| 2019 | Exploring frame segmentation networks for temporal action localization
Ke Yang 0004, Xiaolong Shen, Peng Qiao, Shijie Li 0002, Dongsheng Li 0001, Yong Dou |
J. Vis. Commun. Image Represent. | 6 |
| 2018 | Exploring Temporal Preservation Networks for Precise Temporal Action LocalizationabstractTemporal action localization is an important task of computer vision. Though a variety of methods have been proposed, it still remains an open question how to predict the temporal boundaries of action segments precisely. Most works use segment-level classifiers to select video segments pre-determined by action proposal or dense sliding windows. However, in order to achieve more precise action boundaries, a temporal localization system should make dense predictions at a fine granularity. A newly proposed work exploits Convolutional-Deconvolutional-Convolutional (CDC) filters to upsample the predictions of 3D ConvNets, making it possible to perform per-frame action predictions and achieving promising performance in terms of temporal action localization. However, CDC network loses temporal information partially due to the temporal downsampling operation. In this paper, we propose an elegant and powerful Temporal Preservation Convolutional (TPC) Network that equips 3D ConvNets with TPC filters. TPC network can fully preserve temporal resolution and downsample the spatial resolution simultaneously, enabling frame-level granularity action localization with minimal loss of time information. TPC network can be trained in an end-to-end manner. Experiment results on public datasets show that TPC network achieves significant improvement in both per-frame action prediction and segment-level temporal action localization. Ke Yang 0004, Peng Qiao, Dongsheng Li 0001, Shaohe Lv, Yong Dou |
AAAI | 5 |
| 2018 | Learning Generic Diffusion Processes for Image Restoration
Peng Qiao, Yong Dou, Yunjin Chen, WenSen Feng |
BMVC | 2 |
| 2018 | mmCNN: A Novel Method for Large Convolutional Neural Network on Memory-Limited DevicesabstractDeep learning recently has been widely used in many interactive application fields including but not limited to object recognition, speech recognition, natural language processing and so on. At the same time more and more attractive interactive applications (face recognition and augmented reality) are available on wearable and mobile devices. However, traditional deep learning methods such as CNN cost a lot of memory resources. This challenge makes it difficult to apply the powerful deep learning method on mobile memory limited platforms. In this paper we present a novel memory management strategy called mmCNN to solve this problem. This method helps us deploy a trained large size CNN on an any memory size platform including GPU, FPGA and memory-limited mobile devices. In our experiments, we run a feed-forward CNN process in an extremely small memory size (as low as 5MB) on a GPU platform. The result shows that our method saves more than 98% memory compared to a traditional CNN algorithm and further saves more than 90% compared to the sate-of-the-art related work "vDNN". Our work improve the computing scalability of interaction applications and break the memory bottleneck of using deep learning method on a memory-limited devices. Shijie Li 0002, Yong Dou, Jinwei Xu, Qiang Wang 0006, Xin Niu 0002 |
COMPSAC (1) | 2 |
| 2018 | Deep Image Clustering Using Convolutional Autoencoder Embedding with Inception-Like BlockabstractImage clustering is one of the challenging tasks in machine learning, and has been extensively used in various applications. Recently, various deep clustering methods has been proposed. These methods take a two-stage approach, feature learning and clustering, sequentially or jointly. We observe that these works usually focus on the combination of reconstruction loss and clustering loss, relatively little work has focused on improving the learning representation of the neural network for clustering. In this paper, we propose a deep convolutional embedded clustering algorithm with inception-like block (DCECI). Specifically, an inception-like block with different type of convolution filters are introduced in the symmetric deep convolutional network to preserve the local structure of convolution layers. We simultaneously minimize the reconstruction loss of the convolutional autoencoders with inception-like block and the clustering loss. Experimental results on multiple image datasets exhibit the promising performance of our proposed algorithm compared with other competitive methods. Qiang Wang 0006, Rongchun Li, Peng Qiao, Ke Yang 0004, Shijie Li 0002, Yong Dou |
ICIP | 7 |
| 2018 | Temporal Pyramid Relation Network for Video-Based Gesture RecognitionabstractGesture recognition in video is an important application of computer vision. However, there are few works talked about the temporal order or relation of the frames in video, which is important for model gestures. In this paper, we propose Temporal Pyramid Relation Network (TPRN) which can model the temporal relation of video frames effectively and efficiently. First, we use Temporal Pyramid Pooling (TPP) layer to get temporal feature sequences of multiple scale pyramids. Then, a Temporal Relation Network (TRN) is stacked on the feature sequence of each scale respectively to model the temporal relations of video frames at multiple scales. At last, representations of all scales are aggregated to get the final prediction. TPRN can take video clips of various length as input and is scalable for video length. We evaluate TPRN on a recently released very large video-based gesture recognition dataset - 20BN-Jester dataset v1, and TPRN achieves competitive performance. Ke Yang 0004, Rongchun Li, Peng Qiao, Qiang Wang 0006, Dongsheng Li 0001, Yong Dou |
ICIP | 6 |
| 2018 | Visual Confusion Label Tree for Image ClassificationabstractConvolution neural network models are widely used in image classification tasks. However, the running time of such models is so long that it is not the conforming to the strict real-time requirement of mobile devices. In order to optimize models and meet the requirement mentioned above, we propose a method that replaces the fully-connected layers of convolution neural network models with a tree classifier. Specifically, we construct a Visual Confusion Label Tree based on the output of the convolution neural network models, and use a multi-kernel SVM plus classifier with hierarchical constraints to train the tree classifier. Focusing on those confusion subsets instead of the entire set of categories makes the tree classifier more discriminative and the replacement of the fully-connected layers reduces the original running time. Experiments show that our tree classifier obtains a significant improvement over the state-of-the-art tree classifier by 4.3% and 2.4% in terms of top-l accuracy on CIFAR-100 and ImageNet datasets respectively. Additionally, our method achieves 124× and 115× speedup ratio compared with fully-connected layers on AlexNet and VGG16 without accuracy decline. Yuntao Liu 0004, Yong Dou, Ruochun Jin, Rongchun Li |
ICME | 2 |
| 2018 | Visual Tree Convolutional Neural Network in Image ClassificationabstractIn image classification, Convolutional Neural Net-work(CNN) models have achieved high performance with the rapid development in deep learning. However, some categories in the image datasets are more difficult to distinguished than others. Improving the classification accuracy on these confused categories is benefit to the overall performance. In this paper, we build a Confusion Visual Tree(CVT) based on the confused semantic level information to identify the confused categories. With the information provided by the CVT, we can lead the CNN training procedure to pay more attention on these confused categories. Therefore, we propose Visual Tree Convolutional Neural Networks(VT-CNN) based on the original deep CNN embedded with our CVT. We evaluate our VT-CNN model on the benchmark datasets CIFAR-10 and CIFAR-100. In our experiments, we build up 3 different VT-CNN models and they obtain improvement over their based CNN models by 1.36%, 0.89% and 0.64%, respectively. Yuntao Liu 0004, Yong Dou, Ruochun Jin, Peng Qiao |
ICPR | 2 |
| 2018 | An efficient CPU-GPU hybrid parallel implementation for DVB-RCS2 receiverabstractSummary The second‐generation digital video broadcasting return channel via satellite (DVB‐RCS2) is a promising real‐time wireless protocol that has been widely used in many applications, such as video conferences, video feeds, and video multicasting. However, the receiver end of DVB‐RCS2 is time consuming and should be accelerated by high‐performance processing systems. Today, graphic processing units (GPUs) have been applied in communication systems due to high parallel capability and processing throughput. In this study, we design a novel pipeline of the receiver on the CPU‐GPU platform. Moreover, we propose a CPU‐GPU hybrid strategy to fully utilize resources and reduce communication latency. Compared with the parallel turbo decoder proposed in other work on the same platform, our parallel implementation achieves higher throughput. For the entire DVB‐RCS2 receiver, compared with the non‐pipelined serial and non‐pipelined parallel algorithms, our proposed pipeline obtains 20 times and 6 times speedup, respectively. In addition, the latency of our implementation is lower than that of non‐pipelined CPU‐GPU implementation, which is equal to 1.06 ms. Yueqing Wang, Fang Wang 0004, Rongchun Li, Yong Dou |
Concurr. Comput. Pract. Exp. | 4 |
| 2018 | Local kernel alignment based multi-view clustering using extreme learning machine
Qiang Wang 0006, Yong Dou, Xinwang Liu 0002, Fei Xia 0003, Ke Yang 0004 |
Neurocomputing | 2 |
| 2018 | Distributed sparse bundle adjustment algorithm based on three-dimensional point partition and asynchronous communicationabstractSparse bundle adjustment (SBA) is a key but time- and memory-consuming step in three-dimensional (3D) reconstruction. In this paper, we propose a 3D point-based distributed SBA algorithm (DSBA) to improve the speed and scalability of SBA. The algorithm uses an asynchronously distributed sparse bundle adjustment (A-DSBA) to overlap data communication with equation computation. Compared with the synchronous DSBA mechanism (SDSBA), A-DSBA reduces the running time by 46%. The experimental results on several 3D reconstruction datasets reveal that our distributed algorithm running on eight nodes is up to five times faster than that of the stand-alone parallel SBA. Furthermore, the speedup of the proposed algorithm (running on eight nodes with 48 cores) is up to 41 times that of the serial SBA (running on a single node). Xiaolong Shen, Yong Dou, Steven Mills, David M. Eyers, Huan Feng, Zhiyi Huang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2017 | Multiple Kernel k-Means with Incomplete KernelsabstractMultiple kernel clustering (MKC) algorithms optimally combine a group of pre-specified base kernels to improve clustering performance. However, existing MKC algorithms cannot efficiently address the situation where some rows and columns of base kernels are absent. This paper proposes a simple while effective algorithm to address this issue. Different from existing approaches where incomplete kernels are firstly imputed and a standard MKC algorithm is applied to the imputed kernels, our algorithm integrates imputation and clustering into a unified learning procedure. Specifically, we perform multiple kernel clustering directly with the presence of incomplete kernels, which are treated as auxiliary variables to be jointly optimized. Our algorithm does not require that there be at least one complete base kernel over all the samples. Also, it adaptively imputes incomplete kernels and combines them to best serve clustering. A three-step iterative algorithm with proved convergence is designed to solve the resultant optimization problem. Extensive experiments are conducted on four benchmark data sets to compare the proposed algorithm with existing imputation-based methods. Our algorithm consistently achieves superior performance and the improvement becomes more significant with increasing missing ratio, verifying the effectiveness and advantages of the proposed joint imputation and clustering. Xinwang Liu 0002, Miaomiao Li 0001, Lei Wang 0001, Yong Dou, Jianping Yin, En Zhu |
AAAI | 4 |
| 2017 | Optimal Neighborhood Kernel Clustering with Multiple KernelsabstractMultiple kernel $k$-means (MKKM) aims to improve clustering performance by learning an optimal kernel, which is usually assumed to be a linear combination of a group of pre-specified base kernels. However, we observe that this assumption could: i) cause limited kernel representation capability; and ii) not sufficiently consider the negotiation between the process of learning the optimal kernel and that of clustering, leading to unsatisfying clustering performance. To address these issues, we propose an optimal neighborhood kernel clustering (ONKC) algorithm to enhance the representability of the optimal kernel and strengthen the negotiation between kernel learning and clustering. We theoretically justify this ONKC by revealing its connection with existing MKKM algorithms. Furthermore, this justification shows that existing MKKM algorithms can be viewed as a special case of our approach and indicates the extendability of the proposed ONKC for designing better clustering algorithms. An efficient algorithm with proved convergence is designed to solve the resultant optimization problem. Extensive experiments have been conducted to evaluate the clustering performance of the proposed algorithm. As demonstrated, our algorithm significantly outperforms the state-of-the-art ones in the literature, verifying the effectiveness and advantages of ONKC. Xinwang Liu 0002, Sihang Zhou 0001, Yueqing Wang, Miaomiao Li 0001, Yong Dou, En Zhu, Jianping Yin |
AAAI | 5 |
| 2017 | Platform-Adaptive High-Throughput Surveillance Video Condensation on Heterogeneous Processor Clusters
Peng Qiao, Teng Li 0010, Yong Dou, Yuanwu Lei, Hongbing Luo |
APPT | 3 |
| 2017 | An FPGA-based processor for training convolutional neural networksabstractConvolutional neural networks (CNNs) have gained great success in various computer vision applications. However, training a CNN model is computation-intensive and time-consuming. Hence training is mainly processed on large clusters of high-performance processors like server CPUs and GPUs. In this paper, we propose an FPGA-based processor design to accelerate the training process of CNNs. We first analyze the operations in all types of CNN layers in the training process. A uniform computation engine design is proposed to efficiently carry out all kinds of operations based on the analysis. Then a scalable accelerator framework is presented that exploits the parallelism further by unrolling the loops in two levels. The proposed accelerator design is demonstrated by implementing a processor on the Xilinx ZU19EG FPGA working at 200 MHz. The evaluation results on a group of CNN models show that our processor is 5.7 to 10.7-fold faster than the software implementations on the Intel Core i5-4440 CPU(@3.10GHz). Yong Dou, Jingfei Jiang, Qiang Wang 0006, Paul Chow |
FPT | 2 |
| 2017 | Confusion Graph: Detecting Confusion Communities in Large Scale Image ClassificationabstractFor deep CNN-based image classification models, we observe that confusions between classes with high visual similarity are much stronger than those where classes are visually dissimilar. With these unbalanced confusions, classes can be organized in communities, which is similar to cliques of people in the social network. Based on this, we propose a graph-based tool named "confusion graph" to quantify these confusions and further reveal the community structure inside the database. With this community structure, we can diagnose the model's weaknesses and improve the classification accuracy using specialized expert sub-nets, which is comparable to other state-of-the-art techniques. Utilizing this community information, we can also employ pre-trained models to automatically identify mislabeled images in the large scale database. With our method, researchers just need to manually check approximate 3% of the ILSVRC2012 classification database to locate almost all mislabeled samples. Ruochun Jin, Yong Dou, Yueqing Wang, Xin Niu 0002 |
IJCAI | 2 |
| 2017 | Multiple Kernel Clustering Framework with Improved KernelsabstractMultiple kernel clustering (MKC) algorithms have been successfully applied into various applications. However, these successes are largely dependent on the quality of pre-defined base kernels, which cannot be guaranteed in practical applications. This may adversely affect the clustering performance. To address this issue, we propose a simple while effective framework to adaptively improve the quality of these base kernels. Under our framework, we instantiate three MKC algorithms based on the widely used multiple kernel $k$-means clustering (MKKM), MKKM with matrix-induced regularization (MKKM-MR) and co-regularized multi-view spectral clustering (CRSC). After that, we design the corresponding algorithms with proved convergence to solve the resultant optimization problems. To the best of our knowledge, our framework fills the gap between kernel adaption and clustering procedure for the first time in the literature and is readily extendable. Extensive experimental research has been conducted on 7 MKC benchmarks. As is shown, our algorithms consistently and significantly improve the performance of the base MKC algorithms, indicating the effectiveness of the proposed framework. Meanwhile, our framework shows better performance than compared ones with imperfect kernels. Yueqing Wang, Xinwang Liu 0002, Yong Dou, Rongchun Li |
IJCAI | 3 |
| 2017 | Approximate Large-scale Multiple Kernel k-means Using Deep Neural NetworkabstractMultiple kernel clustering (MKC) algorithms have been extensively studied and applied to various applications. Although they demonstrate great success in both the theoretical aspects and applications, existing MKC algorithms cannot be applied to large-scale clustering tasks due to: i) the heavy computational cost to calculate the base kernels; and ii) insufficient memory to load the kernel matrices. In this paper, we propose an approximate algorithm to overcome these issues, and to make it be applicable to large-scale applications. Specifically, our algorithm trains a deep neural network to regress the indicating matrix generated by MKC algorithms on a small subset, and then obtains the approximate indicating matrix of the whole data set using the trained network, and finally performs the $k$-means on the output of our network. By mapping features into indicating matrix directly, our algorithm avoids computing the full kernel matrices, which dramatically decreases the memory requirement. Extensive experiments show that our algorithm consumes less time than most comparatively similar algorithms, while it achieves comparable performance with MKC algorithms. Yueqing Wang, Xinwang Liu 0002, Yong Dou, Rongchun Li |
IJCAI | 3 |
| 2017 | Learning Non-local Image Diffusion for Image DenoisingabstractImage diffusion plays a fundamental role for the task of image denoising. The recently proposed trainable nonlinear reaction diffusion (TNRD) model defines a simple but very effective framework for image denoising. However, as the TNRD model is a local model, whose diffusion behavior is purely controlled by information of local patches, it is prone to create artifacts in the homogenous regions and over-smooth highly textured regions, especially in the case of strong noise levels. Meanwhile, it is widely known that the non-local self-similarity (NSS) prior stands as an effective image prior for image denoising, which has been widely exploited in many non-local methods. In this work, we are highly motivated to embed the NSS prior into the TNRD model to tackle its weaknesses. In order to preserve the expected property that end-to-end training remains available, we exploit the NSS prior by defining a set of non-local filters, and derive our proposed trainable non-local reaction diffusion (TNLRD) model for image denoising. Together with the local filters and influence functions, the non-local filters are learned by employing loss-specific training. The experimental results show that the trained TNLRD model produces visually plausible recovered images with more textures and less artifacts, compared to its local versions. Moreover, the trained TNLRD model can achieve strongly competitive performance to recent state-of-the-art image denoising methods in terms of peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM). Peng Qiao, Yong Dou, WenSen Feng, Rongchun Li, Yunjin Chen |
ACM Multimedia | 2 |
| 2017 | Robust regularized extreme learning machine for regression using iteratively reweighted least squares
Kai Chen 0020, Yong Dou |
Neurocomputing | 4 |
| 2017 | Multiple kernel clustering with corrupted kernels
Teng Li 0010, Yong Dou, Xinwang Liu 0002, Yang Zhao 0003 |
Neurocomputing | 2 |
| 2017 | A fast and memory saved GPU acceleration algorithm of convolutional neural networks for target detection
Shijie Li 0002, Yong Dou, Xin Niu 0002, Qiang Wang 0006 |
Neurocomputing | 2 |
| 2017 | Heterogeneous blocked CPU-GPU accelerate scheme for large scale extreme learning machine
Shijie Li 0002, Xin Niu 0002, Yong Dou, Yueqing Wang |
Neurocomputing | 3 |
| 2017 | An optimized design of CAN FD for automotive cyber-physical systems
Yong Xie 0003, Ryo Kurachi, Guoqi Xie, Yong Dou, Zhili Zhou 0001 |
J. Syst. Archit. | 5 |
| 2017 | Airport Detection on Optical Satellite Images Using Deep Convolutional Neural NetworksabstractThis letter proposes a method using convolutional neural networks (CNNs) for airport detection on optical satellite images. To efficiently build a deep CNN with limited satellite image samples, a transfer learning approach had been employed by sharing the common image features of the natural images. To decrease the computing cost, an efficient region proposal method had been proposed based on the prior knowledge of the line segments distribution in an airport. The transfer learning ability on deep CNN for airport detection on satellite images had been first evaluated in this letter. The proposed method was tested on an image data set, including 170 different airports and 30 nonairports. The detection rate could reach 88.8% in experiments with seconds' computation time, which showed a great improvement over other the state-of-the-art methods. Peng Zhang 0035, Xin Niu 0002, Yong Dou, Fei Xia 0003 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2017 | Multiple kernel learning with hybrid kernel alignment maximization
Yueqing Wang, Xinwang Liu 0002, Yong Dou |
Pattern Recognit. | 3 |
| 2017 | Variational single image interpolation with time-varying regularization
Peng Qiao, Yunjin Chen, Yong Dou |
Signal Process. Image Commun. | 3 |
| 2017 | Qualitative Action Recognition by Wireless Radio Signals in Human-Machine SystemsabstractHuman-machine systems required a deep understanding of human behaviors. Most existing research on action recognition has focused on discriminating between different actions, however, the quality of executing an action has received little attention thus far. In this paper, we study the quality assessment of driving behaviors and present WiQ, a system to assess the quality of actions based on radio signals. This system includes three key components, a deep neural network based learning engine to extract the quality information from the changes of signal strength, a gradient-based method to detect the signal boundary for an individual action, and an activity-based fusion policy to improve the recognition performance in a noisy environment. By using the quality information, WiQ can differentiate a triple body status with an accuracy of 97%, whereas for identification among 15 drivers, the average accuracy is 88%. Our results show that, via dedicated analysis of radio signals, a fine-grained action characterization can be achieved, which can facilitate a large variety of applications, such as smart driving assistants. Shaohe Lv, Mianxiong Dong, Xiaodong Wang 0002, Yong Dou, Weihua Zhuang |
IEEE Trans. Hum. Mach. Syst. | 5 |
| 2017 | Throughput-Optimized FPGA Accelerator for Deep Convolutional Neural NetworksabstractDeep convolutional neural networks (CNNs) have gained great success in various computer vision applications. State-of-the-art CNN models for large-scale applications are computation intensive and memory expensive and, hence, are mainly processed on high-performance processors like server CPUs and GPUs. However, there is an increasing demand of high-accuracy or real-time object detection tasks in large-scale clusters or embedded systems, which requires energy-efficient accelerators because of the green computation requirement or the limited battery restriction. Due to the advantages of energy efficiency and reconfigurability, Field-Programmable Gate Arrays (FPGAs) have been widely explored as CNN accelerators. In this article, we present an in-depth analysis of computation complexity and the memory footprint of each CNN layer type. Then a scalable parallel framework is proposed that exploits four levels of parallelism in hardware acceleration. We further put forward a systematic design space exploration methodology to search for the optimal solution that maximizes accelerator throughput under the FPGA constraints such as on-chip memory, computational resources, external memory bandwidth, and clock frequency. Finally, we demonstrate the methodology by optimizing three representative CNNs (LeNet, AlexNet, and VGG-S) on a Xilinx VC709 board. The average performance of the three accelerators is 424.7, 445.6, and 473.4GOP/s under 100MHz working frequency, which outperforms the CPU and previous work significantly. Yong Dou, Jingfei Jiang, Jinwei Xu, Shijie Li 0002, Yongmei Zhou, Yingnan Xu |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2016 | Multiple Kernel k-Means Clustering with Matrix-Induced RegularizationabstractMultiple kernel k-means (MKKM) clustering aims to optimally combine a group of pre-specified kernels to improve clustering performance. However, we observe that existing MKKM algorithms do not sufficiently consider the correlation among these kernels. This could result in selecting mutually redundant kernels and affect the diversity of information sources utilized for clustering, which finally hurts the clustering performance. To address this issue, this paper proposes an MKKM clustering with a novel, effective matrix-induced regularization to reduce such redundancy and enhance the diversity of the selected kernels. We theoretically justify this matrix-induced regularization by revealing its connection with the commonly used kernel alignment criterion. Furthermore, this justification shows that maximizing the kernel alignment for clustering can be viewed as a special case of our approach and indicates the extendability of the proposed matrix-induced regularization for designing better clustering algorithms. As experimentally demonstrated on five challenging MKL benchmark data sets, our algorithm significantly improves existing MKKM and consistently outperforms the state-of-the-art ones in the literature, verifying the effectiveness and advantages of incorporating the proposed matrix-induced regularization. Xinwang Liu 0002, Yong Dou, Jianping Yin, Lei Wang 0001, En Zhu |
AAAI | 2 |
| 2016 | Automatic code generation of convolutional neural networks in FPGA implementationabstractConvolutional neural networks (CNNs) have gained great success in various computer vision applications. However, state-of-the-art CNN models are computation-intensive and hence are mainly processed on high performance processors like server CPUs and GPUs. Owing to the advantages of high performance, energy efficiency and reconfigurability, Field-Programmable Gate Arrays (FPGAs) have been widely explored as CNN accelerators. In this paper, we propose parallel structures to exploit the inherent parallelism and efficient computation units to perform operations in convolutional and fully-connected layers. Further, an automatic generator is proposed to generate Verilog HDL source code automatically according to high-level hardware description language. Execution time, DSP consumption and performance are analytically modeled based on some critical design variables. We demonstrate the automatic methodology by implementing two representative CNNs (LeNet and AlexNet) and evaluate the execution time models by comparing estimated and measured values. Our results show that the proposed automatic methodology yields hardware design with good performance and saves much developing round time. Yong Dou, Jingfei Jiang, Jinwei Xu |
FPT | 2 |
| 2016 | Localized region context and object feature fusion for people head detectionabstractPeople head detection in crowded scenes is challenging due to the large variability in clothing and appearance, small scales of people, and strong partial occlusions. Traditional bottom-up proposal methods and existing region proposal network approaches suffer from either poor recall or low precision. In this paper, we propose to improve both the recall and precision of head detection of region proposal models by integrating the local head information. In specific, we first use a region proposal network to predict the bounding boxes and corresponding scores of multiple instances in the region. A local head classifier network is then trained to score the bounding box generated from the region proposal model. After that, we propose an adaptive fusion method by optimally combining both the region and local scores to obtain the final score of each candidate bounding box. Furthermore, our fusion models can automatically learn the optimal hyper-parameters from data. Our algorithm achieves superior people head detection performance on the crowded scenes data set, which significantly outperforms several recent state-of-the-art baselines in the literature. Yule Li, Yong Dou, Xinwang Liu 0002, Teng Li 0010 |
ICIP | 2 |
| 2016 | Hyperspectral image classification via kernel extreme learning machine using local receptive fieldsabstractThis paper proposes a classification approach for hyperspectral image (HSI) using the local receptive fields based kernel extreme learning machine. Extreme learning machine (ELM) has drawn increasing attention in the pattern recognition filed due to its simpleness, speediness and good generalization ability. A kernel method is often used to promote ELM's performance, which is known as kernel ELM. The local receptive field concept originates from research in neuroscience. Considering the local correlations of spectral features, it is promising to improve the performance of HSI classification by combining local receptive fields with kernel ELM. Experimental results on the Pavia University dataset confirm the effectiveness of the proposed HSI classification method. Xin Niu 0002, Yong Dou, Yueqing Wang, Jie Zhou 0007 |
ICIP | 3 |
| 2016 | Multiple Kernel Clustering with Local Kernel Alignment Maximization
Miaomiao Li 0001, Xinwang Liu 0002, Lei Wang 0001, Yong Dou, Jianping Yin, En Zhu |
IJCAI | 4 |
| 2016 | Airport detection from remote sensing images using transferable convolutional neural networksabstractThis paper presents a method for airport detection from optical satellite images using deep convolutional neural networks (CNN). To achieve fast detection with high accuracy, region proposal by searching adjacent parallel line segments has been applied to select candidate fields with potential runways. These proposals were further classified by a CNN model transfer learned from AlexNet to identify the final airport regions from other confusing classes. The proposed method has been tested on a remote sensing dataset consisting of 120 airports. Experiments showed that the proposed method could recognize airports from a large complex area in seconds with an accuracy of 84.1%. Peng Zhang 0035, Xin Niu 0002, Yong Dou, Fei Xia 0003 |
IJCNN | 3 |
| 2016 | ELM based multiple kernel k-means with diversity-induced regularizationabstractMultiple-kernel k-means (MKKM) clustering has demonstrated good clustering performance by combining pre-specified kernels. In this paper, we argue that deep relationships within data and the complementary information among them can improve the performance of MKKM. To illustrate this idea, we propose a diversity-induced MKKM algorithm with extreme learning machine (ELM)-based feature extracting method. First, ELM, which has randomly chosen weights of hidden and output nodes, is applied to thoroughly extract features from data by generating different numbers of hidden nodes and using different functions. Second, an MKKM algorithm with diversity-induced regularization is utilized to explore the complementary information among kernels constructed from features. The problem could be solved efficiently by alternating optimization. Experimental results demonstrate that the proposed method outperforms state-of-the-art kernel methods. Yang Zhao 0003, Yong Dou, Xinwang Liu 0002, Teng Li 0010 |
IJCNN | 2 |
| 2016 | Face Verification Algorithm with Exploiting Feature Distribution
Xuan Li 0002, Yong Dou, Ke Yang 0004 |
PRICAI | 3 |
| 2016 | Joint diversity regularization and graph regularization for multiple kernel k-means clustering via latent variables
Teng Li 0010, Yong Dou, Xinwang Liu 0002 |
Neurocomputing | 2 |
| 2016 | PR-ELM: Parallel regularized extreme learning machine based on cluster
Yueqing Wang, Yong Dou, Xinwang Liu 0002, Yuanwu Lei |
Neurocomputing | 2 |
| 2016 | Multi-view clustering with extreme learning machine
Qiang Wang 0006, Yong Dou, Xinwang Liu 0002, Shijie Li 0002 |
Neurocomputing | 2 |
| 2016 | An efficient and effective convolutional auto-encoder extreme learning machine network for 3d feature learning
Yueqing Wang, Zhige Xie, Kai Xu 0004, Yong Dou, Yuanwu Lei |
Neurocomputing | 4 |
| 2016 | A novel multi-view clustering method via low-rank and matrix-induced regularization
Yang Zhao 0003, Yong Dou, Xinwang Liu 0002, Teng Li 0010 |
Neurocomputing | 2 |
| 2016 | Relative distance features for gait recognition with Kinect
Ke Yang 0004, Yong Dou, Shaohe Lv |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Classification of Hyperspectral Remote Sensing Image Using Hierarchical Local-Receptive-Field-Based Extreme Learning MachineabstractThis letter proposes a novel classification approach for a hyperspectral image (HSI) using a hierarchical local-receptive-field (LRF)-based extreme learning machine (ELM). As a fast and accurate pattern classification algorithm, ELM has been applied in numerous fields, including the HSI classification. The LRF concept originates from research in neuroscience. Considering the local correlations of spectral features, it is promising to improve the performance of HSI classification by introducing the LRFs. Recent research on deep learning has shown that hierarchical architectures with more layers can potentially extract abstract representation and invariant features for better classification performance. Therefore, we further extend the LRF-based ELM method to a hierarchical model for HSI classification. Experimental results on two widely used real hyperspectral data sets confirm the effectiveness of the proposed HSI classification approach. Xin Niu 0002, Yong Dou, Yuanwu Lei |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2016 | Affine-Transformation Parameters Regression for Face AlignmentabstractFace alignment is an important process in facial analysis. Cascaded linear regression approaches have shown the capability to achieve the state-of-the-art accuracy on numerous face alignment datasets. However, most of these approaches only learn to map coordinate offsets of the key points from image features. This regression strategy can be easily trapped in local optima. We propose a novel regression strategy by introducing affine transformation. First, the best affine-transformation parameters between the initial mean shape and the ground truth are estimated by Procrustes analysis. Subsequently, we base the mapping from image features on the best affine-transformation parameters. Experimental results indicate that this strategy can reduce the offsets between two shapes significantly. Combined with coordinate-offset regression strategy, the hybrid approach produces a remarkably performance in term of accuracy, training time, prediction rate, and the model size. Moreover, the affine-transformation parameter regression strategy can be considered as a shape-initialization method that can be combined with other initial shape-based face alignment algorithms to improve the face alignment accuracy. Xuan Li 0002, Yidan Xu, Yong Dou |
IEEE Signal Process. Lett. | 4 |
| 2016 | Coarse-Grained Architecture for Fingerprint MatchingabstractFingerprint matching is a key procedure in fingerprint identification applications. The minutiae-based fingerprint matching algorithm is one of the most typical algorithms achieving a reasonably correct recognition rate. This study proposes a coarse-grained parallel architecture called fingerprint matching core (FMC) to accelerate fingerprint matching. The proposed architecture has a two-level parallel structure (i.e., parallel among groups (PAG) and parallel in group (PIG)). A multirequest controller is added to the PAG structure to obtain a concurrent operation of the multiple processing element group (PEG). The DDR3 controller is used in the PIG structure to read eight minutiae from eight different fingerprints and realize the simultaneous computation of the eight PEs. The whole system is implemented on a Xilinx FPGA board with a Virtex VII XC7VX485T chip. The 16-PEG FMC achieves a throughput of about 9.63 million fingerprint pairs per second, which is larger than that achieved on a Tesla K20c platform. The software execution times are also measured on the 2.93GHz Intel Xeon 5670, 2.3GHz AMD Opteron(tm) Processor 6376, and Tesla K20c platforms. The Intel Xeon 5670 has two processors with 12 cores, and the AMD Opteron(tm) Processor 6376 has two processors with 16 cores. Moreover, the throughput is about 31 times that achieved on a 2.93GHz Intel Xeon 5670 single core. Jinwei Xu, Jingfei Jiang, Yong Dou, Xiaolong Shen |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2015 | Absent Multiple Kernel LearningabstractMultiple kernel learning (MKL) optimally combines the multiple channels of each sample to improve classification performance. However, existing MKL algorithms cannot effectively handle the situation where some channels are missing, which is common in practical applications. This paper proposes an absent MKL (AMKL) algorithm to address this issue. Different from existing approaches where missing channels are firstly imputed and then a standard MKL algorithm is deployed on the imputed data, our algorithm directly classifies each sample with its observed channels. In specific, we define a margin for each sample in its own relevant space, which corresponds to the observed channels of that sample. The proposed AMKL algorithm then maximizes the minimum of all sample-based margins, and this leads to a difficult optimization problem. We show that this problem can be reformulated as a convex one by applying the representer theorem. This makes it readily be solved via existing convex optimization packages. Extensive experiments are conducted on five MKL benchmark data sets to compare the proposed algorithm with existing imputation-based methods. As observed, our algorithm achieves superior performance and the improvement is more significant with the increasing missing ratio. Xinwang Liu 0002, Lei Wang 0001, Jianping Yin, Yong Dou, Jian Zhang 0002 |
AAAI | 4 |
| 2015 | Exploring Relative Motion Features for Gait Recognition with Kinect
Ke Yang 0004, Yong Dou, Shaohe Lv |
ICONIP (4) | 2 |
| 2015 | Optimized deep belief networks on CUDA GPUsabstractA deep belief network (DBN) is an important branch of deep learning models and has been successfully applied in many machine learning and pattern recognition fields such as computer vision and speech recognition. However, the training of billions of parameters in DBN is computationally challenging for modern central processing units (CPUs). Many studies have reported the efficient implementations of the pre-training process of DBNs for graphics processing units (GPUs), but few studies have mentioned the fine-tuning process of DBNs. In this paper, we describe an efficient DBN implementation on the GPU, including the pre-training and fine-tuning processes. Experimental results show that our proposed method on the GPU (NVIDIA Tesla K40c) achieves up to 22 speedups on the pre-training process and 33 speedups on the fine-tuning processes compared with conventional CPU (Intel Core i7-4790K) implementations. Moreover, the performance of our algorithm is superior to that of the OpenBLAS library on the CPU and the CUBLAS library on the GPU. Teng Li 0010, Yong Dou, Jingfei Jiang, Yueqing Wang |
IJCNN | 2 |
| 2015 | Efficient graphics processing unit based layered decoders for quasicyclic low-density parity-check codesabstractSUMMARY Because layered low‐density parity‐check (LDPC) decoding algorithm was proposed, one can exploit the diversity gain to achieve performance comparable to the traditional two‐phase message passing (TPMP) decoding but with about twice faster decoding convergence compared to TPMP. In order to reduce the decoding time of layered LDPC decoder, a graphics processing unit (GPU) is exploited as the modem processor so that the decoding procedure can be processed in parallel using numerous threads in the GPU. In this paper, we present the parallel algorithms and efficient implementations on the GPU for two different layered message passing schemes, the row‐layered and column‐layered decoding. In the experiments, the quasicyclic LDPC codes for WiFi (802.11n) and WiMAX (802.16e) are decoded by the proposed layered LDPC decoders. The experimental results show that our decoder has good bit error ratio (BER) performance comparable to TPMP decoder. The peak throughput is 712 Mbps, which is about two orders of magnitude faster than that of CPU implementation and comparable to the dedicated hardware solutions. Compared to the existing fastest GPU‐based implementation, the presented decoder can achieve a performance improvement of 2.3 times. Copyright © 2013 John Wiley & Sons, Ltd. Rongchun Li, Yong Dou, Dan Zou |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | Classification of land cover based on deep belief networks using polarimetric RADARSAT-2 dataabstractUrban land use and land cover (LULC) classification is one of the core applications in Geographic Information Sys-tem(GIS). In this paper, a novel classification approach based on Deep Belief Network(DBN) for detailed urban mapping is proposed. Deep Belief Network (DBN) is a widely investigated and deployed deep learning model. By applying the DBN model, effective spatio-temporal mapping features can be automatically extracted to improve the classification performance. Six-date RADARSAT-2 Polarimetric SAR (PolSAR) data over the Great Toronto Area were used for evaluation. Experimental results showed that the proposed method can outperform SVM and contextual approaches using adaptive MRF. Yong Dou, Xin Niu 0002, Baoliang Li |
IGARSS | 2 |
| 2014 | 3D pipeline contention: Asymmetric full duplex in wireless networksabstractCoordination among users is an indispensable part in wireless networks for efficient medium access. Alone with the rapid increase of transmission rate, however, coordination time becomes insufferable. We present AFD, namely asymmetric full duplex, to achieve high coordination efficiency at nearly zero overhead. In AFD, channel contention is performed simultaneously with data transmission. We propose a 3D pipeline contention scheme where the contention process is divided into several parallel stages and executed in a pipelined manner in a 3D domain specified by time, frequency and spatial antenna. To mitigate the interference between the data packet and the contention signal, we adopt a singleton PN sequence as a contention pilot. AFD provides a novel network-scale full duplex capability. The performance is evaluated by both simulations and measurements in a testbed. AFD outperforms IEEE 802.11 significantly, i.e., the Jain's fairness index is around 0.95 with a throughput gain up to 120%. Shaohe Lv, Xuan Dong 0002, Xiaoli Du, Xiaodong Wang 0002, Yong Dou, Xingming Zhou |
INFOCOM | 6 |
| 2014 | Efficient parallel implementation of three-point viterbi decoding algorithm on CPU, GPU, and FPGAabstractSUMMARY In wireless communication, Viterbi decoding algorithm (VDA) is the one of most popular channel decoding algorithms, which is widely used in WLAN, WiMAX, or 3G communications. However, the throughput of Viterbi decoder is constrained by the convolutional characteristic. Recently, the three‐point VDA (TVDA) was proposed to solve this problem. In TVDA, the whole procedure can be divided into three phases, the forward, trace‐back, and decoding phases. In this paper, we analyze the parallelism of TVDA and propose parallel TVDA on the multi‐core CPU, graphics processing unit (GPU), and field programmable gate array (FPGA). We demonstrate approaches that fully exploit its performance potential on CPU, GPU, and FPGA computing platforms. For CPU platforms, we perform two optimization methods, single instruction multiple data and multithreading to gain over 145 × speedup over the naive CPU version on a quad‐core CPU platform. For GPU platforms, we propose the combination of cached memory optimization, coalesced global memory accesses, codeword packing scheme, and asynchronous data transition, achieving the throughput of 404.65 Mbps and 12 × speedup over initial GPU versions on an NVIDIA GeForce GTX580 card and 7 × speedup over Intel quad‐core CPU i5‐2300, under the same manufacturing year and both with fully optimized schemes. In addition, for FPGA platforms, we customize a radix‐4 pipelined architecture for the TVDA in a 45‐nm FPGA chip from Xilinx (XC6VLX760). Under 209.15‐MHz clock rate, it achieves a throughput of 418.30 Mbps. Finally, we also discuss the performance evaluation and efficiency comparison of different flexible architectures for real‐time Viterbi decoding in terms of the decoding throughput, power consumption, optimization schemes, programming costs, and price costs.Copyright © 2013 John Wiley & Sons, Ltd. Rongchun Li, Yong Dou, Dan Zou |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | CPU-GPU hybrid parallel strategy for cosmological simulationsabstractSUMMARY Gadget is a simulation application for N‐body and smoothed particle hydrodynamics problems in cosmology, and it is widely applied in solving series of cosmological problems. N‐body focuses on the motion of the interaction of N particles, and smoothed particle hydrodynamics is a fluid simulation algorithm that studies the movement of fluid through particle simulation. Most scholars focus their attention on accelerating Gadget on multi‐core CPU or graphics processing units (GPUs) platforms. However, these research activities failed to achieve CPU–GPU hybrid computing, which resulted in tremendous waste of CPU computing resources. In this paper, we propose a CPU–GPU hybrid parallel strategy to accelerate Gadget‐2, a massively parallel structure formation code for cosmological simulations. This strategy uses CPU and GPU to process the calculation of short‐range force. To ensure CPU and GPU workload balance, a dynamic task allocation scheme is proposed according to the computational performance difference between the CPU and GPU. Experimental results showed that our CPU–GPU hybrid parallel strategy achieved an overall speedup factor of 18.6 and a partial speedup factor for short‐range force calculation of 28.35 compared with a single‐core CPU implementation for particles in million‐size magnitudes. Moreover, compared with a GPU platform that contained 12 CPU cores and one GPU, our hybrid parallel strategy obtained overall speedup and partial speedup factors of 6% and 20%, respectively. Furthermore, the scalability of the hybrid strategy is very fine – its performance will be enhanced when the problem scale is increasing. However, this strategy also has its limitation that the performance enhancement will be decreasing if the ratio(the number of CPU cores divides that of the GPU cards) reduces. Finally, in our hybrid strategy, the CPU coefficient of utilization improved by 17.14% or better. Copyright © 2013 John Wiley & Sons, Ltd. Yueqing Wang, Yong Dou, Song Guo 0003, Yuanwu Lei, Dan Zou |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | Supernodal sparse Cholesky factorization on graphics processing unitsabstractSUMMARY Sparse Cholesky factorization is the most computationally intensive component in solving large sparse linear systems and is the core algorithm of numerous scientific computing applications. A large number of sparse Cholesky factorization algorithms have previously emerged, exploiting architectural features for various computing platforms. The recent use of graphics processing units (GPUs) to accelerate structured parallel applications shows the potential to achieve significant acceleration relative to desktop performance. However, sparse Cholesky factorization has not been explored sufficiently because of the complexity involved in its efficient implementation and the concerns of low GPU utilization. In this paper, we present a new approach for sparse Cholesky factorization on GPUs. We present the organization of the sparse matrix supernode data structure for GPU and propose a queue‐based approach for the generation and scheduling of GPU tasks with dense linear algebraic operations. We also design a subtree‐based parallel method for multi‐GPU system. These approaches increase GPU utilization, thus resulting in substantial computational time reduction. Comparisons are made with the existing parallel solvers by using problems arising from practical applications. The experiment results show that the proposed approaches can substantially improve sparse Cholesky factorization performance on GPUs. Relative to a highly optimized parallel algorithm on a 12‐core node, we were able to obtain speedups in the range 1.59× to 2.31× by using one GPU and 1.80× to 3.21× by using two GPUs. Relative to a state‐of‐the‐art solver based on supernodal method for CPU‐GPU heterogeneous platform, we were able to obtain speedups in the range 1.52× to 2.30× by using one GPU and 2.15× to 2.76× by using two GPUs. Concurrency and Computation: Practice and Experience, 2013. Copyright © 2013 John Wiley & Sons, Ltd. Dan Zou, Yong Dou, Song Guo 0003, Rongchun Li |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | CuSora: Real-time software radio using multi-core graphics processing unit
Rongchun Li, Yong Dou, Jie Zhou 0007 |
J. Syst. Archit. | 2 |
| 2014 | FPGA Implementation of a Special-Purpose VLIW Structure for Double-Precision Elementary FunctionabstractIn the current article, the capability and flexibility of field programmable gate-arrays (FPGAs) to implement IEEE-754 double-precision floating-point elementary functions are explored. To perform various elementary functions on the unified hardware efficiently, we propose a special-purpose very long instruction word (VLIW) processor, called DP_VELP. This processor is equipped with multiple basic units, and its performance is improved through an explicitly parallel technique. Pipelined evaluation of polynomial approximation with Estrin's scheme is proposed, by scheduling basic components in an optimal order to avoid data hazard stalls and achieve minimal latency. The custom VLIW processor can achieve high scalability. Under the control of specific VLIW instructions, the basic units are combined into special-purpose hardware for elementary functions. Common elementary functions are presented as examples to illustrate the design of elementary function in DP_VELP in detail. Minimax approximation scheme is used to reduce degree of polynomial. Compromise between the size of lookup table and the latency is discussed, and the internal precision is carefully planned to guarantee accuracy of the result. Finally, we create a prototype of the DP_VELP unit and an FPGA accelerator based on the DP_VELP unit on a Xilinx XC6VLX760 FPGA chip to implement the SGP4/SDP4 application. Compared with previous researches, the proposed design can achieve low latency with a reasonable amount of resources and evaluate a variety of elementary functions with the unified hardware to satisfy the demands in scientific applications. Experimental results show that the proposed design guarantees more than 99% of correct rounding. Moreover, the SGP4/SDP4 accelerator, which is equipped with 39 DP_VELP units and runs at 200 MHz, outperforms the parallel software approach with hyper-thread technology on an Intel Xeon Quad E5620 CPU at 2.40 GHz by a factor of 7X. Yuanwu Lei, Lei Guo 0029, Yong Dou, Sheng Ma, Jinbo Xu |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2013 | A fully parallel truncated Viterbi decoder for Software Defined Radio on GPUsabstractSoftware-Defined Radio (SDR) on the Graphics Processing Unit (GPU) platform. We exploit a map-reduce strategy based on the three-point Viterbi decoding algorithm (TVDA) due to the high parallelization potential. The trellis of Viterbi decoding algorithm can be divided into sub-trellises in truncation, which can perform independent forward metrics computing and trace-back procedure in parallel. The parallel Viterbi decoding algorithm is mapped on a GPU named NVIDIA GTX580. The experiment shows that our method shows low BER and 36.0x speedup over a C implementation on a CPU with the frequency of 2.0GHz. At the meantime, our method achieves a performance improvement of 1.2x-3.6x times that of the existing GPU-based implementation. Rongchun Li, Yong Dou |
WCNC | 2 |
| 2013 | VLIW coprocessor for IEEE-754 quadruple-precision elementary functionsabstractIn this article, a unified VLIW coprocessor, based on a common group of atomic operation units, for Quad arithmetic and elementary functions (QP_VELP) is presented. The explicitly parallel scheme of VLIW instruction and Estrin's evaluation scheme for polynomials are used to improve the performance. A two-level VLIW instruction RAM scheme is introduced to achieve high scalability and customizability, even for more complex key program kernels. Finally, the Quad arithmetic accelerator (QAA) with the QP_VELP array is implemented on ASIC. Compared with hyper-thread software implementation on an Intel Xeon E5620, QAA with 8 QP_VELP units achieves improvement by a factor of 18X. Yuanwu Lei, Yong Dou, Lei Guo 0029, Jinbo Xu, Jie Zhou 0007, Yazhuo Dong |
ACM Trans. Archit. Code Optim. | 2 |
| 2013 | FPGA implementation of an exact dot product and its application in variable-precision floating-point arithmetic
Yuanwu Lei, Yong Dou, Yazhuo Dong, Jie Zhou 0007, Fei Xia 0003 |
J. Supercomput. | 2 |
| 2012 | A self-organizing and self-adaptive French flag Organism based on lateral activation modelabstractOrganisms possess an amazing self-organizing and self-adaptive abilities compared with the man-made system. These abilities emit a strong attraction to engineers engaged in reliable electronic system design. In the current paper, a bio-inspired approach based on a lateral activation model from biological pattern formation theory is used to design the genotype of a French flag Organism. This lateral activation model represents the intercellular and intracellular interaction of biochemical substances. The designed French flag organism not only has the ability to form arbitrary sizes or a French flag pattern, but also exhibit powerful individual adaptability which is different from the adaptability that can be derived from species evolution. Moreover, this French flag organism shows wonderful dissociated-and-reaggregated capacity as hydrae. The current paper shows that designing reliable electronic systems that are similar to an organism is entirely possible through bio-inspired approach, and provides a new avenue for future reliable electronic systems design. Jie Zhou 0007, Yong Dou |
IEEE Congress on Evolutionary Computation | 4 |
| 2012 | Parallelizing sparse LU decomposition on FPGAsabstractSparse LU decomposition is the core computation in the direct method that solves sparse systems of linear equations. Only little work has been conducted on parallelizing it on FPGAs. In this paper, we study parallelization strategies for sparse LU decomposition on FPGAs. We first analyze how to parallelize the right-looking algorithm and find that this algorithm is not suitable for FPGAs. Then the left-looking algorithm is analyzed and considered as better candidate than the right-looking version. Our design derived from the left-looking algorithm is based on a simple yet efficient parallel computational model for FPGAs. Our design mainly consists of multiple parallel processing elements (PEs). A total of 14 PEs can be integrated into a Xilinx Virtex-5 XC5VLX330. Unlike related work, where their designs are applied to sparse matrices from particular application domains, our hardware design can be applied to any symmetric positive definite or diagonally dominant matrices. Guiming Wu, Xianghui Xie 0001, Yong Dou, Junqing Sun, Yuan Li 0011 |
FPT | 3 |
| 2012 | Optimization schemes and performance evaluation of Smith-Waterman algorithm on CPU, GPU and FPGAabstractSUMMARY With fierce competition between CPU and graphics processing unit (GPU) platforms, performance evaluation has become the focus of various sectors. In this paper, we take a well‐known algorithm in the field of biosequence matching and database searching, the Smith–Waterman (S‐W) algorithm as an example, and demonstrate approaches that fully exploit its performance potentials on CPU, GPU, and field‐programmable gate array (FPGA) computing platforms. For CPU platforms, we perform two optimizations, single instruction, multiple data and multithread, with compiler options, to gain over 70 × speedups over naive CPU versions on quad‐core CPU platforms. For GPU platforms, we propose the combination of coalesced global memory accesses, shared memory tiles, and loop unfolding, achieving 50 × speedups over initial GPU versions on an NVIDIA GeForce GTX 470 card. Experimental results show that the GPU GTX 470 gains 12 × speedups, instead of 100 × reported by some studies, over Intel quadcore CPU Q9400, under the same manufacturing technology and both with fully optimized schemes. In addition, for FPGA platforms, we customize a linear systolic array for the S‐W algorithm in a 45‐nm FPGA chip from Xilinx (XC6VLX760), with up to 1024 processing elements. Under only 133 MHz clock rate, the FPGA platform reaches the highest performance and becomes the most power‐efficient platform, using only 25 W compared with 190 W of the GPU GTX 470. Copyright © 2011 John Wiley & Sons, Ltd. Yong Dou, Fei Xia 0003 |
Concurr. Comput. Pract. Exp. | 2 |
| 2012 | A High Performance and Memory Efficient LU Decomposer on FPGAsabstractLU decomposition for dense matrices is an important linear algebra kernel that is widely used in both scientific and engineering applications. To efficiently perform large matrix LU decomposition on FPGAs with limited local memory, a block LU decomposition algorithm on FPGAs applicable to arbitrary matrix size is proposed. Our algorithm applies a series of transformations, including loop blocking and space-time mapping, onto sequential nonblocking LU decomposition. We also introduce a high performance and memory efficient hardware architecture, which mainly consists of a linear array of processing elements (PEs), to implement our block LU decomposition algorithm. Our design can achieve optimum performance under various hardware resource constraints. Furthermore, our algorithm and design can be easily extended to the multi-FPGA platform by using a block-cyclic data distribution and inter-FPGA communication scheme. A total of 36 PEs can be integrated into a Xilinx Virtex-5 XC5VLX330 FPGA on our self-designed PCI-Express card, reaching a sustained performance of 8.50 GFLOPS at 133 MHz for a matrix size of 16,384, which outperforms several general-purpose processors. For a Xilinx Virtex-6 XC6VLX760, a newer FPGA, we predict that a total of 180 PEs can be integrated, reaching 70.66 GFLOPS at 200 MHz. Compared to the previous work, our design can integrate twice the number of PEs into the same FPGA and has significantly higher performance. Guiming Wu, Yong Dou, Junqing Sun, Gregory D. Peterson |
IEEE Trans. Computers | 2 |
| 2012 | The unified accelerator architecture for RNA secondary structure prediction on FPGA
Fei Xia 0003, Yong Dou, Guoqing Jin |
J. Supercomput. | 2 |
| 2011 | FPGA Implementation of Variable-Precision Floating-Point Arithmetic
Yuanwu Lei, Yong Dou, Song Guo 0003, Jie Zhou 0007 |
APPT | 2 |
| 2011 | VPFPAP: A Special-Purpose VLIW Processor for Variable-Precision Floating-Point ArithmeticabstractMany scientific computing applications require efficient variable-precision floating-point arithmetic. This paper presents a special-purpose Variable-Precision Floating-Point Arithmetic Processor (VPFPAP) based on Very Large Instruction Word (VLIW) structure. The proposed processor uses a unified hardware structure, equipped with multiple custom basic variable-precision arithmetic units, to implement various variable-precision algebraic and transcendental functions. The performance is improved by the explicitly parallel technology of VLIW instruction and by dynamically varying the precision of intermediate computation. Finally, we create a prototype of VPFPAP unit into a Xilinx Virtex-6 XC6VLX760-2FF1760 FPGA chip. The experimental results show that our design, based on FPGA running at 253 MHz, outperforms the approach of a software-based library running on an Intel Core i3 530 CPU at 2.93 GHz by a factor of 5-38X for basic variable precision arithmetic operations and elementary functions. Yuanwu Lei, Yong Dou, Jie Zhou 0007, Sufeng Wang |
FPL | 2 |
| 2011 | Special-purposed VLIW architecture for IEEE-754 quadruple precision elementary functions on FPGAabstractThis work explores the feasibility to implement IEEE-754-2008 standard quadruple precision (Quad) elementary functions on recent FPGAs with plenty of embedded memories and DSP blocks. First, we analysis the implementation algorithm of Quad elementary functions in detail. Then, we present a special-purpose Very Large Instruction Word (VLIW) architecture for Quad elementary function (QE-Processor). The proposed processor uses a unified hardware structure, equipped with multiple basic arithmetic units, to implement various Quad algebraic and transcendental functions, in which several tradeoffs between latency and resource usage are carefully planned to avoid unbalanced resource utilization. The performance is improved through the explicitly parallel technology of custom VLIW instruction. Finally, we create a prototype of QE-Processor into Xilinx Virtex-5 and Virtex-6 FPGA chips. The experimental results show that our design can guarantee that the percentage of correct rounding is more than 99.9%. Moreover, the FPGA implementation on Virtex-6 XC6VLX760-2FF1760 FPGA, running at 220 MHz, outperforms the parallel software approach based on OpenMP running on an Intel Xeon E5620 CPU at 2.40GHz by a factor of 13X-20X for special function applications in Boost library. Yuanwu Lei, Yong Dou, Jie Zhou 0007, Song Guo 0003 |
ICCD | 2 |
| 2011 | FPGA accelerator for protein secondary structure prediction based on the GOR algorithmabstractBACKGROUND: Protein is an important molecule that performs a wide range of functions in biological systems. Recently, the protein folding attracts much more attention since the function of protein can be generally derived from its molecular structure. The GOR algorithm is one of the most successful computational methods and has been widely used as an efficient analysis tool to predict secondary structure from protein sequence. However, the execution time is still intolerable with the steep growth in protein database. Recently, FPGA chips have emerged as one promising application accelerator to accelerate bioinformatics algorithms by exploiting fine-grained custom design. RESULTS: In this paper, we propose a complete fine-grained parallel hardware implementation on FPGA to accelerate the GOR-IV package for 2D protein structure prediction. To improve computing efficiency, we partition the parameter table into small segments and access them in parallel. We aggressively exploit data reuse schemes to minimize the need for loading data from external memory. The whole computation structure is carefully pipelined to overlap the sequence loading, computing and back-writing operations as much as possible. We implemented a complete GOR desktop system based on an FPGA chip XC5VLX330. CONCLUSIONS: The experimental results show a speedup factor of more than 430x over the original GOR-IV version and 110x speedup over the optimized version with multi-thread SIMD implementation running on a PC platform with AMD Phenom 9650 Quad CPU for 2D protein structure prediction. However, the power consumption is only about 30% of that of current general-propose CPUs. Fei Xia 0003, Yong Dou, Guo-Qing Lei, Yusong Tan |
BMC Bioinform. | 2 |
| 2010 | Blocking LU Decomposition for FPGAsabstractTo efficiently perform large matrix LU decomposition on FPGAs with limited local memory, the original algorithm needs to be blocked. In this paper, we propose a block LU decomposition algorithm for FPGAs, which is applicable for matrices of arbitrary size. We introduce a high performance hardware design, which mainly consists of a linear array of processing elements (PEs), to implement our block LU decomposition algorithm. A total of 36 PEs can be integrated into a Xilinx Virtex-5 xc5vlx330 FPGA on our self-designed PCI-Express card, reaching a sustained performance of 8.50 GFLOPS at 133 MHz, which outperforms previous work. Guiming Wu, Yong Dou, Gregory D. Peterson |
FCCM | 2 |
| 2010 | High performance and memory efficient implementation of matrix multiplication on FPGAsabstractWe present a high performance and memory efficient hardware implementation of matrix multiplication for dense matrices of any size on the FPGA devices. By applying a series of transformations and optimizations on the original serial algorithm, we can obtain an I/O and memory optimized block algorithm for matrix multiplication on FPGAs. A linear array of processing elements (PEs) is proposed to implement this block algorithm. We show significant reduction in hardware resources consuming compared to the related work while increasing clock frequency. Moreover, the memory requirement can be reduced to O(S) from O(S2), where S is the block size. Therefore, more PEs can be integrated into the same FPGA devices. Guiming Wu, Yong Dou |
FPT | 2 |
| 2010 | Automatic synthesis of processor arrays with local memories on FPGAsabstractIn this paper, we present an automatic synthesis framework to map loop nests to processor arrays with local memories on FPGAs. An affine transformation approach is firstly proposed to address space-time mapping problem. Then a data-driven architecture model is introduced to enable automatic generation of processor arrays by extracting this data-driven architecture model from transformed loop nests. Some techniques including memory allocation, communication generation and control generation are presented. Synthesizable RTL codes can be easily generated from the architecture model built by these techniques. A preliminary synthesis tool is implemented based on PLUTO, an automatic polyhedral source-to-source transformation and parallelization framework. Guiming Wu, Yong Dou |
FPT | 2 |
| 2010 | FPGA accelerating double/quad-double high precision floating-point applications for ExaScale computingabstractIn this paper we explore the capability and flexibility of FPGA solutions in a sense to accelerate scientific computing applications which require very high precision arithmetic, based on 128-bit or even 256-bit floating-point number representation. Yong Dou, Yuanwu Lei, Guiming Wu, Song Guo 0003, Jie Zhou 0007 |
ICS | 1 |
| 2010 | A Unified Co-Processor Architecture for Matrix Decomposition
Yong Dou, Jie Zhou 0007, Guiming Wu, Jingfei Jiang, Yuanwu Lei, Shi-Ce Ni |
J. Comput. Sci. Technol. | 1 |
| 2010 | Fine-grained parallel RNA secondary structure prediction using SCFGs on FPGA
Fei Xia 0003, Yong Dou |
Parallel Comput. | 2 |
| 2009 | Implementation of Rotation Invariant Multi-View Face Detection on FPGA
Jinbo Xu, Yong Dou, Yuxing Tang, Xiaodong Wang 0002 |
APPT | 2 |
| 2009 | A Fine-Grained Pipelined Implementation for Large-Scale Matrix Inversion on FPGA
Jie Zhou 0007, Yong Dou, Jianxun Zhao, Fei Xia 0003, Yuanwu Lei, Yuxing Tang |
APPT | 2 |
| 2009 | Fine-grained parallel application specific computing for RNA secondary structure prediction using SCFGS on FPGAabstractIn the field of RNA secondary structure prediction, the CYK (Coche-Younger-Kasami) algorithm is a most popular methods using SCFG (stochastic context-free grammars) model. However, general purpose parallel computers including SMP multiprocessors or cluster systems exhibit low parallel efficiency and they are too expensive to be used easily for many research institutes. FPGA chips provide a new approach to accelerate the CYK algorithm by exploiting fine-grained custom design. The CYK algorithm shows complicated data dependence, in which the dependence distance is variable, and the dependence direction is also across two dimensions. We propose a systolic array structure including one master PE and multiple slave PEs for fine grain hardware implementation on FPGA. We partition tasks by columns and assign tasks to PEs for load balance. We exploit data reuse schemes to reduce the need to load matrix from external memory. To our knowledge, our implementation with 16 PEs is the only FPGA accelerator implementing the complete CYK/inside algorithm. The experimental results show a factor of more than 14 speedup over the Infernal-0.55 software running on a PC platform with Pentium 4 2.66GHz CPU. The computational power of our platform with FPGA accelerator is comparable to a PC cluster consisting of 20 Intel-Xeon CPUs for RNA secondary structure prediction using SCFGs, but the hardware cost and power consumption is only about 15% and 10% of the latter respectively. Yong Dou, Fei Xia 0003, Jingfei Jiang |
CASES | 1 |
| 2009 | A Fine-grained Pipelined Implementation of the LINPACK Benchmark on FPGAsabstractPrevious works have projected that the peak performance of FPGAs can outperform that of the general purpose processors. However, no work actually compares the performance between FPGAs and CPUs using the standard benchmarks such as the LINPACK benchmark. We propose and implement an FPGA-based hardware design of the LINPACK benchmark, the key step of which is LU decomposition with pivoting. We introduce a fine-grained pipelined LU decomposition algorithm that enables optimum performance by exploiting fine-grained pipeline parallelism. A scalable linear array of processing elements (PEs), which is the core component of our hardware design, is proposed to implement this algorithm. To the best of our knowledge, this is the first reported FPGA-based pipelined implementation of LU decomposition with pivoting. A total of 19 PEs can be integrated into an Altera Stratix II EP2S130F1020C5 on our self-designed development board. Experimental results show that the speedup up to 6.14 can be achieved relative to a Pentium 4 processor for the LINPACK benchmark. Guiming Wu, Yong Dou, Yuanwu Lei, Jie Zhou 0007, Jingfei Jiang |
FCCM | 2 |
| 2009 | FPGA accelerating three QR decomposition algorithms in the unified pipelined frameworkabstractMany FPGA implementations for QR decomposition have been studied on small-scale matrix and all of them are presented individually. However to the best of our knowledge, there is no FPGA-based accelerator for large-scale QR decomposition. In this paper, we propose a unified FPGA accelerator structure for large-scale QR decomposition. To exploit the computational potential of FPGA, we introduce a fine-grained parallel algorithm for QR decomposition. A scalable linear array processing elements (PEs), which is the core component of the FPGA accelerator, is proposed to implement this algorithm. A total of 15 PEs can integrated into an Altera StratixII EP2S130F1020C5 on our self-designed board. Experimental results show that a factor of 4 speedup and the maximum powerperformance of 60.9 can be achieved compare to Pentium Dual CPU with double SSE thread. Yong Dou, Jie Zhou 0007, Yuanwu Lei, Jinbo Xu |
FPL | 1 |
| 2009 | Fine-grained parallel RNAalifold algorithm for RNA secondary structure prediction on FPGAabstractBACKGROUND: In the field of RNA secondary structure prediction, the RNAalifold algorithm is one of the most popular methods using free energy minimization. However, general-purpose computers including parallel computers or multi-core computers exhibit parallel efficiency of no more than 50%. Field Programmable Gate-Array (FPGA) chips provide a new approach to accelerate RNAalifold by exploiting fine-grained custom design. RESULTS: RNAalifold shows complicated data dependences, in which the dependence distance is variable, and the dependence direction is also across two dimensions. We propose a systolic array structure including one master Processing Element (PE) and multiple slave PEs for fine grain hardware implementation on FPGA. We exploit data reuse schemes to reduce the need to load energy matrices from external memory. We also propose several methods to reduce energy table parameter size by 80%. CONCLUSION: To our knowledge, our implementation with 16 PEs is the only FPGA accelerator implementing the complete RNAalifold algorithm. The experimental results show a factor of 12.2 speedup over the RNAalifold (ViennaPackage - 1.6.5) software for a group of aligned RNA sequences with 2981-residue running on a Personal Computer (PC) platform with Pentium 4 2.6 GHz CPU. Fei Xia 0003, Yong Dou, Xingming Zhou, Xuejun Yang |
BMC Bioinform. | 2 |
| 2009 | A coarse-grained reconfigurable computing architecture with loop self-pipelining
Yong Dou, Guiming Wu, Jinhui Xu 0002, Xingming Zhou |
Sci. China Ser. F Inf. Sci. | 1 |
| 2008 | Collaborative hardware/software partition of coarse-grained reconfigurable system using evolutionary ant colony optimizationabstractThe flexibility, performance and cost effectiveness of reconfigurable architectures have lead to its widespread use for embedded applications. Coarse-grained reconfigurable system design is very complex for multi-fields experts to collaborate on application algorithm design, hardware/software co-design and system decision. However, existing reconfigurable system design methods and environments can only support hardware/software co-design, ignoring the collaboration between multi-field experts. This paper presents a collaborative partition approach of coarse-grained reconfigurable system design using evolutionary ant colony optimization. We create a distributed collaborative design environment for system decision engineers, software designers, hardware designers and application algorithm developers. The method not only utilizes the advantages of ant colony optimization for searching global optimal solutions, but also provides a framework for multi-field experts to work collaboratively. Experimental results show that the method improves the quality and speed of hardware/software partition for coarse-grained reconfigurable system design. Dawei Wang 0020, Sikun Li, Yong Dou |
ASP-DAC | 3 |
| 2008 | Dimensional Bubble Flow Control and Fully Adaptive Routing in the 2-D Mesh Network on ChipabstractIn this paper, the novel flow control strategy called dimensional bubble flow control (DBFC) is presented. The flow control strategy of DBFC builds on virtual cut-through switching and credit-based flow control mechanism and analyzes the credit value of port and the routing information of the packets to realize the point-point flow control. In the 2-D mesh network on chip, when the flow control strategy of DBFC is accepted, the adaptive dimensional bubble routing (ADBR) algorithm designed in this paper can get the goals including deadlock-free and minimal distance even if the cyclic dependencies exist. In this paper, the detail proof is provided for these conclusions. Lastly, we adapt the source code of NOXIM that is a popular simulator of on-chip networks and realize the flow control of DBFC and ADBR algorithm in NOXIM. We test the performance of ADBR on NOXIM. The simulation performance shows our scheme is superior to the usual approach such as XY dimension-order routing, with nearly 17.5% improvement in the packets latency and throughput. Canwen Xiao, Minxuan Zhang, Yong Dou, Zhitong Zhao |
EUC (1) | 3 |
| 2008 | Double Precision Hybrid-Mode Floating-Point FPGA CORDIC Co-processorabstractFPGA chips have become a promising option for accelerating scientific applications, which involve many floating-point transcendental functions, such as sin, log, exp, sqrt and etc. In this paper, we present a 64-bit ANSI/IEEE floating-point CORDIC co-processor on FPGA, providing all known CORDIC functions. And there is no 64-bit CORDIC implementation on FPGA known to us. We propose a hybrid-mode CORDIC algorithm, combining hybrid rotation angle methods with argument reduction algorithm to reduce hardware area usage and meanwhile keep unlimited convergence domain for any floating-point inputs of the functions. Our hybrid-mode CORDIC co-processor is organized into three phases, argument reduction, CORDIC calculation and normalization with 69 pipeline stages for FPGA implementation. The synthesis results show the clock frequency can reach 173 MHz on Xilinx Virtex5 FPGA. Comparing to general-purpose microprocessor in three scientific program kernels, the CORDIC co-processor can achieve a maximum speedup of 49.3 times, 28.7 times in average. Jie Zhou 0007, Yong Dou, Yuanwu Lei, Jinbo Xu, Yazhuo Dong |
HPCC | 2 |
| 2008 | Fine-grained parallel application specific computing for RNA secondary structure prediction on FPGAabstractIn the field of RNA secondary structure prediction, the Zuker algorithm is one of the most popular methods using free energy minimization. However, general-purpose computers including parallel computers or multi-core computers exhibit parallel efficiency of no more than 50% on Zuker. FPGA chips provide a new approach to accelerate the Zuker algorithm by exploiting fine-grained custom design. Zuker shows complicated data dependences, in which the dependence distance is variable, and the dependence direction is also across two dimensions. We propose a systolic array structure including one master PE and multiple slave PEs for fine grain hardware implementation on FPGA. We exploit data reuse schemes to reduce the need to load energy matrices from external memory. We also propose several methods to reduce energy table parameter size by 85%. To our knowledge, our implementation with 16 PEs is the only FPGA accelerator implementing the complete Zuker algorithm. The experimental results show a factor of 14 speedup over the ViennaRNA-1.6.5 software for 2981-residue RNA sequence running on a PC platform with Pentium 4 2.6 GHz CPU. Yong Dou, Fei Xia 0003, Xingming Zhou, Xuejun Yang |
ICCD | 1 |
| 2008 | DMA Performance Analysis and Multi-core Memory Optimization for SWIM Benchmark on the Cell ProcessorabstractThe Cell processor is a typical heterogeneous multi-core processor, which owns powerful computing capability. But we are facing the challenges of 'memory wall' in developing parallel applications, such as, limited capacity of local memory, limited memory bandwidth for multi-cores and the long latency for data communication. The DMA transfer mechanism is often used to hide the long latency and improve the effective usage of memory bandwidth. In the paper, we start with a series of DMA experimental tests in the context of the Cell processor architecture, and perform mathematical analysis to setup a unified formula on the average bandwidth of DMA by means of exponential fitting, which describes that SPE amount and DMA block size take main effects on DMA bandwidth in quantity. With the supports of the DMA performance formula, we perform 4 types of memory optimization in the process of parallelizing the SWIM benchmark program into a multi-core version. We take Sony PlayStation 3 (PS3) as our test-bed. For SWIM benchmark, with 6 SPE cores, we obtain over 13 times of speedup compared to single PPE, and 3.3 to 6.18 times to AMD and Intel CPU. Yong Dou, Jinhui Xu 0002 |
ISPA | 1 |
| 2007 | Reducing Storage Requirements in Accelerating Algorithm of Global BioSequence Alignment on FPGA
Fei Xia 0003, Yong Dou |
APPT | 2 |
| 2007 | FPGA SAR Processor with Window Memory AccessesabstractIn the paper, we present a design of FPGA SAR processor with four 1D FFT processing elements, double internal RAM buffers and double external SDRAM modules. Without traditional corner turn phase, we propose a data layout scheme mapping one row of logical matrix into a rectangular window in physical banks of SDRAM in order to increase the practical I/O throughout between SDRAM modules and SAR processing elements. In addition, we theoretically analyses the optimal window size to minimize the total number of opening/closing pages when performing 2D FFT by balancing the number of handling physical pages between row accesses and column accesses. The experimental results show our window layout approach achieves 650 MB/s of effective bandwidth, reaching nearly 82% of peak bandwidth, with 58.1% increases compared to traditional Corner Turn approaches. The proposed SAR processor has been implemented in an FPGA test-bed, outperforming related works in both of computing speed and image scale. Yong Dou, Jie Zhou 0007, Yuanwu Lei, Xingming Zhou |
ASAP | 1 |
| 2007 | A Parameterized Architecture Model in High Level Synthesis for Image Processing ApplicationsabstractMost image processing applications are computationally intensive and data intensive. Reconfigurable hardware boards provide a convenient and flexible solution to speed up these algorithms. To get a high performance design without going through the time-consuming hardware design process for each different algorithm, we present a universal parameterized architecture in high level synthesis to generate the hardware frames for all image processing applications automatically. The value of the parameters which decide the target architecture can be obtained from the compiler. The algorithm how to get these parameters is also discussed in this paper. Yazhuo Dong, Yong Dou |
ASP-DAC | 2 |
| 2007 | Distributed Collaborative Partition Method of Reconfigurable SoC Using Ant Colony OptimizationabstractThe flexibility, performance and cost effectiveness of reconfigurable architectures have lead to its widespread use for embedded applications. Reconfigurable SoC (system-on-a-chip) design is very complex for multi-fields experts to collaborate on application algorithm design, hardware/software co-design and system decision. However, existing reconfigurable SoC design methods and environments can only support hardware/software co-design, ignoring the collaboration between multi-field experts. This paper presents a distributed collaborative partition method of reconfigurable system design using ant colony optimization. We create a distributed collaborative design environment for multi-field experts. The communication protocol, maintenance of consistency and parallel control are discussed. The method not only utilizes the advantages of ant colony optimization for searching the global best optimized solutions, but also provides a framework for multi-field experts to work collaboratively. Experimental results show that our method effectively improves the quality and efficiency of hardware/software partition for reconfigurable system design. Sikun Li, Dawei Wang 0020, Yong Dou |
CSCWD | 4 |
| 2007 | FPGA Accelerating Algorithms of Active Shape Model in People Tracking ApplicationsabstractAlgorithms of Active Shape Model, as one of the most popular methods for recognizing non-rigid objects, require huge computation power for real time people tracking. After analyzing the parallel characteristics of the algorithm, we propose a deep pipelined structure for accelerating the Active Shape Model algorithm. The computing engine is organized into a deep pipeline network composing of multiple floating-point arithmetic units, including adders, multipliers, dividers and SQRT etc. In the optimization of the memory efficiency for loading random data in large images during the step of local search, we propose an on-chip buffer scheme to eliminate random accesses to off-chip memory. Experimental results show that our FPGA implementation achieves over 15 times of speedup compared with the software implementation in Pentium 4 computer. Jinbo Xu, Yong Dou, Xingming Zhou, Qiang Dou |
DSD | 2 |
| 2007 | FIDP: A Novel Architecture for Lifting-Based 2D DWT in JPEG2000
Baofeng Li, Yong Dou |
MMM (2) | 2 |
| 2006 | Robust and real-time automatic target recognition using partial hausdorff distance measure on reconfigurable hardwareabstractThis paper presents a high performance FPGA-based automatic target recognition system, which matches TV templates with an image efficiently. Theoretically, image matching algorithms based on partial Hausdorff distance (HD) are more tolerant of perturbations in the locations of pixel points than other algorithms, but they are too computationally expensive to be used in embedded systems. In order to solve this problem, we present a robust and real-time implementation of the image matching algorithm based on partial HD, taking advantage of the hardware resources offered by the FPGA chips. A parallel target recognition algorithm under constraints of limited embedded memory and limited memory bandwidth is proposed first. And then the system is organized as a coarse-grained pipeline containing three stages. Each stage is implemented in highly parallel fashion. The implementation of distance transform and template matching are described in detail. Experimental results show that our work outperforms related proposals. A speedup of almost 50 is achieved while compared with the software solution in PC (Pentium 4 2.8 GHz) Jinbo Xu, Yong Dou |
FPT | 2 |
| 2006 | Clustering Multicast on Hypercube Network
Baohua Fan, Yong Dou, Xiaodong Yang 0002 |
HPCC | 3 |
| 2006 | Progress and Challenges in High Performance Computer Technology
Xuejun Yang, Yong Dou, Qingfeng Hu |
J. Comput. Sci. Technol. | 2 |
| 2005 | 64-bit floating-point FPGA matrix multiplicationabstractWe introduce a 64-bit ANSI/IEEE Std 754-1985 floating point design of a hardware matrix multiplier optimized for FPGA implementations. A general block matrix multiplication algorithm, applicable for an arbitrary matrix size is proposed. The algorithm potentially enables optimum performance by exploiting the data locality and reusability incurred by the general matrix multiplication scheme and considering the limitations of the I/O bandwidth and the local storage volume. We implement a scalable linear array of processing elements (PE) supporting the proposed algorithm in the Xilinx Virtex II Pro technology. Synthesis results confirm a superior performance-area ratio compared to related recent works. Assuming the same FPGA chip, the same amount of local memory, and the same I/O bandwidth, our design outperforms related proposals by at least 1.7X and up to 18X consuming the least reconfigurable resources. A total of 39 PEs can be integrated into the xc2vp125-7 FPGA, reaching performance of, e.g., 15.6 GFLOPS with 1600 KB local memory and 400 MB/s external memory bandwidth. Yong Dou, Stamatis Vassiliadis, Georgi Kuzmanov, Georgi Gaydadjiev |
FPGA | 1 |