EDBT 2026 Demo / reviewers in the wild / expert
Jie Tang 0006
dblp:181/2702-6
· DBLP profile ↗
36ranked-venue papers
1as first author
22since 2021 · last 2026
0000-0002-6086-3559ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 17 since 2021Artificial intelligence and machine learning · 15 · 11 since 2021Systems, architecture and hardware · 8 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RASR: Retrieval-Augmented Super Resolution for Practical Reference-based Image Restoration
Shuning Xu, Xiangyu Chen 0006, Dell Zhang, Jiantao Zhou 0001, Jie Tang 0006, Gangshan Wu, Jie Liu 0040 |
ISCAS | 6 |
| 2026 | GLAD: Generative Language-Assisted Visual Tracking for Low-Semantic Templates
Xingyu Luo, Yidong Cai, Jie Liu 0040, Jie Tang 0006, Gangshan Wu, Limin Wang 0002 |
Int. J. Comput. Vis. | 4 |
| 2025 | In-the-wild Audio Spatialization with Flexible Text-guided LocalizationabstractBinaural audio enriches immersive experiences by enabling the perception of the spatial locations of sounding objects in AR, VR, and embodied AI applications. While existing audio spatialization methods can generally map any available monaural audio to binaural audio signals, they often lack the flexible and interactive control needed in complex multi-object user-interactive environments. To address this, we propose a Text-guided Audio Spatialization (TAS) framework that utilizes diverse text prompts and evaluates our model from unified generation and comprehension perspectives. Due to the limited availability of high-quality, large-scale stereo data, we construct the SpatialTAS dataset, which encompasses 376,000 simulated binaural audio samples to facilitate the training of our model. Our model learns binaural differences guided by 3D spatial location and relative position prompts, enhanced with flipped-channel audio. Experimental results show that our model can generate high quality binaural audios for various audio types on both simulated and real-recorded datasets. Besides, we establish an assessment model based on Llama-3.1-8B, which evaluates the semantic accuracy of spatial locations through a spatial reasoning task. Results demonstrate that by utilizing text prompts for flexible and interactive control, we can generate binaural audio with both high quality and semantic consistency in spatial locations. Tianrui Pan, Jie Liu 0040, Zewen Huang, Jie Tang 0006, Gangshan Wu |
ACL (1) | 4 |
| 2025 | CATANet: Efficient Content-Aware Token Aggregation for Lightweight Image Super-ResolutionabstractTransformer-based methods have demonstrated impressive performance in low-level visual tasks such as Image Super-Resolution (SR). However, its computational complexity grows quadratically with the spatial resolution. A series of works attempt to alleviate this problem by dividing Low-Resolution images into local windows, axial stripes, or dilated windows. SR typically leverages the redundancy of images for reconstruction, and this redundancy appears not only in local regions but also in long-range regions. However, these methods limit attention computation to content-agnostic local regions, limiting directly the ability of attention to capture long-range dependency. To address these issues, we propose a lightweight Content-Aware Token Aggregation Network (CATANet). Specifically, we propose an efficient Content-Aware Token Aggregation module for aggregating long-range content-similar tokens, which shares token centers across all image tokens and updates them only during the training phase. Then we utilize intra-group self-attention to enable long-range information interaction. Moreover, we design an inter-group cross-attention to further enhance global information interaction. The experimental results show that, compared with the state-of-the-art cluster-based method SPIN, our method achieves superior performance, with a maximum PSNR improvement of 0.33dB and nearly double the inference speed. Xin Liu 0012, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
CVPR | 3 |
| 2025 | AutoLUT: LUT-Based Image Super-Resolution with Automatic Sampling and Adaptive Residual LearningabstractIn recent years, the increasing popularity of Hi-DPI screens has driven a rising demand for high-resolution images. However, the limited computational power of edge devices poses a challenge in deploying complex super-resolution neural networks, highlighting the need for efficient methods. While prior works have made significant progress, they have not fully exploited pixel-level information. Moreover, their reliance on fixed sampling patterns limits both accuracy and the ability to capture fine details in low-resolution images. To address these challenges, we introduce two plug-and-play modules designed to capture and leverage pixel information effectively in Look-Up Table (LUT) based super-resolution networks. Our method introduces Automatic Sampling (AutoSample), a flexible LUT sampling approach where sampling weights are automatically learned during training to adapt to pixel variations and expand the receptive field without added inference cost. We also incorporate Adaptive Residual Learning (AdaRL) to enhance inter-layer connections, enabling detailed information flow and improving the network’s ability to reconstruct fine details. Our method achieves significant performance improvements on both MuLUT and SPF-LUT while maintaining similar storage sizes. Specifically, for MuLUT, we achieve a PSNR improvement of approximately +0.20 dB improvement on average across five datasets. For SPF-LUT, with more than a 50% reduction in storage space and about a 2/3 reduction in inference time, our method still maintains performance comparable to the original. The code is available at https://github.com/SuperKenVery/AutoLUT. Yuheng Xu, Xin Liu 0012, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
CVPR | 5 |
| 2025 | Towards Practical Real-Time Low-Latency Music Source SeparationabstractIn recent years, significant progress has been made in the field of deep learning for music demixing. However, there has been limited attention on real-time, low-latency music demixing, which holds potential for various applications, such as hearing aids, audio stream remixing, and live performances. Additionally, a notable tendency has emerged towards the development of larger models, limiting their applicability in certain scenarios. In this paper, we introduce a lightweight real-time low-latency model called Real-Time Single-Path TFC-TDF UNET (RT-STT), which is based on the Dual-Path TFC-TDF UNET (DTTNet). In RT-STT, we propose a feature fusion technique based on channel expansion. We also demonstrate the superiority of single-path modeling over dual-path modeling in real-time models. Moreover, we investigate the method of quantization to further reduce inference time. RT-STT exhibits superior performance with significantly fewer parameters and shorter inference times compared to state-of-the-art models. Junyu Wu, Jie Liu 0040, Tianrui Pan, Jie Tang 0006, Gangshan Wu |
ICME | 4 |
| 2025 | FSDM: An efficient video super-resolution method based on Frames-Shift Diffusion Model
Chao Chen 0026, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
Neural Networks | 4 |
| 2024 | Sketch and Refine: Towards Fast and Accurate Lane DetectionabstractLane detection is to determine the precise location and shape of lanes on the road. Despite efforts made by current methods, it remains a challenging task due to the complexity of real-world scenarios. Existing approaches, whether proposal-based or keypoint-based, suffer from depicting lanes effectively and efficiently. Proposal-based methods detect lanes by distinguishing and regressing a collection of proposals in a streamlined top-down way, yet lack sufficient flexibility in lane representation. Keypoint-based methods, on the other hand, construct lanes flexibly from local descriptors, which typically entail complicated post-processing. In this paper, we present a “Sketch-and-Refine” paradigm that utilizes the merits of both keypoint-based and proposal-based methods. The motivation is that local directions of lanes are semantically simple and clear. At the “Sketch” stage, local directions of keypoints can be easily estimated by fast convolutional layers. Then we can build a set of lane proposals accordingly with moderate accuracy. At the “Refine” stage, we further optimize these proposals via a novel Lane Segment Association Module (LSAM), which allows adaptive lane segment adjustment. Last but not least, we propose multi-level feature integration to enrich lane feature representations more efficiently. Based on the proposed “Sketch-and-Refine” paradigm, we propose a fast yet effective lane detector dubbed “SRLane”. Experiments show that our SRLane can run at a fast speed (i.e., 278 FPS) while yielding an F1 score of 78.9%. The source code is available at: https://github.com/passerer/SRLane. Chao Chen 0026, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
AAAI | 4 |
| 2024 | GTPT: Group-Based Token Pruning Transformer for Efficient Human Pose Estimation
Jie Liu 0040, Jie Tang 0006, Gangshan Wu, Yanbing Chou |
ECCV (69) | 3 |
| 2024 | RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual CuesabstractWhile existing Audio-Visual Speech Separation (AVSS) methods primarily concentrate on the audio-visual fusion strategy for two-speaker separation, they demonstrate a severe performance drop in the multi-speaker separation scenarios. Typically, AVSS methods employ guiding videos to sequentially isolate individual speakers from the given audio mixture, resulting in notable missing and noisy parts across various segments of the separated speech. In this study, we propose a simultaneous multi-speaker separation framework that can facilitate the concurrent separation of multiple speakers within a singular process. We introduce speaker-wise interactions to establish distinctions and correlations among speakers. Experimental results on the VoxCeleb2 and LRS3 datasets demonstrate that our method achieves state-of-the-art performance in separating mixtures with 2, 3, 4, and 5 speakers, respectively. Additionally, our model can utilize speakers with complete audio-visual information to mitigate other visual-deficient speakers, thereby enhancing its resilience to missing visual cues. We also conduct experiments where visual information for specific speakers is entirely absent or visual frames are partially missing. The results demonstrate that our model consistently outperforms others, exhibiting the smallest performance drop across all settings involving 2, 3, 4, and 5 speakers. Tianrui Pan, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
ACM Multimedia | 4 |
| 2023 | From Coarse to Fine: Hierarchical Pixel Integration for Lightweight Image Super-resolutionabstractImage super-resolution (SR) serves as a fundamental tool for the processing and transmission of multimedia data. Recently, Transformer-based models have achieved competitive performances in image SR. They divide images into fixed-size patches and apply self-attention on these patches to model long-range dependencies among pixels. However, this architecture design is originated for high-level vision tasks, which lacks design guideline from SR knowledge. In this paper, we aim to design a new attention block whose insights are from the interpretation of Local Attribution Map (LAM) for SR networks. Specifically, LAM presents a hierarchical importance map where the most important pixels are located in a fine area of a patch and some less important pixels are spread in a coarse area of the whole image. To access pixels in the coarse area, instead of using a very large patch size, we propose a lightweight Global Pixel Access (GPA) module that applies cross-attention with the most similar patch in an image. In the fine area, we use an Intra-Patch Self-Attention (IPSA) module to model long-range pixel dependencies in a local patch, and then a spatial convolution is applied to process the finest details. In addition, a Cascaded Patch Division (CPD) strategy is proposed to enhance perceptual quality of recovered images. Extensive experiments suggest that our method outperforms state-of-the-art lightweight SR methods by a large margin. Code is available at https://github.com/passerer/HPINet. Jie Liu 0040, Chao Chen 0026, Jie Tang 0006, Gangshan Wu |
AAAI | 3 |
| 2023 | CPD-GAN: Cascaded Pyramid Deformation GAN for Pose TransferabstractPose-guided person image generation aims to synthesize person images in arbitrary poses. This task requires to perform spatial deformation on source images. Existing work often failed to transfer complex textures to generated images well. To solve this problem, we propose a novel network for this task. The network is called Cascaded Pyramid Deformation GAN(CPD-GAN), which can achieve more realistic results and conform more to the target person. In the extraction sub-network, the multi-scale feature modulation(MFM) blocks are proposed. A MFM block can fuse features at different scales into single scale feature. And in the genration sub-network, the dynamic transferring fusion(DTF) blocks are proposed to perform dynamic deformation on features from extraction sub-network and a cascaded pyramids structure is adopted to improve the quality of resulted complex textures. Experiments prove that our method performs better on pose transfer than other methods and utilizes each part of the network effectively. Yuting Tang, Xiu Zheng, Jie Tang 0006 |
ICASSP | 4 |
| 2023 | Reliable Cluster-Based Framework for Open Set Domain AdaptationabstractExisting Domain Adaptation (DA) has been able to transfer knowledge from a labeled source domain to an unlabeled target domain nicely. However, in a more complex Open Set Domain Adaptation (OSDA) setting containing unknown target domain categories, previous DA methods fail or even negatively transfer. Recently, various OSDA methods have been proposed and achieved good results. Among them, cluster-based methods such as Domain Consensus Clustering have a higher upper boundary because they mine the intrinsic structure of the target domain instead of treating all the unknown classes of the target domain as one category. However, the impact of faulty pseudo-labels suppresses the performance. Instead, we propose a Reliable Cluster-based Framework (RCF), including Coarse Target Clustering, Structured Matching Strategy, and Reliable Pseudo-Label Training modules, as a general framework to solve the impact of faulty pseudo-labels for cluster-based OSDA methods. Experiments on extensive benchmarks demonstrate that RCF significantly outperforms previous state-of-the-arts. Xiu Zheng, Jie Tang 0006 |
ICASSP | 3 |
| 2023 | Robust Object Modeling for Visual TrackingabstractObject modeling has become a core part of recent tracking frameworks. Current popular tackers use Transformer attention to extract the template feature separately or interactively with the search region. However, separate template learning lacks communication between the template and search regions, which brings difficulty in extracting discriminative target-oriented features. On the other hand, interactive template learning produces hybrid template features, which may introduce potential distractors to the template via the cluttered search regions. To enjoy the merits of both methods, we propose a robust object modeling framework for visual tracking (ROMTrack), which simultaneously models the inherent template and the hybrid template features. As a result, harmful distractors can be suppressed by combining the inherent features of target objects with search regions’ guidance. Target-related features can also be extracted using the hybrid template, thus resulting in a more robust object modeling framework. To further enhance robustness, we present novel variation tokens to depict the ever-changing appearance of target objects. Variation tokens are adaptable to object deformation and appearance variations, which can boost overall performance with negligible computation. Experiments show that our ROMTrack sets a new state-of-the-art on multiple benchmarks. Yidong Cai, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
ICCV | 3 |
| 2023 | Video Frame Interpolation with Densely Queried Bilateral CorrelationabstractVideo Frame Interpolation (VFI) aims to synthesize non-existent intermediate frames between existent frames. Flow-based VFI algorithms estimate intermediate motion fields to warp the existent frames. Real-world motions' complexity and the reference frame's absence make motion estimation challenging. Many state-of-the-art approaches explicitly model the correlations between two neighboring frames for more accurate motion estimation. In common approaches, the receptive field of correlation modeling at higher resolution depends on the motion fields estimated beforehand. Such receptive field dependency makes common motion estimation approaches poor at coping with small and fast-moving objects. To better model correlations and to produce more accurate motion fields, we propose the Densely Queried Bilateral Correlation (DQBC) that gets rid of the receptive field dependency problem and thus is more friendly to small and fast-moving objects. The motion fields generated with the help of DQBC are further refined and up-sampled with context features. After the motion fields are fixed, a CNN-based SynthNet synthesizes the final interpolated frame. Experiments show that our approach enjoys higher accuracy and less inference time than the state-of-the-art. Source code is available at https://github.com/kinoud/DQBC. Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
IJCAI | 3 |
| 2023 | Lightweight Super-Resolution Head for Human Pose EstimationabstractHeatmap-based methods have become the mainstream method for pose estimation due to their superior performance. However, heatmap-based approaches suffer from significant quantization errors with downscale heatmaps, which result in limited performance and the detrimental effects of intermediate supervision. Previous heatmap-based methods relied heavily on additional post-processing to mitigate quantization errors. Some heatmap-based approaches improve the resolution of feature maps by using multiple costly upsampling layers to improve localization precision. To solve the above issues, we creatively view the backbone network as a degradation process and thus reformulate the heatmap prediction as a Super-Resolution (SR) task. We first propose the SR head, which predicts heatmaps with a spatial resolution higher than the input feature maps (or even consistent with the input image) by super-resolution, to effectively reduce the quantization error and the dependence on further post-processing. Besides, we propose SRPose to gradually recover the HR heatmaps from LR heatmaps and degraded features in a coarse-to-fine manner. To reduce the training difficulty of HR heatmaps, SRPose applies SR heads to supervise the intermediate features in each stage. In addition, the SR head is a lightweight and generic head that applies to top-down and bottom-up methods. Extensive experiments on the COCO, MPII, and CrowdPose datasets show that SRPose outperforms the corresponding heatmap-based approaches. Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
ACM Multimedia | 3 |
| 2022 | Two Strategies Toward Lightweight Image Super-ResolutionabstractRecent convolution neural networks (CNNs) have achieved remarkable success in lightweight image super-resolution (LISR). The goal of LISR is to restore more accurate details with less model capacity. However, we observe two phenomena in current micro-architectures, one is the lack of consistent learning ability of high-frequency components, the other is large residual problem which does harm to the stability of residual learning. To tackle the two issues, we propose two strategies, namely global-guided attention strategy (GGAS) and channel-wise scaling strategy (CWSS), which can significantly improve the performance of the state-of-the-arts with negligible overheads. Zongcai Du, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
ICASSP | 3 |
| 2022 | Pyramid Fusion Attention Network For Single Image Super-ResolutionabstractRecently, convolutional neural network (CNN) has made a mighty advance in image super-resolution (SR). Most recent models exploit attention mechanism (AM) to focus on high-frequency information. However, these methods exclusively consider interdependencies among channels or spatials, leading to equal treatment of channel-wise or spatial-wise features thus hindering the power of AM. In this paper, we propose a pyramid fusion attention network (PFAN) to tackle this problem. Specifically, a novel pyramid fusion attention (PFA) is developed where stacked residual blocks are employed to model the relationship between pixels among all channels, and pyramid fusion structure is adopted to expand receptive field. Besides, a progressive backward fusion strat-egy is introduced to make full use of hierarchical features, which are beneficial to obtaining more contextual representations. Comprehensive experiments demonstrate the superiority of our proposed PFAN against state-of-the-art methods. Zongcai Du, Jie Tang 0006, Gangshan Wu |
ICASSP | 4 |
| 2022 | Hierarchical Feature Aggregation Network for Deep Image CompressionabstractExisting CNN-based methods for image compression extract features through serially connected high-to-low (encoder) or low-to-high (decoder) resolution stages, leading to insufficient utilization of hierarchical features. To solve this problem, we present a hierarchical feature aggregation network (HFAN) for generating more informative latent representations. In detail, we propose two strategies, namely inter-stage feature aggregation and intra-stage feature aggregation. The inter-stage feature aggregation integrates multi-scale information thereby producing more contextual features. The intra-stage aggregation fuses features within the same stage to enrich representations of one specific resolution. Besides, we incorporate a lightweight pixel-wise attention mechanism to further enhance the discriminative ability of our network. Extensive experiments demonstrate that our HFAN achieves superior performance over state-of-the-art methods without a hyperprior variational autoencoder. Zongcai Du, Jie Tang 0006, Gangshan Wu |
ICASSP | 4 |
| 2022 | IAA-VSR: An iterative alignment algorithm for video super-resolution
Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
Appl. Intell. | 2 |
| 2021 | Learning Discriminative Features for Semi-Supervised Anomaly DetectionabstractAnomaly detection is the task of identifying unusual samples in data. Typically anomaly detection is defined on an unlabeled dataset that is assumed most of the samples are normal and others are anomalies. However, in industrial practice, one may have access to a part of annotated data. This gives us the potential for semi-supervised learning. In addition, existing methods assume all training data is normal and neglect the impact of a small number of anomalous samples. In this paper, we consolidate the model’s discriminative power by introducing a transfer learning scheme to anomaly detection, thereby the model suffers less perturbation caused by pollution. We also propose a novel loss function to further adapt to semi-supervised data scenario. We ensure that the contribution of pollution can be well suppressed and reach a harmonious balance in magnitude of loss/gradient between unlabeled and labeled samples. Experiments on three publicly available datasets show that our method achieves state-of-the-art results. Jie Tang 0006, Yishun Dou, Gangshan Wu |
ICASSP | 2 |
| 2021 | Lightweight Human Pose Estimation under Resource-Limited ScenesabstractRecent research on human pose estimation has achieved significant improvement. However, most existing methods tend to pursue higher scores on benchmark datasets using complex architecture, ignoring the deployment costs in practice. In this paper, we investigate the problem of lightweight human pose estimation under resource-limited scenes.We first redesign a lightweight bottleneck block with two concepts: depthwise convolution and attention mechanism. And then, based on the lightweight block, we present a single-stage Lightweight Pose Network (LPN). Our small network LPN-50 only has 2.7M parameters and 1.0G FLOPs, which is much more lightweight than other popular networks. In order to overcome the training barrier, we propose an iterative training strategy that can give full play to our LPNs’ potential to get more accurate predicted results. We empirically demonstrate the effectiveness and efficiency of our methods on the benchmark dataset: the COCO keypoint detection dataset. Besides, we show the speed superiority of our lightweight network at inference time on a non-GPU platform. Specifically, our LPN-50 can achieve 68.7 in AP score on the COCO test-dev set, with 17 FPS inference speed on an Intel i7-8700K (6 cores) CPU machine. Jie Tang 0006, Gangshan Wu |
ICASSP | 2 |
| 2020 | Residual Feature Aggregation Network for Image Super-ResolutionabstractRecently, very deep convolutional neural networks (CNNs) have shown great power in single image super-resolution (SISR) and achieved significant improvements against traditional methods. Among these CNN-based methods, the residual connections play a critical role in boosting the network performance. As the network depth grows, the residual features gradually focused on different aspects of the input image, which is very useful for reconstructing the spatial details. However, existing methods neglect to fully utilize the hierarchical features on the residual branches. To address this issue, we propose a novel residual feature aggregation (RFA) framework for more efficient feature extraction. The RFA framework groups several residual modules together and directly forwards the features on each local residual branch by adding skip connections. Therefore, the RFA framework is capable of aggregating these informative residual features to produce more representative features. To maximize the power of the RFA framework, we further propose an enhanced spatial attention (ESA) block to make the residual features to be more focused on critical spatial contents. The ESA block is designed to be lightweight and efficient. Our final RFANet is constructed by applying the proposed RFA framework with the ESA blocks. Comprehensive experiments demonstrate the necessity of our RFA framework and the superiority of our RFANet over state-of-the-art SISR methods. Jie Liu 0040, Wenjie Zhang 0006, Yuting Tang, Jie Tang 0006, Gangshan Wu |
CVPR | 4 |
| 2020 | Belief Map Enhancement Network for Accurate Human Pose Estimation
Jie Liu 0040, Yishun Dou, Wenjie Zhang 0006, Jie Tang 0006, Gangshan Wu |
ECAI | 4 |
| 2020 | Low Complexity Single Image Super-Resolution with Channel Splitting and Fusion NetworkabstractRecently, deep convolutional neural networks (CNNs) have made remarkable progress on single image super-resolution (SISR). However, many of these methods use very deep or wide convolutional layers to achieve good performance, which treat all feature channels indiscriminately and neglect the difference among the contribution of each channel to the output results. In this paper, we propose a low complexity solution based on channel splitting and fusion network (CSFN) to address this problem. Our method uses channel splitting and channel fusion to enhance feature maps and make full use of valuable information, and then multiple residual channel splitting and fusion blocks (CSFB) are cascaded to continuously extract more important information for reconstruction. To further minimize redundant parameters and improve efficiency, we adopt group and recursive con-volutional layer strategy in CSFB. Experiments demonstrate that our proposed CSFN could achieve higher performance with low computational complexity than most state-of-the-art methods. Minqiang Zou, Jie Tang 0006, Gangshan Wu |
ICASSP | 2 |
| 2020 | Memory Recursive Network for Single Image Super-ResolutionabstractRecently, extensive works based on convolutional neural network (CNN) have shown great success in single image super-resolution (SISR). In order to improve the SISR performance while reducing the number of model parameters, some methods adopt multiple recursive layers to enhance the intermediate features. However, in the recursive process, these methods only use the output features of current stage as the input of the next stage and neglect the output features of historical stages, which degrades the performance of the recursive blocks. The long-term dependencies can only be learned implicitly during the recursive processes. To address these issues, we propose the memory recursive network (MRNet) to make full use of the output features at each stage. The proposed MRNet utilizes a memory recursive module (MRM) to generate features for each recursive stage, and then these features are fused by our proposed ShuffleConv block. Specifically, MRM adopts a memory updater block to explicitly model the long-term dependencies between the output features of historical recursive stages. The output features from the memory updater will be used as the input of the next recursive stage and will be continuously updated during the recursions. To reduce the number of parameters and ease the training difficulty, we introduce a ShuffleConv module to fuse the features from different recursive stages, which is much more effective than using plain convolutional combinations. Comprehensive experiments demonstrate that the proposed MRNet achieves state-of-the-art SISR performance while using much fewer parameters. Jie Liu 0040, Minqiang Zou, Jie Tang 0006, Gangshan Wu |
ACM Multimedia | 3 |
| 2019 | Personalized Recommendation of Photography Based on Deep Learning
Zhixiang Ji, Jie Tang 0006, Gangshan Wu |
MMM (1) | 2 |
| 2018 | CuAPSS: A Hybrid CUDA Solution for AllPairs Similarity Search
Yilin Feng, Jie Tang 0006, Chong-Jun Wang, Junyuan Xie |
ICA3PP (1) | 2 |
| 2018 | A Parallel Method for All-Pair SimRank Similarity Computation
Xingkun Gao, Jie Tang 0006, Gangshan Wu |
ICA3PP (1) | 3 |
| 2018 | Fast Document Cosine Similarity Self-Join on GPUsabstractSimilarity Search has been studied in many different fields of computer science, including data mining, information retrieval, databases and so on. Document similarity self-join is a crucial part of lots of applications, such as near-duplicate document detection, document clustering and web search. On a collection of documents, document similarity self-join finds out all pairs of documents whose similarity values are no lower than a threshold value. However, similarity search is a computation-intensive procedure and consumes a large amount of time as the dataset size increases. Thus, many serial algorithms focus on speeding up the process by decreasing the possible similarity candidates for each query object on high-dimensional sparse datasets, including documents. However, the efficiency of those serial algorithms degrade badly as the threshold decreases. Parallel implementations based on OpenMP or MapReduce also adopt the pruning policy and do not solve the problem thoroughly. In this context, taking into account features of document datasets, we propose 2Step-SSJ, which solves the document similarity self-join in CUDA environment on GPUs. 2Step-SSJ performs the similarity self-join in two steps, i.e., similarity computing on the inverted list and similarity computing on the forward list, which compromises between the memory visiting and dot-product computation. The experimental results show that 2Step-SSJ could solve the problem much faster than existing methods on three benchmark text corpora, achieving the speedup of 2×-23× against the state-of-the-art parallel algorithm in general, while keep a relatively stable running time with different values of the threshold. Yilin Feng, Jie Tang 0006, Chong-Jun Wang, Junyuan Xie |
ICTAI | 2 |
| 2017 | Deep convolutional neural networks for pedestrian detection with skip poolingabstractWith the big success of deep convolutional neural networks (CNN) in image classification task, many proposal based networks are proposed to detect given objects in an image. Faster R-CNN is such a network that uses a region proposal network (RPN) to generate nearly cost-free region proposals, which has shown excellent performance in ILSVRC and MS COCO datasets. However, Faster R-CNN does not behave so well for the task of pedestrian detection since the images in popular pedestrian detection datasets have more complicated background and contain a lot of small foreground objects. In this work, we leverage the RPN architecture of Faster R-CNN and extend it to a multi-layer version combined with skip pooling to tackle the pedestrian detection problem. Skip pooling is a kind of network connection that combines multiple ROI pooling results from lower layers to form a single input to a higher layer while bypassing intermediate layers. We comprehensively evaluate our network, referred to as SP-CNN, on the Caltech pedestrian detection benchmark and KITTI object detection benchmark. Our method achieves state-of-the-art accuracy on Caltech dataset and presents a comparable result on KITTI dataset while maintaining a good speed. Jie Liu 0040, Xingkun Gao, Nianyuan Bao, Jie Tang 0006, Gangshan Wu |
IJCNN | 4 |
| 2016 | Scalable Single-Source SimRank Computation for Large GraphsabstractSimRank is an effective similarity measure between vertices in a graph, which has become a fundamental technique in graph analytics. Despite its popularity, computation of SimRank is often costly in both space and time, especially with the ever growing scale of graph data nowadays. In this paper, we focus on the computation of Single-Source SimRank: given a query vertex, return the similarities between this vertex and any other vertices in the graph. The traditional centralized SimRank algorithms are not efficient for this problem. To fully utilize the computing power of modern distributed systems, we propose sssSimRank, an efficient distributed algorithm based on the random walk model. Our algorithm achieves scalability via minimizing the total number, the space cost, and the matching time of random walks. We implement our approach on the popular distributed processing platform Spark. Experimental results demonstrate the effectiveness, efficiency and scalability of our method. Xingkun Gao, Nianyuan Bao, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
ICPADS | 4 |
| 2015 | A New Data Replication Scheme for PVFS2
Nianyuan Bao, Jie Tang 0006, Gangshan Wu |
ICA3PP (3) | 2 |
| 2015 | Pre-stack Kirchhoff Time Migration on Hadoop and Spark
Jie Tang 0006, Gangshan Wu |
ICA3PP (3) | 2 |
| 2015 | A Dynamic Extension and Data Migration Method Based on PVFS
Jie Tang 0006, Gangshan Wu |
ICA3PP (2) | 2 |
| 2010 | Interective Point Clouds Fairing on Many-Core SystemabstractThis Paper proposes an interactive point clouds fairing algorithm running on many-core system. The algorithm is composed of four steps. Firstly, a k nearest neighbor searching method was designed which could fully utilize the computing ability of GPU. Secondly, a parallel Gaussian weighted normal estimation was put forward. Thirdly, a weighted fairing method was proposed to get better result especially for the unevenly distributed point clouds. The whole algorithm was implemented on NVIDIA GPU using CUDA. Experimental results show that the algorithm could achieve interactive fairing of large size point clouds with good quality. Jie Tang 0006, Gangshan Wu, Zhongliang Gong |
ISPA | 1 |