VLDB 2026 Research / reviewers in the wild / expert
Wanru Xu
dblp:140/6716
· DBLP profile ↗
60ranked-venue papers
18as first author
39since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 10 first-author · 19 since 2021Artificial intelligence and machine learning · 18 · 3 first-author · 14 since 2021Computer networks · 6 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Debiased Multi-modal Vision-Language Network for Action Quality Assessment
Ruizhao Zhai, Wanru Xu, Zhenjiang Miao, Qinghao Kong, Ruiying Yao |
ICIC (16) | 2 |
| 2026 | Reinforced Multi-expert Ensemble Strategy for Weakly Supervised Video Anomaly Detection
Wanru Xu, Zhenjiang Miao, Ruiying Yao |
ICPR (7) | 2 |
| 2026 | Query-guided predicate decoupling and prototype approximation learning for scene graph generation
Shichao Kan, Yue Zhang 0065, Yi-Gang Cen, Wanru Xu, Yi Jin 0001, Yidong Li |
Expert Syst. Appl. | 5 |
| 2026 | Causal learning with uncertainty-aware transformer for vision-and-language navigation
Wanru Xu, Zhenjiang Miao, Yi-Gang Cen, Wangsheng He |
Neurocomputing | 2 |
| 2026 | AFMT: Adaptive frequency decomposition and multi-scale transformer for time series forecasting
Wanru Xu, Chunqiang Zhu, Jingkai Gao, Lugema Mi, Fan Deng 0003, Jinqi Qu |
Inf. Sci. | 2 |
| 2026 | Multimodal Dance Generation With Multi-Granularity Style Control and Text GuidanceabstractABSTRACT Dance generation is a significant research area in computer arts and artificial intelligence. This study proposes a novel framework to enhance dance controllability and personalization through multimodal and multi‐granularity control. The framework establishes global choreographic control of long sequences via music and dance style factors, while accommodating local style variations. Simultaneously, it enables fine‐grained local control using style, text, and temporal factors for motion refinement. We develop two cross‐modal Transformers: the LS‐M2D model merges music and dance style features for local style‐controllable dance generation, and the LT‐SM2D model integrates textual guidance with music and dance style features for time‐constrained local control. Experimental results demonstrate enhanced motion quality, effective multi‐granularity style control, and precise text‐guided flexibility. This provides valuable technical support for personalized intelligent dance generation systems. Wanru Xu, Shenghui Wang 0003 |
Comput. Animat. Virtual Worlds | 4 |
| 2026 | Vision-Semantics-Label: A New Two-Step Paradigm for Action Recognition With Large Language ModelabstractIn recent years, the rapid advancement of multi-modal large language models has propelled the development of video-based conversation models. Due to their exceptional video understanding capabilities, there is often an expectation that these models can handle all video-related tasks, including action recognition. However, because action recognition datasets typically lack semantic information, limiting the performance of dialogue models. Additionally, as these dialogue models are designed for video understanding, they frequently overlook critical information required for action recognition—continuous motion—in their model architecture and training dataset configurations. To address these challenges, we first propose a novel two-step mapping framework based on large language models, termed “Vision-Semantics-Label” mapping, to better adapt video-based large language models for action recognition. In the first step, we proposed a visual-skeletal collaborative learning large language model (VS-LLM), which utilizes human keypoints to compensate for the missing motion details without increasing the input token length of the large language model. In the second step, we designed two mapping methods: verb noun match (VN-Match) and all text match (ALL-Match), which can effectively extract relevant action descriptions from the text. Finally, we construct semantic action recognition datasets to ensure that the training data inherently contains action details, enabling the model to better achieve action recognition. We evaluate our approach on five benchmark datasets, demonstrating the state-of-the-art performance of large language models in action recognition. The source code and dataset are publicly available at https://github.com/xiaoyu92568/VS-LLM. Wanru Xu, Shichao Kan, Linna Zhang, Yi Jin 0001, Yi-Gang Cen, Yidong Li |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | UAGM: Uncertainty-Aware Geometric Modeling for Multi-Scenario 3-D Object Detection in Autonomous VehiclesabstractVision-based 3-D object detection is a core task for autonomous driving and intelligent transport systems. However, the idealized geometric assumptions and single-view modeling methods relied upon by existing methods are susceptible to geometric uncertainties, caused by external parameter perturbations and depth estimation noise in practical multi-scenario deployments. Building a unified and robust perception model suitable for multi-scenarios of both ego-vehicle and roadside remains challenging. To address these challenges, we propose an Uncertainty-Aware Geometric Modeling (UAGM) method that explicitly handles geometric uncertainty to achieve robust perception across multiple scenarios. We use a dual-branch architecture to establish a robust geometric foundation: the height branch introduces dynamic virtual coordinate calibration to compensate for camera extrinsic parameter perturbations in real time, while explicitly modeling height prediction uncertainty through Multi-Hypothesis Projection (MHP), thereby constructing a robust global geometric representation. Meanwhile, the depth branch integrates a Probabilistic Depth Smoothing (PDS) module that employs Conditional Random Fields (CRF) to model spatial consistency constraints, effectively mitigating geometric discontinuities arising from pixel-level predictions. To facilitate better information fusion and interaction, we first propose a Temporal Pyramid Fusion (TPF) module to effectively capture multi-scale spatio-temporal dynamics to reduce the uncertainty in single-frame estimation, instead of error-prone dynamic ego-motion compensation. Subsequently, our Hierarchical Refinement Decoder (HRD) refines BEV proposal localization by fusing image features with depth embeddings to compensate for spatial distortions caused by forward projection. Experimental results demonstrate that UAGM not only achieves state-of-the-art detection performance on both the nuScenes and DAIR-V2X benchmarks, but more importantly, it successfully demonstrates the strong generalization capability of a single model across different viewpoints and deployment conditions. Zhaojie Sun, Wanru Xu, Lu Shi 0004, Yi-Gang Cen, Yi Jin 0001, Yidong Li |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2026 | Natural Cognizing Video: A Decoupling and Integration Network for General Event Boundary CaptioningabstractVideo captioning is a prominent and challenging research area. Previous studies have focused predominantly on describing entire video segments, often overlooking the more significant status changes within these segments. We propose a novel decoupling and integration network for general event boundary captioning (DIN-GEBC), which is applied to the Kinetics-GEBC dataset, a video dataset with fine-grained status descriptions. DIN-GEBC focuses on generating three types of captions for each video segment boundary: a caption for the dominant subject (subject caption), a caption describing the event prior to the status moment (status-before), and a caption describing the event after the status moment (status-after). To address different descriptive focuses and characteristics, DIN-GEBC is proposed for decoupling and integrating both tasks and features. For task decoupling, DIN-GEBC is designed with a dual branch structure, in which the generation of the subject caption is addressed by the dominant subject branch with a Video Q-former encoder and the generation of the status-before and status-after captions is addressed by the event branch with a Reinventing RNNs for the Transformer Era (RWKV) encoder. For task integration, DIN-GEBC enables the dominant subject branch to guide the event branch in producing detailed information regarding the subject experiencing the change. Feature disentanglement, in which the common features are used by the dominant subject branch to capture the unchanging information, is also performed, and the difference features are applied to the event branch to capture the changing information. The experimental results show that our model outperforms existing models on the Kinetics-GEBC dataset even with fewer parameters. The code is available athttps://github.com/Mrliu001219/DIN-GEBC. Wanru Xu, Zhenjiang Miao, Wangsheng He, Ruiying Yao |
IEEE Trans. Multim. | 2 |
| 2025 | Dual-Branch Diffusion Model for JPEG Artifact Correction
Wenhao Yu 0015, Wanru Xu, Zhenjiang Miao |
ICIC (1) | 3 |
| 2025 | Noise-Guided Predicate Representation Extraction and Diffusion-Enhanced Discretization for Scene Graph GenerationabstractScene Graph Generation (SGG) is a fundamental task in visual understanding, aimed at providing more precise local detail comprehension for downstream applications. Existing SGG methods often overlook the diversity of predicate representations and the consistency among similar predicates when dealing with long-tail distributions. As a result, the model's decision layer fails to effectively capture details from the tail end, leading to biased predictions. To address this, we propose a Noise-Guided Predicate Representation Extraction and Diffusion-Enhanced Discretization (NoDIS) method. On the one hand, expanding the predicate representation space enhances the model's ability to learn both common and rare predicates, thus reducing prediction bias caused by data scarcity. We propose a conditional diffusion model to reconstructs features and increase the diversity of representations for same category predicates. On the other hand, independent predicate representations in the decision phase increase the learning complexity of the decision layer, making accurate predictions more challenging. To address this issue, we introduce a discretization mapper that learns consistent representations among similar predicates, reducing the learning difficulty and decision ambiguity in the decision layer. To validate the effectiveness of our method, we integrate NoDIS with various SGG baseline models and conduct experiments on multiple datasets. The results consistently demonstrate superior performance. Shichao Kan, Fanghui Zhang, Wanru Xu, Yue Zhang 0065, Yi-Gang Cen |
ICML | 4 |
| 2025 | Hierarchical Meta-prototypes Network for Few-shot Action RecognitionabstractExisting few-shot action recognition (FSAR) studies predominantly follow a metric learning framework, where prototypes are generated directly from features extracted by an encoder, and classification is performed via distance-based matching. However, due to the limited number of available samples, significant variations exist between different video features of the same class. As a result, the same query video may yield different classification results when matched against different sets of support videos. To address this issue, we propose a novel Hierarchical Meta-Prototypes Network (HMP-Net). The key innovation of our approach lies in the introduction of a category-agnostic and feature-agnostic meta-prototype module, which guides video feature mapping into a more suitable feature space. To optimize this meta-prototype, we design an alternating meta-prototype training strategy, where the model first learns to transform features under a fixed meta-prototype, and then the meta-prototype is refined to better guide feature mapping. Additionally, to adapt image-based metric learning models to video-based FSAR tasks, we introduce a series of lightweight adaptation modules. Specifically, we integrate an adapter into the encoder to improve video frame feature extraction, design a hierarchical prototype generation mechanism to enhance overall video understanding, and incorporate a task-specific perception module to extract unique features for each task. These adaptations make our model better suited for FSAR, significantly improving performance. We evaluate HMP-Net on five challenging benchmarks, and experimental results demonstrate that our model achieves new state-of-the-art performance on HMDB51, UCF101, Kinetics, and SthSthV2-Small. Extensive empirical evaluations further highlight the effectiveness and robustness of HMP-Net. Yi-Gang Cen, Wanru Xu, Yue Zhang 0065, Yi Jin 0001, Yidong Li, Linna Zhang |
ACM Multimedia | 3 |
| 2025 | InstructStep: Fine-Grained Localization of Step Content and Relation in Instructional VideoabstractExisting methods for video answer localization (VAL) in instructional video focus predominantly on coarse-grained themes, failing to address detailed step content and inter-step relation crucial for effective comprehension. Current datasets, such as MedVidQA, primarily capture video content but lack annotations for step structure and inter-step relation. To address this gap, we introduce InstructStep, a newly proposed VAL task, specifically designed for Instructional Video Step Content and Relation Localization. It extends original VAL task to step-centric content and relation. Accordingly, we create a InstructStep Dataset with fine-grained step content and relation QA pairs. To tackle the challenges of this task, we propose a Step-Centric Multi-Level Knowledge Distillation (SC-MLKD) approach that: (1) A two-stage training strategy that generates step-specific summaries in the first stage and introduces a step branch in the second stage to learn step relations. This is applied only during training, ensuring no additional inference time. (2) Multi-level knowledge distillation, including feature, relation, and response distillation, across visual, text and step branches to capture fine-grained and step-centric features. Comprehensive experiments demonstrate the efficiency of SC-MLKD, with notable gains of up to 5.83% in step content and up to 5.9% in step relation. The dataset have been made publicly available on https://github.com/hewangsh/InstructStep Wangsheng He, Wanru Xu, Zhenjiang Miao |
ACM Multimedia | 2 |
| 2025 | Differential Contrastive Training for Gaze EstimationabstractThe complex application scenarios have raised critical requirements for precise and generalizable gaze estimation methods. Recently, the pre-trained CLIP has achieved remarkable performance on various vision tasks, but its potentials have not been fully exploited in gaze estimation. In this paper, we propose a novel Differential Contrastive Training strategy, which boosts gaze estimation performance with the help of the CLIP. Accordingly, a Differential Contrastive Gaze Estimation network (DCGaze) composed of a Visual Appearance-aware branch and a Semantic Differential-aware branch is introduced. The Visual Appearance-aware branch is essentially a primary gaze estimation network and it incorporates an Adaptive Feature-refinement Unit (AFU) and a Double-head Gaze Regressor (DGR), which both help the primary network to extract informative and gaze-related appearance features. Moreover, the Semantic Difference-aware branch is designed on the basis of the CLIP's text encoder to reveal the semantic difference of gazes. This branch could further empower the Visual Appearance-aware branch with the capability of characterizing the gaze-related semantic information. Extensive experimental results on four challenging datasets over within and cross-domain tasks demonstrate the effectiveness of our DCGaze. The code is available at https://github.com/LinZhang-bjtu/DCGaze. Xiyun Wang, Wanru Xu, Yi Jin 0001 |
ACM Multimedia | 4 |
| 2025 | Causal Debiasing Network for Action Quality Assessment
Ruizhao Zhai, Wanru Xu, Zhenjiang Miao, Qinghao Kong |
PRCV (7) | 2 |
| 2025 | Knowledge-based and Data-driven Fusion for Unsupervised Video Anomaly DetectionabstractVideo Anomaly Detection (VAD) has extensive applications in fields such as intelligent surveillance and autonomous driving. In the field of unsupervised VAD, data-driven methods based on pseudo-label generation have their advantages. They can utilize the characteristics of the data itself to generate pseudo-labels for model learning. However, the unsupervised setting leads to a lack of supervision information, resulting in a low confidence level of the supervision signal. On the other hand, prior knowledge can effectively supplement some supervision information. Nevertheless, prior knowledge is usually general and does not take into account the specific information of each sample. To address these limitations, this paper proposes an unsupervised video anomaly detection method that combines prior knowledge and data. This method takes unlabeled videos as input and learns to predict frame-level anomaly scores. The algorithm consists of a prior knowledge module and a data-driven module. The prior knowledge module calculates the degree of anomaly through normal propagation based on prior knowledge independent of the data. The data-driven module estimates the degree of anomaly through two branches, appearance and motion, and jointly generates pseudo-labels. Experiments conducted on the UCF-Crime and ShanghaiTech datasets, using frame-level AUC as the evaluation metric, show that the proposed method achieves state-of-the-art performance in the unsupervised category, outperforming existing one-class classification methods. Qinghao Kong, Wanru Xu, Zhenjiang Miao, Ruizhao Zhai, Wenhao Yu 0015 |
SMC | 2 |
| 2025 | CroCaps: A CLIP-assisted cross-domain video captioner
Wanru Xu, Yenan Xu, Zhenjiang Miao, Yi-Gang Cen, Xiaole Ma |
Expert Syst. Appl. | 1 |
| 2025 | Counterfactual contrastive learning for weakly supervised temporal sentence grounding
Yenan Xu, Wanru Xu, Zhenjiang Miao |
Neurocomputing | 2 |
| 2025 | Low-resolution human pose estimation and action recognition via pose-driven super-resolution reconstruction
Zhizhuo Zhang, Wanru Xu, Shenghui Wang 0003 |
Mach. Learn. | 3 |
| 2025 | Cross-scene visual context parsing with large vision-language modelabstractRelation analysis is crucial for image-based applications such as visual reasoning and visual question answering . Current relation analysis such as scene graph generation (SGG) only focuses on building relationships among objects within a single image. However, in real-world applications, relationships among objects across multiple images, as seen in video understanding , may hold greater significance as they can capture global information. This is still a challenging and unexplored task. In this paper, we aim to explore the technique of Cross-Scene Visual Context Parsing (CS-VCP) using a large vision-language model. To achieve this, we first introduce a cross-scene dataset comprising 10,000 pairs of cross-scene visual instruction data, with each instruction describing the common knowledge of a pair of cross-scene images. We then propose a Cross-Scene Visual Symbiotic Linkage (CS-VSL) model to understand both cross-scene relationships and objects by analyzing the rationales in each scene. The model is pre-trained on 100,000 cross-scene image pairs and validated on 10,000 image pairs. Both quantitative and qualitative experiments demonstrate the effectiveness of the proposed method. Our method has been released on GitHub: https://github.com/gavin-gqzhang/CS-VSL . Shichao Kan, Lu Shi 0004, Wanru Xu, Gaoyun An, Yi-Gang Cen |
Pattern Recognit. | 4 |
| 2025 | 'Disengage AND Integrate': Personalized Causal Network for Gaze EstimationabstractGaze estimation task aims to predict a 3D gaze direction or a 2D gaze point given a face or eye image. To improve generalization of gaze estimation models to unseen new users, existing methods either disentangle personalized information of all subjects from their gaze features, or integrate unrefined personalized information into blended embeddings. Their methodologies are not rigorous whose performance is still unsatisfactory. In this paper, we put forward a comprehensive perspective named 'Disengage AND Integrate' to deal with personalized information, which elaborates that for specified users, their irrelevant personalized information should be discarded while relevant one should be considered. Accordingly, a novel Personalized Causal Network (PCNet) for generalizable gaze estimation has been proposed. The PCNet adopts a two-branch framework, which consists of a subject-deconfounded appearance sub-network (SdeANet) and a prototypical personalization sub-network (ProPNet). The SdeANet aims to explore causalities among facial images, gazes, and personalized information and extract a subject-invariant appearance-aware feature of each image by means of causal intervention. The ProPNet aims to characterize customized personalization-aware features of arbitrary users with the help of a prototype-based subject identification task. Furthermore, our whole PCNet is optimized in a hybrid episodic training paradigm, which further improve its adaptability to new users. Experiments on three challenging datasets over within-domain and cross-domain gaze estimation tasks demonstrate the effectiveness of our method. Xiyun Wang, Sihui Zhang, Wanru Xu, Yi Jin 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | RESTHT: relation-enhanced spatial-temporal hierarchical transformer for video captioning
Lihuan Zheng, Wanru Xu, Zhenjiang Miao, Xinxiu Qiu, Shanshan Gong |
Vis. Comput. | 2 |
| 2024 | Hypergraph Self-Attention and Channel Topology Specialization Network for Automatic Generation of Labanotation
Wanru Xu, Zhenjiang Miao |
ICPR (15) | 2 |
| 2024 | CLIP-based Semantic Enhancement and Vocabulary Expansion for Video Captioning Using Reinforcement LearningabstractVideo captioning aims to comprehend the content of videos and automatically generate sentences. It necessitates a network with a robust knowledge background to understand complex video events and transform them into coherent sentences. Traditional video captioning is often limited to modeling close-domain videos and the knowledge is fixed after training, which results in generating short and uninformative captions. Different from traditional video captioning, we propose an open-domain video captioning method that incorporates external textual data and expands the current knowledge domain to enhance the model’s ability to generate more nuanced and contextually relevant descriptive sentences. This paper reduces the gap between videos and texts by employing a well-pretrained CLIP network at the lexical level and effectively retrieves pertinent vocabularies from the training corpus of two datasets to serve as prompts. Our model utilizes retrieval words through implicit augmentation and explicit augmentation, providing additional semantic features as implicit knowledge and explicitly updating the word sampling pool in reinforcement learning. The experiments conducted on several benchmark datasets show that our proposed method is effective. Lihuan Zheng, Zhenjiao Miao, Wanru Xu |
IJCNN | 4 |
| 2024 | Probabilistic Distillation Transformer: Modelling Uncertainties for Visual Abductive ReasoningabstractVisual abduction reasoning aims to find the most plausible explanation for incomplete observations, and suffers from inherent uncertainties and ambiguities, which mainly stem from the latent causal relations, incomplete observations, and the reasoning itself. To address this, we propose a probabilistic model named Uncertainty-Guided Probabilistic Distillation Transformer (UPD-Trans) to model uncertainties for Visual Abductive Reasoning. In order to better discover the correct cause-effect chain, we model all the potential causal relations into a unified reasoning framework, thus both the direct relations and latent relations are considered. In order to reduce the effect of the stochasticity and uncertainty for reasoning: 1) we extend the deterministic Transformer to a probabilistic Transformer by considering those uncertain factors as Gaussian random variables and explicitly modeling their distribution; 2) we introduce a distillation mechanism between the posterior branch with complete observations and the prior branch with incomplete observations to transfer posterior knowledge. Evaluation results on the benchmark datasets, consistently demonstrate the commendable performance of our UPD-Trans, with significant improvements after latent relation modeling and uncertainty modeling. Wanru Xu, Zhenjiang Miao, Yi-Gang Cen, Xiaole Ma |
ACM Multimedia | 1 |
| 2024 | Differential Refinement Network for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to recognize novel categories by merely utilizing disjoint seen samples. It is a challenging task as the knowledge of unseen objects is forbidden in the training stage, which easily leads to unseen samples degrading to mismatched categories. In order to alleviate the biased recognition problem, in this article, we propose a differential refinement network (DRNet) for ZSL, which aims to explore robust semantic-to-visual embedding. Our DRNet model consists of two subnetworks: basic network and differential network. The basic network targets to generate initial class-specific visual centers conditioned on corresponding semantic prototypes. The differential network is designed to predict class-unrelated differences between visual centers of arbitrary semantic prototype pairs, which are applied to further polish the initial visual centers. The motivation is that, by comparing different prototypes, interactions between various categories will be characterized, benefiting the generation of authentic and discriminative visual centers. Moreover, a modified episode-based training paradigm is explored to optimize the two subnetworks actively. In the training stage, we form a collection of episodes, each of which is an imitated ZSL task. Our DRNet is optimized by those sampled tasks rather than individual samples, which progressively learns skills to adapt and generalize to novel classes. Experiments on four challenging datasets demonstrate the effectiveness of our method. Wanru Xu, Zhengming Ding |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Weighted hybrid order total variation model using structure tensor for image denoising
Wanru Xu, Haifeng Wu, Ali Abdullah Yahya |
Multim. Tools Appl. | 2 |
| 2022 | Sequential Gesture Learning for Continuous Labanotation Generation Based on the Fusion of Graph Neural NetworksabstractLabanotation is a symbolic recording system for human movements, and also a powerful tool for protecting and spreading folk dances and other performing arts. State-of-the-art automatic Labanotation uses end-to-end methods with sequence-based skeleton representation, which cannot capture the relationship between joints and bones in the skeleton for accurate descriptions of continuous lower limb movements such as dance steps. In this paper, we propose a novel double-stream fusion method of directed graph neural networks (DGNN), combined with connectionist temporal classification (CTC), namely DFGNN-CTC, for sequential fine-grained motion recognition, such as the Labanotation generation of unsegmented dance movement. First, we extract double-stream directed graph feature, employing an orientation-normalized directed acyclic graph (ON-DAG) and an orientation-normalized temporal directed acyclic graph (ON-TDAG), to jointly model spatiotemporal properties of movement recorded in motion capture data. Then, we design a CTC-based fusion-pooling module to fuse the spatial and temporal streams encoded by two DGNNs. It concatenates and fuses the two streams to generate discriminative descriptions of each time step, and concentrates them to make per-time-step predictions of Laban gesture type, from which the CTC searches the optimal Laban symbol sequence, corresponding to elemental motions composing the movement. In this way, the new method enables much finer discrimination for similar Laban gestures with subtle differences in spatial and temporal properties through joint contextual spatiotemporal modeling so that it achieves much superior performance in continuous Labanotation generation to existing methods, which only have single-stream analysis either spatially or temporally. The experiments on two Labanotation-labelled motion capture datasets demonstrate the effectiveness of the components in the proposed method and its superiority comparing with the state-of-the-art methods, especially for lower limb movements. Ningwei Xie, Zhenjiang Miao, Xiao-Ping Zhang 0002, Wanru Xu, Min Li 0025, Jiaji Wang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Bridging Video and Text: A Two-Step Polishing Transformer for Video CaptioningabstractVideo captioning is a joint task of computer vision and natural language processing, which aims to describe the video content using several natural language sentences. Nowadays, most methods cast this task as a mapping problem, which learns a mapping from visual features to natural language and generates captions directly from videos. However, the underlying challenge of video captioning,i.e., sequence to sequence mapping across the different domains, is still not well handled. To address these problems, we introduce the polishing mechanism in an attempt to mimic human polishing process and propose a generate-and-polish framework for video captioning. In this paper, we propose a two-step transformer based polishing network (TSTPN) consisting of two sub-modules: the generation-module is to generate the caption candidate and the polishing-module is to gradually refine the generated candidate. Specifically, the candidate provides a global information of the visual contents in a semantically-meaningful order, where it is firstly considered as a semantic intersnubber to bridge the semantic gap between the text and video, with the cross-modal attention mechanism for better cross-modal modeling; and it secondly provides a global planning ability to maintain the semantic consistency and fluency of the whole sentence for better sequence mapping. In experiments, we present adequate evaluations to show that the proposed TSTPN achieves the comparable and even better performance than the state-of-the-art methods on the benchmark datasets. Wanru Xu, Zhenjiang Miao, Jian Yu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Rhythm-Aware Sequence-to-Sequence Learning for Labanotation Generation With Gesture-Sensitive Graph Convolutional EncodingabstractLabanotation is a professional dance notation system widely used in dance education and choreography preservation. Automatically generatingLabanotation dance scores from motion capture data can save a huge amount of manual time and effort. Recently, the sequence-to-sequence (seq2seq) model is applied to the automatic Labanotation generation. This model is based on an encoder-decoder structure, which encodes the input motion sequence to a fixed-length vector and then decodes it to generate the target sequence. However, the encoding of spatial skeleton structure of motion data is not considered in the existing work. Besides, it is challenging to align between the input motion data and the output Laban symbol sequences due to the severe imbalance of sequence lengths. Therefore, in this paper, we present a new seq2seq model for more effective Labanotation generation. In the encoder, we propose a new gesture-sensitive graph convolutional network with learned adaptive joint weights and non-physical connections to learn both spatial and temporal patterns from motion data sequences. In the decoder, we exploit motion rhythm information and propose a novel rhythm-aware attention mechanism to learn a good alignment between motion sequences and Laban symbol sequences, so that we can focus on relevant parts of the input motion sequence without searching in the whole sequence when predicting a target Laban symbol. Extensive experiments on two real-world datasets show that the proposed method achieves a better performance compared with the state-of-the-art approaches on the task of automatic Labanotation generation. Min Li 0025, Zhenjiang Miao, Xiao-Ping Zhang 0002, Wanru Xu, Cong Ma 0004, Ningwei Xie |
IEEE Trans. Multim. | 4 |
| 2022 | Optimal Charging Oriented Sensor Placement and Flexible Scheduling in Rechargeable WSNsabstractThe recent breakthroughs in Wireless Power Transfer (WPT) facilitate supporting rechargeable sensors to enrich a series of energy-consuming applications. However, most charging scheduling schemes in rechargeable wireless sensor networks (WSNs) focus on sensing tasks instead of charging utility, which leaves a considerably high performance gap in the optimal result. Moreover, the charging scheduling is usually non-flexible, in which a full or nothing charging policy suffers from relatively low charging coverage as well as low efficiency. In this article, we focus on how to efficiently improve charging utility when introducing charging-oriented sensor placement and flexible scheduling policy. We formulate a general maximization optimization problem under a general routing constraint, which generates great difficulty. We utilize area partition and charging discretization methods to transform into the scope of maximizing a submodular function problem. Thus, a constant approximation algorithm is delivered to construct a near optimal charging tour. We analyze the performance loss from the discretization to guarantee that the output of the proposed algorithm has more than (1-ɛ)(1-1/ e )/4 of the optimal solution, where ɛ is an arbitrarily small positive parameter (0 < ɛ < 1). Both simulations and field experiments are conducted to evaluate the performance of our proposed algorithm. Tao Wu 0011, Panlong Yang, Haipeng Dai 0001, Chaocan Xiang, Wanru Xu |
ACM Trans. Sens. Networks | 5 |
| 2021 | An Attention-Seq2Seq Model Based on CRNN Encoding for Automatic Labanotation Generation from Motion Capture DataabstractLabanotation is an important notation system widely used for recording dances. Numerous methods have been proposed for automatic Labanotation generation from motion capture data. Recently, the sequence-to-sequence (seq2seq) model is proposed. However, the encoder of the model only encodes the temporal information of motion data, lacking the encoding for spatial information. And it is challenging for the decoder to align input and output sequences due to the imbalance of the sequence lengths. In this paper, we propose an attention-seq2seq model based on Convolutional Recurrent Neural Network (CRNN). The proposed model employs an encoder based on CRNN to learn the spatial-temporal information of motion data and applies an attention mechanism to align each target Laban symbol with relevant parts of the input motion data in decoding. Experiments show that the proposed method performs favorably against state-of-the-art algorithms in the automatic Labanotation generation task. Min Li 0025, Zhenjiang Miao, Xiao-Ping Zhang 0002, Wanru Xu |
ICASSP | 4 |
| 2021 | Video Anomaly Detection Using Dual Discriminator Based Generative Adversarial NetworkabstractVideo anomaly detection is of great significance due to its wide applications in video surveillance. Recently, there is a trend of using a video prediction framework to tackle this problem. This kind of methods detects anomalies according to the difference between a predicted frame and its ground truth. However, existing prediction methods lack the consideration of multi-scale temporal constraints on generating future frames. Therefore, based on the prediction framework, this paper proposes a temporal enhanced anomaly detection approach, which designs a generative adversarial network with dual discriminator (frame discriminator and sequence discriminator) to predict future frames. In order to obtain more realistic predictions, other than commonly used spatial constraints and the adversarial penalty from the frame discriminator, we also consider both short-range and long-range motions to impose constraints. Specifically, for short-range motion modeling, we utilize the optical flow loss to ensure temporal continuity over two adjacent frames, while for long-range motion modeling, we design a sequence discriminator to identify fake contained sequences from real sequences, making frames more consistent with their previous consecutive frames. Experiments on three datasets, UCSD Ped1, UCSD Ped2 and Avenue, demonstrate the effectiveness of our method in terms of various evaluation criteria for video anomaly detection. Zhenjiang Miao, Wanru Xu, Jiaji Wang, Qiang Zhang 0030, Shaoyue Song |
ICMLA | 3 |
| 2021 | Interactive deformation-driven person silhouette image synthesisabstractAbstract This paper addresses the problem of synthesizing person silhouette images. Previous methods on human‐centric image synthesis deal with normal images taken under consistent lighting conditions. However, a silhouette image has an inconsistent lighting appearance, in which the dark subject and the bright background form a strong contrast. Therefore, it brings great challenges to person silhouette image synthesis. We present a staged method to synthesize realistic person silhouette photos with various poses. The purpose of our method is to give ordinary users the opportunities to interactively adjust the pose of a person in a silhouette photo. The method consists of four main steps: person detection, person analysis, interactive deformation, and image synthesis. To obtain the shape of the target person, we propose a silhouette image segmentation algorithm combined with person detection. Moreover, we also present an effective image inpainting approach to complement a silhouette image with an irregular hole. Experimental results show that the proposed method can generate a set of realistic person silhouette images with interactively changed poses. Di Kong, Zhizhuo Zhang, Wanru Xu, Shenghui Wang 0003 |
Concurr. Comput. Pract. Exp. | 4 |
| 2021 | A CRNN-based attention-seq2seq model with fusion feature for automatic Labanotation generation1
Zhenjiang Miao, Wanru Xu |
Neurocomputing | 3 |
| 2021 | Coupling Adversarial Graph Embedding for transductive zero-shot action recognition
Wanru Xu, Yu Kong 0001 |
Neurocomputing | 3 |
| 2021 | Tolerance-Oriented Wi-Fi Advertisement Scheduling: A Near Optimal Study on Accumulative User Interests
Wanru Xu, Xiaochen Fan, Tao Wu 0011, Panlong Yang |
Mob. Networks Appl. | 1 |
| 2021 | ASPP-DF-PVNet: Atrous Spatial Pyramid Pooling and Distance-Filtered PVNet for occlusion resistant 6D object pose estimation
Yazhi Zhu, Wanru Xu, Shenghui Wang 0003 |
Signal Process. Image Commun. | 3 |
| 2021 | Deep Reinforcement Polishing Network for Video CaptioningabstractThe video captioning task aims to describe video content using several natural-language sentences. Although one-step encoder-decoder models have achieved promising progress, the generations always involve many errors, which are mainly caused by the large semantic gap between the visual domain and the language domain and by the difficulty in long-sequence generation. The underlying challenge of video captioning, i.e., sequence-to-sequence mapping across different domains, is still not well handled. Inspired by the proofreading procedure of human beings, the generated caption can be gradually polished to improve its quality. In this paper, we propose a deep reinforcement polishing network (DRPN) to refine the caption candidates, which consists of a word-denoising network (WDN) to revise word errors and a grammar-checking network (GCN) to revise grammar errors. On the one hand, the long-term reward in deep reinforcement learning benefits the long-sequence generation, which takes the global quality of caption sentences into account. On the other hand, the caption candidate can be considered a bridge between visual and language domains, where the semantic gap is gradually reduced with better candidates generated by repeated revisions. In experiments, we present adequate evaluations to show that the proposed DRPN achieves comparable and even better performance than the state-of-the-art methods. Furthermore, the DRPN is model-irrelevant and can be integrated into any video captioning models to refine their generated caption sentences. Wanru Xu, Jian Yu 0001, Zhenjiang Miao |
IEEE Trans. Multim. | 1 |
| 2020 | Spatio-Temporal Deep Q-Networks for Human Activity LocalizationabstractHuman activity localization aims to recognize category labels and detect the spatio-temporal locations of activities in video sequences. Existing activity localization methods suffer from three major limitations. First, the search space is too large for three-dimensional (3D) activity localization, which requires the generation of a large number of proposals. Second, contextual relations are often ignored in these target-centered methods. Third, locating each frame independently fails to capture the temporal dynamics of human activity. To address the above issues, we propose a unified spatio-temporal deep Q-network (ST-DQN), consisting of a temporal Q-network and a spatial Q-network, to learn an optimized search strategy. Specifically, the spatial Q-network is a novel two-branch sequence-to-sequence deep Q-network, called TBSS-DQN. The network makes a sequence of decisions to search the bounding box for each frame simultaneously and accounts for temporal dependencies between neighboring frames. Additionally, the TBSS-DQN incorporates both the target branch and context branch to exploit contextual relations. The experimental results on the UCF-Sports, UCF-101, ActivityNet, JHMDB, and sub-JHMDB datasets demonstrate that our ST-DQN achieves promising localization performance with a very small number of proposals. The results also demonstrate that exploiting contextual information and temporal dependencies contributes to accurate detection of the spatio-temporal boundary. Wanru Xu, Jian Yu 0001, Zhenjiang Miao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Deep Reinforcement Learning for Weak Human Activity LocalizationabstractHuman activity localization aims at recognizing contents and detecting locations of activities in video sequences. With an increasing number of untrimmed video data, traditional activity localization methods always suffer from two major limitations. First, detailed annotations are needed in most existing methods, i.e., bounding-box annotations in every frame, which are both expensive and time consuming. Second, the search space is too large for 3D activity localization, which requires generating a large number of proposals. In this paper, we propose a unified deep Q-network with weak reward and weak loss (DWRLQN) to address the two problems. Certain weak knowledge and weak constraints involving the temporal dynamics of human activity are incorporated into a deep reinforcement learning framework under sparse spatial supervision, where we assume that only a portion of frames are annotated in each video sequence. Experiments on UCF-Sports, UCF-101 and sub-JHMDB demonstrate that our proposed model achieves promising performance by only utilizing a very small number of proposals. More importantly, our DWRLQN trained with partial annotations and weak information even outperforms fully supervised methods. Wanru Xu, Zhenjiang Miao, Jian Yu 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Bayesian Hierarchical Dynamic Model for Human Action RecognitionabstractHuman action recognition remains as a challenging task partially due to the presence of large variations in the execution of action. To address this issue, we propose a probabilistic model called Hierarchical Dynamic Model (HDM). Leveraging on Bayesian framework, the model parameters are allowed to vary across different sequences of data, which increase the capacity of the model to adapt to intra-class variations on both spatial and temporal extent of actions. Meanwhile, the generative learning process allows the model to preserve the distinctive dynamic pattern for each action class. Through Bayesian inference, we are able to quantify the uncertainty of the classification, providing insight during the decision process. Compared to state-of-the-art methods, our method not only achieves competitive recognition performance within individual dataset but also shows better generalization capability across different datasets. Experiments conducted on data with missing values also show the robustness of the proposed method. Rui Zhao 0015, Wanru Xu, Hui Su |
CVPR | 2 |
| 2019 | Labanotation Generation Based on Bidirectional Gated Recurrent Units with Joint and Line FeaturesabstractLabanotation is an effective carrier for recording and displaying three-dimensional human movements. In the existing methods of Labanotation generation, the spatial characteristics are not fully considered. In addition, Long-term correlation of time series is not reflected. In this paper, we propose a novel method based on Bidirectional Gated Recurrent Units with Joint and Line features which can efficiently convert human movements into Labanotation. Firstly, Joint feature carries the location information; Line feature contains the direction information. These two types of features make full use of the correlation and co-occurrence between adjacent joints, thus embodying good spatial characteristics. Secondly, Bidirectional Gated Recurrent Units are applied to identify human movements. The unique gate-control structure of Bi-GRU can predict current status based on historical and future information. Thus, this method is good in timing modeling especially for long time series. The experimental results show that this method achieved an accuracy of 97.4%. It is higher than the state of the art, which demonstrating the effectiveness of our proposed method. Shanshan Hao, Zhenjiang Miao, Jiaji Wang, Wanru Xu, Qiang Zhang 0030 |
ICIP | 4 |
| 2019 | Charging Oriented Sensor Placement and Flexible Scheduling in Rechargeable WSNsabstractThe recent breakthrough in Wireless Power Transfer (WPT) provides a promising way to support rechargeable sensors to enrich a series of energy-consuming applications. Unfortunately, two major design restrictions hinder the applicability of rechargeable sensor networks. First, most of the sensor placement schemes are focusing on the sensing tasks instead of the charging utility, which leaves a considerably high performance gap towards the optimal result. Second, the charging scheduling is non-flexible, where full or nothing charging policy suffers from the relatively low charging coverage as well as efficiency. In this paper, we focus on how to efficiently improve the charging utility when introducing charging oriented sensor placement and flexible scheduling policy. To this end, we jointly consider optimizing node positions and charging allocations. In particular, we formulate a general convex optimization problem under a general routing constraint, which generates great difficulty. We utilize area partition and charging discretization methods to reformulate a submodular function maximization problem. Thus a constant approximation algorithm is delivered to construct a near optimal charging tour. To this end, we analyze the performance loss from the discretization to guarantee that the output of the proposed algorithm has more than $(1 -\varepsilon)/4 (1 - 1 /e)$ of the optimal solution, where $\varepsilon$ is an arbitrarily small positive parameter $(0 \leq \varepsilon \leq 1)$. Both simulations and field experiments are conducted to evaluate the performance of our proposed algorithm. Tao Wu 0011, Panlong Yang, Haipeng Dai 0001, Wanru Xu, Mingxue Xu |
INFOCOM | 4 |
| 2019 | Collaborated Tasks-driven Mobile Charging and Scheduling: A Near Optimal ResultabstractWireless Power Transfer (WPT) has emerged into an inspiringly commercial and applicable era to charge devices. Existing studies mainly focus on general charging patterns and metrics while overlooking the collaborated task execution, which incurs charging inefficiency among nodes. In this paper we first advocate the collaborated tasks-driven mobile charging and scheduling to respect the energy requirement diversity. Specially, the mobile charging scheduling strategy is considered to maximize the overall task utility which concerns sensor selection and task cooperation. Unfortunately, solving this problem is non-trivial, because it involves solving two coupling NP-hard problems. In tackling with this difficulty, we construct a surrogate function with specific theoretical analysis of its submodularity and gap property. Then, we approximate the traveling cost to transform the formulated problem into an essentially monotone submodular function optimization subject to a general routing constraint, where we propose $a (1-\ 1/e)/4$-approximation algorithm. Extensive simulations are conducted and the results show that our algorithm can achieve a near-optimal solution covering at least S4.9% of the optimal result achieved by the OPT algorithm. Furthermore, field experiments in office room and soccer field environment with 10 and 20 sensors are implemented respectively to validate our proposed algorithm. Tao Wu 0011, Panlong Yang, Haipeng Dai 0001, Wanru Xu, Mingxue Xu |
INFOCOM | 4 |
| 2019 | Prediction-CGAN: Human Action Prediction with Conditional Generative Adversarial NetworksabstractThe underlying challenge of human action prediction, i.e. maintaining prediction accuracy at very beginning of an action execution, is still not well handled. In this paper, we propose a Prediction Conditional Generative Adversarial Network (Prediction-CGAN) for predicting action, which shares information between completely observed and partially observed videos. Instead of generating future frames, we aim at completing visual representations of unfinished video, which can be directly utilized to predict action label no matter at any progress levels. The Prediction-CGAN incorporates the completion constraint to learn a transformation from incomplete actions to complete actions; the adversarial constraint to ensure the generation has similar discriminative power to complete representation; the label consistency constraint to encourage label consistency between each segment and its corresponding complete video; and the confidence monotonically increasing constraint to yield increasingly accurate predictions as observing more frames. Meanwhile, we introduce a novel adversarial criterion especially for prediction task, which requires the generation is more discriminative than its corresponding incomplete representation, while the generation is less discriminative than its real complete representation. In experiments, we present adequate evaluations to show that the proposed Prediction-CGAN outperforms state-of-the-art methods in action prediction. Wanru Xu, Jian Yu 0001, Zhenjiang Miao |
ACM Multimedia | 1 |
| 2019 | Action recognition and localization with spatial and temporal contexts
Wanru Xu, Zhenjiang Miao, Jian Yu 0001 |
Neurocomputing | 1 |
| 2019 | 3D human pose estimation from a single image via exemplar augmentation
Wanru Xu, Shenghui Wang 0003 |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | TIMAO: Time-Sensitive Mobile Advertisement Offloading with Performance GuaranteeabstractMobile advertising has played an important role with the prevalence of smart mobile devices. Most of the previous studies focus on location-based or content-based mobile advertisement propagation and distribution, which are suffered by the low propagation efficiency, because advertisements could not be available to mobile users within limited time span. Conventional offloading schemes could perfectly distribute advertisements according to user's interest, but have not fully respected the time sensitivity in mobile advertisement distribution. In response to this stalemate, we introduce the advertisement platform's expected income maximization problem (EIMP), and prove its NP-hardness. To our knowledge, ours is even harder than conventional 0-1 mutlidimensional and multiple knapsack problem. But inspiringly we find that it could be transformed into a maximizing monotone submodular set function, being subjected to partition matroid constraints. Then a simple but effective greedy algorithm (TIMAO, time-sensitive mobile advertisement offloading)is proposed to solve the EIMP with approximation ratio of 1/3. Finally, the evaluation results show that TIMAO could double the platform's expected income comparing with the random selection method and reach 99.2% of the near optimal values achieved by CPLEX tool-box. At the same time, it increases the time duty cycle by about average 10% compared with the random selection. Wanru Xu, Panlong Yang, Chaocan Xiang |
ICPADS | 1 |
| 2018 | Connection is power: Near optimal advertisement infrastructure placement for vehicular fogs
Wanru Xu, Panlong Yang, Lijing Jiang |
Peer-to-Peer Netw. Appl. | 1 |
| 2017 | Learning a hierarchical spatio-temporal model for human activity recognitionabstractRecent works have shown that hierarchical models lead to significant improvement in human activity recognition, which can not only enhance descriptive capability, but also improve discriminative power. However, most existing methods exploit just one of the two advantages. In this paper, a new hierarchical spatio-temporal model (HSTM) is proposed to integrate feature learning into two-layer hierarchical classification model simultaneously. On the one hand, the two-layer model has sufficient descriptive capability. The bottom layer aims at capturing spatial relations in each frame and learning high-level representations, and the top layer utilizes these learned features to characterize temporal relations in the whole video sequence. On the other hand, the hierarchical model has strong discriminative power. Both spatial similarity and temporal similarity of activities are measured. Experimental results show that the HSTM can successfully recognize human activities with higher accuracies on one-person actions (KTH and UCF), human-human interactions (CASIA), and human-object interactional activities (Gupta). Wanru Xu, Zhenjiang Miao, Xiao-Ping Zhang 0002 |
ICASSP | 1 |
| 2017 | Optimizing the interested area coverage with efficient mobile advertisement user selectionabstractSummary Mobile advertisement distribution effects are vitally important for advertisers as well as users. Status quo studies are focusing on efficient distribution especially when user mobilities are involved. Unfortunately, previous studies have shown the interested area property during mobile advertisement propagation. In achieving efficient and effective mobile advertisement applications, this work advocates the concept of location‐centric mobile crowdsourcing network, where locations are vitally important for advertisement distribution, and mobile users need to be carefully selected for efficiency considerations. Different from traditional user‐centric and platform‐centric crowdsourcing networks, this work focuses on the mobile advertisement user selection problem when interested area coverage is considered. There are several fundamentally important challenges needed to be addressed before developing a location‐centric scheme for mobile advertisement user selection. First of all, we need to deal with the spatio‐temporal features for each user, where the interested area coverage ratio needs to be effectively evaluated. Even worse, budget constraint makes this problem intractable. In tackling aforementioned challenges, this work makes the following efforts: First, a budget‐constrained user selection problem is formulated when location sensitive mobile advertisement applications are considered, which is proved NP‐hard. Second, the submodularity feature is explored, and a simple but efficient heuristic algorithm is presented with guaranteed approximation ratio . Finally, extensive simulation results show that, our scheme could effectively improve the propagation effects for mobile advertisement with 125%. When considering the user's interest to different advertisements, the real user interest data set has also been used to validate that our proposed method could achieve improved performance than the random method. Wanru Xu, Panlong Yang |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | A Hierarchical Spatio-Temporal Model for Human Activity RecognitionabstractThere are two key issues in human activity recognition: spatial dependencies and temporal dependencies. Most recent methods focus on only one of them, and thus do not have sufficient descriptive power to recognize complex activity. In this paper, we propose a hierarchical spatio-temporal model (HSTM) to solve the problem by modeling spatial and temporal constraints simultaneously. The new HSTM is a two-layer hidden conditional random field (HCRF), where the bottom-layer HCRF aims at describing spatial relations in each frame and learning more discriminative representations, and the top-layer HCRF utilizes these high-level features to characterize temporal relations in the whole video sequence. The new HSTM takes advantage of the bottom layer as the building blocks for the top layer and it aggregates evidence from local to global level. A novel learning algorithm is derived to train all model parameters efficiently and its effectiveness is validated theoretically. Experimental results show that the HSTM can successfully classify human activities with higher accuracies on single-person actions (UCF) than other existing methods. More importantly, the HSTM also achieves superior performance on more practical interactions, including human-human interactional activities (UT-Interaction, BIT-Interaction, and CASIA) and human-object interactional activities (Gupta video dataset). Wanru Xu, Zhenjiang Miao, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 1 |
| 2016 | Toward 5G: A Novel Sleeping Strategy for Green Distributed Base Stations in Small Cell NetworksabstractConfronted with the rapidly increasing demand of mobile traffic and heavy energy consumption on base stations (BSs), the base station (BS) sleeping strategy becomes a promising method to promote the system energy efficiency (EE). To switch off the redundant small-cell base stations (s-BSs) without whittling down the network EE in small cell networks, we propose a novel environment-friendly distributed BS sleeping strategy, which consists of two parts, matching and connecting between the s-BSs and the user equipments (UEs), and activating the BS-turning-off procedure for the further promotion of EE. An initializing matching connection algorithm (IMCA) is proposed for the first subproblem and the energy efficiency ratio (EER) obtained by this proposed IMCA surpasses the EE gained from traditional initializing random connection method. Moreover, a turn-off if possible algorithm (TIPA) is proposed as the sleeping strategy to handle the BS-turning-off procedure. Both the EE and the convergence speed obtained by the sleeping deployment using TIPA surpass those of the traditional best response algorithm. Jin Chen 0007, Ducheng Wu, Wanru Xu |
MSN | 4 |
| 2016 | eMAP: Efficient User Selection for Mobile Advertisement PopularizationabstractMobile Advertisement propagation has drawn increasing attention in research and industrial area.In this work, we investigate the mobile advertisement popularization for mobile social networks.Previous studies failed to be applied to mobile social networks because of extremely high overhead and low propagation efficiency.In tackling these difficulties, we propose eMAP, (efficient mobile advertisement popularization), an efficient propagation user selection scheme with local information.Two key technologies enable eMAP to achieve efficient and effective mobile Ads popularization.First, we advocate propagation user selection instead of popular user selection, where mobile users with strong information dissemination ability could be selected.Thus the mobile user could be effectively used.Second, we use local information instead of the global information to achieve near optimal performance for propagation.In that, the information potential is leveraged to find the influential users with local information.With extensive experimental study, we find that, eMAP could effectively improve the mobile Ads delivery ratio.Using the propogation instead of popularization is validated in our experimental studies in different aspects of investigations. Moreover, when the budget is constrained, eMAP could still perform fairly well. Wanru Xu, Panlong Yang, Maotian Zhang, Pengkun Sheng |
VTC Spring | 1 |
| 2016 | A novel mid-level distinctive feature learning for action recognition via diffusion map
Wanru Xu, Zhenjiang Miao |
Neurocomputing | 1 |
| 2015 | Structured feature-graph model for human activity recognitionabstractRecent works have shown that extracting and learning mid-level features lead to significant improvement in human activity recognition. Most existing methods represent activities as a collection of mid-level features and their spatio-temporal relations are completely neglected. Therefore, when scene contains interactional parts or high-level semantic actions, these mid-level features are not able to capture spatial structures as well as high order temporal relationships. In this paper, the activity is represented as a string of structured feature-graphs (SFGs) which models spatial structures and temporal structures simultaneously. A novel temporal graph kernel (TGK) is also proposed to measure similarity between two string representations. Experimental results show that our approach can successfully classify human activities with much higher accuracies for both single-person actions and human-human interactions. Wanru Xu, Zhenjiang Miao, Xiao-Ping Zhang 0002 |
ICIP | 1 |
| 2015 | Context and locality constrained linear coding for human action recognition
Qiuqi Ruan, Gaoyun An, Wanru Xu |
Neurocomputing | 4 |
| 2015 | Projection transform on spatio-temporal context for action recognition
Wanru Xu, Zhenjiang Miao, Qiang Zhang 0030 |
Multim. Tools Appl. | 1 |
| 2014 | Continuous human action recognition in real time
Zhenjiang Miao, Yuan Shen 0004, Wanru Xu, Dianyong Zhang |
Multim. Tools Appl. | 4 |