Wenxin Hu

dblp:163/2853 · DBLP profile ↗
← Back
46ranked-venue papers
2as first author
32since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 1 first-author · 16 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Pose-Free Radiance Reconstruction With Depth and Exposure Priors on a Cloud-Native Hybrid KAN Framework
abstract
Training neural radiance fields without pre-estimated camera poses remains challenging, particularly in the presence of complex illumination and large inter-frame motion. This work presents a self-calibrating radiance-field framework designed to address these difficulties through a unified optimization strategy. Reliable monocular depth priors from the Depth Anything V2 model are incorporated to provide stable geometric constraints, and multiview point-cloud alignment is enforced to prevent pose drift under wide baselines. Illumination-induced photometric inconsistencies are handled by a jointly optimized exposure-compensation module that separates transient brightness variations from the underlying scene radiance. To enhance the representational capacity beyond that of standard multilayer perceptrons, a Kolmogorov-Arnold Network is adopted as the radiance-field backbone, enabling improved recovery of fine geometric and appearance details. The complete pipeline is deployed on a cloud-native infrastructure that supports scalable optimization and interactive access through an LLM-assisted interface. Experiments on indoor and outdoor datasets demonstrate that the proposed method consistently improves pose accuracy and novel-view synthesis quality over existing self-calibrating approaches.
Zhuohao Gong, Xinyang Guo, Wenxin Hu, Weichu Liu, Xiping Hu
IEEE Trans. Cloud Comput.3
2025 NeRF-KAN: An Efficient 3D Reconstruction Method for Neural Radiance Fields Based on Kolmogorov-Arnold Networks
abstract
Neural Radiance Fields (NeRF) have revolutionized 3D scene reconstruction through implicit neural representation, yet face fundamental limitations due to the C° continuity constraints and parameter inefficiency of traditional multilayer perceptrons (MLPs). This paper introduces NeRF-KAN, a novel approach that replaces MLPs with Kolmogorov-Arnold Networks (KANs) for radiance field modeling. Our key innovation lies in the dual-path KAN - Linear unit, which synergistically combines linear transformations for global structure capture with B-spline non-linear paths for local detail fitting, effectively resolving the global-local feature fragmentation inherent in MLPs. Additionally, we design a dual-term regularization strategy comprising activation and entropy regularization, specifically tailored to KAN's B-spline weight characteristics to mitigate overfitting in few-shot scenarios. Comprehensive experiments on both synthetic (Lego) and real-world (Fern) datasets demonstrate that NeRF-KAN achieves PSNR improvements of 0.5 dB (28.7→29.2 dB) and 0.8 dB (25.9→26.7 dB) respectively over original NeRF, while maintaining parameter efficiency. This work not only validates KAN's significant potential in 3D vision tasks but also establishes a new paradigm beyond traditional MLP-based approaches for neural radiance field modeling, with promising applications in cultural heritage digitization and medical imaging reconstruction.
Xinyang Guo, Zhuohao Gong, Wenxin Hu
CloudCom3
2025 Battery Deformation Measurement Based on Electronic Speckle Pattern Interferometry
abstract
To accurately measure the full-field deformation of soft-packed lithium-ion batteries and improve the safety of battery industrial production, an electronic speckle pattern interferometry (ESPI) system was established. After irradiating the object with a laser, uniform speckle patterns were formed on the camera target plane. Quantitative measurement of battery deformation was conducted via charging and discharging experiments on lithium iron phosphate (LFP) lithium-ion batteries. Through the automatic control of the reference surface phase shifting device and the industrial camera, four-step phase shifting was completed, and the wrapped phase was acquired. A phase unwrapping algorithm combining temporal and spatial domains was proposed to obtain the absolute phase variation, thereby achieving battery deformation measurement. The results show that the system exhibits excellent deformation measurement capability.
Wenxin Hu
CloudCom3
2025 Causal Feature Supervision Decoupling: A Novel Method for Clothes-Changing Person Re-identification Algorithm
abstract
Clothes-Changing person re-identification algorithm (Re-ID) is the task of retrieving the query person in the case of the change of pedestrian clothing. Changes in pedestrian clothing lead to an offset in clothing features, resulting in a decrease in identification performance. Simply removing clothing may lead to the loss of contour information. Furthermore, algorithms based on feature decoupling cannot guarantee the accuracy of the positional relationship between clothing features and other features due to the lack of groundtruth. To address these issues, we propose a novel clothes-changing person re-identification algorithm based on causal feature supervision decoupling. Utilizing a multi-scale feature fusion module to extract fine-grained features of clothing and add supervised information labels. This enables the dual-branch network to separately approach overall and clothing features, promoting the extraction of effective identity information by the causal decoupling module, and obtaining unbiased estimations of pedestrians. Experimental results show that the proposed algorithm achieves the highest mAP and Top-1 accuracy on the LTCC and PRCC datasets. The source code is available at https://github.com/zhihu250/CISupNet.
Wenxin Hu, Caidan Zhao, Chenxing Gao, Zhiqiang Wu 0001
ICASSP1
2025 Book2QA: A Framework for Integrating LLMs to Generate High-quality QA Data from Textbooks
abstract
The scarcity of high-quality question answering (QA) data remains a significant bottleneck in the development of intelligent educational systems, as traditional datasets lack the necessary scale and diversity to support personalized model training. To address this challenge, we propose the Book2QA framework, which integrates multiple language models to provide an effective and flexible approach for generating QA datasets derived from textbook content. To further enhance the quality of the generated data, we implement a hierarchical prompting strategy grounded in Bloom’s taxonomy, substantially increasing both the depth and breadth of the QA datasets. In addition, we fine-tune our model using these data, with evaluations conducted by both human reviewers and GPT-4 confirming its strong performance in real-world questioning scenarios. Experimental results demonstrate that our framework excels not only in specific textbook domains but also shows promise for broader applications across diverse fields. We open source our data and code at https://github.com/Curtain2020/Book2QA.
Zhanhao Cui, Xinya Huang, Wenxin Hu
IJCNN5
2025 GSCL-KT: Improving Knowledge Tracing via Intra-Group Similarity Contrastive Learning
abstract
Knowledge tracing models have long grappled with the dual challenges of data sparsity and the limited ability to capture group learning patterns. Current contrastive learning paradigms in knowledge tracing (e.g., CL4KT framework) primarily employ sequence augmentation strategies to alleviate data scarcity constraints. However, such approaches frequently compromise semantic coherence during the augmentation process while failing to account for inherent similarity patterns within learner cohorts. To overcome these limitations, this paper introduces the GSCL-KT model (Group Similarity Contrastive Learning for Knowledge Tracing), which, for the first time, incorporates a group-similarity-aware contrastive learning mechanism into the knowledge tracing domain. Unlike traditional approaches that rely on manual data augmentation, GSCL-KT dynamically identifies positive and negative sample pairs from educationally homogeneous groups, enabling the discovery of group-level cognitive patterns while maintaining semantic coherence. The proposed model incorporates several advanced optimization strategies, including the Talking-Heads attention mechanism for fine-grained interaction modeling, the ContraNorm method for feature distribution regularization, and a correlation network enhanced by label dependencies. Experimental results on four real-world educational datasets demonstrate that GSCL-KT consistently outperforms existing baseline models, achieving the highest AUC and competitive performance across metrics.
Wenxin Hu
SMC3
2025 Knowledge-aware modeling of group commonality for group recommendation
Wen Wu 0006, Guangze Ye, Wenxin Hu, Liang He 0001
Expert Syst. Appl.4
2025 Data Augmentation Aided Automatic Modulation Recognition Using Diffusion Model
abstract
ABSTRACT Automatic modulation recognition enables rapid spectrum access and serves as a key technical component for achieving communication‐aware integration. In recent years, deep learning methods have attracted considerable attention in this field. However, their performance relies heavily on the availability of large‐scale datasets. Data augmentation has proven effective in mitigating data scarcity. To address this challenge, this paper proposes a data augmentation algorithm based on a conditional diffusion model to improve model training under limited data conditions. In the proposed framework, the noise observation model detects noise in the input signal at different time steps. The identified noise is then iteratively removed during the reverse process of the conditional diffusion model to generate the corresponding modulated signals. Generated signals with high confidence are incorporated into the training set to enhance data diversity. Experimental results demonstrate that the proposed algorithm significantly improves the classification network's performance and outperforms existing data augmentation approaches.
Caidan Zhao, Wenxin Hu, Minxin Cai
IET Commun.2
2025 Demographic-Guided Behavior Patterns Contrast for Personality Prediction
abstract
In recent years, personality has been considered as a valuable personal factor being incorporated into the provision of personalized learning. Although some studies have endeavored to obtain learners’ personalities implicitly from their learning behaviors, they failed to achieve satisfactory prediction performance. On the one hand, most existing approaches ignore the imbalanced distribution of personality classes, which causes the personality classifiers to be biased toward the non-extreme personality class. On the other hand, the related methods normally focus on constructing statistical behavior features, while the sequence information of learning behaviors is ignored, but actually it can reflect learners’ behavior patterns more finely. In this paper, inspired by the human learning strategy in the face of small samples, we propose an effective Demographic-Guided Behavior Patterns Contrast (DGBPC) model to classify learners’ personalities through the demographic-guided contrast of learners’ coarse behavior patterns. Besides, we construct and publish the Personality and Learning Behavior Dataset (PLBD), which should be one of the largest public datasets regarding Big-Five personality and learning behavior sequence according to our knowledge. The experimental results on PLBD demonstrate that our DGBPC model could generate learner representations with higher discrimination and outperform the related methods in terms of balanced accuracy.
Wen Wu 0006, Wenxin Hu, Liang Kang, Liang He 0001
IEEE Trans. Affect. Comput.4
2024 A Positive-Unlabeled Metric Learning Framework for Document-Level Relation Extraction with Incomplete Labeling
abstract
The goal of document-level relation extraction (RE) is to identify relations between entities that span multiple sentences. Recently, incomplete labeling in document-level RE has received increasing attention, and some studies have used methods such as positive-unlabeled learning to tackle this issue, but there is still a lot of room for improvement. Motivated by this, we propose a positive-augmentation and positive-mixup positive-unlabeled metric learning framework (P3M). Specifically, we formulate document-level RE as a metric learning problem. We aim to pull the distance closer between entity pair embedding and their corresponding relation embedding, while pushing it farther away from the none-class relation embedding. Additionally, we adapt the positive-unlabeled learning to this loss objective. In order to improve the generalizability of the model, we use dropout to augment positive samples and propose a positive-none-class mixup method. Extensive experiments show that P3M improves the F1 score by approximately 4-10 points in document-level RE with incomplete labeling, and achieves state-of-the-art results in fully labeled scenarios. Furthermore, P3M has also demonstrated robustness to prior estimation bias in incomplete labeled scenarios.
Huazheng Pan, Wen Wu 0006, Wenxin Hu
AAAI5
2024 Build a 50+ Hours Chinese Mandarin Corpus for Children's Speech Recognition
abstract
Children’s speech recognition plays an important role in the education research of children. The usual automatic speech recognition (ASR) systems are not satisfactory in terms of speech recognition for children, mainly due to the lack of child speech corpus. In recent years, there have been a large number of related studies on children’s speech recognition, most of which are improved based on the existing small amount of data through data augmentation, transfer learning, etc, which the promotion is limited. The purpose of our research is to establish a children’s speech corpus for children’s speech recognition. Firstly, we use the children’s language proficiency assessment system to collect weakly labeled children’s speech data from 3 to 12 years old, and obtain transcribed text through a commercial-grade speech recognition interface. Then, this paper proposes a method to quickly verify the transcribed text, and screens out effective label data pairs by combining children’s pronunciation and Chinese pinyin characteristics, and constructs a corpus of 50+ hours Mandarin children’s speech. We experimented with the current mainstream mandarin’s end-to-end pretrained model, finetuning on the data we built, and can achieve a 15% CER reduction on the basis of baseline without using finetuning.
Wenxin Hu
ICASSP4
2024 Enhancing Out-of-Distribution Generalization in VQA through Gini Impurity-guided Adaptive Margin Loss
abstract
In the Visual Question Answering (VQA) task context, most methods are influenced by language bias, resulting in poor performance on out-of-distribution data. Recently, some works attempted to use the adaptive margin loss to address this bias issue. However, these works typically consider only the frequency of answer labels when designing margin loss, leading to some samples being overly emphasized or lacking sufficient attention during model training. To address this issue, we propose a novel margin loss guided by the Gini-impurity for VQA debiasing. By comprehensively considering label distribution and instance complexity, we use Gini impurity to adjust the margin values in margin loss, balancing the attention of the model to different samples. Importantly, our method is plug-and-play and can be directly applied to any baseline. In the VQA-CP v2 task, our evaluation results across various baselines surpass the current state-of-the-art methods.
Tianyu Huai, Anran Wu, Xingjiao Wu, Wenxin Hu, Liang He 0001
ICME5
2024 KD-VSUM: A Vision Guided Models for Multimodal Abstractive Summarization with Knowledge Distillation
abstract
Multimodal abstract summarization is increasingly attracting attention due to its ability to synthesize information from different source modalities and generate high-quality text summaries. Concurrently, there has been significant development in multimodal abstract summarization models for videos. These models are capable of extracting information from multimodal data and generating abstract summaries. Most existing modeling approaches primarily concentrate on instructional videos, such as those teaching sports or life skills, thereby limiting their ability to capture the complexity of dynamic environments in the general world. In this paper, we propose a vision-guided model for multimodal abstractive summarization with knowledge distillation KD-VSUM to address the lack of generalized video domain capabilities in video summarization. This approach includes a vision-guided encoder, which enables the model to better focus on the global spatial and temporal information of video frames. We capitalize on knowledge distillation from multimodal pre-trained video-language models to enhance model performance. We introduce the VersaVision dataset, which includes a broader range of video domains and a higher proportion of medium to long videos. The results demonstrate that our model surpasses existing state-of-the-art models on the VersaVision dataset, achieving ROUGE scores of 1.7 in ROUGE-1, 1.8 in ROUGE-2, and 2 in ROUGE-L. These findings underscore the substantial improvements that the integration of a global vision guided and knowledge distillation can bring to the task of video summary extraction.
Zehong Zheng, Wenxin Hu
IJCNN3
2024 Personality-driven experience storage and retrieval for sentiment classification
Wen Wu 0006, Wenxin Hu, Liang He 0001
J. Supercomput.5
2023 A Two-stage Abdominal Lymph Node Detection Algorithm Based on Multi-scale 2.5D Input
abstract
Abdominal lymph nodes (ALNs) play a pivotal role in tumor metastasis, necessitating their early detection for tumor diagnosis and treatment decisions. However, due to varying ALN sizes and complex abdominal environment, existing methods face challenges, including small ALN absence, high False Positive (FP) rate and low efficiency. This study aims to propose a two-stage ALN detection approach based on multi-scale 2.5D input to overcome these difficulties. It leverages multi-scale images and the Skip-Attention module to generate ALN cadidate regions in the first stage. Then, a new 2.5D input is constructed by extracting slices from different planes and identifing genuine ALNs via the FP reduction algorithm. We validated our method with the dataset containing both public and real patient data. The first stage achieved an 81.1% detection rate with 21 FP/case, while the second stage reached an AUC value of 0.94 after ensemble learning. The final cascade experiment illustrates that our model achieved a 73.5% detection rate, surpassing other state-of-the-art methods including 3D models by at least a margin of 3.3% and significantly improving time efficiency.
Junyi Le, Wenxin Hu
ICPADS2
2023 Unlocking the Potentials of Transcriptomics in Predicting Pan Cancer Immune Therapy Response: A Deep Learning Approach Using PHZNet
abstract
The effectiveness of immune checkpoint blockade (ICB) therapy, a critical strategy in cancer immunotherapy, is often limited by a high rate of primary resistance. Identifying patients likely to respond to this therapy is a significant challenge. To tackle this issue, we introduce a novel deep learning model known as the Permutable Hybrid Zipped Network (PHZNet), which is specifically tailored for classifying patient responses to Immune Checkpoint Blockade (ICB) therapy. We achieved 89.67% accuracy on the transcriptomics public dataset from the Faculty of Pharmacy, Universit'e de Montr'eal, outperforming other excellent models. In addition, to explore interpretability, based on the concept of model distillation, we propose to use XGBoost gradient boosting decision trees to distill PHZNet and thus uncovering key gene features influencing therapy response. Our study provides a reliable basis for clinical decision-making by predicting whether patient benefits from ICB therapy across pan-cancer. Furthermore, it paves the way for understanding the ICB treatment response mechanism and provides critical gene targets for future immune treatment strategies.
Huazheng Pan, Junyi Le, Wenxin Hu, Taojun Jin
TrustCom4
2023 SMAR: Summary-Aware Multi-Aspect Recommendation
Liye Shi, Wen Wu 0006, Jiayi Chen 0002, Wenxin Hu, Liang He 0001
Neurocomputing4
2023 Personality-assisted mood modeling with historical reviews for sentiment classification
Wen Wu 0006, Jiayi Chen 0002, Wenxin Hu, Liang He 0001
Inf. Sci.6
2022 Knowledge-Enhanced Multi-task Learning for Course Recommendation
Qimin Ban, Wen Wu 0006, Wenxin Hu, Liang He 0001
DASFAA (2)3
2022 A Unified Positive-Unlabeled Learning Framework for Document-Level Relation Extraction with Different Levels of Labeling
abstract
Document-level relation extraction (RE) aims to identify relations between entities across multiple sentences. Most previous methods focused on document-level RE under full supervision.However, in real-world scenario, it is expensive and difficult to completely label all relations in a document because the number of entity pairs in document-level RE grows quadratically with the number of entities.To solve the common incomplete labeling problem, we propose a unified positive-unlabeled learning frameworkshift and squared ranking loss positive-unlabeled (SSR-PU) learning.We use positive-unlabeled (PU) learning on documentlevel RE for the first time.Considering that labeled data of a dataset may lead to prior shift of unlabeled data, we introduce a PU learning under prior shift of training data.Also, using none-class score as an adaptive threshold, we propose squared ranking loss and prove its Bayesian consistency with multi-label ranking metrics.Extensive experiments demonstrate that our method achieves an improvement of about 14 F1 points relative to the previous baseline with incomplete labeling.In addition, it outperforms previous state-of-the-art results under both fully supervised and extremely unlabeled settings as well. 1
Wenxin Hu
EMNLP3
2022 Denoising fMRI Message on Population Graph for Multi-site Disease Prediction
Yanyu Lin, Wenxin Hu
ICONIP (6)3
2022 Unexp-DIN: Unexpected Deep Interest Network for Recommendation
abstract
Studies on recommender systems persuaded accuracy of predictions, leading to the filter bubble problem which narrowed the user's interest areas. Unexpected recommendation (UR) is one of the solutions in handling the filter bubble. However, existing research on UR still faces two shortcomings. On the one hand, it is difficult to measure the unexpectedness of items to users accurately. On the other hand, there is a decrease in accuracy when making unexpected item recommendations. In this paper, we propose an Unexpected Deep Interest Network (Unexp-DIN) model to solve these problems. We use mean-shift for multiple clustering and Cluster-attention to maximize the unexpectedness. Then we use the constraint function and the hybrid utility function to simultaneously improve accuracy and unexpectedness. Experiments on three publicly available benchmark datasets show that our model has different degrees of improvement in both accuracy and unexpectedness.
Jiayi Chen 0002, Wen Wu 0006, Wenxin Hu, Liang He 0001
IJCNN4
2022 MTN-Net: A Multi-Task Network for Detection and Segmentation of Thyroid Nodules in Ultrasound Images
Leyao Chen, Wenxin Hu
KSEM (3)3
2022 Prediction of lymph node metastasis in CT based on multi-scale attention fully convolutional network
abstract
With the application of artificial intelligence technology in the medical field, the use of deep learning methods to predict lymph node (LN) metastasis from Computed Tomography (CT) images has become one of the important studies in adjuvant cancer treatment. Although many studies have made some progress, the large difference in LNs size has always limited the performance of traditional convolutional neural network methods. In this paper, we propose a multi-scale attention fully convolutional network for LN metastasis prediction in CT images. We aim to make the network compatible with LNs of different sizes to improve the overall performance of the network. First, the network extracts image features from multiscale input data. Then, we use an attention-based fusion module to study the relationship between features at different scales and adaptively fuse image features. Furthermore, a feature consistency loss is proposed by us to enhance the similarity between image features at different scales. The experimental results show that our proposed network achieves the best performance with the accuracy 92.37%, the sensitivity 90.21% and the specificity 95.71%, which outperforms several state-of the-art methods.
Weiping He, Wenxin Hu, Changquan Lu
SMC2
2022 SENGR: Sentiment-Enhanced Neural Graph Recommender
Liye Shi, Wen Wu 0006, Wang Guo, Wenxin Hu, Jiayi Chen 0002, Liang He 0001
Inf. Sci.4
2022 DualGCN: An Aspect-Aware Dual Graph Convolutional Network for review-based recommender
Liye Shi, Wen Wu 0006, Wenxin Hu, Jie Zhou 0015, Jiayi Chen 0002, Liang He 0001
Knowl. Based Syst.3
2021 LGACN: A Light Graph Adaptive Convolution Network for Collaborative Filtering
Weiguang Jiang, Su Wang 0005, Wenxin Hu
ICANN (3)4
2021 LSTMVAEF: Vivid Layout via LSTM-Based Variational Autoencoder Framework
Xingjiao Wu, Wenxin Hu, Jing Yang 0023
ICDAR (2)3
2021 SSR: Explainable Session-based Recommendation
abstract
In recent years, session-based recommendation has attracted more and more attention. Previous work utilizing deep learning approaches has achieved significant progress in the accuracy of prediction. However, why the user would view the next item is not clear, though attention mechanism assigns different importance on items in a session. Meanwhile, some traditional data mining-based methods are able to provide explanations, but they mainly offer a narrow perspective like measuring item similarity and fail to achieve as good performance as deep learning-based approaches. To this end, we propose an explainable session-based recommendation by considering three factors including sequential patterns, repetition clicks, and item similarity. Concretely, we design a two-stage framework where candidate items are firstly selected according to users' sequential patterns and repeated behaviors, and then they are further ranked by considering short-term interest with the calculation of item similarity. The experimental results on two benchmark datasets show that our model not only achieves competitive recommendation performance with advanced deep learning based models, but also provides reasonable explanations.
Jiayi Chen 0002, Wen Wu 0006, Wenxin Hu, Liang He 0001
IJCNN3
2021 Confidence-Aware Cascaded Network for Fetal Brain Segmentation on MR Images
Xukun Zhang, Zhiming Cui 0001, Changan Chen, Jingjiao Lou, Wenxin Hu, He Zhang 0023, Tao Zhou 0002, Feng Shi 0001, Dinggang Shen
MICCAI (3)6
2021 Attention Guided Filter for Jointly Extracting Entities and Classifying Relations
abstract
Jointly extracting entities and classifying relations aims to detect all possible triples from unstructured text with a single model.Tagging-based method effectively improves the performance of jointly relation extraction.However, some taggingbased approaches ignored that one entity pair may exist multiple relations and others set an empirical threshold value for selecting one or more relevant relations, which becomes the bottlenecks of the model.As a solution, we propose the attention guided filter, namely, AGFRel, which introduces transformer blocks to learn the number of relations for every entity pair to filter out irrelevant relations.Moreover, each module of the model has a multi-head attention guided layer to highlight valuable information.Extensive experimental results show that AGFRel is capable of gaining better performance on various tasks including overlapping triples extraction and multiple triples extraction.On NYT and WebNLG public datasets, our model obtains F1 score 90.8 and 91.9 respectively and achieves a new state-of-the-art performance.
Shaoze Chen, Wenxin Hu
SEKE3
2021 A hybrid deep generative neural model for financial report generation
Yunpeng Ren, Wenxin Hu, Xiaofeng Zhang 0002
Knowl. Based Syst.2
2020 Few-Shot Learning for Medical Image Classification
Aihua Cai, Wenxin Hu
ICANN (1)2
2020 TSCREC: Time-sync Comment Recommendation in Danmu-Enabled Videos
abstract
In recent years, Time-sync Comment (TSC), as known as “Danmaku” or “Danmu”, has been increasingly recognized as a valuable representation of video being incorporated into the process of generating video or highlight recommendation. However, little work has studied how to recommend proper TSC when users are watching the video. Providing some suggestions for users when they want to post TSCs can not only enhance the real-time interactions among users within videos, but also motivate users to post more TSCs which could be useful for video or highlight recommendation in return. In order to accomplish the TSC recommendation task, we extract candidates from existing comments sent by other users. Specifically, we propose a model, namely “TSCREC”, which uses a bidirectional Gated Recurrent Unit to capture the semantic meaning of existing comments and assigns scores for comments by a Multi-Layer Perceptron. Furthermore, considering the importance of correlations among comments, we also take similarities among comments as features. In addition, we design a set of novel evaluation metrics which combines n-grams overlapping and ranking order to measure the quality of the recommendation list. We conduct experiments on realworld datasets, and the results show that our TSCREC model outperforms baseline approaches.
Jiayi Chen 0002, Wen Wu 0006, Wenxin Hu, Liang He 0001
ICTAI3
2020 CPSPNet: Crowd Counting via Semantic Segmentation Framework
abstract
Crowd counting, i.e., estimation number of the pedestrian in crowd images, is emerging as an essential research problem with the public security applications. The density-based method of crowd counting still has some challenges, such as lack of perspective information in density map and background noise. Current models often misjudge background noise as a person and the ground truth density map widely used now is not so accurate. In this paper, we present a novel approach to help generate a higher quality density map. On the one hand, we eliminate the apparent mistakes in the density map with the help of a semantic segmentation model, which provides more information about fine-granted negative samples. On the other hand, we modify the density map to make sure it maintains a natural attribute. The experimental results prove the effectiveness of our method for crowd counting models, especially in uneven distribution monitoring scenario.
Xingjiao Wu, Jing Yang 0023, Wenxin Hu
ICTAI4
2020 Generating Financial Reports from Macro News via Multiple Edits Neural Networks
Wenxin Hu, Xiaofeng Zhang 0002, Yunpeng Ren
ECML/PKDD (3)1
2020 Counting crowds with varying densities via adaptive scenario discovery framework
Xingjiao Wu, Yingbin Zheng, Hao Ye 0005, Wenxin Hu, Tianlong Ma, Jing Yang 0023, Liang He 0001
Neurocomputing4
2020 Two-stage sentiment classification based on user-product interactive information
Wen Wu 0006, Qin Chen 0001, Wenxin Hu, Liang He 0001
Knowl. Based Syst.5
2019 Using Fractional Latent Topic to Enhance Recurrent Neural Network in Text Similarity Modeling
Yang Song 0010, Wenxin Hu, Liang He 0001
DASFAA (2)2
2019 Aggregating Rich Deep Semantic Features for Fine-Grained Place Classification
Tingyu Wei, Wenxin Hu, Xingjiao Wu, Yingbin Zheng, Hao Ye 0005, Jing Yang 0023, Liang He 0001
ICANN (3)2
2019 Adaptive Scenario Discovery for Crowd Counting
abstract
Crowd counting, i.e., estimation number of the pedestrian in crowd images, is emerging as an important research problem with the public security applications. A key component for the crowd counting systems is the construction of counting models which are robust to various scenarios under facts such as camera perspective and physical barriers. In this paper, we present an adaptive scenario discovery framework for crowd counting. The system is structured with two parallel pathways that are trained with different sizes of the receptive field to represent different scales and crowd densities. After ensuring that these components are present in the proper geometric configuration, a third branch is designed to adaptively recalibrate the pathway-wise responses by discovering and modeling the dynamic scenarios implicitly. Our system is able to represent highly variable crowd images and achieves state-of-the-art results in two challenging benchmarks.
Xingjiao Wu, Yingbin Zheng, Hao Ye 0005, Wenxin Hu, Jing Yang 0023, Liang He 0001
ICASSP4
2019 Cascaded Detail-Preserving Networks for Super-Resolution of Document Images
abstract
The accuracy of OCR is usually affected by the quality of the input document image and different kinds of marred document images hamper the OCR results. Among these scenarios, the low-resolution image is a common and challenging case. In this paper, we propose the cascaded networks for document image super-resolution. Our model is composed by the Detail-Preserving Networks with small magnification. The loss function with perceptual terms is designed to simultaneously preserve the original patterns and enhance the edge of the characters. These networks are trained with the same architecture and different parameters and then assembled into a pipeline model with a larger magnification. The low-resolution images can upscale gradually by passing through each Detail-Preserving Network until the final high-resolution images. Through extensive experiments on two scanning document image datasets, we demonstrate that the proposed approach outperforms recent state-of-the-art image super-resolution methods, and combining it with standard OCR system lead to signification improvements on the recognition results.
Zhichao Fu, Yingbin Zheng, Hao Ye 0005, Wenxin Hu, Jing Yang 0023, Liang He 0001
ICDAR5
2019 Chinese Clinical Named Entity Recognition with Word-Level Information Incorporating Dictionaries
abstract
Electronic Medical Records (EMRs) are the digital equivalent of paper records, which include treatment and medical history about a patient. At present, the main research goal of Chinese EMRS is to accurately recognize the body parts, drugs, illnesses and other information in the Chinese medical process. Implementing EMRs can boost both the quality and safety of patient care. In Chinese EMRs, how to accurately recognize named entities is important because it is useful to predict the disease risk, therapeutic method and recovery probability. This paper proposes a novel deep learning framework, which uses character-word joint embedding and combines different feature information based on the dictionary. Compared with the predecessors, we incorporate word-level information based on the basic Bi-LSTM model. In addition, we propose an improved n-gram feature encoding method and compare it with PDET feature and PIET feature. Our experimental results demonstrate that our proposed model performs the best in predicting named entities in Chinese EMRs.
Ningjie Lu, Wen Wu 0006, Yan Yang 0008, Kaiwei Chen, Wenxin Hu
IJCNN6
2019 Enhancing the Healthcare Retrieval with a Self-adaptive Saturated Density Function
Yang Song 0010, Wenxin Hu, Liang He 0001, Liang Dou 0001
PAKDD (1)2
2018 Enhancing the Recurrent Neural Networks with Positional Gates for Sentence Representation
Yang Song 0010, Wenxin Hu, Qin Chen 0001, Qinmin Hu, Liang He 0001
ICONIP (1)2
2016 Feature Selection in Click-Through Rate Prediction Based on Gradient Boosting
Qingsong Yu, Chaomin Shen 0001, Wenxin Hu
IDEAL4