VLDB 2026 Research / reviewers in the wild / expert
Dengshi Li
dblp:134/6330
· DBLP profile ↗
57ranked-venue papers
5as first author
45since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 2 first-author · 24 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 16 since 2021Computer networks · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WavGateMamba: A Frequency-Enhanced and Gated Mamba Model for Multimodal Depression Detection
Haiyang Ye, Dengshi Li, Yulin Wu 0003 |
MMM (1) | 2 |
| 2026 | CMDiff: Clip-guided multi-dimension mamba diffusion model for low light image enhancement
Chengbo Yu, Dengshi Li, Yulin Wu 0003, Aolei Chen |
Image Vis. Comput. | 2 |
| 2026 | Wavelet-driven meta-learning: unifying infrared-visible fusion and semantic segmentation for robust scene perception
Chun Sun, Dengshi Li, Shiwei Hu, Zhiming Zhan |
Vis. Comput. | 3 |
| 2025 | PSRDET: Fast Multimodal Detection Based on Prior Scene Repair for All-Weather Road Sensing
Chengbo Yu, Dengshi Li, Haiyang Ye |
ICANN (2) | 2 |
| 2025 | SE2E: Recognizing Emotion behind Societal BehaviorabstractEmotion recognition, as a core technology in mental health monitoring, has long been constrained by the intrusive nature of data collection methods relying on physiological signals and behavioral cues. Although existing motion-based approaches enable non-intrusive data acquisition, they often overlook the societal dimensions inherent in human behavior. As a result, they often exhibit a significant performance drop in real-world scenarios compared to laboratory settings. In this study, we analyzed the spatial distribution of participants' spatiotemporal trajectories and their visited Points of Interest (POIs), and observed significant differences under varying emotional states. Building on this observation, we propose a novel emotion recognition framework, SE2E, which innovatively incorporates the semantic information of POIs into the emotion recognition task. Specifically, SE2E employs a category-aware semantic embedding mechanism combined with a masked prediction task to ensure that the POI embeddings capture both categorical semantics and contextual information. It then structurally represents individual societal event patterns through a personalized spatiotemporal flow. Finally, a temporal-region consistency attention module is employed to extract continuous representations of societal events, thereby enabling a robust mapping from societal behavior to emotional state. Extensive experimental results demonstrate that SE2E outperforms state-of-the-art methods across multiple benchmarks. To the best of our knowledge, this is the first study to leverage societal event for emotion recognition, offering a new technical direction, benchmark, and insight for future research in the field. Wending Xiong, Ruimin Hu, Lingfei Ren, Dengshi Li |
ACM Multimedia | 5 |
| 2025 | Multimodal and multichannel speech separation using location-guided speech feature mapping network
Yulin Wu 0003, Xiaochen Wang 0001, Dengshi Li, Ruimin Hu |
Neurocomputing | 3 |
| 2024 | Unconventional Face Adversarial Attack
Baojin Huang, Zhen Han 0002, Dengshi Li |
ICANN (2) | 4 |
| 2024 | Noise Adaptive Fine-grained Speech Intelligibility Enhancement With Soft-label Guided DiffusionabstractBackground noise in the listening stage often affects the speech intelligibility and quality of communication devices, such as mobile phones. Traditional approaches like Near-end Listening Enhancement (NELE) aimed at processing speech signals to enhance intelligibility. Recent studies have been conducted to enhance intelligibility by converting normal speech to Lombard speech with varying level. However, overzealous focus on intelligibility improvement in previous research led to over-processing, causing speech distortion and quality degradation. Motivated by soft-label guidance, we propose a noise-adaptive fine-grained speech intelligibility enhancement framework—NELE-Diff. It fine-tunes Lombard intensity based on noise, incorporating multi-metric reinforcement learning into the diffusion model reverse process. Subjective and objective experiments reveal the superiority of NELE-Diff over baselines, presenting a more adaptive fine-grained speech intelligibility enhancement framework for different noise levels. Chenyi Zhu, Dengshi Li, Aolei Chen, Yu Gao 0018 |
ICME | 2 |
| 2024 | Robust Heterophilic Graph Learning against Label Noise for Anomaly Detection
Junhang Wu, Ruimin Hu, Dengshi Li, Lingfei Ren, Yilong Zang |
IJCAI | 3 |
| 2024 | Heterophilic Graph Invariant Learning for Out-of-Distribution of Fraud DetectionabstractGraph-based fraud detection (GFD) has garnered increasing attention due to its effectiveness in identifying fraudsters within multimedia data such as online transactions, product reviews, or telephone voices. However, the prevalent in-distribution (ID) assumption significantly impedes the generalization of GFD approaches to out-of-distribution (OOD) scenarios, which is a pervasive challenge considering the dynamic nature of fraudulent activities. In this paper, we introduce the Heterophilic Graph Invariant Learning Framework (HGIF), a novel approach to bolster the OOD generalization of GFD. HGIF addresses two pivotal challenges: creating diverse virtual training environments and adapting to varying target distributions. Leveraging edge-aware augmentation, HGIF efficiently generates multiple virtual training environments characterized by generalized heterophily distributions, thereby facilitating robust generalization against fraud graphs with diverse heterophily degrees. Moreover, HGIF employs a shared dual-channel encoder with heterophilic graph contrastive learning, enabling the model to acquire stable high-pass and low-pass node representations during training. During the Test-time Training phase, the shared dual-channel encoder is flexibly fine-tuned to adapt to the test distribution through graph contrastive learning. Extensive experiments showcase HGIF's superior performance over existing methods in OOD generalization, setting a new benchmark for GFD in OOD scenarios. Lingfei Ren, Ruimin Hu, Zheng Wang 0007, Yilin Xiao 0002, Dengshi Li, Junhang Wu, Yilong Zang, Jinzhang Hu |
ACM Multimedia | 5 |
| 2024 | Do not ignore heterogeneity and heterophily: Multi-network collaborative telecom fraud detection
Lingfei Ren, Yilong Zang, Ruimin Hu, Dengshi Li, Junhang Wu, Jinzhang Hu |
Expert Syst. Appl. | 4 |
| 2024 | A GNN-based fraud detector with dual resistance to graph disassortativity and imbalance
Junhang Wu, Ruimin Hu, Dengshi Li, Lingfei Ren, Wenyi Hu, Yilong Zang |
Inf. Sci. | 3 |
| 2024 | Improving fraud detection via imbalanced graph structure learning
Lingfei Ren, Ruimin Hu, Yang Liu 0200, Dengshi Li, Junhang Wu, Yilong Zang, Wenyi Hu |
Mach. Learn. | 4 |
| 2024 | STDC-Net: A spatial-temporal deformable convolution network for conference video frame interpolationabstractAbstract Video conference communication can be seriously affected by dropped frames or reduced frame rates due to network or hardware restrictions. Video frame interpolation techniques can interpolate the dropped frames and generate smoother videos. However, existing methods can not generate plausible results in video conferences due to the large motions of the eyes, mouth and head. To address this issue, we propose a Spatial-Temporal Deformable Convolution Network (STDC-Net) for conference video frame interpolation. The STDC-Net first extracts shallow spatial-temporal features by an embedding layer. Secondly, it extracts multi-scale deep spatial-temporal features through Spatial-Temporal Representation Learning (STRL) module, which contains several Spatial-Temporal Feature Extracting (STFE) blocks and downsample layers. To extract the temporal features, each STFE block splits feature maps along the temporal pathway and processes them with Multi-Layer Perceptron (MLP). Similarly, the STFE block splits the temporal features along horizontal and vertical pathways and processes them by another two MLPs to get spatial features. By splitting the feature maps into segments of varying lengths in different scales, the STDC-Net can extract both local details and global spatial features, allowing it to effectively handle large motions. Finally, Frame Synthesis (FS) module predicts weights, offsets and masks using the spatial-temporal features, which are used in deformable convolution to generate the intermediate frames. Experimental results demonstrate the STDC-Net outperforms state-of-the-art methods in both quantitative and qualitative evaluations. Compared to the baseline, the proposed method achieved a PSNR improvement of 0.13 dB and 0.17 dB on the Voxceleb2 and HDTF datasets, respectively. Qianrui Wang, Dengshi Li, Yu Gao 0018 |
Multim. Tools Appl. | 3 |
| 2024 | Fast global tone mapping for high dynamic range compression
Dengshi Li |
Multim. Tools Appl. | 2 |
| 2024 | SVMFI: speaker video multi-frame interpolation with the guidance of audio
Qianrui Wang, Dengshi Li, Yu Gao 0018, Aolei Chen |
Multim. Tools Appl. | 2 |
| 2024 | Beyond the individual: An improved telecom fraud detection approach based on latent synergy graph learning
Junhang Wu, Ruimin Hu, Dengshi Li, Lingfei Ren, Yilong Zang |
Neural Networks | 3 |
| 2023 | User and Interaction Both Matter: Social Relationship Mining Via Interaction Graph PropagatingabstractSocial relationship mining benefits many applications such as leadership analysis and advisor recommendation. Existing methods focus on mining user relationships only from the perspective of user-level. To our knowledge, from this perspective, representing the user interactions by edges is not sufficient for the complex information about interactions between users. In addition, mining users' relationship independently ignores the propagation of social interaction across networks. In this paper, we investigate social relationship mining from a new perspective of interaction-level. We propose an Interaction Graph Propagating(IGP) model which constructs an interaction graph. It not only captures the user interaction information as the union but also exploits the propagation between user interactions. In particular, we utilize the graph attention mechanism to distinguish the contributions of each neighbor union. Experimental results on several public datasets demonstrate that IGP achieves significant improvements over state-of-the-art methods. Yilong Zang, Ruimin Hu, Zheng Wang 0007, Dengshi Li |
ICC | 5 |
| 2023 | ASVFI: Audio-Driven Speaker Video Frame InterpolationabstractDue to limited data transmission, the video frame rate is low during the online conference, severely affecting user experience. Video frame interpolation can solve the problem by interpolating intermediate frames to increase the video frame rate. Generally, most existing video frame interpolation methods are based on the linear motion assumption. However, the mouth motion is nonlinear, and these methods can not generate superior intermediate frames in speaker video. Considering the strong correlation between mouth shape and vocalization, a new method is proposed, named Audio-driven Speaker Video Frame Interpolation(ASVFI). First, we extract the audio feature from Audio Net(ANet). Second, we use Video Net(VNet) encoder to extract the video feature. Finally, we fuse the audio and video features by AVFusion and decode out the intermediate frame in the VNet decoder. The experimental results show that the PSNR is nearly 0.13dB higher than the baseline of interpolating one frame. When interpolating seven frames, the PSNR is 0.33dB higher than the baseline. Qianrui Wang, Dengshi Li, Jing Xiao 0004 |
ICIP | 2 |
| 2023 | Hidden Follower Detection via Refined Gaze and Walking State EstimationabstractHidden following is following behavior with special intentions, and detecting hidden following behavior can prevent many criminal activities in advance. The previous method uses gaze and spacing behaviors to distinguish hidden followers from normal pedestrians. However, they express gaze behaviors in a coarse-grained way with binary values, making it difficult to accurately depict the gaze state of pedestrians. To this end, we propose the Refined Hidden Follower Detection (RHFD) model by choosing a suitable mapping function based on the principle that the closer the gaze direction is to someone, the more likely it is to gaze at someone, which converts the gaze direction into a continuous estimated gaze state representing the complex and variable gaze behavior of pedestrians. Simultaneously, we introduce variations in the magnitude and direction of pedestrian velocity to refine the representation of pedestrian walking states. Experimental results on the surveillance dataset show that RHFD outperforms state-of-the-art methods. Yaxi Chen, Ruimin Hu, Danni Xu, Zheng Wang 0007, Linbo Luo 0001, Dengshi Li |
ICME | 6 |
| 2023 | Noise adaptive speech intelligibility enhancement based on improved StarGAN*abstractWhen people communicate in noisy environments with the phone, it is difficult for listeners to obtain information even if the device outputs clear speech. Previous studies have focused on speech intelligibility enhancement (IENH) via normal speech and different levels of Lombard speech conversion. However, these methods often lead to speech distortion and impair the overall speech quality. We propose an IENH framework based on an improved Star Generative Adversarial Network (StarGAN) named D2StarGAN. It has two main advantages: 1) Inspired by the dual-discriminator idea, we add a speech metric discriminator based on StarGAN to optimize multiple intelligibility-related metrics simultaneously; 2) The framework can adapt to different far-and-near-end noise levels and different noise types. Experimental results using both objective measurements and subjective listening tests indicate that the proposed method outperforms the baseline method. It adaptively converts all the mobile communication scenarios with far-and-near-end noise, thus making IENH more widely used in practice. Lanxin Zhao, Dengshi Li, Jing Xiao 0004, Chenyi Zhu |
ICME | 2 |
| 2023 | Don't Ignore Alienation and Marginalization: Correlating Fraud DetectionabstractThe anonymity of online networks makes tackling fraud increasingly costly. Thanks to the superiority of graph representation learning, graph-based fraud detection has made significant progress in recent years. However, upgrading fraudulent strategies produces more advanced and difficult scams. One common strategy is synergistic camouflage —— combining multiple means to deceive others. Existing methods mostly investigate the differences between relations on individual frauds, that neglect the correlation among multi-relation fraudulent behaviors. In this paper, we design several statistics to validate the existence of synergistic camouflage of fraudsters by exploring the correlation among multi-relation interactions. From the perspective of multi-relation, we find two distinctive features of fraudulent behaviors, i.e., alienation and marginalization. Based on the finding, we propose COFRAUD, a correlation-aware fraud detection model, which innovatively incorporates synergistic camouflage into fraud detection. It captures the correlation among multi-relation fraudulent behaviors. Experimental results on two public datasets demonstrate that COFRAUD achieves significant improvements over state-of-the-art methods. Yilong Zang, Ruimin Hu, Zheng Wang 0007, Danni Xu, Jia Wu 0001, Dengshi Li, Junhang Wu, Lingfei Ren |
IJCAI | 6 |
| 2023 | Collaborative Fraud Detection: How Collaboration Impacts Fraud DetectionabstractCollaborative fraud has become increasingly serious in telecom and social networks, but is hard to detect by traditional fraud detection methods. In this paper, we find a significant positive correlation between the increase of collaborative fraud and the degraded detection performance of traditional techniques, implying that those fraudsters that are difficult to detect with traditional methods are often collaborative in their fraudulent behavior. As we know, multiple objects may contact a single target object over a period of time. We define multiple objects with the same contact target as generalized objects, and their social behaviors can be combined and processed as the social behaviors of one object. We propose Fraud Detection Model based on Second-order and Collaborative Relationship Mining (COFD), exploring new research avenues for collaborative fraud detection. Our code and data are released at https://github.com/CatScarf/COFD-MM https://github.com/CatScarf/COFD-MM. Jinzhang Hu, Ruimin Hu, Zheng Wang 0007, Dengshi Li, Junhang Wu, Lingfei Ren, Yilong Zang |
ACM Multimedia | 4 |
| 2023 | Deformable Spatial-Temporal Attention for Lightweight Video Super-Resolution
Tong Xue, Xinyi Huang 0005, Dengshi Li |
PRCV (10) | 3 |
| 2023 | Self-guided Transformer for Video Super-Resolution
Tong Xue, Qianrui Wang, Xinyi Huang 0005, Dengshi Li |
PRCV (10) | 4 |
| 2023 | Dynamic graph neural network-based fraud detectors against collaborative fraudsters
Lingfei Ren, Ruimin Hu, Dengshi Li, Yang Liu 0200, Junhang Wu, Yilong Zang, Wenyi Hu |
Knowl. Based Syst. | 3 |
| 2023 | Who is your friend: inferring cross-regional friendship from mobility profiles
Lingfei Ren, Ruimin Hu, Dengshi Li, Zheng Wang 0007, Junhang Wu, Wenyi Hu |
Multim. Tools Appl. | 3 |
| 2023 | Where Have You Gone: Category-aware Multigraph Embedding for Missing Point-of-Interest Identification
Junhang Wu, Ruimin Hu, Dengshi Li, Yilin Xiao 0002, Lingfei Ren, Wenyi Hu |
Neural Process. Lett. | 3 |
| 2023 | A Dual Self-Attention mechanism for vehicle re-Identification
Wenqian Zhu, Zhongyuan Wang 0001, Xiaochen Wang 0001, Ruimin Hu, Huikai Liu, Chao Wang 0084, Dengshi Li |
Pattern Recognit. | 8 |
| 2023 | From Collective Attribute Association of Groups to Precise Attribute Association of IndividualsabstractObscured person re-identification (Re-ID) aims to match an obscured image with a complete image of the same person captured by other cameras. As a major challenge in person identification, occlusion severely affects the effectiveness of most traditional person Re-ID methods. To solve this problem, this study proposes a trajectory association method, which, as a pre-processing technique for person Re-ID, can narrow the search range and reduce the problem of degradation caused by mixing. We investigate the method of converting the fuzzy association between sets into the precise association between elements for M video objects and N phone objects (trajectory information) with fuzzy group association relationships at the crime scene. First, we decompose the M-N precise association problem and analyze the similarity of the video objects in the source point and on the trajectories. Then, we define high-similarity points, study their distribution characteristics in different trajectories, and find that there is a significant difference between the distribution of high-similarity points in correct and incorrect matching trajectories. We simplify the full-path association problem into a partial-path high-similarity point distribution difference problem, which effectively reduces the difficulty in accurate association relationship construction. The association experiments in simple and mixed scenarios as well as Re-ID experiments on the PRPW and Market1501 demonstrate the effectiveness of our method. Yun Lan, Ruimin Hu, Xin Xu 0007, Dengshi Li, Chao Wang 0084, Xiaochen Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | A Bi-directional Category-Aware Multi-task Learning Framework for Missing Check-in POI Identification
Junhang Wu, Ruimin Hu, Dengshi Li, Lingfei Ren, Wenyi Hu, Yilong Zang |
ICSOC | 3 |
| 2022 | IDGL: An Imbalanced Disassortative Graph Learning Framework for Fraud Detection
Junhang Wu, Ruimin Hu, Dengshi Li, Lingfei Ren, Wenyi Hu, Yilong Zang |
ICSOC | 3 |
| 2022 | $(\alpha, \ \beta)$-AWCS: $(\alpha, \ \beta)$-Attributed Weighted Community Search on Bipartite GraphsabstractCommunity search on bipartite graphs aims to find a community closely associated with the query vertex for personalized recommendation, fraud detection, and team formation. During community search, considering both the structural closeness and the homogeneity of attributes of nodes is the key to improving the quality of the output community. Traditional work uses the$(\alpha,\beta)$-core model to guarantee structural cohesion of the nodes (i.e., degree of each upper vertex is at least$\alpha$and degree of each lower vertex is at least$\beta$). However, it ignores the attributes of nodes, resulting in an average attribute similarity of only about 0.17 for the node pairs in the output community. In this paper, a framework for$(\alpha,\ \beta)$-Attributed Weighted Community Search ($(\alpha,\ \beta)$-AWCS) was proposed. It output a connected subgraph of$G$containing the query vertex, which satisfies both structurally cohesive (i.e., ($(\alpha,\ \beta){-}$-core) and keyword cohesiveness (i.e., its vertices share common keywords). The framework includes a pruning strategy to strip vertices that do not contain query attributes, thus effectively reducing the search space, and two algorithms improve the attributes cohesiveness of the output community. One of the exact algorithms first obtains a subgraph of attribute cohesion and subsequently keeps the structure cohesive. The other approximate algorithm has higher robustness, which iteratively removes the vertex with the lowest attribute score until the structural cohesion cannot be maintained. We have conducted experiments on real datasets of different sizes. Experiments show that both algorithms can improve the attribute cohesiveness metric by more than 25% compared to the traditional method. Meanwhile, structural cohesion was appropriate. Dengshi Li, Xiaocong Liang, Ruimin Hu, Xiaochen Wang 0001 |
IJCNN | 1 |
| 2022 | ITC: Influential-Truss Community SearchabstractCommunity search is a method of finding a com-munity closely related to a query node. The latest influence community search considers both the structural cohesion of the community and the influence between nodes. It sets the influence threshold to constrain the output community. However, artificially setting the influence threshold makes the output community too large or too small, which leads to low accuracy of the output community. In order to avoid the low accuracy of community search caused by artificially setting influence thresh-old constraints, this paper studies the community search problem based on community influence score. In this paper, an influence-truss community (ITC) model is proposed for community search by combining structural cohesion and community influence score. This model aims to obtain a connected subgraph in a social network containing the query node, which satisfies structural cohesion and satisfies the subgraph's maximum community in-fluence score. In order to obtain ITC, an effective pruning method is proposed, which strips other nodes far away from the query node. Then, the ITCS algorithm is designed, which firstly imposes structural cohesion constraints on query nodes. Then, the search community's influence scores are iteratively calculated until the community has the highest community influence score under the condition of meeting the structural cohesion. Experiments on real-world networks of different scales show that the community search accuracy index of ITCS is improved by about 20% compared with the traditional method. Dengshi Li, Ruimin Hu, Xiaocong Liang, Yilong Zang |
IJCNN | 1 |
| 2022 | User Alignment Across Social Networks Based On ego-Network EmbeddingabstractCross-social network user alignment is to find users with the same identity in multiple social networks. It has important applications in natural and scientific fields, such as link prediction and personality recommendation, and has certain research value in the field of data mining. Most current approaches embed social networks in a low-dimensional vector space and then align users in the low-dimensional space. However, because the social network is extremely complex and large, it is easy to be affected by error propagation and noise of different neighbors in the process of network embedding. Therefore, to obtain better embedding, we first form the user's EGO network, then use the random walk to extract the user node sequence, then use the framework of the natural language model to learn the low-dimensional vector representation of the user, and finally train a matrix to map the two social networks into the same feature space for alignment. Our experiments on real-world data set Foursquare-Twitter and Livejournal-myspace show some improvement over several baseline results. Yu Zhen, Ruimin Hu, Dengshi Li, Yilin Xiao 0002 |
IJCNN | 3 |
| 2022 | Cross-Regional Friendship Inference via Category-Aware Multi-Bipartite Graph EmbeddingabstractThis paper proposes a novel problem of cross-regional friendship inference to solve the geographically restricted friends recommendation. Traditional approaches rely on a fundamental assumption that friends tend to be co-location, which is unrealistic for inferring friendship across regions. By reviewing a large-scale Location-based Social Networks (LBSNs) dataset, we spot that cross-regional users are more likely to form a friendship when their mobility neighbors are of high similarity. To this end, we propose Category-Aware Multi-Bipartite Graph Embedding (CMGE for short) for cross-regional friendship inference. We first utilize multi-bipartite graph embedding to capture users’ Point of Interest (POI) neighbor similarity and activity category similarity simultaneously, then the contributions of each POI and category are learned by a category-aware heterogeneous graph attention network. Experiments on the real-world LBSNs datasets demonstrate that CMGE outperforms state-of-the-art baselines. Linfei Ren, Ruimin Hu, Dengshi Li, Junhang Wu, Yilong Zang, Wenyi Hu |
LCN | 3 |
| 2022 | Gaze- and Spacing-flow Unveil Intentions: Hidden Follower DiscoveryabstractWe raise a new and challenging multimedia application in video surveillance system, i.e., Hidden Follower Discovery (HFD). In contrast to the common abnormal behaviors that are occurring, hidden following is not an ongoing activity, but a preparatory action. Hidden following behavior does not have salient features, making it hard to be discovered. Fortunately, from a socio-cognitive perspective, we found and verified the phenomena that the gaze-flow pattern and the spacing-flow pattern between hidden and normal followers are different. To promote HFD research, we construct two pioneering datasets and devise an HFD baseline network based on the recognition of both gaze-flow and spacing-flow patterns from surveillance videos. Extensive experiments demonstrate their effectiveness. Danni Xu, Ruimin Hu, Zheng Wang 0007, Linbo Luo 0001, Dengshi Li, Wenjun Zeng 0001 |
ACM Multimedia | 5 |
| 2022 | Adaptive Speech Intelligibility Enhancement for Far-and-Near-end Noise Environments Based on Self-attention StarGAN
Dengshi Li, Lanxin Zhao, Jing Xiao 0004, Duanzheng Guan, Qianrui Wang |
MMM (2) | 1 |
| 2022 | Speech Intelligibility Enhancement By Non-Parallel Speech Style Conversion Using CWT and iMetricGAN Based CycleGAN
Jing Xiao 0004, Dengshi Li, Lanxin Zhao, Qianrui Wang |
MMM (1) | 3 |
| 2022 | Where have you been: Dual spatiotemporal-aware user mobility modeling for missing check-in POI identification
Junhang Wu, Ruimin Hu, Dengshi Li, Lingfei Ren, Wenyi Hu, Yilin Xiao 0002 |
Inf. Process. Manag. | 3 |
| 2021 | Multi-level Graph Attention Network based Unsupervised Network AlignmentabstractNetwork alignment is the matching of two networks with corresponding nodes that belong to the same user or entity. The most common application is to analyze which accounts belong to the same user in two social networks. Most of existing techniques rely on matrix factorization so that they cannot be scaled to large-scale networks, are constrained by strict constraints, and cannot learn node embedding without a training set. In this paper, we propose an unsupervised network alignment model based on multi-level graph attention networks. The model uses multi-level graph attention network to learn the embedded representation of nodes, satisfying attribute and structure constraints of alignment. Augmented learning process is proposed to simulate attribute noise and structural noise to improve adaptability of the model. Extensive experiments on real datasets show that the proposed model performs better than the state-of-the-art network alignment model. We also demonstrate the robustness of the proposed model. Yilin Xiao 0002, Ruimin Hu, Dengshi Li, Junhang Wu, Yu Zhen, Lingfei Ren |
LCN | 3 |
| 2021 | Trajectory is not Enough: Hidden Following DetectionabstractIn outdoor crimes such as robbery and kidnapping, suspects generally secretly follow their victims in public places and then look for opportunities to commit crimes. Video anomaly detection (VAD) has achieved fruitful results through deep neural networks (DNN). However, as an abnormal behavior without obvious abnormal physical features, hidden following is highly similar to ordinary walking and accompanying behaviors, so it is difficult to effectively detect hidden dangerous followers using video anomaly detection methods or traditional trajectory analysis methods. We propose "hidden follower'' detection (HFD) task and a HFD model based on gaze pattern extraction. It extracts gaze pattern features of pedestrians from gaze-interval-series and introduces a time series classification model to classify pedestrians with or without hidden following purposes. Based on this model, we propose a hidden follower detection framework (HFDF) to detect hidden followers from normal pedestrians, which utilizes the trajectories and gaze patterns extracted from videos. To cope with the lack of test data, we construct a dataset of 1200 pedestrians from the crowd simulation model to simulate scenes including hidden followers, and we also collected a surveillance video dataset including the hidden following behaviors. The experiments conducted on these two datasets show that HFDF can consistently outperform the state-of-the-art method by a notable margin in the HFD task on the commonly-used F1 benchmark. Danni Xu, Ruimin Hu, Zixiang Xiong, Zheng Wang 0007, Linbo Luo 0001, Dengshi Li |
ACM Multimedia | 6 |
| 2021 | Optimization of sound fields reproduction based Higher-Order Ambisonics (HOA) using the Generative Adversarial Network (GAN)
Lingkun Zhang, Xiaochen Wang 0001, Ruimin Hu, Dengshi Li, Weiping Tu |
Multim. Tools Appl. | 4 |
| 2021 | Estimation of spherical harmonic coefficients in sound field recording using feed-forward neural networks
Lingkun Zhang, Xiaochen Wang 0001, Ruimin Hu, Dengshi Li, Weiping Tu |
Multim. Tools Appl. | 4 |
| 2021 | Trajectory Association for Person Re-identification
Ruimin Hu, Wenxin Huang, Dengshi Li, Xiaochen Wang 0001, Chenhao Hu |
Neural Process. Lett. | 4 |
| 2020 | Social-IFD: Personalized Influential Friends Discovery Based on Semantics in LBSNabstractSocial influence is a hot topic in social network research, and this paper focuses on how to search for the most influential friends for a target user. The key point is to measure the influence between different users, such as adjacent users and the non-adjacent users. However, the traditional method, called the IS model, can only calculate the influence strength between neighboring users based on the inner-product of the Influence vector and the Susceptibility vector. In this paper, the social-IFD algorithm is proposed to compute the influence between different users (not only neighboring users but also non-adjacent users) based on network structure and the semantic information of users in LBSN, which has promoted the development of the IS model. Furthermore, we propose social-IFD ++ algorithm based on dynamic program to reduce the complexity of the social-IFD algorithm. Experiment results on two real large-scale network show that the average precision of the proposed social-IFD algorithm is 30.5% higher than the average precision of the IS model. In addition, the CPU running time of the proposed social-IFD ++ algorithm is nearly ten times lower than that of the IS model. It indicates that the proposed two algorithms have superior performance. Ruimin Hu, Dengshi Li |
ICC | 3 |
| 2020 | Tell The Truth From The Front: Anti-Disguise Vehicle Re-IdentificationabstractRecent efforts have been increasingly made on vehicle reidentification (re-ID), which has huge contributions to intelligent transportation and criminal investigation. However, most existing methods heavily rely on the color and texture features of vehicles to discern their identities, which turn invalid under adversarial social security occasions where vehicles' color and style are always tampered or forged by crime suspects. In this paper, we propose a local feature preservation method to learn the structure-aware features from the position distribution of individual local regions within vehicle front window area, which appears more robust and discriminative upon disguise. We further develop a two-branch deep convolutional network framework to integrate the structure-aware features with vehicle model features for vehicle Re-ID. The experimental results on datasets VehicleID and Vehicle-1M show that our end-to-end framework achieves promising performance and outperforms the state-of-the-art methods proposed so far. Wenqian Zhu, Ruimin Hu, Zhongyuan Wang 0001, Dengshi Li, Xiyue Gao |
ICME | 4 |
| 2020 | HRTF Representation with Convolutional Auto-encoder
Wei Chen 0143, Ruimin Hu, Xiaochen Wang 0001, Dengshi Li |
MMM (1) | 4 |
| 2020 | Perceptual Localization of Virtual Sound Source Based on Loudspeaker Triplet
Duanzheng Guan, Dengshi Li, Xuebei Cai, Xiaochen Wang 0001, Ruimin Hu |
MMM (2) | 2 |
| 2020 | Multi-step Coding Structure of Spatial Audio Object Coding
Chenhao Hu, Ruimin Hu, Xiaochen Wang 0001, Tingzhao Wu, Dengshi Li |
MMM (1) | 5 |
| 2020 | HMM-Based Person Re-identification in Large-Scale Open Scenario
Ruimin Hu, Wenxin Huang, Xiaochen Wang 0001, Dengshi Li |
MMM (1) | 5 |
| 2020 | Loudspeaker triplet selection based on low distortion within head for multichannel conversion of smart 3D home theaterabstractSummary In recent years, with the vigorous development of 3D film industry, the demand for 3D Smart Home Theater, based on Internet of Things (IoT), continues to grow. Theaters are populated with a large number of loudspeakers for more realistic 3D sound effects. However the number of loudspeakers in home is limited. Therefore, multichannel conversion is required to achieve theater 3D sound effects in home. Traditionally, the replaced loudspeaker signal of the original system is assigned to a “loudspeaker triplet” of the converted system. A large amount of subjective evaluations is necessary to judge the consistency of the replaced loudspeaker position with the “phantom source” positions reconstructed by loudspeaker triplets. In this study, after calculating the least‐squares errors of the reproduced sound field within a given region, we explore the constraint between the low distortion of the reproduced sound field and the loudspeaker triplet positions. Using this constraint rule, we present a new loudspeaker triplet selection criteria that can greatly reduce the number and time of subjective evaluations for selecting the optimal loudspeaker triplet. Simulation and subjective evaluation experiments indicate that the proposed selection method outperforms the traditional method, and that the proposed method can be successfully applied to multichannel conversion. Dengshi Li, Ruimin Hu, Xiaochen Wang 0001, Weiping Tu |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | Deep Structural Feature Learning: Re-Identification of simailar vehicles In Structure-Aware Map SpaceabstractVehicle re-identification (re-ID) has received more attention in recent years as a significant work, making huge contribution to the intelligent video surveillance. The complex intra-class and inter-class variation of vehicle images bring huge challenges for vehicle re-ID, especially for the similar vehicle re-ID. In this paper we focus on an interesting and challenging problem, vehicle re-ID of the same/similar model. Previous works mainly focus on extracting global features using deep models, ignoring the individual loa-cal regions in vehicle front window, such as decorations and stickers attached to the windshield, that can be more discriminative for vehicle re-ID. Instead of directly embedding these regions to learn their features, we propose a Regional Structure-Aware model (RSA) to learn structure-aware cues with the position distribution of individual local regions in vehicle front window area, constructing a FW structural map space. In this map sapce, deep models are able to learn more robust and discriminative spatial structure-aware features to improve the performance for vehicle re-ID of the same/similar model. We evaluate our method on a large-scale vehicle re-ID dataset Vehicle-1M. The experimental results show that our method can achieve promising performance and outperforms several recent state-of-the-art approaches. Wenqian Zhu, Ruimin Hu, Zhongyuan Wang 0001, Dengshi Li, Xiyue Gao |
MMAsia | 4 |
| 2017 | The Perceptual Lossless Quantization of Spatial Parameter for 3D Audio Signals
Xiaochen Wang 0001, Ruimin Hu, Dengshi Li |
MMM (2) | 5 |
| 2016 | Multichannel reduction based on sound field within two earsabstractPeople hope to use a small number of loudspeakers to get the experience of the film 3D sound at home. Considering that people use two ears to listen, this paper provides a method which reproduce the sound field within the region of two ears. We develop the fundamental performance limits for the truncated spherical harmonic function expansions of the sound field within the region of ears. Based on this, the low distortion of reproduced sound field within two ears is maintained in the processing of reducing loudspeakers from Q to Q-1. The 22.2 multichannel sound system without two low-frequency effect channels can be simplified to 6 channels automatically and the total of loudspeaker arrangements is ten. The subjective evaluation of the proposed method is better than that of the previous multichannel reduction method with the decrease of the number of loudspeakers. Dengshi Li, Ruimin Hu, Xiaochen Wang 0001, Guo Wu, Weiping Tu |
ICME | 1 |
| 2016 | Adaptive Multichannel Reduction Using Convex Polyhedral Loudspeaker Array
Lingkun Zhang, Ruimin Hu, Dengshi Li, Xiaochen Wang 0001, Weiping Tu |
MMM (1) | 3 |
| 2015 | Spatial perception reproduction of sound events based on sound property coincidencesabstractSound pressure and particle velocity are used to reproduce sound signals in multichannel systems. The two sound properties were estimated step by step and particle velocity was scaled due to ill-conditioned equations in Ando's study. We explore a new system of equations to maintain both sound pressure and particle velocity. The weight equations are solved in a non-traditional way to figure out exact solutions. Based on the proposed method, the perception of the direction of a sound event and the distance to the listening point are both reproduced correctly in a three-dimension reproduction system. The comparison between the proposed method and Ando's method is outlined and the proposed method is more flexible and useful. Objective evaluation shows the wavefront in the proposed method is more accurate than Ando's method and subjective evaluation confirms that the proposed method improves the spatial perception of sound events. Maosheng Zhang, Ruimin Hu, Xiaochen Wang 0001, Dengshi Li |
ICME | 5 |