VLDB 2026 Research / reviewers in the wild / expert
Hao Chen 0051
dblp:175/3324-51
· DBLP profile ↗
22ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0001-5902-1824ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ImCapDA: Fine-tuning CLIP via image captions for unsupervised domain adaptation
Weiwei Xiang, Guangyi Xiao 0001, Shun Peng, Hao Chen 0051, Liming Ding, Lei Yang 0026 |
Expert Syst. Appl. | 4 |
| 2026 | ODPL-CLIP: Open differential prompt learning with CLIP for open-set domain adaptation
Dezhong Li, Guangyi Xiao 0001, Hao Chen 0051 |
Pattern Recognit. | 3 |
| 2025 | Do Adversarial Perturbations Truly Mitigate Gradient Inversion in Federated Learning?abstractFederated Learning (FL) has emerged as a privacy-preserving framework under which multiple participants jointly solve the a collaborative training and user privacy problem. Recent studies find that private training data can still be leaked by the exchanged gradients based on optimization or analytic, i.e., gradient inversion attacks (GIAs). To enhance privacy, adversarial perturbations (AP) are attempted to be applied in FL by introducing carefully crafted noise into the local gradients. However, the effectiveness of adversarial perturbations in strengthening privacy against GIAs remains underexplored. In this work, we empirically evaluate adversarial perturbations on the resistance of GIAs. We show that even adversarial perturbations added to the gradients can still leak training data. We propose an adversarial perturbation gradient inversion attack, APT. Specifically, we design a feature extraction method to extract data features from adversarial perturbation gradients by utilizing the linear layer. Moreover, we design a feature reconstruction method to reconstruct data by a key feature reconstructor. Extensive experiments demonstrate that our method achieves high-quality gradient inversion from adversarial perturbation gradients, surpassing state-of-the-art methods that commonly fail in more challenge scenarios. Overall, our work explores the defense effectiveness as well as reveals the vulnerability of AP under GIAs. We hope this work provide valuable insights into leveraging adversarial perturbations for privacy defense and inspire future research on robust privacy-preserving mechanisms in FL. Hui Zhou 0014, Zheng Qin 0001, Yipeng Zou, Ge Xiao, Hao Chen 0051 |
IJCNN | 7 |
| 2025 | LDInfer: Landmarks Inference-Based for Facial Forgery Detection of Different QualityabstractThe harm caused by deepfakes is becoming increasingly serious, including financial fraud, guiding political public opinion, and more. The vast majority of deepfake detection methods typically perform well on uncompressed deepfake videos, achieving satisfactory results. However, when facing deepfake videos with different compression rates, the effectiveness of certain detection methods will significantly decrease, and they may even be unable to detect compressed videos. To address this challenge, we introduce a new detection mechanism: landmarks inference mechanism. Based on this, we propose a simple and efficient model LDInfer for inferring landmarks to achieve deepfake video detection. Our method utilizes the inference between facial landmarks and their preceding and succeeding frames to improve detection accuracy, which remains effective even in compressed deepfake videos. This method first infers landmarks and verifies their authenticity using the correlation between landmarks in the previous and subsequent frames. Even if the video is compressed, the correlation between facial landmarks in the preceding and following frames remains high. Comparative experiments show that our method outperforms existing methods at various compression rates. Hui Zhou 0014, Tianshuo Jiao, Bohan Tan, Zhuo Zhang 0028, Hao Chen 0051, Zheng Qin 0001 |
TrustCom | 8 |
| 2024 | Unified multi-level neighbor clustering for Source-Free Unsupervised Domain Adaptation
Yuzhe Xiao, Guangyi Xiao 0001, Hao Chen 0051 |
Pattern Recognit. | 3 |
| 2024 | Efficient Cross-Modal Video Retrieval With Meta-Optimized FramesabstractCross-modal video retrieval aims to retrieve semantically relevant videos when given a textual query, and is one of the fundamental multimedia tasks. Most top-performing methods primarily leverage Vision Transformer (ViT) to extract video features (Lei et al., 2021}, (Bain et al., 2021), (Wang et al., 2022). However, they suffer from the high computational complexity of ViT, especially when encoding long videos. A common and simple solution is to uniformly sample a small number (e.g., 4 or 8) of frames from the target video (instead of using the whole video) as ViT inputs. The number of frames has a strong influence on the performance of ViT, e.g., using 8 frames yields better performance than using 4 frames but requires more computational resources, resulting in a trade-off. To get free from this trade-off, this paper introduces an automatic video compression method based on a bilevel optimization program (BOP) consisting of both model-level (i.e., base-level) and frame-level (i.e., meta-level) optimizations. The model-level optimization process learns a cross-modal video retrieval model whose input includes the “compressed frames” learned by frame-level optimization. In turn, frame-level optimization is achieved through gradient descent using the meta loss of the video retrieval model computed on the whole video. We call this BOP method (as well as the “compressed frames”) the Meta-Optimized Frames (MOF) approach. By incorporating MOF, the video retrieval model is able to utilize the information of whole videos (for training) while taking only a small number of input frames in its actual implementation. The convergence of MOF is guaranteed by meta gradient descent algorithms. For evaluation purposes, we conduct extensive cross-modal video retrieval experiments on three large-scale benchmarks: MSR-VTT, MSVD, and DiDeMo. Our results show that MOF is a generic and efficient method that boost multiple baseline methods, and can achieve a new state-of-the-art performance. Ning Han 0005, Xun Yang 0001, Ee-Peng Lim, Hao Chen 0051, Qianru Sun |
IEEE Trans. Multim. | 4 |
| 2024 | BiC-Net: Learning Efficient Spatio-temporal Relation for Text-Video RetrievalabstractThe task of text-video retrieval aims to understand the correspondence between language and vision and has gained increasing attention in recent years. Recent works have demonstrated the superiority of local spatio-temporal relation learning with graph-based models. However, most existing graph-based models are handcrafted and depend heavily on expert knowledge and empirical feedback, which may be unable to mine the high-level fine-grained visual relations effectively. These limitations result in their inability to distinguish videos with the same visual components but different relations. To solve this problem, we propose a novel cross-modal retrieval framework, Bi-Branch Complementary Network (BiC-Net), which modifies Transformer architecture to effectively bridge text-video modalities in a complementary manner via combining local spatio-temporal relation and global temporal information. Specifically, local video representations are encoded using multiple Transformer blocks and additional residual blocks to learn fine-grained spatio-temporal relations and long-term temporal dependency, calling the module a Fine-grained Spatio-temporal Transformer (FST). Global video representations are encoded using a multi-layer Transformer block to learn global temporal features. Finally, we align the spatio-temporal relation and global temporal features with the text feature on two embedding spaces for cross-modal text-video retrieval. Extensive experiments are conducted on MSR-VTT, MSVD, and YouCook2 datasets. The results demonstrate the effectiveness of our proposed model. Our code is public at https://github.com/lionel-hing/BiC-Net . Ning Han 0005, Yawen Zeng, Chuhao Shi, Guangyi Xiao 0001, Hao Chen 0051, Jingjing Chen 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Fine-Grained Alignment for Boundary Samples under Open Set Domain AdaptationabstractOpen set domain adaptation aims to transfer knowledge in the presence of unknown samples in the target domain. Previous approaches use additional classifiers or threshold-based methods to identify unknown samples and try to investigate the information of class diversity within the unknown samples. Despite achieving excellent adaptation results, these methods ignore those samples that lie on the cluster boundaries, especially the clustering-based methods. In this paper, we propose a novel Neighbor Prototype Contrastive Clustering (NPC2) method, which uses the Local Semantic Structure (LSS) to help these low-confidence samples located on the boundary of clusters to return to their own clusters. Further, we propose Local Semantic Consistency (LSC) to evaluate the clustering result and apply it to the domain adaptation process as a metric to assess the reliability of the samples. Results on four benchmarks show that our NPC2significantly outperforms most state-of-the-art methods with higher LSC. Jiang-Lin Wei, Guangyi Xiao 0001, Shun Peng, Hao Chen 0051, Jingzhi Guo, Zhiguo Gong |
ICME | 4 |
| 2023 | Interval-enhanced Graph Transformer solution for session-based recommendation
Huanwen Wang, Yawen Zeng, Jianguo Chen 0001, Ning Han 0005, Hao Chen 0051 |
Expert Syst. Appl. | 5 |
| 2023 | CMFT: Contrastive Memory Feature Transfer for Nonshared-and-Imbalanced Unsupervised Domain AdaptionabstractRecently, nonshared-and-imbalanced unsupervised domain adaption has been proposed to fix domain shift from Big Data source domain with long-tail distribution to specific small target domain with imbalanced distribution, including two challenges: 1) nonshared classes sharing in big data with long-tail distribution; and 2) imbalanced domain adaptation. Prior approaches explore knowledge sharing between classes to improve performance of unsupervised domain adaption methods. However these methods have inductive bias for prior tree or graph. And previous contrastive domain adaptation methods take center-based prototypes as positive samples which only coarsely characterize the domain structure, and fail to depict the local data structure. To fix these problems, we propose a novel framework called contrastive memory feature transfer (CMFT). To solve nonshared data sharing without inductive bias, we build a centroid memory baseddirected memory transfermechanism to enhance imbalanced class features with similar nonshared class centroid. To address the imbalanced domain adaptation, we design a fault-tolerant and fine-grainedneighborhood prototypefor the contrastive learning which can narrow the domain shift. The proposed CMFT outperforms previous methods on most benchmarks. Guangyi Xiao 0001, Shun Peng, Weiwei Xiang, Hao Chen 0051, Jingzhi Guo, Zhiguo Gong |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | Traffic Flow Video Image Recognition and Analysis Based on Multi-Target Tracking Algorithm and Deep LearningabstractTraffic flow parameters are an important data support for the research and development of several technologies in the intelligent transportation system. Therefore, accurate and real-time estimation of traffic flow is particularly important for urban traffic. In this study, a real-time traffic flow detection system framework was constructed based on video image collection and analysis. According to the vehicle detection and tracking results, a traffic flow parameter estimation model and an improved LSTM network are proposed for spatiotemporal counting feature recognition. The results conclude that the developed framework can estimate the traffic flow density and count vehicles, as well as estimate the traffic flow velocity and traffic volume to estimate and optimize traffic flow, respectively. Additionally, the simulation results show that the proposed method can not only counts the two-way traffic vehicles quickly and accurately, but also avoids the use of the complex multi-target tracking method to spatiotemporal correlation of a single target, increases the speed and accuracy of the spatiotemporal information processing procedure, and has stronger scene adaptability. Songshang Zou, Hao Chen 0051, Hui Feng 0005, Guangyi Xiao 0001, Zheng Qin 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | A Spatiotemporal Graph Neural Network for session-based recommendation
Huanwen Wang, Yawen Zeng, Jianguo Chen 0001, Zhouting Zhao, Hao Chen 0051 |
Expert Syst. Appl. | 5 |
| 2022 | Adversarial Multi-Grained Embedding Network for Cross-Modal Text-Video RetrievalabstractCross-modal retrieval between texts and videos has received consistent research interest in the multimedia community. Existing studies follow a trend of learning a joint embedding space to measure the distance between text and video representations. In common practice, video representation is constructed by feeding clips into 3D convolutional neural networks for a coarse-grained global visual feature extraction. In addition, several studies have attempted to align the local objects of video with the text. However, these representations share a drawback of neglecting rich fine-grained relation features capturing spatial-temporal object interactions that benefits mapping textual entities in the real-world retrieval system. To tackle this problem, we propose an adversarial multi-grained embedding network (AME-Net), a novel cross-modal retrieval framework that adopts both fine-grained local relation and coarse-grained global features in bridging text-video modalities. Additionally, with the newly proposed visual representation, we also integrate an adversarial learning strategy into AME-Net, to further narrow the domain gap between text and video representations. In summary, we contribute AME-Net with an adversarial learning strategy for learning a better joint embedding space, and experimental results on MSR-VTT and YouCook2 datasets demonstrate that our proposed framework consistently outperforms the state-of-the-art method. Ning Han 0005, Jingjing Chen 0001, Hao Zhang 0047, Huanwen Wang, Hao Chen 0051 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2021 | Fine-grained Cross-modal Alignment Network for Text-Video RetrievalabstractDespite the recent progress of cross-modal text-to-video retrieval techniques, their performance is still unsatisfactory. Most existing works follow a trend of learning a joint embedding space to measure the distance between global-level or local-level textual and video representation. The fine-grained interactions between video segments and phrases are usually neglected in cross-modal learning, which results in suboptimal retrieval performances. To tackle the problem, we propose a novel Fine-grained Cross-modal Alignment Network (FCA-Net), which considers the interactions between visual semantic units (i.e., sub-actions/sub-events) in videos and phrases in sentences for cross-modal alignment. Specifically, the interactions between visual semantic units and phrases are formulated as a link prediction problem optimized by a graph auto-encoder to obtain the explicit relations between them and enhance the aligned feature representation for fine-grained cross-modal alignment. Experimental results on MSR-VTT, YouCook2, and VATEX datasets demonstrate the superiority of our model as compared to the state-of-the-art method. Ning Han 0005, Jingjing Chen 0001, Guangyi Xiao 0001, Hao Zhang 0047, Yawen Zeng, Hao Chen 0051 |
ACM Multimedia | 6 |
| 2021 | Social-Enhanced Attentive Group RecommendationabstractWith the proliferation of social networks, group activities have become an essential ingredient of our daily life. A growing number of users share their group activities online and invite their friends to join in. This imposes the need of an in-depth study on the group recommendation task, i.e., recommending items to a group of users. Despite its value and significance, group recommendation remains an unsolved problem due to 1) the weights of group members are crucial to the recommendation performance but are rarely learnt from data; 2) social followee information is beneficial to understand users' preferences but is rarely considered; and 3) user-item interactions are helpful to reinforce the performance of group recommendation but are seldom investigated. Toward this end, we devise neural network-based solutions by utilizing the recent developments of attention network and neural collaborative filtering (NCF). First of all, we adopt an attention network to form the representation of a group by aggregating the group members' embeddings, which allows the attention weights of group members to be dynamically learnt from data. Second, the social followee information is incorporated via another attention network to enhance the representation of individual user, which is helpful to capture users' personal preferences. Third, considering that many online group systems also have abundant interactions of individual users on items, we further integrate the modeling of user-item interactions into our method. Through this way, the recommendation for groups and users can be mutually reinforced. Extensive experiments on the scope of both macro-level performance comparison and micro-level analyses justify the effectiveness and rationality of our proposed approaches. Da Cao, Xiangnan He 0001, Lianhai Miao, Guangyi Xiao 0001, Hao Chen 0051, Jia Xu 0005 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2020 | Video-based recipe retrieval
Da Cao, Ning Han 0005, Hao Chen 0051, Xiaochi Wei, Xiangnan He 0001 |
Inf. Sci. | 3 |
| 2020 | A Deep Transfer Learning Solution for Food Material Recognition Using Electronic ScalesabstractIn this article, we present a novel solution to automating the procurement of food materials by using electronic scales, which can automatically identify the food materials along weighing them. Although the CNN model is regarded as one of the most effective solutions to image recognition, the traditional techniques cannot handle the mismatch problem between the lab training data and the real world data. To solve the problem, we propose to embed a partial-and-imbalanced domain adaptation technique (tree adaptation network) in the deep learning model, which can borrow knowledge from sibling classes, to overcome the imbalance problem, and transfer knowledge from the source domain to the target domain, to fight the mismatch problem between the lab training data and the real world data. Experiments show that the proposed approach outperforms state-of-the-art algorithms. Furthermore, the proposed techniques have already been used in practice. Guangyi Xiao 0001, Hao Chen 0051, Da Cao, Jingzhi Guo, Zhiguo Gong |
IEEE Trans. Ind. Informatics | 3 |
| 2018 | Fast auto-clean CNN model for online prediction of food materials
Hao Chen 0051, Jianglong Xu, Guangyi Xiao 0001, Shiqin Zhang |
J. Parallel Distributed Comput. | 1 |
| 2016 | Multi-User Location Correlation Protection with Differential PrivacyabstractIn the big data era, with the rapid development of location-based applications, GPS enabled devices and big data institutions, location correlation privacy raises more and more people's concern. Because adversaries may combine location correlations with their background knowledge to guess users' privacy, such correlation should be protected to preserve users' privacy. In order to deal with the location disclosure problem, location perturbation and generalization have been proposed. However, most proposed approaches depend on syntactic privacy models without rigorous privacy guarantee. Furthermore, many approaches only consider perturbing the locations of one user without considering multi-user location correlations, so these techniques cannot prevent various inference attacks well. Currently, differential privacy has been regarded as a standard for privacy protection, but there are new challenges for applying differential privacy in the location correlations protection. The privacy protection not only should meet the needs of users who request location-based services, but also should protect location correlation among multiple users. In this paper, we propose a systematic solution to protect location correlations privacy among multiple users with rigorous privacy guarantee. First of all, we propose a novel definition, private candidate sets which are obtained by hidden Markov models. Then, we quantify the location correlation between two users by using the similarity of hidden Markov models. Finally, we present a private trajectory releasing mechanism which can preserve the location correlations among users who move under hidden Markov models in a period of time. Experiments on real-world datasets also show that multi-user location correlation protection is efficient. Lu Ou, Zheng Qin 0001, Yonghe Liu, Hui Yin 0001, Yupeng Hu 0004, Hao Chen 0051 |
ICPADS | 6 |
| 2016 | An optimized data integration model based on reverse cleaning for heterogeneous multi-media data
Hao Chen 0051, Yueqi Ouyang |
Multim. Tools Appl. | 1 |
| 2016 | An improved collaborative recommendation algorithm based on optimized user similarity
Hao Chen 0051, Zhongkun Li |
J. Supercomput. | 1 |
| 2005 | Bi-Objective Model for Test-Suite Reduction Based on Modified Condition/Decision CoverageabstractIt is evidence that modified condition/decision coverage (MC/DC) is an effective verification method and can help to detect safety faults despite of its expensive cost. In regression testing, it is quite costly to rerun all of test cases in test suite because new test cases are added to test suite as the software evolves. Therefore, it is necessary to reduce the test suite to improve test efficiency and save test cost. Many existing test-suite reduction techniques are not effective to reduce MC/DC test suite. This paper proposes a new test-suite reduction technique for MC/DC: a bi-objective model that considers both the coverage degree of test case for test requirements and the capability of test cases to reveal error. Our experiment results show that the technique both reduces the size of test suite and better ensures the effectiveness of test suite to reveal error. Lili Pan 0002, Beiji Zou 0001, Hao Chen 0051 |
PRDC | 4 |