VLDB 2026 Research / reviewers in the wild / expert
Kai Zhang 0012
dblp:202/4953
· DBLP profile ↗
27ranked-venue papers
0as first author
16since 2021 · last 2026
0000-0002-5826-3765ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Computer networks · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Software engineering, systems software and programming languages · 3Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PVDSF: A Photovoltaic Generation Forecasting Network With Dynamic-Static Correlation Fusion on Endogenous and Exogenous VariablesabstractPredicting photovoltaic (PV) generation is essential for ensuring grid reliability, optimizing energy allocation, and promoting a green transition in the global energy structure. Previous PV power prediction models only focus on the temporal relationships within exogenous sequences (such as temperature, cloud cover, and humidity) or endogenous sequences (i.e., PV generation), neglecting the impact of exogenous factors on the vulnerable PV generation process. However, the complexity of the physical environment makes it challenging to accurately model the interaction between the two types of variables. To address this issue, we propose a PV generation forecasting network with dynamic-static correlation fusion between endogenous and exogenous variables to improve prediction accuracy. Specifically, we, respectively, encode the exogenous and endogenous variables into static and dynamic components, and accordingly design static and dynamic networks to learn the complex environment, fully capturing the impact of different exogenous components on PV generation. Subsequently, we integrate the dynamic-static enhanced representations through a dual network, achieving lightweight representation prediction. Our model significantly improves PV generation prediction accuracy, achieving state-of-the-art (SOTA) results across four real-world datasets, with average reductions of 12.87% in mean squared error (MSE) and 13.89% in mean absolute error (MAE) compared with competitive baselines. This advancement significantly enhances PV generation forecasting and fosters global sustainability. Tianze Deng, Zengni Zhang, Kai Zhang 0012, Yuhan Dong |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | Relation-Aware Graph Attention Network for Nuclei ClassificationabstractNuclei classification plays a pivotal role in pathological research. Recent advances in graph neural networks (GNNs) have shown great promise in modeling cell-cell interactions. However, many existing methods overlook tissue context, which is crucial for accurate nuclei identification, as nuclei exhibit distinct patterns within specific tissue structures. To address this limitation, we propose a novel Relation-Aware Graph AT-tention network (RAGAT) that effectively leverages nucleus-related features for precise classification. RAGAT constructs a cell graph based on spatial proximity and visual feature similarity, while also introducing a tissue-aware graph by sampling regions around each nucleus to capture the tissue microenvironment and depict local cellular contexts. Furthermore, RAGAT employs a hybrid graph attention module to integrate cell-cell and tissue-cell interactions, enabling a comprehensive understanding of the nuclear context. Experimental results on three benchmark datasets demonstrate that our method significantly outperforms state-of-the-art approaches, offering valuable insight into the analysis of nuclear microenvironments. Our code is available at https://github.com/lingboboo/RAGAT. Lingbo Zhang, Ye Zhang 0043, Linghan Cai, Xianchao Guan, Kai Zhang 0012, Yongbing Zhang 0002 |
ICME | 5 |
| 2025 | FLAP: Fully-controllable Audio-driven Portrait Video Generation through 3D head conditioned diffusion modelabstractDiffusion-based video generation techniques have significantly improved zero-shot talking-head avatar generation, enhancing the naturalness of both head motion and facial expressions. However, existing methods suffer from poor controllability, making them less applicable to real-world scenarios such as filmmaking and live streaming for e-commerce. To address this limitation, we propose FLAP, a novel approach that integrates explicit 3D intermediate parameters (head poses and facial expressions) into the diffusion model for end-to-end generation of realistic portrait videos. The proposed architecture allows the model to generate vivid portrait videos from audio while simultaneously incorporating additional control signals, such as head rotation angles and eye-blinking frequency. Furthermore, the decoupling of head pose and facial expression allows for independent control of each, offering precise manipulation of both the avatar's pose and facial expressions. We also demonstrate its flexibility in integrating with existing 3D head generation methods, bridging the gap between 3D model-based approaches and end-to-end diffusion techniques. Extensive experiments show that our method outperforms recent audio-driven portrait video models in both naturalness and controllability. Lingzhou Mu, Baiji Liu, Guiming Mo, Jiawei Jin, Kai Zhang 0012, Hao-Zhi Huang 0001 |
ACM Multimedia | 6 |
| 2025 | Counting by Points: Density-Guided Weakly-Supervised Nuclei Segmentation in Histopathological Images
Lingbo Zhang, Bingqian Sun, Linghan Cai, Yifeng Wang 0001, Ye Zhang 0043, Songhan Jiang, Kai Zhang 0012, Yongbing Zhang 0002 |
ACM Multimedia | 7 |
| 2025 | Hierarchical task network-enhanced multi-agent reinforcement learning: Toward efficient cooperative strategies
Xuechen Mu, Hankui Zhuo, Chen Chen 0077, Kai Zhang 0012, Chao Yu 0004, Jianye Hao |
Neural Networks | 4 |
| 2024 | Fraud Detection for Financial Transactions: Leveraging Big Data Analytics and Machine LearningabstractRecently, transaction fraud costs billions of dollars to card issuers. With the increase in fraud rates, it is important to establish a comprehensive monitoring mechanism for detecting abnormal accounts that conform to the characteristics of online fraudulent activities. This paper primarily utilizes statistical methods and data mining techniques to analyze new signals that can represent the latest fraudulent transaction methods based on account static information, account transaction data and fraud blacklists collected by the financial big data system. To safeguard the financial security of customers, interpretable machine learning algorithms are also proposed to construct fraud detection model. Eventually, extensive experiments demonstrate that the proposed framework achieves state-of-the-art results with the highest F1 score, Recall and Accuracy. Yixuan Chen 0007, Kai Zhang 0012 |
SMC | 3 |
| 2023 | ARA-GAN: Adaptive Residual Attention Generative Adversarial Network for Retinal Vessel SegmentationabstractAutomatic segmentation of retinal vessels is a critical task in fundoscopic image analysis. The emergence of deep learning has shown promising abilities of feature representation, particularly with Convolutional Neural Networks (CNNs). However, the fixed receptive field in CNNs limits their ability to adapt to the scale variation of natural vascular networks and capture nonlocal context dependencies across feature maps. To address these limitations, we propose a novel model called ARA-GAN that can adaptively extract nonlocal feature contexts and aggregate multi-scale information for retinal vessel segmentation. The proposed model comprises a novel Generative Adversarial Network (GAN) as the overall framework to obtain global information and strong robustness. Additionally, we integrate Residual Nonlocal Attention (RNA) Module into the framework to adaptively capture nonlocal context dependencies across the input features. Finally, we add a Pyramid Pooling Module (PPM) to extract the morphological characteristics of natural retinal vessels at multiple scales. Our experimental results demonstrate that our method outperforms state-of-the-art approaches on both the DRIVE and STARE datasets. Yixuan Chen 0007, Yuhan Dong, Kai Zhang 0012 |
SMC | 4 |
| 2023 | Investment Value Evaluation of Listed Companies Based on Machine LearningabstractIn this paper, we investigate a heterogeneous set of listed companies and extract significant features that can accurately evaluate the investment value of them. Specifically, we analyze both financial data and non-financial data, including: corporate annual report, commercial information, industrial information, land acquisition information, financial information, tax report, intellectual property report, etc. In order to effectively handle a large number of categorical features and mitigate over-fitting problem, CatBoost, LightGBM and ensemble learning framework are adopted to output precise value for each enterprise. Furthermore, extensive experiments demonstrate that the proposed framework achieves state-of-the-art result with RMSE as low as 2.97. Finally, we find that in addition to financial features, many non-financial factors such as Number of Patents (NOP) and Number of Qualification Certifications (NOQC) also play important roles in company investment value evaluation. Yixuan Chen 0007, Kai Zhang 0012 |
SMC | 3 |
| 2023 | Breaking Filter Bubble: A Reinforcement Learning Framework of Controllable Recommender SystemabstractIn the information-overloaded era of the Web, recommender systems that provide personalized content filtering are now the mainstream portal for users to access Web information. Recommender systems deploy machine learning models to learn users’ preferences from collected historical data, leading to more centralized recommendation results due to the feedback loop. As a result, it will harm the ranking of content outside the narrowed scope and limit the options seen by users. In this work, we first conduct data analysis from a graph view to observe that the users’ feedback is restricted to limited items, verifying the phenomenon of centralized recommendation. We further develop a general simulation framework to derive the procedure of the recommender system, including data collection, model learning, and item exposure, which forms a loop. To address the filter bubble issue under the feedback loop, we then propose a general and easy-to-use reinforcement learning-based method, which can adaptively select few but effective connections between nodes from different communities as the exposure list. We conduct extensive experiments in the simulation framework based on large-scale real-world datasets. The results demonstrate that our proposed reinforcement learning-based control method can serve as an effective solution to alleviate the filter bubble and the separated communities induced by it. We believe the proposed framework of controllable recommendation in this work can inspire not only the researchers of recommender systems, but also a broader community concerned with artificial intelligence algorithms’ impact on humanity, especially for those vulnerable populations on the Web. Yancheng Dong, Chen Gao 0001, Dong Li 0016, Jianye Hao, Kai Zhang 0012, Yong Li 0008, Zhi Wang 0001 |
WWW | 7 |
| 2022 | Adaptive Range Guided Multi-view Depth Estimation with Normal Ranking Loss
Yikang Ding, Dihe Huang, Kai Zhang 0012, Zhiheng Li 0001, Wensen Feng |
ACCV (1) | 4 |
| 2022 | ClusterGNN: Cluster-based Coarse-to-Fine Graph Neural Network for Efficient Feature MatchingabstractGraph Neural Networks (GNNs) with attention have been successfully applied for learning visual feature matching. However, current methods learn with complete graphs, resulting in a quadratic complexity in the number of features. Motivated by a prior observation that self- and cross- attention matrices converge to a sparse representation, we propose ClusterGNN, an attentional GNN architecture which operates on clusters for learning the feature matching task. Using a progressive clustering module we adaptively divide keypoints into different subgraphs to reduce redundant connectivity, and employ a coarse-to-fine paradigm for mitigating miss-classification within images. Our approach yields a 59.7% reduction in runtime and 58.4% reduction in memory consumption for dense detection, compared to current state-of-the-art GNN-based matching, while achieving a competitive performance on various computer vision tasks. Junxiong Cai, Yoli Shavit, Tai-Jiang Mu, Wensen Feng, Kai Zhang 0012 |
CVPR | 6 |
| 2022 | Enhancing Multi-View Stereo with Contrastive Matching and Weighted Focal LossabstractLearning-based multi-view stereo (MVS) methods have made impressive progress and surpassed traditional methods in recent years. However, their accuracy and completeness are still struggling. In this paper, we propose a new method to enhance the performance of existing networks inspired by contrastive learning and feature matching. First, we propose a Contrast Matching Loss (CML), which treats the correct matching points in depth-dimension as positive sample and other points as negative samples, and computes the contrastive loss based on the similarity of features. We further propose a Weighted Focal Loss (WFL) for better classification capability, which weakens the contribution of low-confidence pixels in unimportant areas to the loss according to predicted confidence. Extensive experiments performed on DTU, Tanks and Temples and BlendedMVS datasets show our method achieves state-of-the-art performance and significant improvement over baseline network. Yikang Ding, Dihe Huang, Zhiheng Li 0001, Kai Zhang 0012 |
ICIP | 5 |
| 2022 | GCD-PKAug: A Gradient Consistency Discriminator-Based Augmentation Method for Pharmacokinetics Time Courses
Pingping Song, Yuhan Dong, Kai Zhang 0012 |
ICONIP (5) | 3 |
| 2022 | Adversarial Learning of Hard Positives for Place RecognitionabstractImage retrieval methods for place recognition learn global image descriptors that are used for fetching geo-tagged images at inference time. Recent works have suggested employing weak and self-supervision for mining hard positives and hard negatives in order to improve localization accuracy and robustness to visibility changes (e.g. in illumination or view point). However, generating hard positives, which is essential for obtaining robustness, is still limited to hard-coded or global augmentations. In this work we propose an adversarial method to guide the creation of hard positives for training image retrieval networks. Our method learns local and global augmentation policies which will increase the training loss, while the image retrieval network is forced to learn more powerful features for discriminating increasingly difficult examples. This approach allows the image retrieval network to generalize beyond the hard examples presented in the data and learn features that are robust to a wide range of variations. Our method achieves state-of-the-art recalls on the Pitts250 and Tokyo 24/7 benchmarks and outperforms recent image retrieval methods on the rOxford and rParis datasets by a noticeable margin. Kai Zhang 0012, Yoli Shavit, Wensen Feng |
IJCNN | 2 |
| 2022 | WT-MVSNet: Window-based Transformers for Multi-view StereoabstractRecently, Transformers have been shown to enhance the performance of multi-view stereo by enabling long-range feature interaction. In this work, we propose Window-based Transformers (WT) for local feature matching and global feature aggregation in multi-view stereo. We introduce a Window-based Epipolar Transformer (WET) which reduces matching redundancy by using epipolar constraints. Since point-to-line matching is sensitive to erroneous camera pose and calibration, we match windows near the epipolar lines. A second Shifted WT is employed for aggregating global information within cost volume. We present a novel Cost Transformer (CT) to replace 3D convolutions for cost volume regularization. In order to better constrain the estimated depth maps from multiple views, we further design a novel geometric consistency loss (Geo Loss) which punishes unreliable areas where multi-view consistency is not satisfied. Our WT multi-view stereo method (WT-MVSNet) achieves state-of-the-art performance across multiple datasets and ranks $1^{st}$ on Tanks and Temples benchmark. Code will be available upon acceptance. Jinli Liao, Yikang Ding, Yoli Shavit, Dihe Huang, Shihao Ren, Wensen Feng, Kai Zhang 0012 |
NeurIPS | 8 |
| 2021 | Nonlinear HPA Impact on Artificial Noise Aided MISO Secure SystemsabstractIn this paper, we consider the artificial noise (AN) aided multiple-input single-output (MISO) secure systems in which a multi-antenna transmitter simultaneously transmits the combination of information-bearing signal and artificial noise to a single-antenna legitimate receiver and eavesdropper. At the transmitter front-ends, the signal and/or AN will be amplified by high-power amplifiers (HPAs) which however may work in a nonlinear region and introduce nonlinear distortion. We calculate the received signal-to-noise ratio (SNR) at legitimate receiver and eavesdropper under HPA nonlinearity, and derive the approximated closed-form expression of secrecy rate. Numerical results suggest that the approximated theoretical secrecy performance is very close to Monte Carlo simulations for relatively lower SNRs, and is severely degraded by the nonlinear distortion caused by HPAs as the transmit power increases. Moreover, we further provide a power allocation strategy between information-bearing signal and AN to alleviate the nonlinear distortion. Fan Yang 0086, Kai Zhang 0012, Yongzhi Zhai, Yuhan Dong |
ICC | 2 |
| 2020 | A Uniform Spatial Channel Model for Underwater Wireless Optical Communication LinksabstractIn underwater wireless optical communications(UWOC), absorption and scattering cause attenuation and dispersion in both time and space domains. There have been many studies on temporal channel investigation and modeling such as impulse response but fewer on spatial behavior. In this paper, we consider the characteristics of actual light sources, i.e., laser diodes (LDs) and light-emitting diodes (LEDs), and propose a uniform spatial channel (USC) model for UWOC system employing single or multiple light sources. We first simulate the irradiance distributions of single-source UWOC systems based on Monte Carlo method and fit them with a closed-form expression for both types of light sources. Then we generalize this single-source model to multiple-source link geometry considering uniform linear array and circular array of light sources. Numerical results suggest that the proposed model fits well with the irradiance distributions of UWOC links regardless of the type, number and array geometry of light sources for various water types. Kai Zhang 0012, Yuhan Dong |
GLOBECOM | 2 |
| 2020 | PCANet: Pyramid Context-aware Network for Retinal Vessel SegmentationabstractAutomated retinal vessel segmentation plays an important role in the diagnosis of some diseases such as diabetes, arteriosclerosis and hypertension. Recent works attempt to improve segmentation performance by exploring either global or local contexts. However, the context demands are varying from regions in each image and different levels of network. To address these problems, we propose Pyramid Context-aware Network (PCANet), which can adaptively capture multi-scale context representations. Specifically, PCANet is composed of multiple Adaptive Context-aware (ACA) blocks arranged in parallel, each of which can adaptively obtain the context-aware features by estimating affinity coefficients at a specific scale under the guidance of global contextual dependencies. Meanwhile, we import ACA blocks with specific scales in different levels of the network to obtain a coarse-to-fine result. Furthermore, an integrated test-time augmentation method is developed to further boost the performance of PCANet. Finally, extensive experiments demonstrate the effectiveness of the proposed PCANet, and state-of-the-art performances are achieved with AUCs of 0.9866, 0.9886 and F1 Scores of 0.8274, 0.8371 on two public datasets, DRIVE and STARE, respectively. Yixuan Chen 0007, Kai Zhang 0012 |
ICPR | 3 |
| 2020 | RNA-Net: Residual Nonlocal Attention Network for Retinal Vessel SegmentationabstractAutomatic segmentation of retinal vessels is an important step in fundoscopic image analysis. Recently, convolutional-neural-network-based methods have been widely explored in this vision task. However, the local fixed receptive field makes network unable to collect global information and adapt to scale variation of retinal vessels. In this paper, we propose a novel RNA-Net which can capture nonlocal context dependencies across the inputs and extract multi-scale features for segmentation task. Firstly, we build a Residual Nonlocal Attention (RNA) Module, which can guide the network to pay more attention to task-related regions of the whole feature map. Secondly, to better capture the morphological characteristics of natural blood vessels, Pyramid Pooling Module (PPM) is added to capture features at multiple scales. Experimental results on two public datasets DRIVE and STARE clearly demonstrate that our method outperforms the current state-of-the-art approaches. Yixuan Chen 0007, Yuhan Dong, Kai Zhang 0012 |
SMC | 4 |
| 2017 | Improved joint antenna selection and user scheduling for massive MIMO systemsabstractMassive multi-input multi-output (MIMO) technology is promising by employing a large number of antennas at the base station to support a large amount of users. However, due to the limitation of analog front-ends at the base station, the antenna selection and user scheduling strategies are essential to achieve spatial diversity and reduce hardware cost at the same time. In this work, we consider the strategy of joint antenna selection and user scheduling (JASUS) for uplink massive MIMO systems and propose a greedy two-step JASUS algorithm referred to as largest minimum singular value based JASUS (LMSVJASUS). In its first step, a simplified downward branch and bound based JASUS is used to find a near-optimal antenna and user sets whose channel matrix has the near-largest MSV. In its second step, a swapping-based algorithm is proposed to find a better solution by swapping antennas and users between the selected and the discarded. The numerical results suggest that the proposed algorithm outperforms traditional approaches in terms of system sum-rate and computational complexity. Yuhan Dong, Kai Zhang 0012 |
ICIS | 3 |
| 2017 | A spatial-temporal model to improve PM2.5 inferenceabstractPM2.5 is one of the major indicators of ambient air quality which has become a focus of public attention. Urban PM2.5 can be measured by air quality monitoring stations which are costly and not sufficiently installed in a city. In this paper, we aim to infer the PM2.5 information at the place where there is no air quality monitoring station. As PM2.5 concentration varies over time and space domains, we propose a joint topic model to jointly model the spatial and temporal patterns of PM2.5. Numerical results suggest that the proposed model achieves better inference based on five related datasets compared with traditional methods. Yuhan Dong, Kai Zhang 0012 |
ICIS | 3 |
| 2017 | Analysis and evaluation of driving behavior recognition based on a 3-axis accelerometer using a random forest approach: poster abstractabstractUnderstanding human drivers' behavior is critical for the self-driving cars, and has been intensively studied in the past decade. We exploit the widely available camera and motion sensor data from car recorders, and propose a hybrid method of recognizing driving events based on the random forest approach. The classification results are analyzed by comparing different features, classifiers and filters. A high accuracy of 98.1% on driving behavior classification is obtained and the robustness is verified on a dataset including 2400 driving events. Wangjing Cao, Kai Zhang 0012, Yuhan Dong, Shao-Lun Huang, Lin Zhang 0001 |
IPSN | 3 |
| 2017 | Range-based localization in underwater wireless sensor networks using deep neural network: poster abstractabstractIn underwater wireless sensor networks (USWNs), localizing unknown nodes is essential for most applications while is more complex than that of terrestrial WSNs. In this paper, we propose a range-based localization scheme using deep neural network (DNN). Numerical results suggest that the proposed DNN localization algorithm outperforms traditional schemes using least squares support vector machines (LS-SVM) or generalized least squares (GLS) in terms of localization accuracy and efficiency. Moreover, the proposed algorithm requires a small number of anchor nodes, which is plausible for practical applications. Yuhan Dong, Kai Zhang 0012 |
IPSN | 4 |
| 2017 | Improved reverse localization schemes for underwater wireless sensor networks: poster abstractabstractLocalization of sensor nodes is important for wireless sensor networks (WSNs) especially for underwater WSNs (UWSNs). Among the existing UWSN localization approaches, reverse localization scheme (RLS) is an event-driven method suitable for underwater surveillance. RLS adopts the strongest arrival or the first arrival as the direct path, which is however not accurate enough due to the severe multipath effect in underwater acoustic channels. We propose median RLS (MRLS) to select the median path and weighted RLS (WRLS) by assigning each path with a possibility to be the direct path for UWSNs. Numerical results have validated the proposed schemes and further suggested that WRLS outperforms MRLS and traditional approaches and is capable to diminish the multipath effect. Yuhan Dong, Kai Zhang 0012 |
IPSN | 5 |
| 2017 | A novel passenger hotspots searching algorithm for taxis in urban areaabstractPassenger hotspots searching is essential to increase profits for taxis drivers in urban area. In this paper, we propose a two-step approach for pick-up hotspots searching. In the first step, a traveling similarity model is built to quantify the similarity of traveling behaviors. In the second step, we utilize affinity propagation and simulated annealing to identify the daily passenger hotspots in a selected period. Numerical results based on GPS data of Manhattan taxis suggest that the proposed approach outperforms the traditional spatio-temporal clustering regardless of buffer radius. Yuhan Dong, Siyuan Qian, Kai Zhang 0012, Yongzhi Zhai |
SNPD | 3 |
| 2016 | Optimal placement of charging stations for electric taxis in urban area with profit maximizationabstractThe deployment of charging infrastructures is a key factor for the operation of electric taxis in urban area. This paper focuses on the operational efficiency and the charging convenience of electric taxis, and introduces a two-step optimization process of the charging station location for electric taxis. We propose a modified k-means clustering method to divide the urban area into multiple service regions according to the market demand distribution. Then we utilize the optimal location model in the service region to calculate the optimal sites of the charging stations to maximize the operational efficiency and charging convenience. We collect GPS data of electric taxis in Shenzhen and examine the behavior of our proposed approach. Numerical results suggest that the proposed approach outperforms the traditional particle swarm optimization (PSO) method and has improved the operational efficiency by 14.03% and reduced the average distance for the charging service by 12.30% compared with the actual physical sites. Yuhan Dong, Siyuan Qian, Lin Zhang 0001, Kai Zhang 0012 |
SNPD | 5 |
| 2016 | An improved model for PM2.5 inference based on support vector machineabstractPM2.5 is one of the major ambient air pollutants to threaten our health in urban area. However, there are only a few air monitoring stations in a city, which make it difficult to precisely measure PM2.5 concentration at the place without installation of air monitor. In this paper, we consider the PM2.5 inference problem and propose a model taking the PM2.5 nonlinear characteristic into account based on support vector machine (SVM). We collect the features of meteorology, geographical locations and PM2.5 indexes observed by air monitoring stations. Unlike the previous work, we add points of interest (POIs) feature and reduce its dimension by latent Dirichlet allocation (LDA) model due to its sparsity and adopt wavelet decomposition to improve the inference accuracy. Numerical results show that the proposed approach outperforms other related methods in terms of RMSE, correlation coefficient and mean absolute error (MAE) with the real value. Yuhan Dong, Lin Zhang 0001, Kai Zhang 0012 |
SNPD | 4 |