VLDB 2026 Research / reviewers in the wild / expert
Jiakai Zhang
dblp:179/2299
· DBLP profile ↗
14ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prescribed-time neurodynamic ADMM for Tikhonov regularization: Algorithms, circuits and application
Gehao Zhang, Zhuoqin Yang, Jiakai Zhang |
Neurocomputing | 3 |
| 2025 | CryoFastAR: Fast Cryo-EM AB Initio Reconstruction Made EasyabstractPose estimation from unordered images is fundamental for 3D reconstruction, robotics, and scientific imaging. Recent geometric foundation models, such as DUSt3R, enable end-to-end dense 3D reconstruction but remain underexplored in scientific imaging fields like cryo-electron microscopy (cryo-EM) for near-atomic protein reconstruction. In cryo-EM, pose estimation and 3D reconstruction from unordered particle images still depend on time-consuming iterative optimization, primarily due to challenges such as low signal-to-noise ratios (SNR) and distortions from the contrast transfer function (CTF). We introduce CryoFastAR, the first geometric foundation model that can directly predict poses from Cryo-EM noisy images for Fast ab initio Reconstruction. By integrating multi-view features and training on large-scale simulated cryo-EM data with realistic noise and CTF modulations, CryoFastAR enhances pose estimation accuracy and generalization. To enhance training stability, we propose a progressive training strategy that first allows the model to extract essential features under simpler conditions before gradually increasing difficulty to improve robustness. Experiments show that CryoFastAR achieves comparable quality while significantly accelerating inference over traditional iterative approaches on both synthetic and real datasets. Jiakai Zhang, Shouchen Zhou, Haizhao Dai, Xinhang Liu, Peihao Wang, Zhiwen Fan, Yuan Pei, Jingyi Yu 0001 |
ICCV | 1 |
| 2025 | Temperature-Aware Adaptive Federated Distillation for Energy-Constrained AIoT with Non-IID DataabstractFederated Learning (FL) can help multiple Internet of Things (IoT) devices to collaboratively train a machine learning model to provide intelligent services and applications (FL-AIoT). Due to IoT devices' limited storage capacity and energy, fresh data collected by devices often overwrites outdated data and establishes a heterogeneous data distribution. This causes the global model to forget outdated data's characteristics (i.e., catastrophic forgetting). Existing methods incorporate knowledge distillation into FL (i.e., federated distillation, FD) to extract and integrate characteristics from both fresh and outdated data, but they use fixed distillation temperatures for different devices, which overlooks that fixed distillation temperatures cannot match the heterogeneous data distribution on different devices and degrades global model accuracy. To this end, we propose a Federated Dynamic Decoupled Distillation method based on Logits distribution (Fed3DL). Specifically, Fed3DL utilizes decoupled distillation to mitigate catastrophic forgetting. To alleviate the impact of heterogeneous data distributions, Fed3DL novelly builds an adaptive temperature-aware mechanism to dynamically adjust the distillation temperature of each device based on the distribution of Logits. Additionally, Fed3DL introduces a regularization term into the local distillation loss to reduce inter-class characteristics disparity and improve model accuracy. Experiments on two datasets show that compared with the best of the 5 baselines, Fed3DL can improve the global model accuracy by an average of 3.40 %, reduce the forgetting rate by an average of 4.98 %, and achieve the lowest inter-class accuracy disparity. Yingchi Mao, Jiakai Zhang, Litao Qu, Benteng Zhang, Xiaoming He 0004 |
VTC2025-Spring | 2 |
| 2024 | Dual Adaptive Compression for Efficient Communication in Heterogeneous Federated LearningabstractIn federated learning, multiple rounds of communication are involved between clients and the server to train a global model. The extensive model updates transmitted during the training lead to significant communication costs. Previous methods usually employ quantization or sparsification to compress model updates. However, the lossy compression leads to a decline in accuracy, it is challenging to strike a balance between communication efficiency and model accuracy. Meanwhile, due to the data heterogeneity, local updates among different clients are biased towards each other. Employing the same compression ratios for each local updates will further degrade the model accuracy. To achieve the trade-off between communication efficiency and model accuracy, we propose FedDAC, a Dual Adaptive Compression method in heterogeneous federated learning. In the local computation phase, the loss queue is adopted to detect the convergence trends within each client. FedDAC can then dynamically quantify model updates and allow for various compression ratios among heterogeneous clients. In the global aggregation phase, FedDAC can determine the fluctuations in training based on the similarity between clients and the server, thereby adjusting the sparsity ratio flexibly. To alleviate the reduction in model accuracy caused by lossy compression, we introduce residual updates in the local computation and global aggregation phases to maintain model accuracy. Experiment results show that compared with one-way compression methods NAGC and AdaQuantFL, FedDAC can maintain comparable accuracy while the accumulated communication volume is reduced by about 29.6 times, and 22.8 times, respectively. Moreover, the global model accuracy of FedDAC surpasses the two-way compression method T-FedAvg by about 2.4%, and the accumulated communication volume is about 2.5 times lower than T-FedAvg. Yingchi Mao, Chenxin Li, Jiakai Zhang, Shufang Xu, Jie Wu 0001 |
CCGrid | 4 |
| 2024 | NILM-LANN: A Lightweight Attention-based Neural Network in Non-Intrusive Load MonitoringabstractNon-Intrusive Load Monitoring (NILM) has attracted much attention as a promising method of identifying electrical appliances. It discriminates electrical devices based on changes in load characteristics to enable electrical device scheduling strategies for optimal energy utilization. Existing NILM methods mainly use high-frequency electrical-specific signals and have high computational and memory requirements, which are difficult to implement on resource-constrained devices. Therefore, we propose NILM-LANN, a novel lightweight neural network with an attention mechanism, in which several tricks including Convolutional Long Short-Term Memory (ConvLSTM), 1-D convolutional operations and DenseNet are comprehensively utilized and balanced to reduce the number of model parameters, deepen the network while avoiding gradient explosion or vanishing. The Squeeze and Excitation block (SE-block) is also used to capture channel-wise dependencies based on the aggregated information. The NILM-LANN model is evaluated on public datasets of UK-DALE and LIT, and a self-built CAE dataset, with a maximum classification accuracy of 99.9% and F1-Score of 99.9%, while reducing the number of model parameters by over 90% compared to other existing methods such as Alex-Net and Light-LSTM Finally, the NILM-LANN is deployed in NVIDIA Jetson Nano B01 embedded device to verily its lightweight and efficiency. Yanjing Lei, Zehui Feng, Xiangqing Lin, Jiakai Zhang |
CSCWD | 5 |
| 2024 | DRACO: A Denoising-Reconstruction Autoencoder for Cryo-EMabstractFoundation models in computer vision have demonstrated exceptional performance in zero-shot and few-shot tasks by extracting multi-purpose features from large-scale datasets through self-supervised pre-training methods. However, these models often overlook the severe corruption in cryogenic electron microscopy (cryo-EM) images by high-level noises. We introduce DRACO, a Denoising-Reconstruction Autoencoder for CryO-EM, inspired by the Noise2Noise (N2N) approach. By processing cryo-EM movies into odd and even images and treating them as independent noisy observations, we apply a denoising-reconstruction hybrid training scheme. We mask both images to create denoising and reconstruction tasks. For DRACO's pre-training, the quality of the dataset is essential, we hence build a high-quality, diverse dataset from an uncurated public database, including over 270,000 movies or micrographs. After pre-training, DRACO naturally serves as a generalizable cryo-EM image denoiser and a foundation model for various cryo-EM downstream tasks. DRACO demonstrates the best performance in denoising, micrograph curation, and particle picking tasks compared to state-of-the-art baselines. Yingjun Shen, Haizhao Dai, Qihe Chen, Jiakai Zhang, Yuan Pei, Jingyi Yu 0001 |
NeurIPS | 5 |
| 2024 | CryoGEM: Physics-Informed Generative Cryo-Electron MicroscopyabstractIn the past decade, deep conditional generative models have revolutionized the generation of realistic images, extending their application from entertainment to scientific domains. Single-particle cryo-electron microscopy (cryo-EM) is crucial in resolving near-atomic resolution 3D structures of proteins, such as the SARS-COV-2 spike protein. To achieve high-resolution reconstruction, a comprehensive data processing pipeline has been adopted. However, its performance is still limited as it lacks high-quality annotated datasets for training. To address this, we introduce physics-informed generative cryo-electron microscopy (CryoGEM), which for the first time integrates physics-based cryo-EM simulation with a generative unpaired noise translation to generate physically correct synthetic cryo-EM datasets with realistic noises. Initially, CryoGEM simulates the cryo-EM imaging process based on a virtual specimen. To generate realistic noises, we leverage an unpaired noise translation via contrastive learning with a novel mask-guided sampling scheme. Extensive experiments show that CryoGEM is capable of generating authentic cryo-EM images. The generated dataset can be used as training data for particle picking and pose estimation models, eventually improving the reconstruction resolution. Jiakai Zhang, Qihe Chen, Wenyuan Gao, Xuming He 0001, Jingyi Yu 0001 |
NeurIPS | 1 |
| 2024 | MRS-Net: Brain tumour segmentation network based on feature fusion and attention mechanismabstractAbstract Accurate segmentation of brain tumor magnetic resonance imaging (MRI) is crucial for treatment planning. Addressing the challenges of complex tumor structures and inadequate cross‐channel information utilization in Unet‐based segmentation, this paper proposes the multi‐scale residual brain tumor MRI segmentation network (MRS‐Net) incorporating an attention mechanism to enhance segmentation accuracy. First, the double residual feature fusion module is utilized to enhance the fusion of feature information between different levels. Second, the Atrous Spatial Pyramid Pooling is introduced as a bridging module of the network to capture the features at different scales of the image, so as to enhance the extraction capability of the network for detailed features. Finally, the inverted residual coordinate attention module replaces the direct splicing in Unet to fuse the large feature information at each level and scale, thus enhancing the model's ability to recognize the spatial location information of brain tumors. The Dice coefficients, positive predictive values (PPVs), sensitivities (Sensitivity) and Hausdorff distance (HD), which are the four evaluation indexes, reach 84.54%, 87.43%, 88.37% and 2.248, respectively, which are improved by 1.85%, 2.11%, 2.88% and 6.0%, respectively, compared with Unet. The experimental results show that MRS‐Net achieves better brain tumor image segmentation. Xiaoyan Shen, Jiakai Zhang, Hongming Shen |
IET Image Process. | 6 |
| 2022 | Fourier PlenOctrees for Dynamic Radiance Field Rendering in Real-timeabstractImplicit neural representations such as Neural Radiance Field (NeRF) have focused mainly on modeling static objects captured under multi-view settings where real-time rendering can be achieved with smart data structures, e.g., PlenOctree. In this paper, we present a novel Fourier PlenOctree (FPO) technique to tackle efficient neural mod-eling and real-time rendering of dynamic scenes captured under the free-view video (FVV) setting. The key idea in our FPO is a novel combination of generalized NeRF, PlenOctree representation, volumetric fusion and Fourier transform. To accelerate FPO construction, we present a novel coarse-to-fine fusion scheme that leverages the gen-eralizable NeRF technique to generate the tree via spatial blending. To tackle dynamic scenes, we tailor the implicit network to model the Fourier coefficients of time-varying density and color attributes. Finally, we construct the FPO and train the Fourier coefficients directly on the leaves of a union PlenOctree structure of the dynamic sequence. We show that the resulting FPO enables compact memory overload to handle dynamic objects and supports efficient fine-tuning. Extensive experiments show that the proposed method is 3000 times faster than the original NeRF and achieves over an order of magnitude acceleration over SOTA while preserving high visual quality for the free-viewpoint rendering of unseen dynamic scenes. Jiakai Zhang, Xinhang Liu, Fuqiang Zhao, Yanshun Zhang, Yingliang Zhang, Minye Wu, Jingyi Yu 0001, Lan Xu 0003 |
CVPR | 2 |
| 2022 | HumanNeRF: Efficiently Generated Human Radiance Field from Sparse InputsabstractRecent neural human representations can produce high-quality multi-view rendering but require using dense multi-view inputs and costly training. They are hence largely limited to static models as training each frame is infeasible. We present HumanNeRF - a neural representation with efficient generalization ability - for high-fidelity free-view synthesis of dynamic humans. Analogous to how IBRNet assists NeRF by avoiding perscene training, HumanNeRF employs an aggregated pixel-alignment feature across multi-view inputs along with a pose embedded non-rigid deformation field for tackling dynamic motions. The raw Human-NeRF can already produce reasonable rendering on sparse video inputs of unseen subjects and camera settings. To further improve the rendering quality, we augment our solution with in-hour scene-specific fine-tuning, and an appearance blending module for combining the benefits of both neural volumetric rendering and neural texture blending. Extensive experiments on various multi-view dynamic hu-man datasets demonstrate effectiveness of our approach in synthesizing photo-realistic free-view humans under challenging motions and with very sparse camera view inputs. Fuqiang Zhao, Wei Yang 0034, Jiakai Zhang, Pei Lin, Yingliang Zhang, Jingyi Yu 0001, Lan Xu 0003 |
CVPR | 3 |
| 2022 | Human Performance Modeling and Rendering via Neural Animated MeshabstractWe have recently seen tremendous progress in the neural advances for photo-real human modeling and rendering. However, it's still challenging to integrate them into an existing mesh-based pipeline for downstream applications. In this paper, we present a comprehensive neural approach for high-quality reconstruction, compression, and rendering of human performances from dense multi-view videos. Our core intuition is to bridge the traditional animated mesh workflow with a new class of highly efficient neural techniques. We first introduce a neural surface reconstructor for high-quality surface generation in minutes. It marries the implicit volumetric rendering of the truncated signed distance field (TSDF) with multi-resolution hash encoding. We further propose a hybrid neural tracker to generate animated meshes, which combines explicit non-rigid tracking with implicit dynamic deformation in a self-supervised framework. The former provides the coarse warping back into the canonical space, while the latter implicit one further predicts the displacements using the 4D hash encoding as in our reconstructor. Then, we discuss the rendering schemes using the obtained animated meshes, ranging from dynamic texturing to lumigraph rendering under various bandwidth settings. To strike an intricate balance between quality and bandwidth, we propose a hierarchical solution by first rendering 6 virtual views covering the performer and then conducting occlusion-aware neural texture blending. We demonstrate the efficacy of our approach in a variety of mesh-based applications and photo-realistic free-view experiences on various platforms, i.e., inserting virtual human performances into real environments through mobile AR or immersively watching talent shows with VR headsets. Fuqiang Zhao, Yuheng Jiang, Kaixin Yao, Jiakai Zhang, Haizhao Dai, Yuhui Zhong, Yingliang Zhang, Minye Wu, Lan Xu 0003, Jingyi Yu 0001 |
ACM Trans. Graph. | 4 |
| 2021 | Editable free-viewpoint video using a layered neural representationabstractGenerating free-viewpoint videos is critical for immersive VR/AR experience, but recent neural advances still lack the editing ability to manipulate the visual perception for large dynamic scenes. To fill this gap, in this paper, we propose the first approach for editable free-viewpoint video generation for large-scale view-dependent dynamic scenes using only 16 cameras. The core of our approach is a new layered neural representation, where each dynamic entity, including the environment itself, is formulated into a spatio-temporal coherent neural layered radiance representation called ST-NeRF. Such a layered representation supports manipulations of the dynamic scene while still supporting a wide free viewing experience. In our ST-NeRF, we represent the dynamic entity/layer as a continuous function, which achieves the disentanglement of location, deformation as well as the appearance of the dynamic entity in a continuous and self-supervised manner. We propose a scene parsing 4D label map tracking to disentangle the spatial information explicitly and a continuous deform module to disentangle the temporal motion implicitly. An object-aware volume rendering scheme is further introduced for the re-assembling of all the neural layers. We adopt a novel layered loss and motion-aware ray sampling strategy to enable efficient training for a large dynamic scene with multiple performers, Our framework further enables a variety of editing functions, i.e., manipulating the scale and location, duplicating or retiming individual neural layers to create numerous visual effects while preserving high realism. Extensive experiments demonstrate the effectiveness of our approach to achieve high-quality, photo-realistic, and editable free-viewpoint video generation for dynamic scenes. Jiakai Zhang, Xinhang Liu, Fuqiang Zhao, Yanshun Zhang, Minye Wu, Yingliang Zhang, Lan Xu 0003, Jingyi Yu 0001 |
ACM Trans. Graph. | 1 |
| 2020 | LGNN: A Context-aware Line Segment DetectorabstractWe present a novel real-time line segment detection scheme called Line Graph Neural Network (LGNN). Existing approaches require a computationally expensive verification or postprocessing step. Our LGNN employs a deep convolutional neural network (DCNN) for proposing line segment directly, with a graph neural network (GNN) module for reasoning their connectivities. Specifically, LGNN exploits a new quadruplet representation for each segment where the GNN module takes the predicted candidates as vertexes and constructs a sparse graph to enforce structural context. Compared with the state-of-the-art, LGNN achieves near real-time performance without compromising accuracy. LGNN further enables time-sensitive 3D applications. When a 3D point cloud is accessible, we present a multi-modal line segment classification technique for extracting a 3D wireframe of the environment robustly and efficiently. Quan Meng, Jiakai Zhang, Qiang Hu 0003, Xuming He 0001, Jingyi Yu 0001 |
ACM Multimedia | 2 |
| 2017 | Query-Efficient Imitation Learning for End-to-End Simulated DrivingabstractOne way to approach end-to-end autonomous driving is to learn a policy that maps from a sensory input, such as an image frame from a front-facing camera, to a driving action, by imitating an expert driver, or a reference policy. This can be done by supervised learning, where a policy is tuned to minimize the difference between the predicted and ground-truth actions. A policy trained in this way however is known to suffer from unexpected behaviours due to the mismatch between the states reachable by the reference policy and trained policy. More advanced algorithms for imitation learning, such as DAgger, addresses this issue by iteratively collecting training examples from both reference and trained policies. These algorithms often require a large number of queries to a reference policy, which is undesirable as the reference policy is often expensive. In this paper, we propose an extension of the DAgger, called SafeDAgger, that is query-efficient and more suitable for end-to-end autonomous driving. We evaluate the proposed SafeDAgger in a car racing simulator and show that it indeed requires less queries to a reference policy. We observe a significant speed up in convergence, which we conjecture to be due to the effect of automated curriculum learning. Jiakai Zhang, Kyunghyun Cho |
AAAI | 1 |