VLDB 2026 Research / reviewers in the wild / expert
Zhenyuan Chen
dblp:121/0032
· DBLP profile ↗
8ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RSCC: A Large-Scale Remote Sensing Change Caption Dataset for Disaster EventsabstractRemote sensing is critical for disaster monitoring, yet existing datasets lack temporal image pairs and detailed textual annotations. While single-snapshot imagery dominates current resources, it fails to capture dynamic disaster impacts over time. To address this gap, we introduce the Remote Sensing Change Caption (RSCC) dataset, a large-scale benchmark comprising 62,351 pre-/post-disaster image pairs (spanning earthquakes, floods, wildfires, and more) paired with rich, human-like change captions. By bridging the temporal and semantic divide in remote sensing data, RSCC enables robust training and evaluation of vision-language models for disaster-aware bi-temporal understanding. Our results highlight RSCC’s ability to facilitate detailed disaster-related analysis, paving the way for more accurate, interpretable, and scalable vision-language applications in remote sensing. Code and dataset are available at https://github.com/Bili-Sakura/RSCC. Zhenyuan Chen, Ningyu Zhang 0001 |
NeurIPS | 1 |
| 2025 | Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You ThinkabstractREPA and its variants effectively mitigate training challenges in diffusion models by incorporating external visual representations from pretrained models, through alignment between the noisy hidden projections of denoising networks and foundational clean image representations. We argue that the external alignment, which is absent during the entire denoising inference process, falls short of fully harnessing the potential of discriminative representations. In this work, we propose a straightforward method called $\textit{$\textbf{R}$epresentation $\textbf{E}$ntanglement for $\textbf{G}$eneration}$ ($\textbf{REG}$), which entangles low-level image latents with a single high-level class token from pretrained foundation models for denoising.
REG acquires the capability to produce coherent image-class pairs directly from pure noise, substantially improving both generation quality and training efficiency.
This is accomplished with negligible additional inference overhead, requiring only one single additional token for denoising (<0.5\% increase in FLOPs and latency).
The inference process concurrently reconstructs both image latents and their corresponding global semantics, where the acquired semantic knowledge actively guides and enhances the image generation process.
On ImageNet 256$\times$256, SiT-XL/2 + REG demonstrates remarkable convergence acceleration, achieving $\textbf{63}\times$ and $\textbf{23}\times$ faster training than SiT-XL/2 and SiT-XL/2 + REPA, respectively.
More impressively, SiT-L/2 + REG trained for merely 400K iterations outperforms SiT-XL/2 + REPA trained for 4M iterations ($\textbf{10}\times$ longer). Code is available at: https://github.com/Martinser/REG. Ruijing Shi, Shanghua Gao, Zhenyuan Chen, Lei Wang 0118, Zhaowei Chen, Hongcheng Gao, Jian Yang 0003, Ming-Ming Cheng, Xiang Li 0041 |
NeurIPS | 5 |
| 2025 | S2R-CMI: A Robust Deep Reinforcement Learning method based on counterfactual estimation and state importance evaluation under additive noise disturbance
Zhenyuan Chen |
Neurocomputing | 1 |
| 2023 | Reduced-search guessing random additive noise decoding of polar codes
Kefan Wang, Yuejun Wei, Zhenyuan Chen, Huarui Yin, Wenyi Zhang 0001 |
Sci. China Inf. Sci. | 3 |
| 2023 | ORBGRAND Is Almost Capacity-AchievingabstractDecoding via sequentially guessing the error pattern in a received noisy sequence has received attention recently, and ORBGRAND has been proposed as one such decoding algorithm that is capable of utilizing the soft information embedded in the received noisy sequence. An information theoretic study is conducted for ORBGRAND, and it is shown that the achievable rate of ORBGRAND using independent and identically distributed random codebooks almost coincides with the channel capacity, for an additive white Gaussian noise channel under antipodal input. For finite-length codes, improved guessing schemes motivated by the information theoretic study are proposed that attain lower error rates than ORBGRAND, especially in the high signal-to-noise ratio regime. Mengxiao Liu, Yuejun Wei, Zhenyuan Chen, Wenyi Zhang 0001 |
IEEE Trans. Inf. Theory | 3 |
| 2022 | Class Activation Map Refinement via Semantic Affinity Exploration for Weakly Supervised Object DetectionabstractWeakly Supervised Object Detection (WSOD) aims to train a detector to specify the interesting targets in an image by only using image-level labels. An important trend of current WSOD methods is to integrate object detection with Weakly Supervised Semantic Segmentation (WSSS), so that more discriminative regions can be obtained and more accurate detection can be achieved. However, due to the unreliable segmentation supervision generated by WSOD, their performance is still very limited. To address this problem, in this paper, we propose a novel end-to-end framework termed Class activation map Guided Detection Network (CGDN), where the detection process is guided by Class Activation Map (CAM) rather than the segmentation results. The proposed CGDN is composed of a detection branch and a CAM refinement branch, where the CAM refinement branch critically refines the CAMs generated by the detection branch, and then the refined CAMs are deployed to provide more reliable fore-ground cues for the detection branch in turn. Therefore, the two branches interact which leads to progressively improved detection and CAM outputs. Extensive experiments on PASCAL VOC 2007 and 2012 datasets verify the effectiveness of our proposed network. Zhenyuan Chen, Chen Gong 0002 |
ICIP | 2 |
| 2012 | Downlink multicasting beamforming with imperfect CSI on both transceiver sidesabstractThis paper deals with the issue of robust beamforming design of multicast downlink system with bounded imperfect channel state information (CSI) on both transceiver sides. Two optimization objectives are considered: the minimization of transmit power that subject to the quality of service (QoS) constraints of all receivers; the maximization of worst case effective signal-to-interference-plus-noise ratio (eSINR) bounded by the transmit power. Both of the optimization problems encounter non-convex constraints. By utilizing mathematical approaches like S-lemma and semidefinite relaxations (SDR), these stubborn obstacles can be eliminated and the problems can therefore be solved efficiently. Numerical results are provided to demonstrate the performance of proposed designs. Qiushi Gong, Zhenyuan Chen |
PIMRC | 2 |
| 2012 | Robust Transmit Beamforming for Multigroup MulticastingabstractThis paper addresses a robust downlink beamforming optimization problem for the multigroup multicast scenario, when only imperfect channel state information (CSI) is available at the transmitter. We consider two different optimization criteria: minimizing the total transmit power subject to quality of service (QoS) constraints at each receiver; max-min fair (MMF) signal-to-interference- plus-noise ratio (SINR) subject to total power constraint. With the aid of S-lemma, the infinite non-convex QoS constraints of robust downlink beamforming problem are transformed into finite linear matrix inequalities (LMI). By applying the semidefinite relaxation (SDR) method, the robust downlink beamforming problem can be relaxed and solved efficiently. Simulation results are presented to corroborate our design. Zhenyuan Chen, Wenyi Zhang 0001, Guo Wei 0001 |
VTC Fall | 1 |