EDBT 2026 Demo / reviewers in the wild / expert
Ziteng Cui
dblp:150/5743
· DBLP profile ↗
19ranked-venue papers
11as first author
14since 2021 · last 2026
0000-0002-6126-2712ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 10 since 2021Computer networks · 4 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Perception-Inspired Color Space Design for Photo White Balance EditingabstractWhite balance (WB) is a key step in the image signal processor (ISP) pipeline that mitigates color casts caused by varying illumination and restores the scene’s true colors. Currently, sRGB-based WB editing for post-ISP WB correction is widely used [2], [18] to address color constancy failures in the ISP pipeline when the original camera RAW is unavailable. However, additive color models (e.g., sRGB) are inherently limited by fixed nonlinear transformations and entangled color channels, which often impede their generalization to complex lighting conditions.To address these challenges, we propose a novel framework for WB correction that leverages a perception-inspired Learnable HSI (LHSI) color space. Built upon a cylindrical color model that naturally separates luminance from chromatic components, our framework further introduces dedicated parameters to enhance this disentanglement and learnable mapping to adaptively refine the flexibility. Moreover, a new Mamba-based network is introduced, which is tailored to the characteristics of the proposed LHSI color space.Experimental results on benchmark datasets demonstrate the superiority of our method, highlighting the potential of perception-inspired color space design in computational photography. The source code is avail-able at https://github.com/YangCheng58/WB_Color_Space. Yang Cheng 0008, Ziteng Cui, Lin Gu 0003, Shenghan Su, Zenghui Zhang |
WACV | 2 |
| 2025 | Luminance-GS: Adapting 3D Gaussian Splatting to Challenging Lighting Conditions with View-Adaptive Curve AdjustmentabstractCapturing high-quality photographs under diverse real-world lighting conditions is challenging, as both natural lighting (e.g., low-light) and camera exposure settings (e.g., exposure time) significantly impact image quality. This challenge becomes more pronounced in multi-view scenarios, where variations in lighting and image signal processor (ISP) settings across viewpoints introduce photometric inconsistencies. Such lighting degradations and view-dependent variations pose substantial challenges to novel view synthesis (NVS) frameworks based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS).To address this, we introduce Luminance-GS, a novel approach to achieving high-quality novel view synthesis results under diverse challenging lighting conditions using 3DGS. By adopting per-view color matrix mapping and view adaptive curve adjustments, Luminance-GS achieves state-of-the-art (SOTA) results across various lighting conditions—including low-light, overexposure, and varying exposure—while not altering the original 3DGS explicit representation. Compared to previous NeRF- and 3DGS-based baselines, Luminance-GS provides real-time rendering speed with improved reconstruction quality. The source code is available at1. Ziteng Cui, Xuangeng Chu, Tatsuya Harada |
CVPR | 1 |
| 2025 | Dr. RAW: Towards General High-Level Vision from RAW with Efficient Task ConditioningabstractWe introduce Dr. RAW, a unified and tuning-efficient framework for high-level computer vision tasks directly operating on camera RAW data. Unlike previous approaches that optimize image signal processing (ISP) pipelines and fully fine-tune networks for each task, Dr. RAW achieves state-of-the-art performance with minimal parameter updates. At the input stage, we apply lightweight pre-processing modules, sensor and illumination mapping, followed by re-mosaicing, to mitigate data inconsistencies stemming from sensor variation and lighting. At the network level, we introduce task-specific adaptation through two modules: Sensor Prior Prompts (SPP) and Low-Rank Adaptation (LoRA). SPP injects sensor-aware conditioning into the network via learnable prompts derived from imaging priors, while LoRA enables efficient task-specific tuning by updating only low-rank matrices in key backbone layers. Despite minimal tuning, our method delivers superior results across four RAW-based tasks (object detection, semantic segmentation, instance segmentation, and pose estimation) on nine datasets encompassing low-light and over-exposed conditions. By harnessing the intrinsic physical cues of RAW data alongside parameter-efficient techniques, our method advances RAW-based vision systems, achieving both high accuracy and computational economy. We will release our source code. Wenjun Huang 0001, Ziteng Cui, Yinqiang Zheng, Yirui He, Tatsuya Harada, Mohsen Imani |
NeurIPS | 2 |
| 2025 | I2-NeRF: Learning Neural Radiance Fields Under Physically-Grounded Media InteractionsabstractParticipating in efforts to endow generative AI with the 3D physical world perception, we propose I2-NeRF, a novel neural radiance field framework that enhances isometric and isotropic metric perception under media degradation. While existing NeRF models predominantly rely on object-centric sampling, I2-NeRF introduces a reverse-stratified upsampling strategy to achieve near-uniform sampling across 3D space, thereby preserving isometry. We further present a general radiative formulation for media degradation that unifies emission, absorption, and scattering into a particle model governed by the Beer–Lambert attenuation law. By matting direct and media-induced in-scatter radiance, this formulation extends naturally to complex media environments such as underwater, haze, and even low-light scenes. By treating light propagation uniformly in both vertical and horizontal directions, I2-NeRF enables isotropic metric perception and can even estimate medium properties such as water depth. Experiments on real-world datasets demonstrate that our method significantly improves both reconstruction fidelity and physical plausibility compared to existing approaches. The source code is available at https://github.com/ShuhongLL/I2-NeRF. Shuhong Liu, Lin Gu 0003, Ziteng Cui, Xuangeng Chu, Tatsuya Harada |
NeurIPS | 3 |
| 2025 | ARTalk: Speech-Driven 3D Head Animation via Autoregressive ModelabstractSpeech-driven 3D facial animation aims to generate realistic lip movements and facial expressions for 3D head models from arbitrary audio clips. Although existing diffusion-based methods are capable of producing natural motions, their slow generation speed limits their application potential. In this paper, we introduce a novel autoregressive model that achieves real-time generation of highly synchronized lip movements and realistic head poses and eye blinks by learning a mapping from speech to a multi-scale motion codebook. Furthermore, our model can adapt to unseen speaking styles, enabling the creation of 3D talking avatars with unique personal styles beyond the identities seen during training. Extensive evaluations and user studies demonstrate that our method outperforms existing approaches in lip synchronization accuracy and perceived quality. Demos and codes are available at https://xg-chu.site/project_artalk/. Xuangeng Chu, Nabarun Goswami, Ziteng Cui, Hanqin Wang, Tatsuya Harada |
SIGGRAPH Asia | 3 |
| 2024 | Aleth-NeRF: Illumination Adaptive NeRF with Concealing Field AssumptionabstractThe standard Neural Radiance Fields (NeRF) paradigm employs a viewer-centered methodology, entangling the aspects of illumination and material reflectance into emission solely from 3D points. This simplified rendering approach presents challenges in accurately modeling images captured under adverse lighting conditions, such as low light or over-exposure. Motivated by the ancient Greek emission theory that posits visual perception as a result of rays emanating from the eyes, we slightly refine the conventional NeRF framework to train NeRF under challenging light conditions and generate normal-light condition novel views unsupervisedly. We introduce the concept of a ``Concealing Field," which assigns transmittance values to the surrounding air to account for illumination effects. In dark scenarios, we assume that object emissions maintain a standard lighting level but are attenuated as they traverse the air during the rendering process. Concealing Field thus compel NeRF to learn reasonable density and colour estimations for objects even in dimly lit situations. Similarly, the Concealing Field can mitigate over-exposed emissions during rendering stage. Furthermore, we present a comprehensive multi-view dataset captured under challenging illumination conditions for evaluation. Our code and proposed dataset are available at https://github.com/cuiziteng/Aleth-NeRF. Ziteng Cui, Lin Gu 0003, Xiao Sun 0001, Xianzheng Ma, Yu Qiao 0001, Tatsuya Harada |
AAAI | 1 |
| 2024 | Discovering an Image-Adaptive Coordinate System for Photography Processing
Ziteng Cui, Lin Gu 0003, Tatsuya Harada |
BMVC | 1 |
| 2024 | RAW-Adapter: Adapting Pre-trained Visual Model to Camera RAW Images
Ziteng Cui, Tatsuya Harada |
ECCV (4) | 1 |
| 2023 | MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionabstractMonocular 3D object detection has long been a challenging task in autonomous driving. Most existing methods follow conventional 2D detectors to first localize object centers, and then predict 3D attributes by neighboring features. However, only using local visual features is insufficient to understand the scene-level 3D spatial structures and ignores the long-range inter-object depth relations. In this paper, we introduce the first DETR framework for Monocular DEtection with a depth-guided TRansformer, named MonoDETR. We modify the vanilla transformer to be depth-aware and guide the whole detection process by contextual depth cues. Specifically, concurrent to the visual encoder that captures object appearances, we introduce to predict a foreground depth map, and specialize a depth encoder to extract non-local depth embeddings. Then, we formulate 3D object candidates as learnable queries and propose a depth-guided decoder to conduct object-scene depth interactions. In this way, each object query estimates its 3D attributes adaptively from the depth-guided regions on the image and is no longer constrained to local visual features. On KITTI benchmark with monocular images as input, MonoDETR achieves state-of-the-art performance and requires no extra dense depth annotations. Besides, our depth-guided modules can also be plug-and-play to enhance multi-view 3D object detectors on nuScenes dataset, demonstrating our superior generalization capacity. Code is available at https://github.com/ZrrSkywalker/MonoDETR. Renrui Zhang, Han Qiu 0010, Ziteng Cui, Yu Qiao 0001, Hongsheng Li 0001, Peng Gao 0007 |
ICCV | 5 |
| 2022 | You Only Need 90K Parameters to Adapt Light: a Light Weight Transformer for Image Enhancement and Exposure Correction
Ziteng Cui, Kunchang Li 0002, Lin Gu 0003, Shenghan Su, Peng Gao 0007, Zhengkai Jiang 0001, Yu Qiao 0001, Tatsuya Harada |
BMVC | 1 |
| 2022 | Exploring Resolution and Degradation Clues as Self-supervised Signal for Low Quality Object Detection
Ziteng Cui, Yingying Zhu 0004, Lin Gu 0003, Guo-Jun Qi, Renrui Zhang, Zenghui Zhang, Tatsuya Harada |
ECCV (9) | 1 |
| 2022 | Explainable Analysis of Deep Learning Methods for Sar Image ClassificationabstractDeep learning methods exhibit outstanding performance in synthetic aperture radar (SAR) image interpretation tasks. However, these are black box models that limit the com-prehension of their predictions. Therefore, to meet this challenge, we have utilized explainable artificial intelli-gence (XAI) methods for the SAR image classification task. Specifically, we trained state-of-the-art convolutional neural networks for each polarization format on OpenSARUrban dataset and then investigate eight explanation methods to analyze the predictions of the CNN classifiers of SAR images. These XAI methods are also evaluated qualitatively and quantitatively which shows that Occlusion achieves the most reliable interpretation performance in terms of Max-Sensitivity but with a low-resolution explanation heatmap. The explanation results provide some insights into the in-ternal mechanism of black-box decisions for SAR image classification. Shenghan Su, Ziteng Cui, Weiwei Guo, Zenghui Zhang, Wenxian Yu |
IGARSS | 2 |
| 2021 | Multitask AET with Orthogonal Tangent Regularity for Dark Object DetectionabstractDark environment becomes a challenge for computer vision algorithms owing to insufficient photons and undesirable noise. To enhance object detection in a dark environment, we propose a novel multitask auto encoding transformation (MAET) model which is able to explore the intrinsic pattern behind illumination translation. In a self-supervision manner, the MAET learns the intrinsic visual structure by encoding and decoding the realistic illumination-degrading transformation considering the physical noise model and image signal processing (ISP). Based on this representation, we achieve the object detection task by decoding the bounding box coordinates and classes. To avoid the over-entanglement of two tasks, our MAET disentangles the object and degrading features by imposing an orthogonal tangent regularity. This forms a parametric manifold along which multitask predictions can be geometrically formulated by maximizing the orthogonality between the tangents along the outputs of respective tasks. Our framework can be implemented based on the mainstream object detection architecture and directly trained end-to-end using normal target detection datasets, such as VOC and COCO. We have achieved the state-of-the-art performance using synthetic and real-world datasets. Codes will be released at https://github.com/cuiziteng/MAET. Ziteng Cui, Guo-Jun Qi, Lin Gu 0003, Shaodi You, Zenghui Zhang, Tatsuya Harada |
ICCV | 1 |
| 2021 | Self-Supervised Auto-Encoding Multi-Transformations for Airplane ClassificationabstractIn this paper, we present a self-supervised learning method of Auto-Encoding Multi-Transformations (AEMT) for airplane classification. In this method, the image features are learned in an unsupervised way by simultaneously estimating multiple image transformations from the features of original and transformed images instead of reconstructing the input images. Besides, we propose two structure variants of the AEMT method: composite and parallel modes of which the former transforms the images in a composite fashion while the latter does it in parallel. The experimental results demonstrate that the proposed method outperforms the state-of-the-art self-supervised learning methods for the airplane classification task. Ziteng Cui, Weiwei Guo, Zenghui Zhang, Wenxian Yu |
IGARSS | 2 |
| 2020 | Ellipse-FCN: Oil Tanks Detection from Remote Sensing Images with Fully Convolution NetworkabstractOil is an essential asset for every country, and plays a key role in world trade system. The detection of oil tanks is a very important task for both military and commerce. Recently, researchers have shown an increasing interest in oil tanks detection in remote sensing imagery. However, the previous works almost used the methods of circle detection, but the real oil tanks in remote sensing imagery are more close to ellipses. In this paper, we propose an oil tanks detector base on an U-shape Fully Convolutional Network(FCN) in optical remote sensing images. The structure of our network consists of three parts: feature extraction part, feature merge part and the output layer. The output layer consists of two output branches, one branch is a score map branch, which generates confidence score to indicate the region of oil tanks at pixel wise, and the other ends up with several channels which regress the ellipse geometric parameters (center, horizontal axis and vertical axis). In addition, we also design a novel loss function adapted to our network. The experimental results conducted on our dataset collected from Google Earth show that this method achieves promising performance on oil tanks detection in terms of both efficiency and accuracy in high-resolution optical remote sensing images. Ziteng Cui, Weiwei Guo, Zenghui Zhang, Huiyuan Chen, Wenxian Yu |
IGARSS | 1 |
| 2016 | An Approach to Improve the Cooperation between Heterogeneous SDN OverlaysabstractThe overlay network has been widely developed in recent years. There may be various overlays that co-exist with each other upon the same underlying network. These overlays have heterogeneous performance goals, and they will compete for the physical resources, so that a sub-optimal performance of the overlays may be achieved. Moreover, the heterogeneity of the overlays makes them difficult to coordinate with each other to improve their performance. We introduce the concept of SDN to the deployment of overlay network and propose an approach to make the overlays cooperate with each other. A cooperative solution is proposed for co-existing overlays to improve their performance while leveraging their heterogeneous performance goals. Simulations are performed to evaluate the cooperative solution. Ziteng Cui, Jianxin Liao, Jingyu Wang 0001, Qi Qi 0001, Jing Wang 0039 |
LCN | 1 |
| 2015 | Cooperative traffic management for co-existing overlaysabstractThe overlay network has been widely deployed by Service Providers to provide services. Since there are multiple SPs built upon the same ISP, their overlays are co-existing and may interfere with each other. The selfishness of overlay may lead to sub-optimal performance and traffic arrangement dilemma for overlays. To optimize the performances of overlays and maximize the benefit of SPs, we proposed a cooperative traffic management framework. Several models are applied to analyze and solve the overlay routing problem, the revenue allocation problem, and the coalition formation problem in the framework. Simulations are performed to evaluate the framework. Ziteng Cui, Jianxin Liao, Jingyu Wang 0001, Qi Qi 0001, Jing Wang 0039 |
LCN | 1 |
| 2015 | A coalitional game approach on improving interactions in multiple overlay environments
Jianxin Liao, Ziteng Cui, Jingyu Wang 0001, Tonghong Li, Qi Qi 0001, Jing Wang 0039 |
Comput. Networks | 2 |
| 2014 | Cooperative overlay routing in a multiple overlay environmentabstractOverlay networks have been widely developed over the past few years. More and more overlays are deployed on the top of the same native network, and share the same physical resources. Competing for these physical resources, co-existing overlays may affect each other adversely. It has been showed that by using selfish overlay routing, co-existing overlays would be likely to converge to a Nash equilibrium which is sub-optimal. However, to achieve the global optimal may also cause the performance degradation of certain overlays, which make it hard to realize. Inspired by the Nash bargaining solution, a cooperative method is proposed for two co-existing overlays to achieve a near Pareto optimal. Simulations are performed to evaluate the proposed approach. The results show that the approach is effective and efficient in the multiple overlay networks environment. Ziteng Cui, Jianxin Liao, Jingyu Wang 0001, Qi Qi 0001, Jing Wang 0039 |
ICC | 1 |