Zichong Chen

dblp:57/8656 · DBLP profile ↗
← Back
26ranked-venue papers
14as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 8 since 2021Computer networks · 3 · 3 first-authorSystems, architecture and hardware · 2Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Multi-feature collaboration with spatial-frequency learning guided by vision foundation model for remote sensing image captioning
Jian Cheng 0003, Ziying Xia, Siyu Liu 0003, Changjian Deng, Zunni Zhu, Zichong Chen
Neurocomputing7
2025 Attention Distillation: A Unified Approach to Visual Characteristics Transfer
abstract
Recent advances in generative diffusion models have shown a notable inherent understanding of image style and semantics. In this paper, we leverage the self-attention features from pretrained diffusion networks to transfer the visual characteristics from a reference to generated images. Unlike previous work that uses these features as plug-and-play attributes, we propose a novel attention distillation loss calculated between the ideal and current stylization results, based on which we optimize the synthesized image via backpropagation in latent space. Next, we propose an improved Classifier Guidance that integrates attention distillation loss into the denoising sampling process, further accelerating the synthesis and enabling a broad range of image generation applications. Extensive experiments have demonstrated the extraordinary performance of our approach in transferring the examples’ style, appearance, and texture to new images in synthesis. Code is available at https://github.com/xugao97/AttentionDistillation.
Yang Zhou 0007, Zichong Chen, Hui Huang 0004
CVPR3
2025 Rectified Mixed-Label Learning for Semi-Supervised Medical Image Segmentation
abstract
Semi-supervised medical image segmentation (SSMIS) has gained increasing attention due to its potential to alleviate the manual annotation burden. However, existing works face two key challenges: i) how to deal with the information loss caused by learning labeled and unlabeled data in an inconsistent manner and ii) how to reduce the impact of label noise derived by the model’s cognitive bias. To address these challenges, we propose the Rectified Mixed-label Learning (RML) method for SSMIS. First, we mix labeled and unlabeled images, and then encourage the model to learn common semantics from mixed images straightly. More importantly, we conduct mixed-label learning in both the image and feature levels to fully exploit mixed images. In detail, a rectified supervision objective is implemented at the image level, which adaptively enhances high-quality pseudo-labels while weakening unreliable pseudo-labels. Moreover, for ambiguous voxels, we guide them to acquire reliable semantic information from high-quality prototypes in the feature space, thereby improving the identification of unreliable regions. Numerous experimental results demonstrate the superiority of our proposed method over previous SoTA methods.
Zeyu An, Zichong Chen
ICME2
2025 Region Confidence Refinement with Progressive Semantic Mining for Source-Free Domain Adaptive Object Detection
abstract
Source-Free Domain Adaptive (SFDA) object detection addresses the challenges of detection in scenarios where both source domain data and target domain labels are unavailable. Due to the lack of data supervision, pseudo label learning has become the key to SFDA object detection. However, prevailing SFDA methods primarily concentrate on pseudo labels exhibiting exceptionally high or low confidence, without simultaneously considering false negative samples in high confidence and false positive samples in low confidence. We summarize this issue as the double-sided problem of pseudo labels. To address this issue, we propose the Region Confidence Refinement (RCR) aimed at refining the quality of pseudo labels via progressive semantic mining. Specifically, we bolster the semantic representation capacity of the detector across both pixel and image levels. Firstly, we design the Multi-Channel Style Filter (MSF) module to enrich pixel-level semantic representation by eliminating background-induced noise. Secondly, we design the Cross-Modal Semantic Enhancement (CSE) to enhance the classification efficacy of the detector amidst supervised information scarcity by aligning textual and image features, thereby amplifying image-level semantic representation. Finally, we design a Semantic Aggregation Strategy (SAS) for reconstructing region-level confidence. Extensive experiments demonstrate our proposed RCR achieves the state-of-the-art (SOTA) performance.
Zichong Chen, Zeyu An, Jian Cheng 0003
ICME1
2025 A reinforced final belief divergence for mass functions and its application in target recognition
Fuxiao Zhang, Zichong Chen, Rui Cai 0001
Appl. Intell.2
2025 StyleBlend: Enhancing Style-Specific Content Creation in Text-to-Image Diffusion Models
abstract
Abstract Synthesizing visually impressive images that seamlessly align both text prompts and specific artistic styles remains a significant challenge in Text‐to‐Image (T2I) diffusion models. This paper introduces StyleBlend, a method designed to learn and apply style representations from a limited set of reference images, enabling content synthesis of both text‐aligned and stylistically coherent. Our approach uniquely decomposes style into two components, composition and texture, each learned through different strategies. We then leverage two synthesis branches, each focusing on a corresponding style component, to facilitate effective style blending through shared features without affecting content generation. StyleBlend addresses the common issues of text misalignment and weak style representation that previous methods have struggled with. Extensive qualitative and quantitative comparisons demonstrate the superiority of our approach.
Zichong Chen, Yang Zhou 0007
Comput. Graph. Forum1
2025 SPG: Style-Prompting Guidance for Style-Specific Content Creation
abstract
Abstract Although recent text‐to‐image (T2I) diffusion models excel at aligning generated images with textual prompts, controlling the visual style of the output remains a challenging task. In this work, we propose Style‐Prompting Guidance (SPG), a novel sampling strategy for style‐specific image generation. SPG constructs a style noise vector and leverages its directional deviation from unconditional noise to guide the diffusion process toward the target style distribution. By integrating SPG with Classifier‐Free Guidance (CFG), our method achieves both semantic fidelity and style consistency. SPG is simple, robust, and compatible with controllable frameworks like ControlNet and IPAdapter, making it practical and widely applicable. Extensive experiments demonstrate the effectiveness and generality of our approach compared to state‐of‐the‐art methods. Code is available at https://github.com/Rumbling281441/SPG .
Zichong Chen, Yang Zhou 0007, Hui Huang 0004
Comput. Graph. Forum2
2025 High-Fidelity Texture Transfer Using Multi-Scale Depth-Aware Diffusion
abstract
Abstract Textures are a key component of 3D assets. Transferring textures from one shape to another, without user interaction or additional semantic guidance, is a classical yet challenging problem. It can enhance the diversity of existing shape collections, augmenting their application scope. This paper proposes an innovative 3D texture transfer framework that leverages the generative power of pre‐trained diffusion models. While diffusion models have achieved significant success in 2D image generation, their application to 3D domains faces great challenges in preserving coherence across different viewpoints. Addressing this issue, we designed a multi‐scale generation framework to optimize the UV maps coarse‐to‐fine. To ensure multi‐view consistency, we use depth info as geometric guidance; meanwhile, a novel consistency loss is proposed to further constrain the color coherence and reduce artifacts. Experimental results demonstrate that our multi‐scale framework not only produces high‐quality texture transfer results but also excels in handling complex shapes while preserving correct semantic correspondences. Compared to existing techniques, our method achieves improvements in both consistency and texture clarity, as well as time efficiency.
Rongzhen Lin, Zichong Chen, Xiaoyong Hao, Yang Zhou 0007, Hui Huang 0004
Comput. Graph. Forum2
2025 Focusing on feature-level domain alignment with text semantic for weakly-supervised domain adaptive object detection
Zichong Chen, Jian Cheng 0003, Ziying Xia, Yongxiang Hu 0004, Zhicheng Dong 0003, Nyima Tashi
Neurocomputing1
2025 CITAL: Counterfactual intervention for temporal action localization with point-level annotation
Yongxiang Hu 0004, Ziying Xia, Zichong Chen, Thupten Tsering, Jian Cheng 0003, Nyima Tashi
Neurocomputing3
2024 Deformable One-Shot Face Stylization via DINO Semantic Guidance
abstract
This paper addresses the complex issue of one-shot face stylization, focusing on the simultaneous consideration of appearance and structure, where previous methods have fallen short. We explore deformation-aware face stylization that diverges from traditional single-image style reference, opting for a real-style image pair instead. The cornerstone of our method is the utilization of a self-supervised vision transformer, specifically DINO-ViT, to establish a robust and consistent facial structure representation across both real and style domains. Our stylization process begins by adapting the StyleGAN generator to be deformation-aware through the integration of spatial transformers (STN). We then introduce two innovative constraints for generator fine-tuning under the guidance of DINO semantics: i) a directional deformation loss that regulates directional vectors in DINO space, and ii) a relative structural consistency constraint based on DINO token self-similarities, ensuring diverse generation. Additionally, style-mixing is employed to align the color generation with the reference, minimizing inconsistent correspondences. This framework delivers enhanced deformability for general one-shot face stylization, achieving notable efficiency with a fine-tuning duration of approximately 10 minutes. Extensive qualitative and quantitative comparisons demonstrate our superiority over state-of-the-art one-shot face stylization methods. Code is available at https://github.com/zichongc/DoesFS.
Yang Zhou 0007, Zichong Chen, Hui Huang 0004
CVPR2
2024 TS-ILM: Class Incremental Learning for Online Action Detection
abstract
Online action detection aims to identify ongoing actions within untrimmed video streams, with extensive applications in real-life scenarios. However, in practical applications, video frames are received sequentially over time and new action categories continually emerge, giving rise to the challenge of catastrophic forgetting - a problem that remains inadequately explored. Generally, in the field of video understanding, researchers address catastrophic forgetting through class-incremental learning. Nevertheless, online action detection is based solely on historical observations, thus demanding higher temporal modeling capabilities for class-incremental learning methods. In this paper, we conceptualize this task as Class-Incremental Online Action Detection (CIOAD) and propose a novel framework, TS-ILM, to address it. Specifically, TS-ILM consists of two components: task-level temporal pattern extractor and temporal-sensitive exemplar selector. The former extracts the temporal patterns of actions in different tasks and saves them, allowing the data to be comprehensively observed on a temporal level before it is input into the backbone. The latter selects a set of frames with the highest causal relevance and minimum information redundancy for subsequent replay, enabling the model to learn the temporal information of previous tasks more effectively. We benchmark our approach against SoTA class-incremental learning methods applied in the image and video domains on THUMOS'14 and TVSeries datasets. Our method outperforms the previous approaches.
Jian Cheng 0003, Ziying Xia, Zichong Chen, Junhao Shi, Zhicheng Dong 0003, Nyima Tashi
ACM Multimedia4
2024 Semantic consistency knowledge transfer for unsupervised cross domain object detection
Zichong Chen, Ziying Xia, Junhao Shi, Nyima Tashi, Jian Cheng 0003
Appl. Intell.1
2024 A belief interval euclidean distance entropy of the mass function and its application in multi-sensor data fusion
Fuxiao Zhang, Zichong Chen, Rui Cai 0001
Appl. Intell.2
2024 An adaptive optimization machine of mass function for conflict management
Zichong Chen, Rui Cai 0001
Eng. Appl. Artif. Intell.1
2024 Symmetric Renyi-Permutation divergence and conflict management for random permutation set
Zichong Chen, Rui Cai 0001
Expert Syst. Appl.1
2022 Updating incomplete framework of target recognition database based on fuzzy gap statistic
Zichong Chen, Rui Cai 0001
Eng. Appl. Artif. Intell.1
2022 A novel divergence measure of mass function for conflict management
abstract
Dempster–Shafer evidence theory, which is an extension of Bayesian probability theory, is a useful approach to realize multisensor data fusion. It uses mass functions to represent uncertainty, which can produce a satisfactory fusion result. However, when the evidence is highly conflicting, using Dempster–Shafer evidence theory fusion rule to combine the evidence will generate the result contrary to common sense. To solve this issue, we propose a new method for conflict management based on Renyi divergence (RD). Then, by combining RD with the mass function, we develop Renyi-Belief divergence (RBD). To expand its utility, we modify it and define the modified Renyi-Belief divergence (MRBD). Our method MRBD integrates the characteristics of mass functions and can handle conflict by measuring the differences between mass functions. Experiments show that MRBD can effectively deal with conflicts. After dealing with the conflicting evidence, we realize multisensor data fusion based on the Dempster–Shafer combination rule. Moreover, we also consider the information quality and belief entropy to reinforce the credibility of evidence. A large number of examples show that the proposed method is feasible and efficient. Finally, in the application of fault diagnosis, our method can effectively determine the fault type.
Zichong Chen, Rui Cai 0001
Int. J. Intell. Syst.1
2018 Map-based Deep Imitation Learning for Obstacle Avoidance
abstract
Making an optimal decision to avoid obstacles while heading to the goal is one of the fundamental challenges for mobile robots equipped with limited computational resources. In this paper, we present a deep imitation learning algorithm that develops a computationally efficient obstacle avoidance policy based on egocentric local occupancy maps. The trained model embedded with a variant of the value iteration networks is able to provide near-optimal continuous action commands through fast feed-forward inferences and generalize well to unseen planning-based scenarios. To improve the policy robustness, we augment the training data set with artificially generated maps, which effectively alleviates the shortage of catastrophic samples in normal demonstrations. Extensive experiments on a Segway robot show the effectiveness of the proposed approach in terms of solution optimality, robustness as well as computation time.
Yuejiang Liu, An Xu, Zichong Chen
IROS3
2017 Depth enhanced visual-inertial odometry based on Multi-State Constraint Kalman Filter
abstract
There have been increasing demands for developing robotic system combining camera and inertial measurement unit in navigation task, due to their low-cost, lightweight and complementary properties. In this paper, we present a Visual Inertial Odometry (VIO) system which can utilize sparse depth to estimate 6D pose in GPS-denied and unstructured environments. The system is based on Multi-State Constraint Kalman Filter (MSCKF), which benefits from low computation load when compared to optimization-based method, especially on resource-constrained platform. Features are enhanced with depth information forming 3D landmark position measurements in space, which reduces uncertainty of position estimate. And we derivate measurement model to access compatibility with both 2D and 3D measurements. In experiments, we evaluate the performance of the system in different in-flight scenarios, both cluttered room and industry environment. The results suggest that the estimator is consistent, substantially improves the accuracy compared with original monocular-based MSKCF and achieves competitive accuracy with other research.
Fumin Pang, Zichong Chen, Li Pu, Tianmiao Wang
IROS2
2015 DASS: Distributed Adaptive Sparse Sensing
abstract
Wireless sensor networks are often designed to perform two tasks: sensing a physical field and transmitting the data to end-users. A crucial design aspect of a WSN is the minimization of the overall energy consumption. Previous researchers aim at optimizing the energy spent for the communication, while mostly ignoring the energy cost of sensing. Recently, it has been shown that considering the sensing energy cost can be beneficial for further improving the overall energy efficiency. More precisely, sparse sensing techniques were proposed to reduce the amount of collected samples and recover the missing data using data statistics. While the majority of these techniques use fixed or random sampling patterns, we propose adaptively learning the signal model from the measurements and using the model to schedule when and where to sample the physical field. The proposed method requires minimal on-board computation, no inter-node communications, and achieves appealing reconstruction performance. With experiments on real-world datasets, we demonstrate significant improvements over both traditional sensing schemes and the state-of-the-art sparse sensing schemes, particularly when the measured data is characterized by a strong intra-sensor (temporal) or inter-sensors (spatial) correlation.
Zichong Chen, Juri Ranieri, Runwei Zhang, Martin Vetterli
IEEE Trans. Wirel. Commun.1
2012 Event-driven video coding for outdoor wireless monitoring cameras
abstract
Reducing communication cost is crucial for outdoor wireless monitoring cameras which are constrained by limited energy budgets. From event detection point of view, traditional video coding schemes such as H.264 are inefficient as they ignore the “meaning” of video content and thus waste many bits to convey irrelevant information. To take advantage of the powerful computing resource on cameras, we propose a novel event-driven video coding scheme. Unlike previous approach that attempts to find anomalous image frame with potential events, we propose to detect salient regions in each image and transmit the image fragments marked with saliency to the receiver. This scheme rarely drops an event as it transmits all image fragments with potential events, and also requires no training procedure. The experimental results show that it performs substantially better than conventional video coding schemes for outdoor monitoring task.
Zichong Chen, Guillermo Barrenetxea, Martin Vetterli
ICIP1
2012 Howis the weather: Automatic inference from images
abstract
Low-cost monitoring cameras/webcams provide unique visual information. To take advantage of the vast image dataset captured by a typical webcam, we consider the problem of retrieving weather information from a database of still images. The task is to automatically label all images with different weather conditions (e.g., sunny, cloudy, and overcast), using limited human assistance. To address the drawbacks in existing weather prediction algorithms, we first apply image segmentation to the raw images to avoid disturbance of the non-sky region. Then, we propose to use multiple kernel learning to gather and select an optimal subset of image features from a certain feature pool. To further increase the recognition performance, we adopt multi-pass active learning for selecting the training set. The experimental results show that our weather recognition system achieves high performance.
Zichong Chen, Albrecht J. Lindner, Guillermo Barrenetxea, Martin Vetterli
ICIP1
2012 Share risk and energy: Sampling and communication strategies for multi-camera wireless monitoring networks
abstract
In the context of environmental monitoring, outdoor wireless cameras are vulnerable to natural hazards. To benefit from the inexpensive imaging sensors, we introduce a multi-camera monitoring system to share the physical risk. With multiple cameras focusing at a common scenery of interest, we propose an interleaved sampling strategy to minimize per-camera consumption by distributing sampling tasks among cameras. To overcome the uncertainties in the sensor network, we propose a robust adaptive synchronization scheme to build optimal sampling configuration by exploiting the broadcast nature of wireless communication. The theory as well as simulation results verify the fast convergence and robustness of the algorithm. Under the interleaved sampling configuration, we propose three video coding methods to compress correlated video streams from disjoint cameras, namely, distributed/independent/joint coding schemes. The energy profiling on a two-camera system shows that independent and joint coding perform substantially better. The comparison between two-camera and single-camera system shows 30%-50% per-camera consumption reduction. On top of these, we point out that MIMO technology can be potentially utilized to push the communication consumption even lower.
Zichong Chen, Guillermo Barrenetxea, Martin Vetterli
INFOCOM1
2012 Sensorcam: an energy-efficient smart wireless camera for environmental monitoring
abstract
Reducing energy cost is crucial for energy-constrained smart wireless cameras. Existing platforms impose two main challenges: First, most commercial smart phones have a closed platform, which makes it impossible to manage low-level circuits. Since the sampling frequency is moderate in environmental monitoring context, any improper power management in idle period will incur significant energy leak. Secondly, low-end cameras tailored for wireless sensor networks usually have limited processing power or communication range, and thus are not capable of outdoor monitoring task under low data rate. To tackle these issues, we develop Sensorcam, a long-range, smart wireless camera running a Linux-base open system. Through better power management in idle period and the "intelligence" of the camera itself, we demonstrate an energy-efficient wireless monitoring system in a real deployment.
Zichong Chen, Paolo Prandoni, Guillermo Barrenetxea, Martin Vetterli
IPSN1
2012 Distributed Successive Refinement of Multiview Images Using Broadcast Advantage
abstract
In environmental monitoring applications, having multiple cameras focus on common scenery increases robustness of the system. To save energy based on user demand, successive refinement image coding is important, as it allows us to progressively request better image quality. By exploiting the broadcast nature and correlation between multiview images, we investigate a two-camera setup and propose a novel two-encoder successive refinement scheme which imitates a ping-pong game. For the bivariate Gaussian case, we prove that this scheme is successively refinable on the theoretical rate-distortion limit of distributed coding (Wagner surface) under arbitrary settings. For stereo-view images, we develop a practical successive refinement coding algorithm using the same idea. The simulation results show that this scheme operates close to the distributed coding bound.
Zichong Chen, Guillermo Barrenetxea, Martin Vetterli
IEEE Trans. Image Process.1