Kun Zhan

dblp:46/8462 · DBLP profile ↗
← Back
55ranked-venue papers
13as first author
38since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 8 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 3 first-author · 29 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 DriveLiDAR4D: Sequential and Controllable LiDAR Scene Generation for Autonomous Driving
abstract
The generation of realistic LiDAR point clouds plays a crucial role in the development and evaluation of autonomous driving systems. Although recent methods for 3D LiDAR point cloud generation have shown significant improvements, they still face notable limitations, including the lack of sequential generation capabilities and the inability to produce accurately positioned foreground objects and realistic backgrounds. These shortcomings hinder their practical applicability. In this paper, we introduce DriveLiDAR4D, a novel LiDAR generation pipeline consisting of multimodal conditions and a novel sequential noise prediction model LiDAR4DNet, capable of producing temporally consistent LiDAR scenes with highly controllable foreground objects and realistic backgrounds. To the best of our knowledge, this is the first work to address the sequential generation of LiDAR scenes with full scene manipulation capability in an end-to-end manner. We evaluated DriveLiDAR4D on the nuScenes and KITTI datasets, where we achieved an FRD score of 743.13 and an FVD score of 16.96 on the nuScenes dataset, surpassing the current state-of-the-art (SOTA) method, UniScene, with an performance boost of 37.2% in FRD and 24.1% in FVD, respectively.
Kaiwen Cai, Hengtong Hu, Xueyang Zhang, Kun Zhan, Yifei Zhan, Xianpeng Lang
AAAI8
2026 CorrectAD: A Self-Correcting Agentic System to Improve End-to-end Planning in Autonomous Driving
abstract
End-to-end planning methods are the de-facto standard of the current autonomous driving system, while the robustness of the data-driven approaches suffers due to the notorious long-tail problem (i.e., rare but safety-critical failure cases). In this work, we explore whether recent diffusion-based video generation methods (a.k.a. world models), paired with structured 3D layouts, can enable a fully automated pipeline to self-correct such failure cases. We first introduce an agent to simulate the role of product manager, dubbed PM-Agent, which formulates data requirements to collect data similar to the failure cases. Then, we use a generative model that can simulate both data collection and annotation. However, existing generative models struggle to generate high-fidelity data conditioned on 3D layouts. To address this, we propose DriveSora, which can generate spatiotemporally consistent videos aligned with the 3D annotations requested by PM-Agent. We integrate these components into our self-correcting agentic system, CorrectAD. Importantly, our pipeline is end-to-end model agnostic and can be applied to improve any end-to-end planner. Evaluated on both nuScenes and a more challenging in-house dataset across multiple end-to-end planners, CorrectAD corrects 62.5% and 49.8% of failure cases, reducing collision rates by 39% and 27%, respectively.
Enhui Ma, Junpeng Jiang, Kun Zhan, Xueyang Zhang, Xianpeng Lang, Di Lin 0002, Kaicheng Yu
AAAI8
2026 WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous Driving
abstract
Latent World Models enhance scene representation through temporal self-supervised learning, presenting a perception annotation-free paradigm for end-to-end autonomous driving. However, the reconstruction-oriented representation learning tangles perception with planning tasks, leading to suboptimal optimization for planning. To address this challenge, we propose WorldRFT, a planning-oriented latent world model framework that aligns scene representation learning with planning via a hierarchical planning decomposition and local-aware interactive refinement mechanism, augmented by reinforcement learning fine-tuning (RFT) to enhance safety-critical policy performance. Specifically, WorldRFT integrates a vision-geometry foundation model to improve 3D spatial awareness, employs hierarchical planning task decomposition to guide representation optimization, and utilizes local-aware iterative refinement to derive a planning-oriented driving policy. Furthermore, we introduce Group Relative Policy Optimization (GRPO), which applies trajectory Gaussianization and collision-aware rewards to fine-tune the driving policy, yielding systematic improvements in safety. WorldRFT achieves state-of-the-art (SOTA) performance on both open-loop nuScenes and closed-loop NavSim benchmarks. On nuScenes, it reduces collision rates by 83% (0.30% → 0.05%). On NavSim, using camera-only sensors input, it attains competitive performance with the LiDAR-based SOTA method DiffusionDrive (87.8 vs. 88.1 PDMS).
Pengxuan Yang, Ben Lu, Zhongpu Xia, Yinfeng Gao, Kun Zhan, Xianpeng Lang, Yupeng Zheng
AAAI7
2026 Street Gaussians: Modeling Dynamic Urban Scenes With Gaussian Primitives
abstract
This paper aims to tackle the problem of modeling dynamic urban streets for autonomous driving scenes. Recent methods extend NeRF by incorporating tracked vehicle poses to animate vehicles, enabling photo-realistic view synthesis of dynamic urban street scenes. However, significant limitations are their slow training and rendering speed. We introduce Street Gaussians, a new explicit scene representation that tackles these limitations. Specifically, the dynamic urban scene is represented as a set of point clouds equipped with semantic logits and Gaussian primitives, each associated with either a foreground object or the background. To model the dynamics of foreground objects, each object point cloud is optimized with optimizable tracked poses, along with a 4D spherical harmonics model for the dynamic appearance. The explicit representation allows easy composition of objects and background, which in turn allows for scene editing operations and rendering at 135 FPS (1066 * 1600 resolution) within half an hour of training. The proposed method is evaluated on multiple challenging benchmarks, including KITTI and Waymo Open datasets. Experiments show that the proposed method consistently outperforms state-of-the-art methods across all datasets.
Sida Peng, Yushi Long, Yunzhi Yan, Haotong Lin, Chenxu Zhou, Kun Zhan, Xianpeng Lang, Hujun Bao, Xiaowei Zhou 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2025 Autonomous LLM-Enhanced Adversarial Attack for Text-to-Motion
abstract
Human motion generative models have enabled promising applications, but the ability of text-to-motion (T2M) models to produce realistic motions raises security concerns if exploited maliciously. Despite growing interest in T2M, limited research focus on safeguarding these models against adversarial attacks, with existing work on text-to-image models proving insufficient for the unique motion domain. In the paper, we propose ALERT-Motion, an autonomous framework that leverages large language models (LLMs) to generate targeted adversarial attacks against black-box T2M models. Unlike prior methods that modify prompts through predefined rules, ALERT-Motion uses the knowledge of LLMs of human motion to autonomously generate subtle yet powerful adversarial text descriptions. It comprises two key modules: an adaptive dispatching module that constructs an LLM-based agent to iteratively refine and search for adversarial prompts; and a multimodal information contrastive module that extracts semantically relevant motion information to guide the agent's search. Through this LLM-driven approach, ALERT-Motion produces adversarial prompts querying victim models to produce outputs closely matching targeted motions, while avoiding obvious perturbations. Evaluations across popular T2M models demonstrate ALERT-Motion's superiority over previous methods, achieving higher attack success rates with stealthier adversarial prompts. This pioneering work on T2M adversarial attacks highlights the urgency of developing defensive measures as motion generation technology advances, urging further research into safe and responsible deployment.
Honglei Miao, Fan Ma, Ruijie Quan, Kun Zhan, Yi Yang 0001
AAAI4
2025 BEV-TSR: Text-Scene Retrieval in BEV Space for Autonomous Driving
abstract
The rapid development of the autonomous driving industry has led to a significant accumulation of autonomous driving data. Consequently, there comes a growing demand for retrieving data to provide specialized optimization. However, directly applying previous image retrieval methods faces several challenges, such as the lack of global feature representation and inadequate text retrieval ability for complex driving scenes. To address these issues, firstly, we propose the BEV-TSR framework which leverages descriptive text as an input to retrieve corresponding scenes in the Bird’s Eye View (BEV) space. Then to facilitate complex scene retrieval with extensive text descriptions, we employ a large language model (LLM) to extract the semantic features of the text inputs and incorporate knowledge graph embeddings to enhance the semantic richness of the language embedding. To achieve feature alignment between the BEV feature and language embedding, we propose Shared Cross-modal Embedding with a set of shared learnable embeddings to bridge the gap between these two modalities, and employ a caption generation task to further enhance the alignment. Furthermore, there lack of well-formed retrieval datasets for effective evaluation. To this end, we establish a multi-level retrieval dataset, nuScenes-Retrieval, based on the widely adopted nuScenes dataset. Experimental results on the multi-level nuScenes-Retrieval show that BEV-TSR achieves state-of-the-art performance, e.g., 85.78% and 87.66% top-1 accuracy on scene-to-test and text-to-scene retrieval respectively.
Dafeng Wei, Zhengyu Jia, Changwei Cai, Chengkai Hou, Peng Jia 0007, Kun Zhan, Jingchen Fan, Yixing Zhao, Xiaodan Liang, Xianpeng Lang
AAAI8
2025 BrainGuard: Privacy-Preserving Multisubject Image Reconstructions from Brain Activities
abstract
Reconstructing perceived images from human brain activity forms a crucial link between human and machine learning through Brain-Computer Interfaces. Early methods primarily focused on training separate models for each individual to account for individual variability in brain activity, overlooking valuable cross-subject commonalities. Recent advancements have explored multisubject methods, but these approaches face significant challenges, particularly in data privacy and effectively managing individual variability. To overcome these challenges, we introduce BrainGuard, a privacy-preserving collaborative training framework designed to enhance image reconstruction from multisubject fMRI data while safeguarding individual privacy. BrainGuard employs a collaborative global-local architecture where personalized models are trained on each subject's data and operate in conjunction with a shared commonality model that captures and leverages cross-subject patterns. This architecture eliminates the need to aggregate fMRI data across subjects, thereby ensuring privacy preservation. To tackle the complexity of fMRI data, BrainGuard integrates a hybrid synchronization strategy, enabling individual models to dynamically incorporate parameters from the global model. By establishing a secure and collaborative training environment, BrainGuard not only protects sensitive brain activity data but also improves the accuracy of image reconstructions. Extensive experiments demonstrate that BrainGuard sets a new benchmark in both high-level and low-level metrics, advancing the state-of-the-art in brain decoding through its innovative design.
Zhibo Tian, Ruijie Quan, Fan Ma, Kun Zhan, Yi Yang 0001
AAAI4
2025 Hierarchical Consensus Network for Multiview Feature Learning
abstract
Multiview feature learning aims to learn discriminative features by integrating the distinct information in each view. However, most existing methods still face significant challenges in learning view-consistency features, which are crucial for effective multiview learning. Motivated by the theories of CCA and contrastive learning in multiview feature learning, we propose the hierarchical consensus network (HCN) in this paper. The HCN derives three consensus indices for capturing the hierarchical consensus across views, which are classifying consensus, coding consensus, and global consensus, respectively. Specifically, classifying consensus reinforces class-level correspondence between views from a CCA perspective, while coding consensus closely resembles contrastive learning and reflects contrastive comparison of individual instances. Global consensus aims to extract consensus information from two perspectives simultaneously. By enforcing the hierarchical consensus, the information within each view is better integrated to obtain more comprehensive and discriminative features. The extensive experimental results obtained on four multiview datasets demonstrate that the proposed method significantly outperforms several state-of-the-art methods.
Chengwei Xia, Chaoxi Niu, Kun Zhan
AAAI3
2025 ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration
abstract
Closed-loop simulation is crucial for end-to-end autonomous driving. Existing sensor simulation methods (e.g., NeRF and 3DGS) reconstruct driving scenes based on conditions that closely mirror training data distributions. However, these methods struggle with rendering novel trajectory, such as lane changes. Recent works have demonstrated that integrating world model knowledge alleviates these issues. Despite their efficiency, these approaches still encounter difficulties in the accurate representation of more complex maneuvers, with multi-lane shifts being a notable example. Therefore, we introduce ReconDreamer, which enhances driving scene reconstruction through incremental integration of world model knowledge. Specifically, DriveRestorer is proposed to mitigate artifacts via online restoration. This is complemented by a progressive data update strategy designed to ensure high-quality rendering for more complex maneuvers. To the best of our knowledge, ReconDreamer is the first method to effectively render in large maneuvers. Experimental results demonstrate that ReconDreamer outperforms Street Gaussians in the NTA-IoU, NTL-IoU, and FID, with relative improvements by 24.87%, 6.72%, and 29.97%. Furthermore, ReconDreamer surpasses DriveDreamer4D with PVG during large maneuver rendering, as verified by a relative improvement of 195.87% in the NTA-IoU metric and a user study.
Chaojun Ni, Guosheng Zhao, Wenkang Qin, Guan Huang 0003, Yuyin Chen, Xueyang Zhang, Yifei Zhan, Kun Zhan, Peng Jia 0007, Xianpeng Lang, Xingang Wang 0003, Wenjun Mei
CVPR12
2025 StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
abstract
This paper aims to tackle the problem of photorealistic view synthesis from vehicle sensor data. Recent advancements in neural scene representation have achieved notable success in rendering high-quality autonomous driving scenes, but the performance significantly degrades as the viewpoint deviates from the training trajectory. To mitigate this problem, we introduce StreetCrafter, a novel controllable video diffusion model that utilizes LiDAR point cloud renderings as pixel-level conditions, which fully exploits the generative prior for novel view synthesis, while preserving precise camera control. Moreover, the utilization of pixel-level LiDAR conditions allows us to make accurate pixel-level edits to target scenes. In addition, the generative prior of StreetCrafter can be effectively incorporated into dynamic scene representations to achieve real-time rendering. Experiments on Waymo Open Dataset and PandaSet demonstrate that our model enables flexible control over viewpoint changes, enlarging the view synthesis regions for satisfying rendering, which outperforms existing methods. The code is available at https://zju3dv.github.io/streetcrafter.
Yunzhi Yan, Zhen Xu 0008, Haotong Lin, Haian Jin, Kun Zhan, Xianpeng Lang, Hujun Bao, Xiaowei Zhou 0001, Sida Peng
CVPR7
2025 DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation
abstract
Autonomous driving evaluation requires simulation environments that closely replicate actual road conditions, including real-world sensory data and responsive feedback loops. However, many existing simulations need to predict waypoints along fixed routes on public datasets or synthetic photorealistic data, i.e., open-loop simulation usually lacks the ability to assess dynamic decision-making. While the recent efforts of closed-loop simulation offer feedback-driven environments, they cannot process visual sensor inputs or produce outputs that differ from real-world data. To address these challenges, we propose DrivingSphere, a realistic and closed-loop simulation framework. Its core idea is to build 4D world representation and generate real-life and controllable driving scenarios. In specific, our framework includes a Dynamic Environment Composition module that constructs a detailed 4D driving world with a format of occupancy equipping with static backgrounds and dynamic objects, and a Visual Scene Synthesis module that transforms this data into high-fidelity, multi-view video outputs, ensuring spatial and temporal consistency. By providing a dynamic and realistic simulation environment, DrivingSphere enables comprehensive testing and validation of autonomous driving algorithms, ultimately advancing the development of more reliable autonomous cars. The benchmark will be publicly released.
Tianyi Yan, Dongming Wu 0005, Wencheng Han, Junpeng Jiang, Kun Zhan, Cheng-Zhong Xu 0001, Jianbing Shen
CVPR6
2025 3DRealCar: An In-the-Wild RGB-D Car Dataset with 360-Degree Views
abstract
3D cars are commonly used in self-driving systems, virtual/augmented reality, and games. However, existing 3D car datasets are either synthetic or low-quality, limiting their applications in practical scenarios and presenting a significant gap toward high-quality real-world 3D car datasets. In this paper, we propose the first large-scale 3D real car dataset, termed 3DRealCar, offering three distinctive features. (1) \textbf{High-Volume}: 2,500 cars are meticulously scanned by smartphones, obtaining car images and point clouds with real-world dimensions; (2) \textbf{High-Quality}: Each car is captured in an average of 200 dense, high-resolution 360-degree RGB-D views, enabling high-fidelity 3D reconstruction; (3) \textbf{High-Diversity}: The dataset contains various cars from over 100 brands, collected under three distinct lighting conditions, including reflective, standard, and dark. Additionally, we offer detailed car parsing maps for each instance to promote research in car parsing tasks. Moreover, we remove background point clouds and standardize the car orientation to a unified axis for the reconstruction only on cars and controllable rendering without background. We benchmark 3D reconstruction results with state-of-the-art methods across different lighting conditions in 3DRealCar. Extensive experiments demonstrate that the standard lighting condition part of 3DRealCar can be used to produce a large number of high-quality 3D cars, improving various 2D and 3D tasks related to cars. Notably, our dataset brings insight into the fact that recent 3D reconstruction methods face challenges in reconstructing high-quality 3D cars under reflective and dark lighting conditions. \textcolor{red}{\href{https://xiaobiaodu.github.io/3drealcar/}{Our dataset is here.}}
Xiaobiao Du, Zhuojie Wu, Hongwei Sheng, Jiaying Ying, Ming Lu 0002, Tianqing Zhu, Kun Zhan, Xin Yu 0002
ICCV10
2025 Hierarchy UGP: Hierarchy Unified Gaussian Primitive for Large-Scale Dynamic Scene Reconstruction
Hongyang Sun 0005, Qinglin Yang, Zhen Xu 0008, Chen Liu 0028, Kun Zhan, Hujun Bao, Xiaowei Zhou 0001, Sida Peng
ICCV7
2025 RoboPearls: Editable Video Simulation for Robot Manipulation
Tang Tao, Likui Zhang, Youpeng Wen, Kaidong Zhang, Jiawang Bian, Tianyi Yan, Kun Zhan, Peng Jia 0007, Hefeng Wu, Xiaodan Liang
ICCV8
2025 HiNeuS: High-Fidelity Neural Surface Mitigating Low-Texture and Reflective Ambiguity
abstract
Neural surface reconstruction faces persistent challenges in reconciling geometric fidelity with photometric consistency under complex scene conditions. We present HiNeuS, a unified framework that holistically addresses three core limitations in existing approaches: multi-view radiance inconsistency, missing keypoints in textureless regions, and structural degradation from over-enforced Eikonal constraints during joint optimization. To resolve these issues through a unified pipeline, we introduce: 1) Differential visibility verification through SDF-guided ray tracing, resolving reflection ambiguities via continuous occlusion modeling; 2) Planar-conformal regularization via ray-aligned geometry patches that enforce local surface coherence while preserving sharp edges through adaptive appearance weighting; and 3) Physically-grounded Eikonal relaxation that dynamically modulates geometric constraints based on local radiance gradients, enabling detail preservation without sacrificing global regularity. Unlike prior methods that handle these aspects through sequential optimizations or isolated modules, our approach achieves cohesive integration where appearance-geometry constraints evolve synergistically throughout training. Comprehensive evaluations across synthetic and real-world datasets demonstrate state-of-the-art performance, including a 21.4% reduction in Chamfer distance over reflection-aware baselines and 2.32 dB PSNR improvement against neural rendering counterparts. Qualitative analyses reveal superior capability in recovering specular instruments, urban layouts with centimeter-scale infrastructure, and low-textured surfaces without local patch collapse. The method's generalizability is further validated through successful application to inverse rendering tasks, including material decomposition and view-consistent relighting.
Xueyang Zhang, Kun Zhan, Peng Jia 0007, Xianpeng Lang
ICCV3
2025 S2-Track: A Simple yet Strong Approach for End-to-End 3D Multi-Object Tracking
abstract
3D multiple object tracking (MOT) plays a crucial role in autonomous driving perception. Recent end-to-end query-based trackers simultaneously detect and track objects, which have shown promising potential for the 3D MOT task. However, existing methods are still in the early stages of development and lack systematic improvements, failing to track objects in certain complex scenarios, like occlusions and the small size of target object’s situations. In this paper, we first summarize the current end-to-end 3D MOT framework by decomposing it into three constituent parts: query initialization, query propagation, and query matching. Then we propose corresponding improvements, which lead to a strong yet simple tracker: S2-Track. Specifically, for query initialization, we present 2D-Prompted Query Initialization, which leverages predicted 2D object and depth information to prompt an initial estimate of the object’s 3D location. For query propagation, we introduce an Uncertainty-aware Probabilistic Decoder to capture the uncertainty of complex environment in object prediction with probabilistic attention. For query matching, we propose a Hierarchical Query Denoising strategy to enhance training robustness and convergence. As a result, our S2-Track achieves state-of-the-art performance on nuScenes benchmark, i.e., 66.3% AMOTA on test split, surpassing the previous best end-to-end solution by a significant margin of 8.9% AMOTA. We achieve 1st place on the nuScenes tracking task leaderboard.
Pengkun Hao, Kalok Ho, Shuo Gu, Zhihui Hao, Kun Zhan, Peng Jia 0007, Xianpeng Lang, Xiaodan Liang
ICML9
2025 Generalizing Motion Planners with Mixture of Experts for Autonomous Driving
abstract
Large real-world driving datasets have sparked significant research into various aspects of learning-based motion planners for autonomous driving. These include data augmentation, model architecture, reward design, training strategies, and planner pipelines. In this paper, we review and benchmark previous methods. Experiments show that many of these approaches have limited generalization abilities in planning performance due to overly complex designs or training paradigms. Experiments further reveal that as models are appropriately scaled, many designs become redundant. Therefore, we introduce StateTransformer-2 (STR2), a scalable, decoder-only motion planner. STR2uses a Vision Transformer (ViT) encoder and a mix-of-experts (MoE) causal transformer architecture. The MoE backbone addresses modality collapse and reward balancing by expert routing during training. Extensive experiments on the NuPlan dataset show that our method generalizes better than previous approaches across different test sets and closed-loop simulations. We evaluate its scalability on billions of real-world urban driving scenarios, demonstrating consistent accuracy improvements as both data and model size grow.
Qiao Sun 0001, Jiahao Zhan, Fan Nie, Leimeng Xu, Kun Zhan, Peng Jia 0007, Xianpeng Lang, Hang Zhao 0021
ICRA7
2025 PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth
abstract
Recent advancements in autonomous driving (AD) systems have highlighted the potential of world models in achieving robust and generalizable performance across both ordinary and challenging driving conditions. However, a key challenge remains: precise and flexible camera pose control, which is crucial for accurate viewpoint transformation and realistic simulation of scene dynamics. In this paper, we introduce PosePilot, a lightweight yet powerful framework that significantly enhances camera pose controllability in generative world models. Drawing inspiration from self-supervised depth estimation, PosePilot leverages structure-from-motion principles to establish a tight coupling between camera pose and video generation. Specifically, we incorporate self-supervised depth and pose readouts, allowing the model to infer depth and relative camera motion directly from video sequences. These outputs drive pose-aware frame warping, guided by a photometric warping loss that enforces geometric consistency across synthesized frames. To further refine camera pose estimation, we introduce a reverse warping step and a pose regression loss, improving viewpoint precision and adaptability. Extensive experiments on autonomous driving and general-domain video datasets demonstrate that PosePilot significantly enhances structural understanding and motion reasoning in both diffusion-based and auto-regressive world models. By steering camera pose with self-supervised depth, PosePilot sets a new benchmark for pose controllability, enabling physically consistent, reliable viewpoint synthesis in generative world models.
Bu Jin, Weize Li 0001, Baihan Yang, Zhenxin Zhu, Junpeng Jiang, Huan-ang Gao, Kun Zhan, Hengtong Hu, Xueyang Zhang, Peng Jia 0007, Hao Zhao 0002
IROS8
2025 OmniGen: Unified Multimodal Sensor Generation for Autonomous Driving
abstract
Autonomous driving has seen remarkable advancements, largely driven by extensive real-world data collection. However, acquiring diverse and corner-case data remains costly and inefficient. Generative models have emerged as a promising solution by synthesizing realistic sensor data. However, existing approaches primarily focus on single-modality generation, leading to inefficiencies and misalignment in multimodal sensor data. To address these challenges, we propose OminiGen, which generates aligned multimodal sensor data in a unified framework. Our approach leverages a shared Bird's Eye View (BEV) space to unify multimodal features and designs a novel generalizable multimodal reconstruction method, UAE, to jointly decode LiDAR and multi-view camera data. UAE achieves multimodal sensor decoding through volume rendering, enabling accurate and flexible reconstruction. Furthermore, we incorporate a Diffusion Transformer (DiT) with a ControlNet branch to enable controllable multimodal sensor generation. Our comprehensive experiments demonstrate that OminiGen achieves desired performances in unified multimodal sensor data generation with multimodal consistency and flexible sensor adjustments.
Enhui Ma, Tianyi Yan, Xueyang Zhang, Kun Zhan, Peng Jia 0007, Xianpeng Lang, Jiawang Bian, Kaicheng Yu, Xiaodan Liang
ACM Multimedia7
2025 Frequency Regulation for Exposure Bias Mitigation in Diffusion Models
abstract
Diffusion models exhibit impressive generative capabilities but are significantly impacted by exposure bias. In this paper, we make a key observation: the energy of predicted noisy samples in the reverse process continuously declines compared to perturbed samples in the forward process. Building on this, we identify two important findings: 1) The reduction in energy follows distinct patterns in the low-frequency and high-frequency subbands; 2) The subband energy of reverse-process reconstructed samples is consistently lower than that of forward-process ones, and both are lower than the original data samples. Based on the first finding, we introduce a dynamic frequency regulation mechanism utilizing wavelet transforms, which separately adjusts the low- and high-frequency subbands. Leveraging the second insight, we derive the rigorous mathematical form of exposure bias. It is worth noting that, our method is training-free and plug-and-play, significantly improving the generative quality of various diffusion models and frameworks with negligible computational cost. The source code is available at https://github.com/kunzhan/wpp.
Kun Zhan
ACM Multimedia2
2025 RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation
abstract
Synthetic data is crucial for advancing autonomous driving (AD) systems, yet current state-of-the-art video generation models, despite their visual realism, suffer from subtle geometric distortions that limit their utility for downstream perception tasks. We identify and quantify this critical issue, demonstrating a significant performance gap in 3D object detection when using synthetic versus real data. To address this, we introduce Reinforcement Learning with Geometric Feedback (RLGF), RLGF uniquely refines video diffusion models by incorporating rewards from specialized latent-space AD perception models. Its core components include an efficient Latent-Space Windowing Optimization technique for targeted feedback during diffusion, and a Hierarchical Geometric Reward (HGR) system providing multi-level rewards for point-line-plane alignment, and scene occupancy coherence. To quantify these distortions, we propose GeoScores. Applied to models like DiVE on nuScenes, RLGF substantially reduces geometric errors (e.g., VP error by 21\%, Depth error by 57\%) and dramatically improves 3D object detection mAP by 12.7\%, narrowing the gap to real-data performance. RLGF offers a plug-and-play solution for generating geometrically sound and reliable synthetic videos for AD development.
Tianyi Yan, Wencheng Han, Xueyang Zhang, Kun Zhan, Cheng-Zhong Xu 0001, Jianbing Shen
NeurIPS5
2025 Domain generalization plant leaf disease recognition: Toward from laboratory to field
Kun Zhan, Yingqiong Peng, Muxin Liao
Eng. Appl. Artif. Intell.1
2024 TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
Bu Jin, Yupeng Zheng, Pengfei Li 0007, Weize Li 0001, Yuhang Zheng 0004, Sujie Hu, Zhijie Yan, Kun Zhan, Peng Jia 0007, Xiaoxiao Long, Hao Zhao 0002
ECCV (18)11
2024 Street Gaussians: Modeling Dynamic Urban Scenes with Gaussian Splatting
Yunzhi Yan, Haotong Lin, Chenxu Zhou, Weijie Wang 0014, Kun Zhan, Xianpeng Lang, Xiaowei Zhou 0001, Sida Peng
ECCV (73)6
2024 InfoMatch: Entropy Neural Estimation for Semi-Supervised Image Classification
Zhibo Tian, Chengwei Xia, Kun Zhan
IJCAI4
2024 Identity Semantic Correspondence for Cloth-Changing Person Re-Identification
abstract
Cloth-changing Person Re-Identification (CC-ReID) aims at retrieving the same person who might change clothes across different locations. Although remarkable progress has been achieved in recent studies, most of the recent methods still lack sufficient emphasis on identity-related regions. To address these issues, we propose a novel Identity Semantic Correspondence framework (ISC) to fully utilize human semantic information, which includes dual-stream identity semantic correspondance networks, i.e., a Clothing-invariant Identity (CI) stream and a Fine-grained Identity Semantic Highlighting (FISH) stream. The CI mitigates the interference of human dressing and enhances clothing-invariant identity information by erasing clothing information. And, the FISH exploits fine-grained identity information by highlighting contribution of different body parts to identity with a part-aware weighting module. Additionally, an identity semantic consistency module is further proposed to extract the most representative and discriminative semantic features for each identity. Besides, we employ a mutual loss to transfer identity-related knowledge between components, which enables the original appearance module to be deployed independently during the inference stage. Extensive experiments on two CC-ReID benchmarks, including PRCC and VC-Clothes, are conducted to demonstrate the effectiveness of the proposed ISC method.
Yongtang Bao, Kun Zhan, Peng Zhang 0057
IJCNN4
2024 Non-local sparse attention based swin transformer V2 for image super-resolution
Ningning Lv, Yufei Xie, Kun Zhan, Fuxiang Lu
Signal Process.4
2024 Benchmarking deep models on retinal fundus disease diagnosis and a large-scale dataset
abstract
Retinal fundus imaging contributes to monitoring the vision of patients by providing views of the interior surface of the eyes. Machine learning models greatly aided ophthalmologists in detecting retinal disorders from color fundus images. Hence, the quality of the data is pivotal for enhancing diagnosis algorithms, which ultimately benefits vision care and maintenance. To facilitate further research in this domain, we introduce the Eye Disease Diagnosis and Fundus Synthesis (EDDFS) dataset, comprising 28,877 fundus images. These include 15,000 healthy samples and a diverse range of images depicting various disorders such as diabetic retinopathy, age-related macular degeneration, glaucoma, pathological myopia, hypertension retinopathy, retinal vein occlusion, and Laser photocoagulation. In addition to providing the dataset, we propose a Transformer-joint convolution network for automated eye disease screening. Firstly, a co-attention structure is integrated to capture long-range attention information along with local features. Secondly, a cross-stage feature fusion module is designed to extract multi-level and disease-related information. By leveraging the dataset and our proposed network, we establish benchmarks for disease screening and grading tasks. Our experimental results underscore the network’s proficiency in both multi-label and single-label disease diagnosis, while also showcasing the dataset’s capability in supporting fundus synthesis. (The dataset and code will be available on https://github.com/xia-xx-cv/EDDFS_dataset).
Xue Xia 0005, Guobei Xiao, Kun Zhan, Jinhua Yan, Yuming Fang 0001, Guofu Huang
Signal Process. Image Commun.4
2023 Curriculum Knowledge Switching for Pancreas Segmentation
abstract
Pancreas segmentation is challenging due to the small proportion and highly changeable anatomical structure. It motivates us to propose a novel segmentation framework, namely Curriculum Knowledge Switching (CKS) framework, which decomposes detecting pancreas into three phases with different difficulty extent: straightforward, difficult, and challenging. The framework switches from straightforward to challenging phases and thereby gradually learns to detect pancreas. In addition, we adopt the momentum update parameter updating mechanism during switching, ensuring the loss converges gradually when the input dataset changes. Experimental results show that different neural network backbones with the CKS framework achieved state-of-the-art performance on the NIH dataset as measured by the DSC metric. The code is available at https://github.com/kunzhan/CKS_Pancreas
Yumou Tang, Kun Zhan, Zhibo Tian, Saisai Wang, Xueming Wen
ICIP2
2023 CLIP-FG:Selecting Discriminative Image Patches by Contrastive Language-Image Pre-Training for Fine-Grained Image Classification
abstract
Fine-grained Visual Classification (FGVC), which aims to identify objects from subcategories, presents great challenges for classification due to large intra-class differences and subtle inter-class differences. To address these issues of FGVC, this paper proposes a patch selection model referenced from CLIP for Fine Grained Visual Classification, namely CLIP-FG. Specifically, unlike the previous CLIP, which focused only on the level of text and image, we calculate the similarity between labels and image patches. Top k image patches are selected and their indexes fed into the Vision Transformer to select discriminative areas to improve the performance of fine grained image classification. Quantitative evaluations show CLIP-FG’s competitive performance against mainstream methods.
Ningning Lv, Yufei Xie, Fuxiang Lu, Kun Zhan
ICIP5
2023 Self-Contrastive Graph Diffusion Network
abstract
Augmentation techniques and sampling strategies are crucial in contrastive learning, but in most existing works, augmentation techniques require careful design, and their sampling strategies can only capture a small amount of intrinsic supervision information. Additionally, the existing methods require complex designs to obtain two different representations of the data. To overcome these limitations, we propose a novel framework called the Self-Contrastive Graph Diffusion Network (SCGDN). Our framework consists of two main components: the Attentional Module (AttM) and the Diffusion Module (DiFM). AttM aggregates higher-order structure and feature information to get an excellent embedding, while DiFM balances the state of each node in the graph through Laplacian diffusion learning and allows the cooperative evolution of adjacency and feature information in the graph. Unlike existing methodologies, SCGDN is an augmentation-free approach that avoids "sampling bias" and semantic drift, without the need for pre-training. We conduct a high-quality sampling of samples based on structure and feature information. If two nodes are neighbors, they are considered positive samples of each other. If two disconnected nodes are also unrelated on kNN graph, they are considered negative samples for each other. The contrastive objective reasonably uses our proposed sampling strategies, and the redundancy reduction term minimizes redundant information in the embedding and can well retain more discriminative information. In this novel framework, the graph self-contrastive learning paradigm gives expression to a powerful force. The results manifest that SCGDN can consistently generate out performance over both the contrastive methods and the classical methods. The source code is available at https://github.com/kunzhan/SCGDN.
Yixuan Ma, Kun Zhan
ACM Multimedia2
2023 Entropy Neural Estimation for Graph Contrastive Learning
abstract
Contrastive learning on graphs aims at extracting distinguishable high-level representations of nodes. We theoretically illustrate that the entropy of a dataset is approximated by maximizing the lower bound of the mutual information across different views of a graph, i.e., entropy is estimated by a neural network. Based on this finding, we propose a simple yet effective subset sampling strategy to contrast pairwise representations between views of a dataset. In particular, we randomly sample nodes and edges from a given graph to build the input subset for a view. Two views are fed into a parameter-shared Siamese network to extract the high-dimensional embeddings and estimate the information entropy of the entire graph. For the learning process, we propose to optimize the network using two objectives, simultaneously. Concretely, the input of the contrastive loss consists of positive and negative pairs. Our selection strategy of pairs is different from previous works and we present a novel strategy to enhance the representation ability by selecting nodes based on cross-view similarities. We enrich the diversity of the positive and negative pairs by selecting highly similar samples and totally different data with the guidance of cross-view similarity scores, respectively. We also introduce a cross-view consistency constraint on the representations generated from the different views. We conduct experiments on seven graph benchmarks, and the proposed approach achieves competitive performance compared to the current state-of-the-art methods. The source code is available at https://github.com/kunzhan/M-ILBO.
Yixuan Ma, Peng Zhang 0057, Kun Zhan
ACM Multimedia4
2023 Improving Semi-Supervised Semantic Segmentation with Dual-Level Siamese Structure Network
abstract
Semi-supervised semantic segmentation (SSS) is an important task that utilizes both labeled and unlabeled data to reduce expenses on labeling training examples. However, the effectiveness of SSS algorithms is limited by the difficulty of fully exploiting the potential of unlabeled data. To address this, we propose a dual-level Siamese structure network (DSSN) for pixel-wise contrastive learning. By aligning positive pairs with a pixel-wise contrastive loss using strong augmented views in both low-level image space and high-level feature space, the proposed DSSN is designed to maximize the utilization of available unlabeled data. Additionally, we introduce a novel class-aware pseudo-label selection strategy for weak-to-strong supervision, which addresses the limitations of most existing methods that do not perform selection or apply a predefined threshold for all classes. Specifically, our strategy selects the top high-confidence prediction of the weak view for each class to generate pseudo labels that supervise the strong augmented views. This strategy is capable of taking into account the class imbalance and improving the performance of long-tailed classes. Our proposed method achieves state-of-the-art results on two datasets, PASCAL VOC 2012 and Cityscapes, outperforming other SSS algorithms by a significant margin. The source code is available at https://github.com/kunzhan/DSSN.
Zhibo Tian, Peng Zhang 0057, Kun Zhan
ACM Multimedia4
2023 View-Consistent Heterogeneous Network on Graphs With Few Labeled Nodes
abstract
Performing transductive learning on graphs with very few labeled data, that is, two or three samples for each category, is challenging due to the lack of supervision. In the existing work, self-supervised learning via a single view model is widely adopted to address the problem. However, recent observation shows multiview representations of an object share the same semantic information in high-level feature space. For each sample, we generate heterogeneous representations and use view-consistency loss to make their representations consistent with each other. Multiview representation also inspires to supervise the pseudolabels generation by the aid of mutual supervision between views. In this article, we thus propose a view-consistent heterogeneous network (VCHN) to learn better representations by aligning view-agnostic semantics. Specifically, VCHN is constructed by constraining the predictions between two views so that the view pairs can supervise each other. To make the best use of cross-view information, we further propose a novel training strategy to generate more reliable pseudolabels, which thus enhances predictions of the VCHN. Extensive experimental results on three benchmark datasets demonstrate that our method achieves superior performance over state-of-the-art methods under very low label rates.
Zhuolin Liao, Wei Su 0008, Kun Zhan
IEEE Trans. Cybern.4
2022 Stationary Diffusion State Neural Estimation for Multiview Clustering
abstract
Although many graph-based clustering methods attempt to model the stationary diffusion state in their objectives, their performance limits to using a predefined graph. We argue that the estimation of the stationary diffusion state can be achieved by gradient descent over neural networks. We specifically design the Stationary Diffusion State Neural Estimation (SDSNE) to exploit multiview structural graph information for co-supervised learning. We explore how to design a graph neural network specially for unsupervised multiview learning and integrate multiple graphs into a unified consensus graph by a shared self-attentional module. The view-shared self-attentional module utilizes the graph structure to learn a view-consistent global graph. Meanwhile, instead of using auto-encoder in most unsupervised learning graph neural networks, SDSNE uses a co-supervised strategy with structure information to supervise the model learning. The co-supervised strategy as the loss function guides SDSNE in achieving the stationary state. With the help of the loss and the self-attentional module, we learn to obtain a graph in which nodes in each connected component fully connect by the same weight. Experiments on several multiview datasets demonstrate effectiveness of SDSNE in terms of six clustering evaluation metrics.
Chenghua Liu, Zhuolin Liao, Yixuan Ma, Kun Zhan
AAAI4
2022 Eye Disease Diagnosis and Fundus Synthesis: A Large-Scale Dataset and Benchmark
abstract
As one of the most common imaging modalities, retinal fundus imaging offers images of interior surface of eyes for initial examination of disorders. Data-driven machine learning methods, especially deep learning models in recent years, provide automatic ophthalmological disease diagnosis techniques from color fundus images. Data with high quality, diversity and balanced distribution supports deep model-based eye disease diagnosis. However, many existing datasets focus on a specific kind of eye disease, and some suffer from label noise or quality degeneration, which hinders automatic screening algorithms from dealing with multiple eye diseases. To solve this, we propose a high-quality dataset containing 28877 color fundus images for deep learning-based diagnosis. Except for 15000 healthy samples, the dataset consists of 8 eye disorders including diabetic retinopathy, agerelated macular degeneration, glaucoma, pathological myopia, hypertension, retinal vein occlusion, LASIK spot and others. Based on this, we propose a co-attention network for disease diagnosis, establish benchmark on screening and grading tasks, and demonstrate that the proposed dataset supports generative adversarial network-based image synthesis. The dataset will be made publicly available.
Xue Xia 0005, Kun Zhan, Guobei Xiao, Jinhua Yan, Zhuxiang Huang, Guofu Huang, Yuming Fang 0001
MMSP2
2022 Texture-aware Network for Smoke Density Estimation
abstract
Smoke density estimation, also termed as soft segmentation, was developed from pixel-wise smoke (hard) segmen-tation and it aims at providing transparency and segmentation confidence for each pixel. The key difference between them lies in that segmentation focuses on classifying pixels into smoke and non-smoke ones, while density estimation obtains inner transparency of smoke component rather than treat all smoke pixels as an equal value. Based on this, we propose a texture-aware network being able to capture inner transparency of smoke components rather than merely focus on general smoke distribution for pixel-wise smoke density estimation. Besides, we adapt the Squeeze-and-Excitation (SE) layer for smoke feature extraction by involving max values for robustness. In order to represent inhomogeneous smoke pixels, we proposed a simple yet efficient attention-based texture-aware module that involves both gradient and semantic information. Experimental results show that our method outperforms others in both single image density estimation or segmentation and video smoke detection.
Xue Xia 0005, Kun Zhan, Yajing Peng, Yuming Fang 0001
VCIP2
2021 Mutual teaching for graph convolutional networks
Kun Zhan, Chaoxi Niu
Future Gener. Comput. Syst.1
2019 Zero-shot event detection via event-adaptive concept relevance mining
Zhihui Li 0001, Lina Yao 0001, Xiaojun Chang, Kun Zhan, Jiande Sun 0001, Huaxiang Zhang 0001
Pattern Recognit.4
2019 Adaptive Structure Discovery for Multimedia Analysis Using Multiple Features
abstract
Multifeature learning has been a fundamental research problem in multimedia analysis. Most existing multifeature learning methods exploit graph, which must be computed beforehand, as input to uncover data distribution. These methods have two major problems confronted. First, graph construction requires calculating similarity based on nearby data pairs by a fixed function, e.g., the RBF kernel, but the intrinsic correlation among different data pairs varies constantly. Therefore, feature learning based on such predefined graphs may degrade, especially when there is dramatic correlation variation between nearby data pairs. Second, in most existing algorithms, each single-feature graph is computed independently and then combine them for learning, which ignores the correlation between multiple features. In this paper, a new unsupervised multifeature learning method is proposed to make the best utilization of the correlation among different features by jointly optimizing data correlation from multiple features in an adaptive way. As opposed to computing the affinity weight of data pairs by a fixed function, the weight of affinity graph is learned by a well-designed optimization problem. Additionally, the affinity graph of data pairs from different features is optimized in a global level to better leverage the correlation among different channels. In this way, the adaptive approach correlates the features of all features for a better learning process. Experimental results on real-world datasets demonstrate that our approach outperforms the state-of-the-art algorithms on leveraging multiple features for multimedia analysis.
Kun Zhan, Xiaojun Chang, Junpeng Guan, Ling Chen 0006, Zhigang Ma, Yi Yang 0001
IEEE Trans. Cybern.1
2019 Multiview Consensus Graph Clustering
abstract
A graph is usually formed to reveal the relationship between data points and graph structure is encoded by the affinity matrix. Most graph-based multiview clustering methods use predefined affinity matrices and the clustering performance highly depends on the quality of graph. We learn a consensus graph with minimizing disagreement between different views and constraining the rank of the Laplacian matrix. Since diverse views admit the same underlying cluster structure across multiple views, we use a new disagreement cost function for regularizing graphs from different views toward a common consensus. Simultaneously, we impose a rank constraint on the Laplacian matrix to learn the consensus graph with exactly connected components where is the number of clusters, which is different from using fixed affinity matrices in most existing graph-based methods. With the learned consensus graph, we can directly obtain the cluster labels without performing any post-processing, such as -means clustering algorithm in spectral clustering-based methods. A multiview consensus clustering method is proposed to learn such a graph. An efficient iterative updating algorithm is derived to optimize the proposed challenging optimization problem. Experiments on several benchmark datasets have demonstrated the effectiveness of the proposed method in terms of seven metrics.
Kun Zhan, Feiping Nie 0001, Jing Wang 0023, Yi Yang 0001
IEEE Trans. Image Process.1
2019 Graph Structure Fusion for Multiview Clustering
abstract
Most existing multiview clustering methods take graphs, which are usually predefined independently in each view, as input to uncover data distribution. These methods ignore the correlation of graph structure among multiple views and clustering results highly depend on the quality of predefined affinity graphs. We address the problem of multiview clustering by seamlessly integrating graph structures of different views to fully exploit the geometric property of underlying data structure. The proposed method is based on the assumption that the intrinsic underlying graph structure would assign corresponding connected component in each graph to the same cluster. Different graphs from multiple views are integrated by using the Hadamard product since different views usually together admit the same underlying structure across multiple views. Specifically, these graphs are integrated into a global one and the structure of the global graph is adaptively tuned by a well-designed objective function so that the number of components of the graph is exactly equal to the number of clusters. It is worth noting that we directly obtain cluster indicators from the graph itself without performing further graph-cut or k-means clustering algorithms. Experiments show the proposed method obtains better clustering performance than the state-of-the-art methods.
Kun Zhan, Chaoxi Niu, Changlu Chen, Feiping Nie 0001, Changqing Zhang 0002, Yi Yang 0001
IEEE Trans. Knowl. Data Eng.1
2018 Adaptive Structure Concept Factorization for Multiview Clustering
abstract
Most existing multiview clustering methods require that graph matrices in different views are computed beforehand and that each graph is obtained independently. However, this requirement ignores the correlation between multiple views. In this letter, we tackle the problem of multiview clustering by jointly optimizing the graph matrix to make full use of the data correlation between views. With the interview correlation, a concept factorization-based multiview clustering method is developed for data integration, and the adaptive method correlates the affinity weights of all views. This method differs from nonnegative matrix factorization-based clustering methods in that it can be applicable to data sets containing negative values. Experiments are conducted to demonstrate the effectiveness of the proposed method in comparison with state-of-the-art approaches in terms of accuracy, normalized mutual information, and purity.
Kun Zhan, Jinhui Shi, Jing Wang 0023, Yuange Xie
Neural Comput.1
2018 Diverse Non-Negative Matrix Factorization for Multiview Data Representation
abstract
Non-negative matrix factorization (NMF), a method for finding parts-based representation of non-negative data, has shown remarkable competitiveness in data analysis. Given that real-world datasets are often comprised of multiple features or views which describe data from various perspectives, it is important to exploit diversity from multiple views for comprehensive and accurate data representations. Moreover, real-world datasets often come with high-dimensional features, which demands the efficiency of low-dimensional representation learning approaches. To address these needs, we propose a diverse NMF (DiNMF) approach. It enhances the diversity, reduces the redundancy among multiview representations with a novel defined diversity term and enables the learning process in linear execution time. We further propose a locality preserved DiNMF (LP-DiNMF) for more accurate learning, which ensures diversity from multiple views while preserving the local geometry structure of data in each view. Efficient iterative updating algorithms are derived for both DiNMF and LP-DiNMF, along with proofs of convergence. Experiments on synthetic and real-world datasets have demonstrated the efficiency and accuracy of the proposed methods against the state-of-the-art approaches, proving the advantages of incorporating the proposed diversity term into NMF.
Jing Wang 0023, Feng Tian 0006, Hongchuan Yu, Chang Hong Liu, Kun Zhan, Xiao Wang 0017
IEEE Trans. Cybern.5
2018 Graph Learning for Multiview Clustering
abstract
Most existing graph-based clustering methods need a predefined graph and their clustering performance highly depends on the quality of the graph. Aiming to improve the multiview clustering performance, a graph learning-based method is proposed to improve the quality of the graph. Initial graphs are learned from data points of different views, and the initial graphs are further optimized with a rank constraint on the Laplacian matrix. Then, these optimized graphs are integrated into a global graph with a well-designed optimization procedure. The global graph is learned by the optimization procedure with the same rank constraint on its Laplacian matrix. Because of the rank constraint, the cluster indicators are obtained directly by the global graph without performing any graph cut technique and the k-means clustering. Experiments are conducted on several benchmark datasets to verify the effectiveness and superiority of the proposed graph learning-based multiview clustering algorithm comparing to the state-of-the-art methods.
Kun Zhan, Changqing Zhang 0002, Junpeng Guan
IEEE Trans. Cybern.1
2018 Two-Stream Multirate Recurrent Neural Network for Video-Based Pedestrian Reidentification
abstract
Video-based pedestrian reidentification is an emerging task in video surveillance and is closely related to several real-world applications. Its goal is to match pedestrians across multiple nonoverlapping network cameras. Despite the recent effort, the performance of pedestrian reidentification needs further improvement. Hence, we propose a novel two-stream multirate recurrent neural network for video-based pedestrian reidentification with two inherent advantages: First, capturing the static spatial and temporal information; Second,Author: Figure II is not cited in the text. Please cite it at the appropriate place. dealing with motion speed variance. Given video sequences of pedestrians, we start with extracting spatial and motion features using two different deep neural networks. Then, we explore the feature correlation which results in a regularized fusion network integrating the two aforementioned networks. Considering that pedestrians, sometimes even the same pedestrian, move in different speeds across different camera views, we extend our approach by feeding the two networks into a multirate recurrent network to exploit the temporal correlations. Extensive experiments have been conducted on two real-world video-based pedestrian reidentification benchmarks: iLIDS-VID and PRID 2011 datasets. The experimental results confirm the efficacy of the proposed method. Our code will be released upon acceptance.
Zhihui Li 0001, De Cheng, Huaxiang Zhang 0001, Kun Zhan, Yi Yang 0001
IEEE Trans. Ind. Informatics5
2017 Linking synaptic computation for image enhancement
Kun Zhan, Jinhui Shi, Jicai Teng, Qiaoqiao Li, Mingying Wang, Fuxiang Lu
Neurocomputing1
2017 Graph-regularized concept factorization for multi-view document clustering
abstract
We propose a novel multi-view document clustering method with the graph-regularized concept factorization (MVCF). MVCF makes full use of multi-view features for more comprehensive understanding of the data and learns weights for each view adaptively. It also preserves the local geometrical structure of the manifolds for multi-view clustering. We have derived an efficient optimization algorithm to solve the objective function of MVCF and proven its convergence by utilizing the auxiliary function method. Experiments carried out on three benchmark datasets have demonstrated the effectiveness of MVCF in comparison to several state-of-the-art approaches in terms of accuracy, normalized mutual information and purity.
Kun Zhan, Jinhui Shi, Jing Wang 0023, Feng Tian 0006
J. Vis. Commun. Image Represent.1
2016 Unsupervised discriminative hashing
Kun Zhan, Junpeng Guan, Yi Yang 0001
J. Vis. Commun. Image Represent.1
2016 Feature-Linking Model for Image Enhancement
abstract
Inspired by gamma-band oscillations and other neurobiological discoveries, neural networks research shifts the emphasis toward temporal coding, which uses explicit times at which spikes occur as an essential dimension in neural representations. We present a feature-linking model (FLM) that uses the timing of spikes to encode information. The first spiking time of FLM is applied to image enhancement, and the processing mechanisms are consistent with the human visual system. The enhancement algorithm achieves boosting the details while preserving the information of the input image. Experiments are conducted to demonstrate the effectiveness of the proposed method. Results show that the proposed method is effective.
Kun Zhan, Jicai Teng, Jinhui Shi, Qiaoqiao Li, Mingying Wang
Neural Comput.1
2015 Image segmentation using fast linking SCM
abstract
Spiking cortical model (SCM) is applied to image segmentation. A natural image is processed to produce a series of spike images by SCM, and the segmented result is obtained by the integration of the series of spike images. An appropriate maximum iterative times is selected to achieve an optimal threshold of SCM. In each iteration, neurons that produced spikes correspond to pixels with an intensity of the input natural image approximately. SCM synchronizes the output spikes via the fast linking synaptic modulation, which makes objects in the image as homogeneous as possible. Experimental results show that the output image not only separates objects and background well, but also pixels in each object are homogeneous. The proposed method performs well over other methods and the quantitative metrics are consistent with the visual performance.
Kun Zhan, Jinhui Shi, Qiaoqiao Li, Jicai Teng, Mingying Wang
IJCNN1
2014 Spiking cortical model for multifocus image fusion
Nianyi Wang, Yide Ma, Kun Zhan
Neurocomputing3
2014 Joint tracking and classification based on aerodynamic model and radar cross section
Hong Jiang 0004, Kun Zhan
Pattern Recognit.3
2010 Pulse-coupled neural networks and one-class support vector machines for geometry invariant texture retrieval
Yide Ma, Kun Zhan, Yongqing Wu
Image Vis. Comput.3
2009 New Spiking Cortical Model for Invariant Texture Retrieval and Image Processing
abstract
Based on the studies of existing local-connected neural network models, in this brief, we present a new spiking cortical neural networks model and find that time matrix of the model can be recognized as a human subjective sense of stimulus intensity. The series of output pulse images of a proposed model represents the segment, edge, and texture features of the original image, and can be calculated based on several efficient measures and forms a sequence as the feature of the original image. We characterize texture images by the sequence for an invariant texture retrieval. The experimental results show that the retrieval scheme is effective in extracting the rotation and scale invariant features. The new model can also obtain good results when it is used in other image processing applications.
Kun Zhan, Hongjuan Zhang, Yide Ma
IEEE Trans. Neural Networks1