Haozhe Chi

dblp:330/9145 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0007-5695-0557ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 60% Video understanding and tracking · 17% Vision and language · 17%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
normalizing flow
0.912025
Enhancing Consistency of Flow-Based Image Editing through Kalman Control · NeurIPS 2025
Visual content generation and editing
image editing
0.912025
Enhancing Consistency of Flow-Based Image Editing through Kalman Control · NeurIPS 2025
Visual content generation and editing › image editing › text-guided image editing
instruction-based image editing
0.912025
Enhancing Consistency of Flow-Based Image Editing through Kalman Control · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.812024
RectifID: Personalizing Rectified Flow with Anchored Classifier Guidance · NeurIPS 2024
Computer vision › Video understanding and tracking
long video understanding
0.812024
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding · CVPR 2024
Machine learning › Generative modeling › diffusion model
personalized image generation
0.812024
RectifID: Personalizing Rectified Flow with Anchored Classifier Guidance · NeurIPS 2024
Computer vision › Vision and language
video-language model
0.812024
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding · CVPR 2024
Natural language and speech › Question answering and dialogue systems
long-context memory
0.212024
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding · CVPR 2024
Machine learning › Generative modeling › diffusion model
rectified flow
0.212024
RectifID: Personalizing Rectified Flow with Anchored Classifier Guidance · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

vector field velocity estimation · 1.7kalman filter · 1.7control theory · 1.7transformer · 0.8memory mechanism · 0.8fixed-point solution · 0.8classifier guidance · 0.8anchored reference flow · 0.8
YearPublicationVenuePosition
2025 Enhancing Consistency of Flow-Based Image Editing through Kalman Control
abstract
Flow-based generative models have gained popularity for image generation and editing. For instruction-based image editing, it is critical to ensure that modifications are confined to the targeted regions. Yet existing methods often fail to maintain consistency in non-targeted regions between the original / edited images. Our primary contribution is to identify the cause of this limitation as the error accumulation across individual editing steps and to address it by incorporating the historical editing trajectory. Specifically, we formulate image editing as a control problem and leverage the Kalman filter to integrate the historical editing trajectory. Our proposed algorithm, dubbed Kalman-Edit, reuses early-stage details from the historical trajectory to enhance the structural consistency of the editing results. To speed up editing, we introduce a shortcut technique based on approximate vector field velocity estimation. Extensive experiments on several datasets demonstrate its superior performance compared to previous state-of-the-art methods.
Haozhe Chi, Zhicheng Sun 0001, Yadong Mu
NeurIPS1
2024 MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
abstract
Recently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet, existing systems can only handle videos with very few frames. For long videos, the computation complexity, memory cost, and long-term temporal connection impose additional challenges. Taking advantage of the Atkinson-Shiffrin memory model, with tokens in Transformers being employed as the carriers of memory in combination with our specially designed memory mechanism, we propose the MovieChat to overcome these challenges. MovieChat achieves state-of-the-art performance in long video understanding, along with the released MovieChat-1K benchmark with 1K long video and 14K manual annotations for validation of the effectiveness of our method. The code, models and data can be found in https://reself.github.io/MovieChat.
Enxin Song, Wenhao Chai, Guanhong Wang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo 0002, Tian Ye 0001, Yanting Zhang 0001, Yan Lu 0001, Jenq-Neng Hwang, Gaoang Wang
CVPR7
2024 RectifID: Personalizing Rectified Flow with Anchored Classifier Guidance
abstract
Customizing diffusion models to generate identity-preserving images from user-provided reference images is an intriguing new problem. The prevalent approaches typically require training on extensive domain-specific images to achieve identity preservation, which lacks flexibility across different use cases. To address this issue, we exploit classifier guidance, a training-free technique that steers diffusion models using an existing classifier, for personalized image generation. Our study shows that based on a recent rectified flow framework, the major limitation of vanilla classifier guidance in requiring a special classifier can be resolved with a simple fixed-point solution, allowing flexible personalization with off-the-shelf image discriminators. Moreover, its solving procedure proves to be stable when anchored to a reference flow trajectory, with a convergence guarantee. The derived method is implemented on rectified flow with different off-the-shelf image discriminators, delivering advantageous personalization results for human faces, live subjects, and certain objects. Code is available at https://github.com/feifeiobama/RectifID.
Zhicheng Sun 0001, Zhenhao Yang, Haozhe Chi, Kun Xu 0005, Hao Jiang 0032, Yang Song 0008, Kun Gai, Yadong Mu
NeurIPS4
2024 Segment anything model for medical images?
Yuhao Huang 0001, Xin Yang 0009, Ao Chang, Rusi Chen, Junxuan Yu, Jiongquan Chen, Chaoyu Chen, Sijing Liu, Haozhe Chi, Xindi Hu, Kejuan Yue, Lei Li 0020, Vicente Grau, Deng-Ping Fan, Fajin Dong, Dong Ni 0001
Medical Image Anal.12
2023 Fourier Test-Time Adaptation with Multi-level Consistency for Robust Classification
Yuhao Huang 0001, Xin Yang 0009, Xiaoqiong Huang, Haozhe Chi, Haoran Dou, Xindi Hu, Jian Wang 0099, Xuedong Deng, Dong Ni 0001
MICCAI (3)5