Haozhe Feng

dblp:241/9604 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0002-5900-356XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Language models and text generation · 26% Transfer learning and domain adaptation · 22% Efficient and distributed learning · 16%
Computer graphics and multimedia
5 papers
Visualization and visual analytics · 100%
Databases, data mining, and information retrieval
2 papers
Data mining · 40% Machine learning and data management · 30% Data integration and cleaning · 30%
Network and information security
3 papers
Security and privacy of machine learning · 55% Privacy and data protection · 45%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 22 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visualization and visual analytics › visualization generation › automated visualization generation
chart generation
1.012026
Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports from Scratch with Agentic Framework · AAAI 2026
Data integration and cleaning
data preprocessing
0.912025
DataLab: A Unified Platform for LLM-Powered Business Intelligence · ICDE 2025
Visualization and visual analytics › visualization generation
automated visualization generation
0.912025
DataLab: A Unified Platform for LLM-Powered Business Intelligence · ICDE 2025
Machine learning › Efficient and distributed learning
federated learning
0.812024
BAFFLE: A Baseline of Backpropagation-Free Federated Learning · ECCV (75) 2024
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.812024
Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning · ACL (1) 2024
Visualization and visual analytics › interaction techniques
interactive annotation
0.712023
ChartNavigator: An Interactive Pattern Identification and Annotation Framework for Charts · IEEE Trans. Knowl. Data Eng. 2023
Visualization and visual analytics › information visualization
privacy-preserving visualization
0.712023
DPVisCreator: Incorporating Pattern Constraints to Privacy-preserving Visualizations via Differential Privacy · IEEE Trans. Vis. Comput. Graph. 2023
Privacy and data protection
differential privacy
0.712023
DPVisCreator: Incorporating Pattern Constraints to Privacy-preserving Visualizations via Differential Privacy · IEEE Trans. Vis. Comput. Graph. 2023
Machine learning › Graph learning › graph neural network
dynamic graph neural network
0.612022
Multivariate Time-Series Forecasting with Temporal Polynomial Graph Neural Networks · NeurIPS 2022
Machine learning › Graph learning
graph neural network
0.612022
Multivariate Time-Series Forecasting with Temporal Polynomial Graph Neural Networks · NeurIPS 2022
Data mining › time series analysis › time series forecasting
multivariate time series forecasting
0.612022
Multivariate Time-Series Forecasting with Temporal Polynomial Graph Neural Networks · NeurIPS 2022
Data mining › time series analysis
time series forecasting
0.612022
Multivariate Time-Series Forecasting with Temporal Polynomial Graph Neural Networks · NeurIPS 2022
Machine learning › Efficient and distributed learning
federated and distributed training
0.512021
KD3A: Unsupervised Multi-Source Decentralized Domain Adaptation via Knowledge Distillation · ICML 2021
Machine learning › Transfer learning and domain adaptation › domain adaptation
multi-source domain adaptation
0.512021
KD3A: Unsupervised Multi-Source Decentralized Domain Adaptation via Knowledge Distillation · ICML 2021
Machine learning › Trustworthy machine learning › privacy › privacy-preserving machine learning
privacy-preserving distributed learning
0.512021
KD3A: Unsupervised Multi-Source Decentralized Domain Adaptation via Knowledge Distillation · ICML 2021
Machine learning › Learning paradigms
semi-supervised learning
0.512021
SHOT-VAE: Semi-supervised Deep Generative Models With Label-aware ELBO Approximations · AAAI 2021
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.512021
KD3A: Unsupervised Multi-Source Decentralized Domain Adaptation via Knowledge Distillation · ICML 2021
Machine learning › Generative modeling
variational autoencoder
0.512021
SHOT-VAE: Semi-supervised Deep Generative Models With Label-aware ELBO Approximations · AAAI 2021
Visualization and visual analytics
visual analytics
0.522025
InsightLens: Augmenting LLM-Powered Data Analysis With Interactive Insight Management and Navigation · IEEE Trans. Vis. Comput. Graph. 2025
DPVisCreator: Incorporating Pattern Constraints to Privacy-preserving Visualizations via Differential Privacy · IEEE Trans. Vis. Comput. Graph. 2023
Natural language and speech › Language models and text generation
large language model safety
0.312026
A Causal Perspective for Enhancing Jailbreak Attack and Defense · NDSS 2026
Machine learning › Learning theory
classification
0.112021
SHOT-VAE: Semi-supervised Deep Generative Models With Label-aware ELBO Approximations · AAAI 2021
Privacy and data protection
privacy-preserving data analysis
0.112021
KD3A: Unsupervised Multi-Source Decentralized Domain Adaptation via Knowledge Distillation · ICML 2021

Methods — techniques the papers use, named apart from their topics

large language model · 3.5retrieval-augmented generation · 2.0formal description of visualization · 2.0causal inference · 2.0agentic framework · 2.0adversarial prompting · 2.0interactive visualization · 1.7inter-agent communication · 1.7domain knowledge incorporation · 1.7agent framework · 1.7LLM agents · 1.7bayesian network · 1.3matrix polynomial · 1.1zeroth-order optimization · 0.8self-distillation · 0.8distribution matching · 0.8backpropagation-free training · 0.8differential privacy · 0.7
YearPublicationVenuePosition
2026 Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports from Scratch with Agentic Framework
abstract
Visualizations play a crucial part in effective communication of concepts and information. Recent advances in reasoning and retrieval augmented generation have enabled Large Language Models (LLMs) to perform deep research and generate comprehensive reports. Despite its progress, existing deep research frameworks primarily focus on generating text-only content, leaving the automated generation of interleaved texts and visualizations underexplored. This novel task poses key challenges in designing informative visualizations and effectively integrating them with text reports. To address these challenges, we propose Formal Description of Visualization (FDV), a structured textual representation of charts that enables LLMs to learn from and generate diverse, high-quality visualizations. Building on this representation, we introduce Multimodal DeepResearcher, an agentic framework that decomposes the task into four stages: (1) researching, (2) exemplar report textualization, (3) planning and (4) multimodal report generation. For the evaluation of the generated reports, we develop MultimodalReportBench which contains 100 diverse topics as inputs, and a set of dedicated metrics for report and chart evaluation. Extensive experiments across models and evaluation methods demonstrate the effectiveness of Multimodal DeepResearcher. Notably, utilizing the same Claude 3.7 Sonnet model, Multimodal DeepResearcher achieves an 82% overall win rate over the baseline method.
Zhaorui Yang 0001, Bo Pan 0004, Yiyao Wang, Xingyu Liu 0003, Luoxuan Weng, Yingchaojie Feng, Haozhe Feng, Minfeng Zhu 0001, Wei Chen 0001
AAAI8
2026 A Causal Perspective for Enhancing Jailbreak Attack and Defense
Licheng Pan, Yunsheng Lu, Jiexi Liu 0005, Jialing Tao, Haozhe Feng, Hui Xue 0001, Zhixuan Chu, Kui Ren 0001
NDSS5
2025 DataLab: A Unified Platform for LLM-Powered Business Intelligence
abstract
Business intelligence (BI) transforms large volumes of data within modern organizations into actionable insights for informed decision-making. Recently, large language model (LLM)-based agents have streamlined the BI workflow by automatically performing task planning, reasoning, and actions in executable environments based on natural language (NL) queries. However, existing approaches primarily focus on individual BI tasks such as NL2SQL and NL2VIS. The fragmentation of tasks across different data roles and tools lead to inefficiencies and potential errors due to the iterative and collaborative nature of BI. In this paper, we introduce DataLab, a unified BI platform that integrates a one-stop LLM-based agent framework with an augmented computational notebook interface. DataLab supports various BI tasks for different data roles in data preparation, analysis, and visualization by seamlessly combining LLM assistance with user customization within a single environment. To achieve this unification, we design a domain knowledge incorporation module tailored for enterprise-specific BI tasks, an inter-agent communication mechanism to facilitate information sharing across the BI workflow, and a cell-based context management strategy to enhance context utilization efficiency in BI notebooks. Extensive experiments demonstrate that DataLab achieves state-of-the-art performance on various BI tasks across popular research benchmarks. Moreover, DataLab maintains high effectiveness and efficiency on real-world datasets from Tencent, achieving up to a 58.58% increase in accuracy and a 61.65 % reduction in token cost on enterprise-specific BI tasks.
Luoxuan Weng, Yinghao Tang, Yingchaojie Feng, Zhuo Chang, Ruiqin Chen, Haozhe Feng, Chen Hou, Danqing Huang, Yang Li 0106, Huaming Rao, Canshi Wei, Xiuqi Huang, Minfeng Zhu 0001, Yuxin Ma 0001, Bin Cui 0001, Peng Chen 0021, Wei Chen 0001
ICDE6
2025 FedCare: towards interactive diagnosis of federated learning systems
Tian-Ye Zhang, Haozhe Feng, Wenqi Huang 0002, Lingyu Liang, Huanming Zhang, Zexian Chen, Anthony K. H. Tung, Wei Chen 0001
Frontiers Comput. Sci.2
2025 InsightLens: Augmenting LLM-Powered Data Analysis With Interactive Insight Management and Navigation
abstract
The proliferation of large language models (LLMs) has revolutionized the capabilities of natural language interfaces (NLIs) for data analysis. LLMs can perform multi-step and complex reasoning to generate data insights based on users' analytic intents. However, these insights often entangle with an abundance of contexts in analytic conversations such as code, visualizations, and natural language explanations. This hinders efficient recording, organization, and navigation of insights within the current chat-based LLM interfaces. In this paper, we first conduct a formative study with eight data analysts to understand their general workflow and pain points of insight management during LLM-powered data analysis. Accordingly, we introduce InsightLens, an interactive system to overcome such challenges. Built upon an LLM-agent-based framework that automates insight recording and organization along with the analysis process, InsightLens visualizes the complex conversational contexts from multiple aspects to facilitate insight navigation. A user study with twelve data analysts demonstrates the effectiveness of InsightLens, showing that it significantly reduces users' manual and cognitive effort without disrupting their conversational data analysis workflow, leading to a more efficient analysis experience.
Luoxuan Weng, Xingbo Wang 0001, Yingchaojie Feng, Haozhe Feng, Danqing Huang, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.6
2024 Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning
abstract
Zhaorui Yang, Tianyu Pang, Haozhe Feng, Han Wang, Wei Chen, Minfeng Zhu, Qian Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zhaorui Yang 0001, Tianyu Pang, Haozhe Feng, Wei Chen 0001, Minfeng Zhu 0001, Qian Liu 0033
ACL (1)3
2024 BAFFLE: A Baseline of Backpropagation-Free Federated Learning
Haozhe Feng, Tianyu Pang, Wei Chen 0001, Shuicheng Yan
ECCV (75)1
2024 GraphFederator: Federated Visual Analysis for Multi-party Graphs
abstract
This paper presents GraphFederator, a novel approach to construct federated representations of multi-party graphs and supports privacy-preserving visual analysis of graphs. Inspired by the concept of federated learning, we reformulate the analysis of multi-party graphs into a decentralization process. The new federation framework consists of a shared module that is responsible for federated modeling and analysis, and a set of local modules that run on respective graph data. Specifically, we propose a Federated Graph Representation Model (FGRM) that is learned from encrypted characteristics of multi-party graphs in local modules. We also design multiple visualization tools for federated visualization, exploration, and analysis of multi-party graphs. Experimental results on two datasets demonstrate the effectiveness of our approach.
Dongming Han, Wei Chen 0001, Rusheng Pan, Yijing Liu 0003, Jiehui Zhou, Haozhe Feng, Tian-Ye Zhang, Xumeng Wang, Minfeng Zhu 0001, Jianrong Tao, Changjie Fan, Xiaolong Zhang 0001
PacificVis7
2023 ChartNavigator: An Interactive Pattern Identification and Annotation Framework for Charts
abstract
Patterns in charts refer to interesting visual features or forms. Identifying patterns not only helps analysts understand the ‘shape’ of the data but also supports better and faster decision-making. Existing solutions for identifying patterns in charts require a large number of labeled data instances, making it intractable without user supervision. In this paper, we propose ChartNavigator, an interactive pattern identification and annotation framework for unlabeled visualization charts. ChartNavigator leverages a novel chart-sensitive deep factor model to map patterns into a low-dimensional factor representation space, and facilitates rich analysis with the derived representations. We design and implement a visual interface to support efficient identification and annotation of potential patterns in charts. Evaluations with multiple datasets show that our approach outperforms the baseline models in identifying and annotating patterns
Tian-Ye Zhang, Haozhe Feng, Wei Chen 0001, Zexian Chen, Wenting Zheng, Wenqi Huang 0002, Anthony K. H. Tung
IEEE Trans. Knowl. Data Eng.2
2023 DPVisCreator: Incorporating Pattern Constraints to Privacy-preserving Visualizations via Differential Privacy
abstract
Data privacy is an essential issue in publishing data visualizations. However, it is challenging to represent multiple data patterns in privacy-preserving visualizations. The prior approaches target specific chart types or perform an anonymization model uniformly without considering the importance of data patterns in visualizations. In this paper, we propose a visual analytics approach that facilitates data custodians to generate multiple private charts while maintaining user-preferred patterns. To this end, we introduce pattern constraints to model users' preferences over data patterns in the dataset and incorporate them into the proposed Bayesian network-based Differential Privacy (DP) model PriVis. A prototype system, DPVisCreator, is developed to assist data custodians in implementing our approach. The effectiveness of our approach is demonstrated with quantitative evaluation of pattern utility under the different levels of privacy protection, case studies, and semi-structured expert interviews.
Jiehui Zhou, Xumeng Wang, Jason K. Wong, Huanliang Wang, Xiaoran Yan, Haozhe Feng, Huamin Qu, Haochao Ying, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.8
2022 Multivariate Time-Series Forecasting with Temporal Polynomial Graph Neural Networks
abstract
Modeling multivariate time series (MTS) is critical in modern intelligent systems. The accurate forecast of MTS data is still challenging due to the complicated latent variable correlation. Recent works apply the Graph Neural Networks (GNNs) to the task, with the basic idea of representing the correlation as a static graph. However, predicting with a static graph causes significant bias because the correlation is time-varying in the real-world MTS data. Besides, there is no gap analysis between the actual correlation and the learned one in their works to validate the effectiveness. This paper proposes a temporal polynomial graph neural network (TPGNN) for accurate MTS forecasting, which represents the dynamic variable correlation as a temporal matrix polynomial in two steps. First, we capture the overall correlation with a static matrix basis. Then, we use a set of time-varying coefficients and the matrix basis to construct a matrix polynomial for each time step. The constructed result empirically captures the precise dynamic correlation of six synthetic MTS datasets generated by a non-repeating random walk model. Moreover, the theoretical analysis shows that TPGNN can achieve perfect approximation under a commutative condition. We conduct extensive experiments on two traffic datasets with prior structure and four benchmark datasets. The results indicate that TPGNN achieves the state-of-the-art on both short-term and long-term MTS forecastings.
Yijing Liu 0003, Qinxian Liu, Jianwei Zhang 0015, Haozhe Feng, Zihan Zhou 0009, Wei Chen 0001
NeurIPS4
2021 SHOT-VAE: Semi-supervised Deep Generative Models With Label-aware ELBO Approximations
abstract
Semi-supervised variational autoencoders (VAEs) have obtained strong results, but have also encountered the challenge that good ELBO values do not always imply accurate inference results.In this paper, we investigate and propose two causes of this problem: (1) The ELBO objective cannot utilize the label information directly. (2) A bottleneck value exists, and continuing to optimize ELBO after this value will not improve inference accuracy. On the basis of the experiment results, we propose SHOT-VAE to address these problems without introducing additional prior knowledge. The SHOT-VAE offers two contributions: (1) A new ELBO approximation named smooth-ELBO that integrates the label predictive loss into ELBO. (2) An approximation based on optimal interpolation that breaks the ELBO value bottleneck by reducing the margin between ELBO and the data likelihood. The SHOT-VAE achieves good performance with 25.30% error rate on CIFAR-100 with 10k labels and reduces the error rate to 6.11% on CIFAR-10 with 4k labels.
Haozhe Feng, Kezhi Kong, Tian-Ye Zhang, Minfeng Zhu 0001, Wei Chen 0001
AAAI1
2021 KD3A: Unsupervised Multi-Source Decentralized Domain Adaptation via Knowledge Distillation
abstract
Conventional unsupervised multi-source domain adaptation (UMDA) methods assume all source domains can be accessed directly. However, this assumption neglects the privacy-preserving policy, where all the data and computations must be kept decentralized. There exist three challenges in this scenario: (1) Minimizing the domain distance requires the pairwise calculation of the data from the source and target domains, while the data on the source domain is not available. (2) The communication cost and privacy security limit the application of existing UMDA methods, such as the domain adversarial training. (3) Since users cannot govern the data quality, the irrelevant or malicious source domains are more likely to appear, which causes negative transfer. To address the above problems, we propose a privacy-preserving UMDA paradigm named Knowledge Distillation based Decentralized Domain Adaptation (KD3A), which performs domain adaptation through the knowledge distillation on models from different source domains. The extensive experiments show that KD3A significantly outperforms state-of-the-art UMDA approaches. Moreover, the KD3A is robust to the negative transfer and brings a 100x reduction of communication cost compared with other decentralized UMDA methods.
Haozhe Feng, Zhaoyang You, Tian-Ye Zhang, Minfeng Zhu 0001, Fei Wu 0001, Chao Wu 0001, Wei Chen 0001
ICML1
2020 Lung adenocarcinoma diagnosis in one stage
Pengyi Hao, Kun You, Haozhe Feng, Xinnan Xu, Fan Zhang 0056, Fuli Wu, Peng Zhang 0043, Wei Chen 0001
Neurocomputing3