Haoran Guo

dblp:124/2047 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Cellflow: Advancing pathological image augmentation from spatial views to temporal trajectories
Zeyu Liu 0013, Haoran Guo, Peng Zhang 0078, Chenbin Ma, Shangqing Lyu, Yunlu Feng, Yueming Jin, Dachun Zhao, Guanglei Zhang
Medical Image Anal.5
2026 Implicit hierarchical temporal-spatial residual model for long-term video prediction
Guiqin Wang, Peng Zhao 0001, Haoran Guo, Cong Zhao 0001, Qinghai Guo, Shusen Yang
Neural Networks3
2026 DHPT: Dual-Modality Heterogeneous Prompt Tuning for Online Test-Time Adaption in Vision-Language Models
abstract
Test-Time Adaptation (TTA) has recently emerged as a promising research direction, enabling vision-language models (VLMs) to adapt to unlabeled test data in zero-shot settings. Among TTA approaches, test-time prompt tuning has shown great potential for enhancing the practical applicability of VLMs. However, existing methods typically either focus on adapting a single modality or apply uniform optimization to both modalities, without explicitly defining modality-specific optimization objectives. Such a one-size-fits-all strategy often results in suboptimal performance under test-time conditions. To address this limitation, we propose Dual-modality Heterogeneous Prompt Tuning (DHPT), a novel framework designed to simultaneously capture fine-grained textual semantics and alleviate domain shift noise in the visual modality. Specifically, we leverage a large language model to provide textual cognition guidance for the text encoder, while on the vision side, we develop a lightweight calibration module that adaptively mitigates domain shift noise across different scales. Furthermore, we introduce a cluster-tight optimization objective that enhances the stability and generalizability of prompt tuning under distribution shifts. Extensive experiments conducted on 11 benchmark datasets demonstrate that DHPT consistently and significantly outperforms existing TTA methods for VLMs.
Guiqin Wang, Peng Zhao 0001, Haoran Guo, Shusen Yang, Qinghai Guo
IEEE Trans. Circuits Syst. Video Technol.4
2026 StainExpert: A Unified Multi-Expert Diffusion Framework for Multi-Target Pathological Stain Translation
abstract
Histopathological analysis constitutes the diagnostic cornerstone in disease characterization, employing diverse staining methodologies to elucidate tissue architecture. While hematoxylin and eosin (H&E) remains the foundational technique, ancillary modalities, including specialized histochemical stains, immune-histochemistry (IHC), and multiplex immune-fluorescence (mpIF), yield critical complementary data essential for comprehensive diagnosis. Nevertheless, sequential implementation of these techniques necessitates protracted processing times, substantial labor investment, and significant tissue consumption, often requiring serial sectioning with iterative staining procedures that compromise sample integrity. To address these challenges, we propose StainExpert, a unified multimodal diffusion framework for source-to-multi-target pathological stain translation. Unlike existing approaches that require separate models for each staining pair, StainExpert establishes the first multi-expert system where specialized networks collaboratively learn staining principles while maintaining domain-specific expertise. Through multi-expert and multi-objective optimization, it enables efficient translation from a single source to multiple targets. Additionally, our multimodal diffusion architecture integrates textual guidance with visual features, achieving superior accuracy and pathology-informed translation. Leveraging parameter-efficient design and model distillation, StainExpert matches GAN-level efficiency while delivering superior generation quality. We validate StainExpert across three datasets spanning H&E, special stains, IHC, and mpIF modalities. Extensive evaluation demonstrates that StainExpert generates high-quality virtual stains that preserve critical pathological features for accurate diagnosis. Beyond robust cross-domain generalization, StainExpert offers a transformative platform for efficient multi-target stain translation, advancing toward streamlined, tissue-conserving, and resource-efficient diagnostic workflows in computational pathology. The code is available at https://rowerliu.github.io/StainExpert.
Zeyu Liu 0013, Chenbin Ma, Huijie Wu, Ruxin Cai, Haoran Guo, Peng Zhang 0078, Dachun Zhao, Guanglei Zhang
IEEE Trans. Medical Imaging8
2025 OptiPathD: A Capacity-Optimized Diffusion Foundation Model for Pathology Image Generation
abstract
Generative models hold promise in addressing data scarcity and imbalance in computational pathology, yet current approaches often suffer from limited generalization due to either overfitting on narrow domains or reliance on pre-trained models from unrelated natural image distributions. In this work, we introduce OptiPathD, the first pathology-specific generative foundation model optimized for scalable and generalizable image synthesis. Leveraging our curated dataset CPIA comprising over 148 million multi-scale, multi-organ whole-slide image patches, we pre-train a transformer-based diffusion model with pathology-aware design. To enhance both fidelity and generalization, we propose a principled capacity optimization strategy that aligns model complexity with data scale. Extensive evaluations demonstrate that OptiPathD achieves state-of-the-art performance in conditional image generation, outperforming present generative models across fidelity, diversity, and transferability metrics. Further experiments using downstream classification task on ROSE dataset confirm the efficacy of our generated images. Our work provides a foundation for generative pathology modeling, offering a scalable, domain-specialized, and transferable solution to support data-driven clinical research and diagnostic applications.
Zeyu Liu 0013, Peng Zhang 0078, Chenbin Ma, Haoran Guo, Nan Ying, Shangqing Lyu, Guanglei Zhang
BIBM7
2025 Think Thrice Before You Act: Progressive Thought Refinement in Large Language Models
abstract
Recent advancements in large language models (LLMs) have demonstrated that progressive refinement, rather than providing a single answer, results in more accurate and thoughtful outputs. However, existing methods often rely heavily on supervision signals to evaluate previous responses, making it difficult to effectively assess output quality in more open-ended scenarios. Additionally, these methods are typically designed for specific tasks, which limits their generalization to new domains. To address these limitations, we propose Progressive Thought Refinement (PTR), a framework that enables LLMs to progressively refine their responses. PTR operates in two phases: (1) Thought data construction stage: We propose a weak and strong model collaborative selection strategy to build a high-quality progressive refinement dataset to ensure logical consistency from thought to answers, and the answers are gradually refined in each round. (2) Thought-Mask Fine-Tuning Phase: We design a training structure to mask the "thought" and adjust loss weights to encourage LLMs to refine prior thought, teaching them to implicitly understand "how to improve" rather than "what is correct." Experimental results show that PTR significantly enhances LLM performance across ten diverse tasks (avg. from 49.6% to 53.5%) without task-specific fine-tuning. Notably, in more open-ended tasks, LLMs also demonstrate substantial improvements in the quality of responses beyond mere accuracy, suggesting that PTR truly teaches LLMs to self-improve over time. Our work is now open-source. https://github.com/cydu24/Progressive-Thought-Refinement
Chengyu Du, Jinyi Han, Yizhou Ying, Aili Chen, Qianyu He, Haokun Zhao, Haoran Guo, Sirui Xia, Jiaqing Liang, Zulong Chen, Liangyue Li, Yanghua Xiao
ICLR7
2025 CoSER: Coordinating LLM-Based Persona Simulation of Established Roles
abstract
Role-playing language agents (RPLAs) have emerged as promising applications of large language models (LLMs). However, simulating established characters presents a challenging task for RPLAs, due to the lack of authentic character datasets and nuanced evaluation methods using such data. In this paper, we present CoSER, a collection of a high-quality dataset, open models, and an evaluation protocol towards effective RPLAs of established characters. The CoSER dataset covers 17,966 characters from 771 renowned books. It provides authentic dialogues with real-world intricacies, as well as diverse data types such as character experiences and internal thoughts. Drawing from acting methodology, we introduce given-circumstance acting for training and evaluating role-playing LLMs, where LLMs sequentially portray multiple characters in book scenes. Using our dataset, we develop CoSER 8B and CoSER 70B, i.e., advanced open role-playing LLMs built on LLaMA-3.1 models. Extensive experiments demonstrate the value of the CoSER dataset for RPLA training, evaluation and retrieval. Moreover, CoSER 70B exhibits state-of-the-art performance surpassing or matching GPT-4o on our evaluation and three existing benchmarks, i.e., achieving 75.80% and 93.47% accuracy on the InCharacter and LifeChoice benchmarks respectively. Our code, dataset and models are available at: https://github.com/Neph0s/CoSER.
Xintao Wang 0001, Xinfeng Yuan, Rui Xu 0026, Jen-tse Huang 0001, Haoran Guo, Jiangjie Chen, Shuchang Zhou 0003, Wei Wang 0009, Yanghua Xiao
ICML8
2025 Exactly Tight Information-theoretic Generalization Bounds via Binary Jensen-Shannon Divergence
abstract
Information-theoretic bounds, while achieving significant success in analyzing the generalization of randomized learning algorithms, have been criticized for their slow convergence rates and overestimation. This paper presents novel bounds that bridge the expected empirical and population risks through a binarized variant of the Jensen-Shannon divergence. Leveraging our foundational lemma that characterizes the interaction between an arbitrary and a binary variable, we derive hypothesis-based bounds that enhance existing conditional mutual information bounds by reducing the number of conditioned samples from $2$ to $1$. We additionally establish prediction-based bounds that surpass prior bounds based on evaluated loss mutual information measures. Thereafter, through a new binarization technique for the evaluated loss variables, we obtain exactly tight generalization bounds broadly applicable to general randomized learning algorithms for any bounded loss functions. Our results effectively address key limitations of previous results in analyzing certain stochastic convex optimization problems, without requiring additional stability or compressibility assumptions about the learning algorithm.
Yuxin Dong 0003, Haoran Guo, Tieliang Gong, Wen Wen 0013, Chen Li 0011
ICML2
2025 Bio-Skin: A Cost-Effective Thermostatic Tactile Sensor with Multi-Modal Force and Temperature Detection
abstract
Tactile sensors can significantly enhance the perception of humanoid robotics systems by providing contact information that facilitates human-like interactions. However, existing commercial tactile sensors focus on improving the resolution and sensitivity of single-modal detection with high-cost components and densely integrated design, incurring complex manufacturing processes and unaffordable prices. In this work, we present Bio-Skin, a cost-effective multi-modal tactile sensor that utilizes single-axis Hall-Effect sensors for planar normal force measurement and bar-shape piezo resistors for 2D shear force measurement. A thermistor coupling with a heating wire is integrated into a silicone body to achieve temperature sensation and thermostatic function analogous to human skin. We also present a cross-reference framework to validate the two modalities of the force sensing signal, improving the sensing fidelity in a complex electromagnetic environment. Bio-Skin has a multi-layer design, and each layer is manufactured sequentially and subsequently integrated, thereby offering a fast production pathway. After calibration, Bio-Skin demonstrates performance metrics—including signal-to-range ratio, sampling rate, and measurement range—comparable to current commercial products, with one-tenth of the cost. The sensor’s real-world performance is evaluated using an Allegro hand in object grasping tasks, while its temperature regulation functionality was assessed in a material detection task.
Haoran Guo, Zhengxiong Li, Lingfeng Tao
IROS1
2025 Adaptive Anomaly Recovery for Telemanipulation: A Diffusion Model Approach to Vision-Based Tracking
abstract
Dexterous telemanipulation critically relies on the continuous and stable tracking of the human operator’s commands to ensure robust operation. Vison-based tracking methods are widely used but have low stability due to anomalies such as occlusions, inadequate lighting, and loss of sight. Traditional filtering, regression, and interpolation methods are commonly used to compensate for explicit information such as angles and positions. These approaches are restricted to low-dimensional data and often result in information loss compared to the original high-dimensional image and video data. Recent advances in diffusion-based approaches, which can operate on high-dimensional data, have achieved remarkable success in video reconstruction and generation. However, these methods have not been fully explored in continuous control tasks in robotics. This work introduces the Diffusion-Enhanced Telemanipulation (DET) framework, which incorporates the Frame-Difference Detection (FDD) technique to identify and segment anomalies in video streams. These anomalous clips are replaced after reconstruction using diffusion models, ensuring robust telemanipulation performance under challenging visual conditions. We validated this approach in various anomaly scenarios and compared it with the baseline methods. Experiments show that DET achieves an average RMSE reduction of 17.2% compared to the cubic spline and 51.1% compared to FFT-based interpolation for different occlusion durations.
Haoran Guo, Zhengxiong Li, Lingfeng Tao
IROS2
2024 InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews
abstract
Xintao Wang, Yunze Xiao, Jen-tse Huang, Siyu Yuan, Rui Xu, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang, Jiangjie Chen, Cheng Li, Yanghua Xiao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Xintao Wang 0001, Yunze Xiao, Jen-tse Huang 0001, Rui Xu 0026, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang 0009, Jiangjie Chen, Yanghua Xiao
ACL (1)6
2024 MCFP: A multi-target 3D perception method with weak dependence on 2D detectors
Haoran Guo, Mingyun He, Fan Li 0033, Kexin He, Lina Chen
Pattern Recognit. Lett.1
2023 Neighborhood Attention-based Transformer Line Segment Detector with Edge Computing
abstract
Line segment detection is increasingly being widely applied in visual tasks. Traditional methods for line segment detection are known for their speed and accuracy, but they lack robustness in handling noisy images. CNN-based line segment detection methods achieve impressive detection results by extracting more complex features. Nowadays, with the development of edge computing, these methods are being optimized for edge devices to enable on-device processing without the need for extensive cloud resources. However, current CNN-based methods fail to effectively utilize the global features of an image, and they usually require the stacking of a large number of convolutional layers to achieve the best performance, resulting in high computation costs. Transformer-based line segment detection methods have gradually emerged as promising solutions for edge computing applications. They leverage the power of distributed attention mechanisms to capture both global and local features, making them suitable for edge devices with limited computational resources. In this paper, we introduce the concept of edge-enhanced Transformer-based line segment detection, specifically the Neighborhood Attention-Based Transformer (NBAT). By incorporating neighborhood attention into the Transformer architecture, NBAT efficiently processes image data on the edge, focusing on local regions of interest. This localized approach not only ensures real-time responsiveness but also maintains accuracy in detecting line segments. Extensive experiments demonstrate that the proposed NBAT achieves higher detection accuracy.
Haoran Guo, Mingyue Shi
ICPADS1
2023 User story clustering in agile development: a framework and an empirical study
Xiuyin Ma, Haoran Guo, Huai Liu
Frontiers Comput. Sci.4
2023 Evaluation and assessment of machine learning based user story grouping: A framework and empirical studies
Haoran Guo, Huai Liu
Sci. Comput. Program.2
2022 CO-SNE: Dimensionality Reduction and Visualization for Hyperbolic Data
abstract
Hyperbolic space can naturally embed hierarchies that often exist in real-world data and semantics. While high-dimensional hyperbolic embeddings lead to better representations, most hyperbolic models utilize low-dimensional embeddings, due to non-trivial optimization and visualization of high-dimensional hyperbolic data. We propose CO-SNE, which extends the Euclidean space visualization tool, t-SNE, to hyperbolic space. Like t-SNE, it converts distances between data points to joint probabilities and tries to minimize the Kullback-Leibler divergence between the joint probabilities of high-dimensional data$X$and low-dimensional embedding$Y$. However, unlike Euclidean space, hyperbolic space is inhomogeneous: A volume could contain a lot more points at a location far from the origin. CO-SNE thus uses hyperbolic normal distributions for$X$and hyperbolic Cauchy instead of t-SNE's Student's t-distribution for$Y$, and it additionally seeks to preserve$X$'s individual distances to the Origin in$Y$. We apply CO-SNE to naturally hyperbolic data and supervisedly learned hyperbolic features. Our results demonstrate that CO-SNE deflates high-dimensional hyperbolic data into a low-dimensional space without losing their hyperbolic characteristics, significantly outperforming popular visualization tools such as PCA, t-SNE, UMAP, and HoroPCA which is also designed for hyperbolic data.
Yunhui Guo, Haoran Guo, Stella X. Yu
CVPR2
2016 Exploiting Path Diversity for Thwarting Pollution Attacks in Named Data Networking
abstract
With information becoming a first-class citizen on the Internet, information-centric networking (ICN) is considered as a promising direction for the future Internet. Named data networking (NDN) is a prominent example of emerging ICN architectures. Unfortunately, NDN is vulnerable to various attacks targeting its in-network caching mechanism. In this paper, we focus on the false-locality pollution attack, in which an adversary repeatedly requests a number of unpopular data objects to waste the precious cache space on the NDN router and to reduce normal users' hit ratios. With simulation experiments, we show that such an attack can cause considerable damage to the NDN network. To detect and mitigate such an attack, we introduce an algorithm that exploits the diversity of the Interest traversing paths within an Internet service provider's point-of-presence network. We also propose inexpensive methodologies based on the probabilistic counting and Bloom filter techniques to implement the algorithm on an NDN router. The experimental results indicate that our proposed algorithm is effective in thwarting false-locality pollution. We also experiment with strategies that the adversary may utilize against our antipollution algorithm and demonstrate that such strategies are either ineffective or impractical in the real world.
Haoran Guo, Xiaodong Wang 0013, Kun Chang, Ye Tian 0004
IEEE Trans. Inf. Forensics Secur.1