Hao Shao

dblp:66/3089 · DBLP profile ↗
← Back
38ranked-venue papers
16as first author
15since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 9 first-author · 12 since 2021Databases, data management, data science and information retrieval · 10 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 2 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding
abstract
Visual Document Understanding has become essential with the increase of text-rich visual content. This field poses significant challenges due to the need for effective integration of visual perception and textual comprehension, particularly across diverse document types with complex lay-outs. Moreover, existing fine-tuning datasets for this domain often fall short in providing the detailed contextual information for robust understanding, leading to hallucinations and limited comprehension of spatial relationships among visual elements. To address these challenges, we propose an innovative pipeline that utilizes adaptive generation of markup languages, such as Markdown, JSON, HTML, and TiKZ, to build highly structured document representations and deliver contextually-grounded responses. We intro-duce two fine-grained structured datasets: DocMark-Pile, comprising approximately 3.8M pretraining data pairs for document parsing, and DocMark-Instruct, featuring 624k fine-tuning data annotations for grounded instruction following. Extensive experiments demonstrate that our pro-posed model significantly outperforms existing state-of-the-art MLLMs across a range of visual document understanding benchmarks, facilitating advanced reasoning and comprehension capabilities in complex visual scenarios. Our code and models are released at https://github.com/Euphoria16/DocMark.
Han Xiao 0010, Yina Xie, Guanxin Tan, Ke Wang 0036, Aojun Zhou, Hao Li 0069, Hao Shao, Peng Gao 0007, Yafei Wen, Xiaoxin Chen 0001, Shuai Ren 0002, Hongsheng Li 0001
CVPR9
2025 SmartPretrain: Model-Agnostic and Dataset-Agnostic Representation Learning for Motion Prediction
abstract
Predicting the future motion of surrounding agents is essential for autonomous vehicles (AVs) to operate safely in dynamic, human-robot-mixed environments. However, the scarcity of large-scale driving datasets has hindered the development of robust and generalizable motion prediction models, limiting their ability to capture complex interactions and road geometries. Inspired by recent advances in natural language processing (NLP) and computer vision (CV), self-supervised learning (SSL) has gained significant attention in the motion prediction community for learning rich and transferable scene representations. Nonetheless, existing pre-training methods for motion prediction have largely focused on specific model architectures and single dataset, limiting their scalability and generalizability. To address these challenges, we propose SmartPretrain, a general and scalable SSL framework for motion prediction that is both model-agnostic and dataset-agnostic. Our approach integrates contrastive and reconstructive SSL, leveraging the strengths of both generative and discriminative paradigms to effectively represent spatiotemporal evolution and interactions without imposing architectural constraints. Additionally, SmartPretrain employs a dataset-agnostic scenario sampling strategy that integrates multiple datasets, enhancing data volume, diversity, and robustness. Extensive experiments on multiple datasets demonstrate that SmartPretrain consistently improves the performance of state-of-the-art prediction models across datasets, data splits and main metrics. For instance, SmartPretrain significantly reduces the MissRate of Forecast-MAE by 10.6\%. These results highlight SmartPretrain's effectiveness as a unified, scalable solution for motion prediction, breaking free from the limitations of the small-data regime.
Yang Zhou 0054, Hao Shao, Steven Lake Waslander, Hongsheng Li 0001, Yu Liu 0015
ICLR2
2025 EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM
abstract
Significant achievements in personalization of diffusion models have been witnessed. Conventional tuning-free methods mostly encode multiple reference images by averaging or concatenating their image embeddings as the injection condition, but such an image-independent operation cannot perform interaction among images to capture consistent visual elements within multiple references. Although tuning-based approaches can effectively extract consistent elements within multiple images through the training process, it necessitates test-time finetuning for each distinct image group. This paper introduces EasyRef, a plug-and-play adaption method that empowers diffusion models to condition consistent visual elements (e.g., style and human facial identity, etc.) across multiple reference images under instruction controls. To effectively exploit consistent visual elements within multiple images, we leverage the multi-image comprehension and instruction-following capabilities of the multimodal large language model (MLLM), prompting it to capture consistent visual elements based on the instruction. Besides, injecting the MLLM’s representations into the diffusion process through adapters can easily generalize to unseen domains. To mitigate computational costs and enhance fine-grained detail preservation, we introduce an efficient reference aggregation strategy and a progressive training scheme. Finally, we introduce MRBench, a new multi-reference image generation benchmark. Experimental results demonstrate EasyRef surpasses both tuning-free and tuning-based methods, achieving superior aesthetic quality and robust zero-shot generalization across diverse domains.
Zhuofan Zong, Dongzhi Jiang, Bingqi Ma, Guanglu Song, Hao Shao, Dazhong Shen, Yu Liu 0015, Hongsheng Li 0001
ICML5
2025 VividFace: A Robost and High-Fidelity Video Face Swapping Framework
abstract
Video face swapping has seen increasing adoption in diverse applications, yet existing methods primarily trained on static images struggle to address temporal consistency and complex real-world scenarios. To overcome these limitations, we propose the first video face swapping framework, VividFace, a robust and high-fidelity diffusion-based framework. VividFace employs a novel hybrid training strategy that leverages abundant static image data alongside temporal video sequences, enabling it to effectively model temporal coherence and identity consistency in videos. Central to our approach is a carefully designed diffusion model integrated with a specialized VAE, capable of processing image-video hybrid data efficiently. To further enhance identity and pose disentanglement, we introduce and release the Attribute-Identity Disentanglement Triplet (AIDT) dataset, comprising a large-scale collection of triplets where each set contains three face images—two sharing the same pose and two sharing the same identity. Augmented comprehensively with occlusion scenarios, AIDT significantly boosts the robustness of VividFace against occlusions. Moreover, we incorporate advanced 3D reconstruction techniques as conditioning inputs to address significant pose variations effectively. Extensive experiments demonstrate that VividFace achieves state-of-the-art performance in identity preservation, temporal consistency, and visual realism, surpassing existing methods while requiring fewer inference steps. Our framework notably mitigates common challenges such as temporal flickering, identity loss, and sensitivity to occlusions and pose variations. The AIDT dataset, source code, and pre-trained weights will be released to support future research. The code and pretrained weights are available on the [project page](https://hao-shao.com/projects/vividface.html).
Hao Shao, Shulun Wang, Yang Zhou 0054, Guanglu Song, Dailan He, Zhuofan Zong, Yu Liu 0015, Hongsheng Li 0001
NeurIPS1
2024 Polyper: Boundary Sensitive Polyp Segmentation
abstract
We present a new boundary sensitive framework for polyp segmentation, termed Polyper.Our method is motivated by a clinical approach that seasoned medical practitioners often leverage the inherent features of interior polyp regions to tackle blurred boundaries.Inspired by this, we propose to explicitly leverages boundary regions to bolster the model's boundary discrimination capability while minimizing computational resource wastage. Our approach first extracts low-confidence boundary regions and high-confidence prediction regions from an initial segmentation map through differentiable morphological operators.Then, we design the boundary sensitive attention that concentrates on augmenting the features near the boundary regions using the high-confidence prediction region's characteristics to generate good segmentation results.Our proposed method can be seamlessly integrated with classical encoder networks, like ResNet-50, MiT-B1, and Swin Transformer.To evaludate the effectiveness of Polyper, we conduct experiments on five publicly available challenging datasets, and receive state-of-the-art performance on all of them. Code is available at https://github.com/haoshao-nku/medical_seg.git.
Hao Shao, Qibin Hou
AAAI1
2024 LMDrive: Closed-Loop End-to-End Driving with Large Language Models
abstract
Despite significant recent progress in the field of autonomous driving, modern methods still struggle and can incur serious accidents when encountering long-tail unfore-seen events and challenging urban scenarios. On the one hand, large language models (LLM) have shown impres-sive reasoning capabilities that approach “Artificial Gen-eral Intelligence”. On the other hand, previous autonomous driving methods tend to rely on limited-format inputs (e.g., sensor data and navigation waypoints), restricting the vehi-cle's ability to understand language information and inter-act with humans. To this end, this paper introduces LM-Drive, a novel language-guided, end-to-end, closed-loop autonomous driving framework. LMDrive uniquely processes and integrates multimodal sensor data with naturallanguage instructions, enabling interaction with humans and navigation software in realistic instructional settings. To facilitate research in language-based closed-loop autonomous driving, we also publicly release the corresponding dataset which includes approximately 64K instruction-following data clips, and the LangAuto benchmark that tests the system's ability to handle complex instructions and challenging driving scenarios. Extensive closed-loop experiments are conducted to demonstrate LMDrive's effectiveness. To the best of our knowledge, we're the very first work to leverage LLMs for closed-loop end-to-end autonomous driving. Code is available on our webpage.
Hao Shao, Guanglu Song, Steven Lake Waslander, Yu Liu 0015, Hongsheng Li 0001
CVPR1
2024 SmartRefine: A Scenario-Adaptive Refinement Framework for Efficient Motion Prediction
abstract
Predicting the future motion of surrounding agents is essential for autonomous vehicles (AVs) to operate safely in dy-namic, human-robot-mixed environments. Context information, such as road maps and surrounding agents' states, provides crucial geometric and semantic information for motion behavior prediction. To this end, recent works explore two-stage prediction frameworks where coarse trajectories are first proposed, and then used to select critical context information for trajectory refinement. However, they either incur a large amount of computation or bring limited improvement, if not both. In this paper, we introduce a novel scenario-adaptive refinement strategy, named SmartRefine, to refine prediction with minimal additional computation. Specifically, SmartRefine can comprehensively adapt refinement configurations based on each scenario's properties, and smartly chooses the number of refinement iterations by introducing a quality score to measure the prediction quality and remaining refinement potential of each scenario. SmartRefine is designed as a generic and flexible approach that can be seamlessly integrated into most state-of-the-art motion prediction models. Experiments on Argoverse (1 & 2) show that our method consistently improves the prediction accuracy of multiple state-of-the-art prediction models. Specifically, by adding SmartRefine to QCNet, we outper-form all published ensemble-free works on the Argoverse 2 leaderboard (single agent track) at submission11November 2023.: Compre-hensive studies are also conducted to ablate design choices and explore the mechanism behind multi-iteration refinement. Codes are available at our webpage.
Yang Zhou 0054, Hao Shao, Steven Lake Waslander, Hongsheng Li 0001, Yu Liu 0015
CVPR2
2024 SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
abstract
We propose SPHINX-X, an extensive Multi-modality Large Language Model (MLLM) series developed upon SPHINX. To improve the architecture and training efficiency, we modify the SPHINX framework by removing redundant visual encoders, bypassing fully-padded sub-images with skip tokens, and simplifying multi-stage training into a one-stage all-in-one paradigm. To fully unleash the potential of MLLMs, we assemble a comprehensive multi-domain and multi-modal dataset covering publicly available resources in language, vision, and vision-language tasks. We further enrich this collection with our curated OCR intensive and Set-of-Mark datasets, extending the diversity and generality. By training over different base LLMs including TinyLlama-1.1B, InternLM2-7B, LLaMA2-13B, and Mixtral-8$\times$7B, we obtain a spectrum of MLLMs that vary in parameter size and multilingual capabilities. Comprehensive benchmarking reveals a strong correlation between the multi-modal performance with the data and parameter scales. Code and models are released at https://github.com/Alpha-VLLM/LLaMA2-Accessory.
Renrui Zhang, Longtian Qiu, Siyuan Huang 0004, Weifeng Lin, Shitian Zhao, Shijie Geng, Kaipeng Zhang, Wenqi Shao, Conghui He, Junjun He, Hao Shao, Pan Lu, Yu Qiao 0001, Hongsheng Li 0001, Peng Gao 0007
ICML15
2024 Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
abstract
Multi-Modal Large Language Models (MLLMs) have demonstrated impressive performance in various VQA tasks. However, they often lack interpretability and struggle with complex visual inputs, especially when the resolution of the input image is high or when the interested region that could provide key information for answering the question is small. To address these challenges, we collect and introduce the large-scale Visual CoT dataset comprising 438k question-answer pairs, annotated with intermediate bounding boxes highlighting key regions essential for answering the questions. Additionally, about 98k pairs of them are annotated with detailed reasoning steps. Importantly, we propose a multi-turn processing pipeline that dynamically focuses on visual inputs and provides interpretable thoughts. We also introduce the related benchmark to evaluate the MLLMs in scenarios requiring specific local region identification.Extensive experiments demonstrate the effectiveness of our framework and shed light on better inference strategies. The Visual CoT dataset, benchmark, and pre-trained models are available on this website to support further research in this area.
Hao Shao, Shengju Qian, Han Xiao 0010, Guanglu Song, Zhuofan Zong, Yu Liu 0015, Hongsheng Li 0001
NeurIPS1
2024 MoVA: Adapting Mixture of Vision Experts to Multimodal Context
abstract
As the key component in multimodal large language models (MLLMs), the ability of the visual encoder greatly affects MLLM's understanding on diverse image content. Although some large-scale pretrained vision encoders such as vision encoders in CLIP and DINOv2 have brought promising performance, we found that there is still no single vision encoder that can dominate various image content understanding, e.g., the CLIP vision encoder leads to outstanding results on general image understanding but poor performance on document or chart content. To alleviate the bias of CLIP vision encoder, we first delve into the inherent behavior of different pre-trained vision encoders and then propose the MoVA, a powerful and novel MLLM, adaptively routing and fusing task-specific vision experts with a coarse-to-fine mechanism. In the coarse-grained stage, we design a context-aware expert routing strategy to dynamically select the most suitable vision experts according to the user instruction, input image, and expertise of vision experts. This benefits from the powerful model function understanding ability of the large language model (LLM). In the fine-grained stage, we elaborately conduct the mixture-of-vision-expert adapter (MoV-Adapter) to extract and fuse task-specific knowledge from various experts. This coarse-to-fine paradigm effectively leverages representations from experts based on multimodal context and model expertise, further enhancing the generalization ability. We conduct extensive experiments to evaluate the effectiveness of the proposed approach. Without any bells and whistles, MoVA can achieve significant performance gains over current state-of-the-art methods in a wide range of challenging multimodal benchmarks.
Zhuofan Zong, Bingqi Ma, Dazhong Shen, Guanglu Song, Hao Shao, Dongzhi Jiang, Hongsheng Li 0001, Yu Liu 0015
NeurIPS5
2024 An Adaptive Semantic Mining Framework for Heterogeneous Information Network Embedding
abstract
Heterogeneous information network (HIN) embedding aims to map heterogeneous nodes to the low-dimensional vector space. The existing embedding models cannot determine the optimal length of semantics automatically and reveal full semantic information adaptively for different heterogeneous networks. To address this challenge, an HIN embedding model with adaptive semantic mining is proposed. First, we project heterogeneous nodes into the same space and aggregate the features of target types in the first-degree range. Then, the semantics of different node types is combined through the attention mechanism, and latent meta-paths are mined using the attention coefficients. Finally, multiple feature aggregation layers are stacked with residual blocks. The residual weights control the proportion of semantics transferred between layers to aggregate more distal features selectively. In addition, we designed the Selected DropLink unit to remove links which transfer negative information, which can further improve the resistance of model to over-smoothing. Experiments show that our model can obtain more accurate embedding results and can automatically mine complex semantic connections between heterogeneous nodes without prior definition of meta-path and semantic depth.
Hao Shao, Rangang Zhu, Lunwen Wang
IEEE Trans. Comput. Soc. Syst.1
2023 ReasonNet: End-to-End Driving with Temporal and Global Reasoning
abstract
The large-scale deployment of autonomous vehicles is yet to come, and one of the major remaining challenges lies in urban dense traffic scenarios. In such cases, it remains challenging to predict the future evolution of the scene and future behaviors of objects, and to deal with rare adverse events such as the sudden appearance of occluded objects. In this paper, we present ReasonNet, a novel end-to-end driving framework that extensively exploits both temporal and global information of the driving scene. By reasoning on the temporal behavior of objects, our method can effectively process the interactions and relationships among features in different frames. Reasoning about the global information of the scene can also improve overall perception performance and benefit the detection of adverse events, especially the anticipation of potential danger from occluded objects. For comprehensive evaluation on occlusion events, we also release publicly a driving simulation benchmark DriveOcclusionSim consisting of diverse occlusion events. We conduct extensive experiments on multiple CARLA benchmarks, where our model outperforms all prior methods, ranking first on the sensor track of the public CARLA Leaderboard [53].
Hao Shao, Ruobing Chen 0005, Steven Lake Waslander, Hongsheng Li 0001, Yu Liu 0015
CVPR1
2022 Leaning compact and representative features for cross-modality person re-identification
Guangwei Gao, Hao Shao, Fei Wu 0004, Meng Yang 0001, Yi Yu 0001
World Wide Web2
2021 Blending Anti-Aliasing into Vision Transformer
abstract
The transformer architectures, based on self-attention mechanism and convolution-free design, recently found superior performance and booming applications in computer vision. However, the discontinuous patch-wise tokenization process implicitly introduces jagged artifacts into attention maps, arising the traditional problem of aliasing for vision transformers. Aliasing effect occurs when discrete patterns are used to produce high frequency or continuous information, resulting in the indistinguishable distortions. Recent researches have found that modern convolution networks still suffer from this phenomenon. In this work, we analyze the uncharted problem of aliasing in vision transformer and explore to incorporate anti-aliasing properties. Specifically, we propose a plug-and-play Aliasing-Reduction Module (ARM) to alleviate the aforementioned issue. We investigate the effectiveness and generalization of the proposed method across multiple tasks and various vision transformer families. This lightweight design consistently attains a clear boost over several famous structures. Furthermore, our module also improves data efficiency and robustness of vision transformers.
Shengju Qian, Hao Shao, Yi Zhu 0001, Mu Li 0003, Jiaya Jia
NeurIPS2
2021 Kansei evaluation for group of users: A data-driven approach using dominance-based rough sets
Fu Guo, Mingcai Hu, Vincent G. Duffy, Hao Shao, Zenggen Ren
Adv. Eng. Informatics4
2020 Temporal Interlacing Network
abstract
For a long time, the vision community tries to learn the spatio-temporal representation by combining convolutional neural network together with various temporal models, such as the families of Markov chain, optical flow, RNN and temporal convolution. However, these pipelines consume enormous computing resources due to the alternately learning process for spatial and temporal information. One natural question is whether we can embed the temporal information into the spatial one so the information in the two domains can be jointly learned once-only. In this work, we answer this question by presenting a simple yet powerful operator – temporal interlacing network (TIN). Instead of learning the temporal features, TIN fuses the two kinds of information by interlacing spatial representations from the past to the future, and vice versa. A differentiable interlacing target can be learned to control the interlacing process. In this way, a heavy temporal model is replaced by a simple interlacing operator. We theoretically prove that with a learnable interlacing target, TIN performs equivalently to the regularized temporal convolution network (r-TCN), but gains 4% more accuracy with 6x less latency on 6 challenging benchmarks. These results push the state-of-the-art performances of video understanding by a considerable margin. Not surprising, the ensemble model of the proposed TIN won the 1st place in the ICCV19 - Multi Moments in Time challenge. Code is made available to facilitate further research.1
Hao Shao, Shengju Qian
AAAI1
2020 How User's First Impression Forms on Mobile user Interface?: An ERPs Study
abstract
In an era of mobile Internet, the using of mobile interfaces is ineluctable in our daily life. To some extent, the first impression of the mobile user interfaces determines users’ downloading and using behavior, even affects the overall user experience. In fact, the underlying neural mechanism of users’ first impression formation of mobile user interface is worthy to be investigated. Considering the process of first impression formation is always unconscious and transient, event-related potentials (ERPs), which are used to track users’ cerebral activities could be an appropriate method to explore this process. In the present study, the perceived usability and esthetics are taken as the evaluation dimensions of user’ first impression formation, and ERPs were used to investigate which dimensions were automatically activated when users browsed the mobile user interfaces passively without any explicit guidance. The ERPs results showed that N2 was dominated by perceived usability and esthetics, which reflected in enhanced N2 for high esthetics than for low esthetics, and larger N2 for high-perceived usability/high esthetics compared to low-perceived usability/low esthetics. Moreover, there was a larger late positive potential (LPP) for low esthetics interfaces with respect to high esthetics ones. The findings revealed that users could make an implicit evaluation spontaneously to the mobile user interfaces with different levels of perceived usability and esthetics. Furthermore, users’ first impression is significantly affected by the differences in esthetics but marginally influenced by the difference in perceived usability. The suggestions from this study for mobile application designers might be that the esthetics design of the mobile user interface should be prioritized in the early design stage.
Fu Guo, Xueshuang Wang, Hao Shao, Xiao-Rong Wang
Int. J. Hum. Comput. Interact.3
2019 Subject Recognition in Chinese Sentences for Chatbots
Huanhuan Wei, Qiangda Hao, Ruihong Zeng, Hao Shao, Wenliang Chen
NLPCC (2)5
2019 Syntax-aware entity representations for neural relation extraction
Zhengqiu He, Wenliang Chen, Zhenghua Li, Wei Zhang 0027, Hao Shao, Min Zhang 0005
Artif. Intell.5
2019 Query by diverse committee in transfer active learning
Hao Shao
Frontiers Comput. Sci.1
2017 Active Learning for Text Mining from Crowds
Hao Shao
IEA/AIE (2)1
2014 Transfer active learning by querying committee
abstract
In real applications of inductive learning for classification, labeled instances are often deficient, and labeling them by an oracle is often expensive and time-consuming. Active learning on a single task aims to select only informative unlabeled instances for querying to improve the classification accuracy while decreasing the querying cost. However, an inevitable problem in active learning is that the informative measures for selecting queries are commonly based on the initial hypotheses sampled from only a few labeled instances. In such a circumstance, the initial hypotheses are not reliable and may deviate from the true distribution underlying the target task. Consequently, the informative measures will possibly select irrelevant instances. A promising way to compensate this problem is to borrow useful knowledge from other sources with abundant labeled information, which is called transfer learning. However, a significant challenge in transfer learning is how to measure the similarity between the source and the target tasks. One needs to be aware of different distributions or label assignments from unrelated source tasks; otherwise, they will lead to degenerated performance while transferring. Also, how to design an effective strategy to avoid selecting irrelevant samples to query is still an open question. To tackle these issues, we propose a hybrid algorithm for active learning with the help of transfer learning by adopting a divergence measure to alleviate the negative transfer caused by distribution differences. To avoid querying irrelevant instances, we also present an adaptive strategy which could eliminate unnecessary instances in the input space and models in the model space. Extensive experiments on both the synthetic and the real data sets show that the proposed algorithm is able to query fewer instances with a higher accuracy and that it converges faster than the state-of-the-art methods.
Hao Shao, Feng Tao 0002, Rui Xu 0004
J. Zhejiang Univ. Sci. C1
2014 Transfer dimensionality reduction by Gaussian process in parallel
Bin Tong, Junbin Gao, Thach Huy Nguyen, Hao Shao, Einoshin Suzuki
Knowl. Inf. Syst.4
2013 An Estimation of Distribution Algorithm for the 3D Bin Packing Problem with Various Bin Sizes
Yaxiong Cai, Huaping Chen 0001, Rui Xu 0004, Hao Shao, Xueping Li 0002
IDEAL4
2013 Gaussian Process for Transfer Learning through Minimum Encoding
Hao Shao, Rui Xu 0004, Feng Tao 0002
IDEAL1
2013 An Effective Ant Colony Approach for Scheduling Parallel Batch-Processing Machines
Rui Xu 0004, Huaping Chen 0001, Hao Shao
IDEAL3
2013 Transfer learning by centroid pivoted mapping in noisy environment
Thach Huy Nguyen, Bin Tong, Hao Shao, Einoshin Suzuki
J. Intell. Inf. Syst.3
2013 A feature-free and parameter-light multi-task clustering framework
Thach Huy Nguyen, Hao Shao, Bin Tong, Einoshin Suzuki
Knowl. Inf. Syst.2
2013 Extended MDL principle for feature-based inductive transfer learning
Hao Shao, Bin Tong, Einoshin Suzuki
Knowl. Inf. Syst.1
2012 Query by Committee in a Heterogeneous Environment
Hao Shao, Bin Tong, Einoshin Suzuki
ADMA1
2012 Linear semi-supervised projection clustering by transferred centroid regularization
Bin Tong, Hao Shao, Bin-Hui Chou, Einoshin Suzuki
J. Intell. Inf. Syst.2
2011 A Compression-Based Dissimilarity Measure for Multi-task Clustering
Thach Huy Nguyen, Hao Shao, Bin Tong, Einoshin Suzuki
ISMIS2
2011 Compact Coding for Hyperplane Classifiers in Heterogeneous Environment
Hao Shao, Bin Tong, Einoshin Suzuki
ECML/PKDD (3)1
2011 Feature-based Inductive Transfer Learning through Minimum Encoding
abstract
This paper proposes an Extended Minimum Description Length Principle (EMDLP) for feature-based inductive transfer learning, in which both the source and the target data sets contain class labels and relevant features are transferred from the source domain to the target one. Despite numerous works on this topic, few of them have a solid theoretical framework and are parameter-free. Our EMDLP overcomes these flaws and allows us to evaluate the inferiority of the results of transfer learning with the add-sum of the code lengths of five components: the corresponding two hypotheses, the two data sets with the help of the hypotheses, and the set of the transferred features. We design a code book to build the connections between the source and the target tasks. Extensive experiments using both real and artificial data sets show that EMDLP is robust against noise and performs better on the classification accuracy than the state-of-the-art methods.
Hao Shao, Einoshin Suzuki
SDM1
2010 Semi-supervised Projection Clustering with Transferred Centroid Regularization
Bin Tong, Hao Shao, Bin-Hui Chou, Einoshin Suzuki
ECML/PKDD (3)2
2008 A chaotic Ant Colony Optimization method for scheduling a single batch-processing machine with non-identical job sizes
abstract
The problem of minimizing makespan on a single batch-processing machine with non-identical job sizes is strongly NP-hard. This paper proposes an ant colony optimization (ACO) algorithm with chaotic control to solve the problem. The metropolis criterion is adopted to select the paths of ants to escape immature convergence. In order to improve the solutions of ACO, a chaotic optimizer is designed and integrated into ACO to reinforce the capacity of global optimization. Batch first fit is introduced to decode the paths into feasible solutions of the problem. In the experiment, the instances of 24 levels are simulated and the results show that the proposed CACO outperforms genetic algorithm and simulated annealing on all the instances.
Bayi Cheng, Huaping Chen 0001, Hao Shao, Rui Xu 0004, George Q. Huang
IEEE Congress on Evolutionary Computation3
2003 Fast indexing for image retrieval based on local appearance with re-ranking
abstract
This paper describes an approach to retrieve images containing specific objects, scenes or buildings. The image content is captured by a set of local features. More precisely, we use so-called invariant regions. These are features with shapes that self-adapt to the viewpoint. The physical parts on the object surface that they carve out are the same in all views, even though the extraction proceeds from a single view only. The surface patterns within the regions are then characterized by a feature vector of moment invariants. Invariance is under affine geometric deformations and scaled color bands with an offset added. This allows regions from different views to be matched efficiently. An indexing technique based on vantage point tree organizes the feature vectors in such a way that a naive sequential search can be avoided. This results in sublinear computation times to retrieve images from a database. In order to get sufficient certainty about the correctness of the retrieved images, a method to increase the number of matched regions is introduced. This way, the system is both efficient and discriminant. It is demonstrated how scenes or buildings are recognized, even in case of partial visibility and under a large variety of viewing condition changes.
Hao Shao, Tomás Svoboda, Vittorio Ferrari, Tinne Tuytelaars, Luc Van Gool
ICIP (3)1
1999 Development & Application of Wall-Climbing Robots
abstract
In this paper we introduce wall-climbing robots of two different absorption schemes: one with single suction cup and the other with permanent magnetic crawlers. The former adopts an omnidirectional vehicle and can perform remote-control inspection of nuclear storage tanks. Moreover, it is further developed for cleaning both the ceramic tile and glass surfaces of high-rise buildings. The latter consists of two types of robots: one is for maintenance automation of storage tanks in petrochemical enterprises, which can perform operations of sand-blasting, spray-painting and inspection; and the other is for monitoring the boiler wall water tubing.
Yan Wang 0001, Dianguo Xu 0001, Yanzheng Zhao, Hao Shao, Xueshan Gao
ICRA5