EDBT 2026 Demo / reviewers in the wild / expert
Hao Jiang 0014
dblp:38/6049-14
· DBLP profile ↗
29ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0002-7323-4751ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 10 since 2021Systems, architecture and hardware · 8 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MAU-GPT: Enhancing Multi-type Industrial Anomaly Understanding via Anomaly-aware and Generalist Experts AdaptationabstractAs industrial manufacturing scales, automating fine-grained product image analysis has become critical for quality control. However, existing approaches are hindered by limited dataset coverage and poor model generalization across diverse and complex anomaly patterns. To address these challenges, we introduce MAU-Set, a comprehensive dataset for Multi-type industrial Anomaly Understanding. It spans multiple industrial domains and features a hierarchical task structure, ranging from binary classification to complex reasoning. Alongside this dataset, we establish a rigorous evaluation protocol to facilitate fair and comprehensive model assessment. Building upon this foundation, we further present MAU-GPT, a domain-adapted multimodal large model specifically designed for industrial anomaly understanding. It incorporates a novel AMoE-LoRA mechanism that unifies anomaly-aware and generalist experts adaptation, enhancing both understanding and reasoning across diverse defect classes. Extensive experiments show that MAU-GPT consistently outperforms prior state-of-the-art methods across all domains, demonstrating strong potential for scalable and automated industrial inspection. Zhuonan Wang, Zhenxuan Fan, Siwen Tan, Yuqian Yuan, Haoyuan Li 0002, Hao Jiang 0014, Wenqiao Zhang, Feifei Shao, Jun Xiao 0001 |
AAAI | 7 |
| 2026 | CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory AugmentationabstractWhile previous multimodal slow-thinking methods have demonstrated remarkable success in single-image understanding scenarios, their effectiveness becomes fundamentally constrained when extended to more complex multi-image comprehension tasks. This limitation stems from their predominant reliance on text-based intermediate reasoning processes. While for human, when engaging in sophisticated multi-image analysis, they typically perform two complementary cognitive operations: (1) continuous cross-image visual comparison through region-of-interest matching, and (2) dynamic memorization of critical visual concepts throughout the reasoning chain. Motivated by these observations, we propose the Complex Multi-Modal Chain-of-Thought (CMMCoT) framework, a multi-step reasoning framework that mimics human-like "slow thinking" for multi-image understanding. Our approach incorporates two key innovations: (1) The construction of interleaved multimodal multi-step reasoning chains, which utilize critical visual region tokens, extracted from intermediate reasoning steps, as supervisory signals. This mechanism not only facilitates comprehensive cross-modal understanding but also enhances model interpretability. (2) The introduction of a test-time memory augmentation module that expands the model’s reasoning capacity during inference while preserving parameter efficiency. Furthermore, to facilitate research in this direction, we have curated a novel multi-image slow-thinking dataset. Extensive experiments demonstrate the effectiveness of our model. Yan Xia 0006, Mushui Liu, Zhelun Yu, Haoyuan Li 0002, Wanggui He, Dong She, Yi Wang 0068, Hao Jiang 0014 |
AAAI | 10 |
| 2025 | MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image SynthesisabstractAuto-regressive models have made significant progress in the realm of text-to-image synthesis, yet devising an appropriate model architecture and training strategy to achieve a satisfactory level remains an important avenue of exploration. In this work, we introduce MARS, a novel framework for T2I generation that incorporates a specially designed Semantic Vision-Language Integration Expert (SemVIE). This innovative component integrates pre-trained LLMs by independently processing linguistic and visual information—freezing the textual component while fine-tuning the visual component. This methodology preserves the NLP capabilities of LLMs while imbuing them with exceptional visual understanding. Building upon the powerful base of the pre-trained Qwen-7B, MARS stands out with its bilingual generative capabilities corresponding to both English and Chinese language prompts and the capacity for joint image and text generation. The flexibility of this framework lends itself to migration towards any-to-any task adaptability. Furthermore, MARS employs a multi-stage training strategy that first establishes robust image-text alignment through complementary bidirectional tasks and subsequently concentrates on refining the T2I generation process, significantly augmenting text-image synchrony and the granularity of image details. Notably, MARS requires only 9% of the GPU days needed by SD1.5, yet it achieves remarkable results across a variety of benchmarks, illustrating the training efficiency and the potential for swift deployment in various applications. Wanggui He, Siming Fu, Mushui Liu, Xierui Wang, Wenyi Xiao, Fangxun Shu, Yi Wang 0068, Lei Zhang 0006, Zhelun Yu, Haoyuan Li 0002, Ziwei Huang 0005, Leilei Gan, Hao Jiang 0014 |
AAAI | 13 |
| 2025 | Resolving Multi-Condition Confusion for Finetuning-Free Personalized Image GenerationabstractPersonalized text-to-image generation methods can generate customized images based on the reference images, which have garnered wide research interest. Recent methods propose a finetuning-free approach with a decoupled cross-attention mechanism to generate personalized images requiring no test-time finetuning. However, when multiple reference images are provided, the current decoupled cross-attention mechanism encounters the object confusion problem and fails to map each reference image to its corresponding object, thereby seriously limiting its scope of application. To address the object confusion problem, in this work we investigate the relevance of different positions of the latent image features to the target object in diffusion model, and accordingly propose a weighted-merge method to merge multiple reference image features into the corresponding objects. Next, we integrate this weighted-merge method into existing pre-trained models and continue to train the model on a multi-object dataset constructed from the open-sourced SA-1B dataset. To mitigate object confusion and reduce training costs, we propose an object quality score to estimate the image quality for the selection of high-quality training samples. Furthermore, our weighted-merge training framework can be employed on single-object generation when a single object has multiple reference images. The experiments verify that our method achieves superior performance to the state-of-the-arts on the Concept101 dataset and DreamBooth dataset of multi-object personalized image generation, and remarkably improves the performance on single-object personalized image generation. Qihan Huang, Siming Fu, Hao Jiang 0014, Yipeng Yu, Jie Song 0011 |
AAAI | 4 |
| 2025 | Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI FeedbackabstractThe rapidly developing Large Vision Language Models (LVLMs) still face the hallucination phenomena where the generated responses do not align with the given contexts, significantly restricting the usages of LVLMs. Most previous work detects and mitigates hallucination at the coarse-grained level or requires expensive annotation (e.g., labeling by human experts or proprietary models). To address these issues, we propose detecting and mitigating hallucinations in LVLMs via fine-grained AI feedback. The basic idea is that we generate a small-size sentence-level hallucination annotation dataset by proprietary models, whereby we train a detection model which can perform sentence-level hallucination detection. Then, we propose a detect-then-rewrite pipeline to automatically construct preference dataset for hallucination mitigation training. Furthermore, we propose differentiating the severity of hallucinations, and introducing a Hallucination Severity-Aware Direct Preference Optimization (HSA-DPO) which prioritizes the mitigation of critical hallucination in LVLMs by incorporating the severity of hallucinations into preference learning. Extensive experiments on hallucination detection and mitigation benchmarks demonstrate that our method sets a new state-of-the-art in hallucination detection on MHaluBench, surpassing GPT-4V and Gemini, and reduces the hallucination rate by 36.1% on AMBER and 76.3% on Object HalBench compared to the base model. Wenyi Xiao, Ziwei Huang 0005, Leilei Gan, Wanggui He, Haoyuan Li 0002, Zhelun Yu, Fangxun Shu, Hao Jiang 0014, Linchao Zhu |
AAAI | 8 |
| 2025 | T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive ConceptsabstractZiwei Huang, Wanggui He, Quanyu Long, Yandi Wang, Haoyuan Li, Zhelun Yu, Fangxun Shu, Weilong Dai, Hao Jiang, Fei Wu, Leilei Gan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ziwei Huang 0005, Wanggui He, Quanyu Long, Yandi Wang, Haoyuan Li 0002, Zhelun Yu, Fangxun Shu, Weilong Dai, Hao Jiang 0014, Fei Wu 0001, Leilei Gan |
ACL (1) | 9 |
| 2025 | TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and CompetitionabstractWhile Parameter-Efficient Fine-Tuning (PEFT) methods like Low-Rank Adaptation (LoRA) effectively address resource constraints during fine-tuning, their performance often falls short, especially in multidimensional task scenarios. To address this issue, one straightforward solution is to introduce task-specific LoRA as domain experts, leveraging the modeling of multiple capabilities of experts and thus enhancing the general capability of multi-task learning.Although promising, these additional components often add complexity to the training and inference process, contravening the efficiency that PEFT is designed to deliver. Considering this, we introduce an innovative PEFT method, TeamLoRA, consisting of a collaboration and competition module for LoRA experts, thus achieving the right balance of effectiveness and efficiency:(i) For collaboration, we introduce a novel knowledge sharing and organization mechanism designed to optimize hierarchical learning while enhancing the efficiency of model training and inference.(ii) For competition, we propose leveraging a game-theoretic interaction mechanism for experts, encouraging experts to transfer their domain-specific knowledge while facing diverse downstream tasks, thus enhancing the performance.By doing so, TeamLoRA elegantly connects the experts as a “Team” with internal collaboration and competition, enabling a faster and more accurate PEFT paradigm. Meanwhile, we curate a Comprehensive Multi-Task Evaluation (CME) benchmark to thoroughly assess the capability of multi-task learning. Experiments conducted on our CME and other benchmarks indicate the effectiveness and efficiency of TeamLoRA. Our project is available at https://github.com/DCDmllm/TeamLoRA. Tianwei Lin 0001, Wenqiao Zhang, Haoyuan Li 0002, Zhelun Yu, Wanggui He, Juncheng Li 0006, Jiannan Guo 0003, Hao Jiang 0014, Siliang Tang, Yueting Zhuang |
ACL (1) | 10 |
| 2025 | Streaming Video Question-Answering with In-context Video KV-Cache RetrievalabstractWe propose ReKV, a novel training-free approach that enables efficient streaming video question-answering (StreamingVQA), by seamlessly integrating with existing Video Large Language Models (Video-LLMs). Traditional VideoQA systems struggle with long videos, as they must process entire videos before responding to queries, and repeat this process for each new question. In contrast, our approach analyzes long videos in a streaming manner, allowing for prompt responses as soon as user queries are received. Building on a common Video-LLM, we first incorporate a sliding-window attention mechanism, ensuring that input frames attend to a limited number of preceding frames, thereby reducing computational overhead. To prevent information loss, we store processed video key-value caches (KV-Caches) in RAM and disk, reloading them into GPU memory as needed. Additionally, we introduce a retrieval method that leverages an external retriever or the parameters within Video-LLMs to retrieve only query-relevant KV-Caches, ensuring both efficiency and accuracy in question answering. ReKV enables the separation of video analyzing and question-answering across different processes and GPUs, significantly enhancing the efficiency of StreamingVQA. Through comprehensive experimentation, we validate the efficacy and practicality of our approach, which significantly boosts efficiency and enhances applicability over existing VideoQA models. Shangzhe Di, Zhelun Yu, Haoyuan Li 0002, Bolin Li, Wanggui He, Fangxun Shu, Hao Jiang 0014 |
ICLR | 10 |
| 2025 | HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge AdaptationabstractWe present **HealthGPT**, a powerful Medical Large Vision-Language Model (Med-LVLM) that integrates medical visual comprehension and generation capabilities within a unified autoregressive paradigm. Our bootstrapping philosophy is to progressively adapt heterogeneous comprehension and generation knowledge to pre-trained Large Language Models (LLMs). This is achieved through a novel heterogeneous low-rank adaptation **(H-LoRA)** technique, which is complemented by a tailored hierarchical visual perception **(HVP)** approach and a three-stage learning strategy **(TLS)**. To effectively learn the HealthGPT, we devise a comprehensive medical domain-specific comprehension and generation dataset called **VL-Health**. Experimental results demonstrate exceptional performance and scalability
of HealthGPT in medical visual unified tasks. Our project can be accessed at https://github.com/DCDmllm/HealthGPT. Tianwei Lin 0001, Wenqiao Zhang, Sijing Li, Yuqian Yuan, Binhe Yu, Haoyuan Li 0002, Wanggui He, Hao Jiang 0014, Mengze Li 0001, Siliang Tang, Jun Xiao 0001, Yueting Zhuang, Beng Chin Ooi |
ICML | 8 |
| 2024 | On the Evaluation Consistency of Attribution-Based Explanations
Jiarui Duan, Haoling Li, Haofei Zhang, Hao Jiang 0014, Mengqi Xue, Mingli Song, Jie Song 0011 |
ECCV (70) | 4 |
| 2024 | LiDUT-Depth: A Lightweight Self-supervised Depth Estimation Model Featuring Dynamic Upsampling and Triplet Loss Optimization
Hao Jiang 0014, Zhijun Fang 0001, Xuan Shao, Jenq-Neng Hwang |
ICPR (16) | 1 |
| 2024 | LG-CAV: Train Any Concept Activation Vector with Language GuidanceabstractConcept activation vector (CAV) has attracted broad research interest in explainable AI, by elegantly attributing model predictions to specific concepts. However, the training of CAV often necessitates a large number of high-quality images, which are expensive to curate and thus limited to a predefined set of concepts. To address this issue, we propose Language-Guided CAV (LG-CAV) to harness the abundant concept knowledge within the certain pre-trained vision-language models (e.g., CLIP). This method allows training any CAV without labeled data, by utilizing the corresponding concept descriptions as guidance. To bridge the gap between vision-language model and the target model, we calculate the activation values of concept descriptions on a common pool of images (probe images) with vision-language model and utilize them as language guidance to train the LG-CAV. Furthermore, after training high-quality LG-CAVs related to all the predicted classes in the target model, we propose the activation sample reweighting (ASR), serving as a model correction technique, to improve the performance of the target model in return. Experiments on four datasets across nine architectures demonstrate that LG-CAV achieves significantly superior quality to previous CAV methods given any concept, and our model correction method achieves state-of-the-art performance compared to existing concept-based methods. Our code is available at https://github.com/hqhQAQ/LG-CAV. Qihan Huang, Jie Song 0011, Mengqi Xue, Haofei Zhang, Bingde Hu, Huiqiong Wang, Hao Jiang 0014, Xingen Wang, Mingli Song |
NeurIPS | 7 |
| 2023 | Quality Assessment for High Dynamic Range Stereoscopic Omnidirectional Image System
Liuyan Cao, Hao Jiang 0014, Zhidi Jiang, Jihao You, Mei Yu 0001, Gangyi Jiang |
ACIVS | 2 |
| 2022 | Multi-Angle Projection Based Blind Omnidirectional Image Quality AssessmentabstractMost of the existing blind omnidirectional image quality assessment (BOIQA) methods are based on data-driven approach where the end-to-end neural network or deep learning tools are mainly used for feature extraction. However, it usually lacks interpretability and is difficult to discover the perceptual mechanism behind. In this paper, from the perspective of perception modeling, we propose a novel multi-angle projection based BOIQA (MP-BOIQA) method. Considering the omnibearing and near eye display characteristics with head mounted display, multiple color cubemap projection images with respect to different viewpoints are grouped as the color omnidirectional distortion (COD) units so as to simulate the user’s viewing behavior in subjective quality assessment. In the designed multi-angle projection based feature extractor, tensor decomposition is implemented on each COD unit for dimensionality reduction, and piecewise exponential fitting is used to get the distribution of mean subtracted contrast normalized coefficients of the unit’s feature matrices in tensor domain. Finally, the extracted features are pooled with random forest. The experimental results on three omnidirectional image quality datasets show that the MP-BOIQA method can deliver highly competitive performance compared with some representative full-reference quality assessment methods, as well as some state-of-the-art BOIQA methods. Hao Jiang 0014, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Haiyong Xu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Electromagnetic Imaging of Uniaxial Objects by Artificial Intelligence TechnologyabstractThe electromagnetic imaging of uniaxial objects by the Artificial Intelligence (AI) technology is presented in this paper. We study the two-dimensional inverse scattering problem from uniaxial objects illuminated by the TM (Transverse Magnetic) and TE (Transverse Electric) polarized incident waves. As the uniaxial objects have different components of permittivity along different transverse directions, the problem of TE polarization will be more severe than that of TM polarization. We use the Dominant Current Scheme (DCS) and Back Propagation Scheme (BPS) to calculate the preliminary permittivity distribution. By combining with deep learning and neural networks, the permittivity distribution of those uniaxial objects can be reconstructed more accurately. U-Net is used to reconstruct the permittivity distribution, because U-Net has shared the weights and biases, which can effectively reduce the network complexity and is very suitable for solving image processing problems. In the numerical results, we added different noises to compare the reconstruction results of the DCS and BPS initial estimations through the U-Net. Numerical results show that the reconstruction permittivity for the DCS initial estimation is better than that for the BPS’s. Our diversity is that we have reconstructed the uniaxial objects by neural network successfully with less time-consuming effort and real-time imaging. Chien-Ching Chiu, Po-Hsiang Chen, Hao Jiang 0014 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Cubemap-Based Perception-Driven Blind Quality Assessment for 360-degree Imagesabstractimage can be represented with different formats, such as the equirectangular projection (ERP) image, viewport images or spherical image, for its different processing procedures and applications. Accordingly, the 360-degree image quality assessment (360-IQA) can be performed on these different formats. However, the performance of 360-IQA with the ERP image is not equivalent with those with the viewport images or spherical image due to the over-sampling and the resulted obvious geometric distortion of ERP image. This imbalance problem brings challenge to ERP image based applications, such as 360-degree image/video compression and assessment. In this paper, we propose a new blind 360-IQA framework to handle this imbalance problem. In the proposed framework, cubemap projection (CMP) with six inter-related faces is used to realize the omnidirectional viewing of 360-degree image. A multi-distortions visual attention quality dataset for 360-degree images is firstly established as the benchmark to analyze the performance of objective 360-IQA methods. Then, the perception-driven blind 360-IQA framework is proposed based on six cubemap faces of CMP for 360-degree image, in which human attention behavior is taken into account to improve the effectiveness of the proposed framework. The cubemap quality feature subset of CMP image is first obtained, and additionally, attention feature matrices and subsets are also calculated to describe the human visual behavior. Experimental results show that the proposed framework achieves superior performances compared with state-of-the-art IQA methods, and the cross dataset validation also verifies the effectiveness of the proposed framework. In addition, the proposed framework can also be combined with new quality feature extraction method to further improve the performance of 360-IQA. All of these demonstrate that the proposed framework is effective in 360-IQA and has a good potential for future applications. Hao Jiang 0014, Gangyi Jiang, Mei Yu 0001, Yun Zhang 0002, You Yang 0002, Zongju Peng |
IEEE Trans. Image Process. | 1 |
| 2020 | Towards Context-aware Distributed Learning for CNN in Mobile ApplicationsabstractIntelligent mobile applications have been ubiquitous on mobile devices. These applications keep collecting new and sensitive data from different users while being expected to have the ability to continually adapt the embedded machine learning model to these newly collected data. To improve the quality of service while protecting users' privacy, distributed mobile learning (e.g., Federated Learning (FedAvg) [1]) has been proposed to offload model training from the cloud to the mobile devices, which enables multiple devices collaboratively train a shared model without leaking the data to the cloud. However, this design becomes impracticable when training the machine learning model (e.g., Convolutional Neural Network (CNN)) on mobile devices with diverse application context. For example, in conventional distributed training schemes, different devices are assumed to have integrated training datasets and train identical CNN model structures. Distributed collaboration between devices is implemented by a straightforward weight average of each identical local models. While, in mobile image classification tasks, different mobile applications have dedicated classification targets depending on individual users' preference and application specificity. Therefore, directly averaging the model weight of each local model will result in a significant reduction of the test accuracy. To solve this problem, we proposed CAD: a context-aware distributed learning framework for mobile applications, where each mobile device is deployed with a context-adaptive submodel structure instead of the entire global model structure. Zhuwei Qin, Hao Jiang 0014 |
SEC | 2 |
| 2020 | Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark StudyabstractExisting enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions. Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin |
IEEE Trans. Image Process. | 17 |
| 2018 | Pulse-Width Modulation based Dot-Product Engine for Neuromorphic Computing System using Memristor Crossbar ArrayabstractThe Dot-Product Engine (DPE) is a critical circuit for implementing neural networks in hardware. The recent-developed memristor crossbar array technology, which is able to efficiently carry out dot-product multiplication and update its weights in real time, has been considered as one of the viable technologies to build a high-efficient neural network computing system. In this paper, the Pulse-Width-Modulation (PWM) based DPE has been presented and analyzed. Here, the PWM based signal, instead of the traditional amplitude modulated (AM) signal, is used as the computation variable. Comparing to the existing AM based system, this PWM counterpart provides an alternative approach to reduce the power consumption and chip area of its peripheral circuits. Power and area saving becomes more prominent when the size and/or the number of arrays increase. This new approach also provides the critically needed scalability to accommodate the computation variable with higher precision. In this paper, a 4-bit (can be easily expanded to 8-bit) feed forward neural network with 3-bit weights (memristor's conductance) is constructed using the proposed PWM DPE to identify digits from the MNIST data set. The circuit system is implemented in 130 nm standard CMOS technology. The entire circuit system consumes about 53mW with more than 86% recognition accuracy in average. Hao Jiang 0014, Kevin Yamada, Zizhe Ren, Thomas Kwok, Fu Luo, Qing Yang 0011, J. Joshua Yang, Qiangfei Xia, Yiran Chen 0001, Hai Li 0001, Qing Wu 0002, Mark Barnell |
ISCAS | 1 |
| 2018 | 3D visual discomfort predictor based on subjective perceived-constraint sparse representation in 3D display system
Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Zongju Peng, Feng Shao 0001, Hao Jiang 0014 |
Future Gener. Comput. Syst. | 7 |
| 2017 | A memristor-based neuromorphic engine with a current sensing scheme for artificial neural network applicationsabstractBy following the big data revolution, neuromorphic computing makes a comeback for its great potential in information processing capability. Despite of many types of architectures reported in conventional CMOS domain, memristor, as an example of emerging devices, demonstrates an intrinsic support of parallel matrix-vector multiplication operation that is widely used in artificial neural network applications. However, its computation accuracy and speed are far from satisfactory, mainly constrained by the features of memristor crossbar array and peripheral circuitry. In this work, we propose a new memristor crossbar based computing engine design by leveraging a current sensing scheme. High parallelism in operation and therefore fast computation can be achieved via simultaneously supplying analog voltages into a memristor crossbar and directly converting the weighted current through a current-to-voltage converter. We implemented and compared the feed-forward neural networks with different array sizes and layer numbers. Our design demonstrates a good computation accuracy, e.g., 96.6% classification accuracy for MNIST handwritten digit in a two-layer design. Qing Yang 0011, Hao Jiang 0014, Qing Wu 0002, Hai Li 0001 |
ASP-DAC | 4 |
| 2017 | A new tone-mapped image quality assessment approach for high dynamic range imaging systemabstractTone-mapping operators are designed to apply high dynamic range (HDR) images on widely-used low dynamic range (LDR) devices. Developing well-performed tone-mapped image quality assessment (IQA) method is highly desired because traditional IQA method cannot be adopted in cross dynamic range quality measuring. To this end, we proposed a quality assessment method based on image exposure property. Specifically, an image exposure property determination model is utilized to segment HDR image into different exposure region. Then, quality features are extracted according to the distortion characteristics of each exposure region. Finally, the quality of tone-mapped image can be acquired by a trained regression model. Validation experiments on public database show that the proposed method can accurately predict the quality of tone-mapped image. Yang Song 0015, Gangyi Jiang, Hao Jiang 0014, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
ICIP | 3 |
| 2017 | A Mismatch Detection Method Based on Affine Transformation for Stereo Light Microscopy Stereo MatchingabstractFor the light microscopy images that have the characteristics of shallow depth of field, serious distortion and poor resolution, mismatch is a ubiquitous phenomenon. The paper presents a mismatch detection method for the stereo light microscopy stereo matching. Affine transformation matrix and matching constraint condition are calibrated by the calibration board which has the precision solid dots and the motorized stage. Bias vector of affine transformation of each matching pair is taken as the criteria to apply mismatch detection. The experimental results show that the method can detect more mismatching pairs and preserve more matching pairs than the traditional RANSAC method and the epipolar rectification method. Shengli Fan, Mei Yu 0001, Gangyi Jiang, Yigang Wang, Hao Jiang 0014 |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2016 | Built-in selectors self-assembled into memristorsabstractWe demonstrate an approach to build a selector into ReRAM (memristors) using engineered materials. In this approach, a segment(s) of “nonlinear material” is self-assembled into the conduction channel (s) (filament) of a memristor. The nonlinear material exhibits a highly nonlinear current-voltage characteristic, which gives rise to a nonlinear i-v characteristic of the memristor in the ON state. Somnath Chakraborty 0002, Saumil Joshi, Qiangfei Xia, Hai Li 0001, Yiran Chen 0001, Hao Jiang 0014, Qing Wu 0002, Mark Barnell, J. Joshua Yang |
ISCAS | 6 |
| 2016 | Cyclical sensing integrate-and-fire circuit for memristor array based neuromorphic computingabstractThe brain-inspired, spike-based neuromorphic system is highly anticipated in the artificial intelligence community due to its high computational efficiency. The recently developed memristor-crossbar-array technology, which is able to efficiently emulate the plasticity of biological synapses and accommodate matrix multiplication, has demonstrated its potential for neuromorphic computing. To facilitate the computation, a high-speed integrate-and-fire circuit (IFC) and a counter were previously developed to efficiently convert the current from the memristor array into rate-coded spikes. However, the linear dynamic range of the circuit, which is limited by its responding speed, is challenged when the input intensity and the conductance of the memristor array are both high simultaneously. In this paper, a novel cyclical sensing scheme is developed that can significantly extend the linear dynamic range of the original IFC. Meanwhile, the power efficiency of the IFC can also be increased. The circuit simulation results indicated that the cyclical sensing IFC was able to efficiently and accurately facilitate the matrix multiplication when it was integrated with a 32×32 memristor crossbar array. With the optimized crossbar array structure and its peripheral circuits, the developed cyclical sensing IFC has shown great promise in accelerating matrix multiplication in spike-based computing systems. Hao Jiang 0014, Fu Luo, Kangjun Bai, J. Joshua Yang, Qiangfei Xia, Yiran Chen 0001, Qing Wu 0002 |
ISCAS | 1 |
| 2016 | MyoHMI: A low-cost and flexible platform for developing real-time human machine interface for myoelectric controlled applicationsabstractEMG pattern recognition has been studied for control of prostheses and rehabilitation systems for decades. Existing research platforms for developing EMG pattern recognition algorithms are typically based on MATLAB and the collection of EMG signals is often done by expensive, non-portable data acquisition systems. The requirement of these resources usually limits the use of these platforms in the lab environments and prohibits their widespread to other fields and applications. To address this limitation, this paper presents a low-cost, easy to use, and flexible platform called MyoHMI for developing real-time human machine interfaces for myoelectric controlled applications. MyoHMI facilitates the interface with a commercial EMG-based armband Myo, which costs less than $200 and can be easily worn by the user without the need of special preparation. MyoHMI also provides a highly modular and customizable C/C++ based software engine which seamlessly integrates a variety of interfacing and signal processing modules, from data acquisition through signal processing and pattern recognition, to real-time evaluation and control. The experimental results on able-bodied human subjects for controlling two evaluation platforms in real time verified the merit of the MyoHMI platform and demonstrated the feasibility of a low-cost solution for the development of myoelectric controlled applications. Ian Donovan, Kevin Valenzuela, Alejandro Ortiz, Sergey Dusheyko, Hao Jiang 0014, Kazunori Okada |
SMC | 5 |
| 2015 | Spiking-based matrix computation by leveraging memristor crossbar arrayabstractAs process technology continues scaling down, the memory barrier becomes more severe. Thus, spiking neuromorphic computing that can significantly enhance computing and communication efficiencies has been widely studied. Both conventional CMOS technology and emerging devices have been used in hardware implementation of spiking neuromorphic computing. Particularly, the memristor technology that can naturally emulate plasticity and energy efficiency of biological synapses have gained a lot of attention. However, the use of memristors in high density computation, such as matrix-vector operation, is still missing. In this work, a spiking (pulse-based) computing component that leverages memristor crossbar array is proposed for matrix-vector operation. We adopt the rate coding model and count the produced spike number during a given time period of T as the computation output. We carefully design the crossbar array structure and the integrate-and-fire circuit. The linear relationship between output spike numbers and the sum-of-production of input vector and matrix entries is observed in our simulation results. The proposed spiking computing design realizes matrix computation successfully and demonstrates good adaptability in neural network. Hai Li 0001, Bonan Yan, Chaofei Yang, Linghao Song, Yiran Chen 0001, Qing Wu 0002, Hao Jiang 0014 |
CISDA | 10 |
| 2015 | RENO: a high-efficient reconfigurable neuromorphic computing accelerator designabstractNeuromorphic computing is recently gaining significant attention as a promising candidate to conquer the well-known von Neumann bottleneck. In this work, we propose RENO -- a efficient reconfigurable neuromorphic computing accelerator. RENO leverages the extremely efficient mixed-signal computation capability of memristor-based crossbar (MBC) arrays to speedup the executions of artificial neural networks (ANNs). The hierarchically arranged MBC arrays can be configured to a variety of ANN topologies through a mixed-signal interconnection network (M-Net). Simulation results on seven ANN applications show that compared to the baseline general-purpose processor, RENO can achieve on average 178.4x (27.06x) performance speedup and 184.2x (25.23x) energy savings in high-efficient multilayer perception (high-accurate auto-associative memory) implementation. Moreover, in the comparison to a pure digital neural processing unit (D-NPU) and a design with MBC arrays co-operating through a digital interconnection network, RENO still achieves the fastest execution time and the lowest energy consumption with similar computation accuracy. Xiaoxiao Liu 0001, Mengjie Mao, Beiye Liu, Hai Li 0001, Yiran Chen 0001, Boxun Li, Yu Wang 0002, Hao Jiang 0014, Mark Barnell, Qing Wu 0002, J. Joshua Yang |
DAC | 8 |
| 2015 | A spiking neuromorphic design with resistive crossbarabstractNeuromorphic systems recently gained increasing attention for their high computation efficiency. Many designs have been proposed and realized with traditional CMOS technology or emerging devices. In this work, we proposed a spiking neuromorphic design built on resistive crossbar structures and implemented with IBM 130nm technology. Our design adopts a rate coding scheme where pre- and post-neuron signals are represented by digitalized pulses. The weighting function of pre-neuron signals is executed on the resistive crossbar in analog format. The computing result is transferred into digitalized output spikes via an integrate-and-fire circuit (IFC) as the post-neuron. We calibrated the computation accuracy of the entire system through circuit simulations. The results demonstrated a good match to our analytic modeling. Furthermore, we implemented both feedforward and Hopfield networks by utilizing the proposed neuromorphic design. The system performance and robustness were studied through massive Monte-Carlo simulations based on the application of digital image recognition. Comparing to the previous crossbar-based computing engine that represents data with voltage amplitude, our design can achieve >50% energy savings, while the average probability of failed recognition increase only 1.46% and 5.99% in the feedforward and Hopfield implementations, respectively. Bonan Yan, Chaofei Yang, Linghao Song, Beiye Liu, Yiran Chen 0001, Hai Li 0001, Qing Wu 0002, Hao Jiang 0014 |
DAC | 10 |