Jiahao Tang

dblp:254/3250 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Charts to Code: A Hierarchical Benchmark for Multimodal Models
abstract
Jiahao Tang, Henry Hengyuan Zhao, Lijian Wu, Zijian Zhang, Yifei Tao, Dongxing Mao, Yang Wan, Jingru Tan, Min Zeng, Min Li, Alex Jinpeng Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiahao Tang, Hengyuan Zhao, Lijian Wu, Yifei Tao, Dongxing Mao, Yang Wan, Jingru Tan, Min Zeng 0004, Min Li 0007, Alex Jinpeng Wang
ACL (1)1
2026 GeoDiffuser: A geometry-aware extension of pretrained diffusion models for consistent multi-view synthesis
Jiahao Tang, Mingxuan Chen, Ying Li 0020, Zuolei Sun, Yongbin Gao
Neurocomputing1
2026 EDTC: Exact Triangle Counting for Dynamic Graphs on GPU
abstract
In the process of updating a dynamic graph, an update to one edge may result in the addition or deletion of multiple triangles, while an update to multiple edges may only result in the addition or deletion of a single triangle. Consequently, accurately counting triangles on a dynamic graph is a challenging undertaking. As dynamic graphs are continuously updated, the GPU's memory may be insufficient to accommodate the storage of larger graphs. This presents a challenge when the graph, which is constantly growing, cannot be stored. The hash-based and binary search-based triangle counting algorithm is regarded as the most efficient for static graphs. However, when vertices with high degrees are encountered, the hash-based triangle counting method results in significant memory wastage due to the traditional construction of a hash table, leading to a shortage of memory. This issue remains unresolved. In this paper a triangle counting system EDTC is developed for dynamic graphs while ensuring the accuracy of counting. The system addresses three main problems: (1) An efficient EHTC algorithm is introduced to rapidly and accurately count the number of triangles in a graph. (2) The concept of an Update Activation CSR(UA-CSR) is introduced, along with a data structure to facilitate its implementation. This structure loads only the subgraph portion affected by the updated edge into the GPU, allowing calculations to be performed on this specific subgraph. (3) A compressed hash table is designed to reduce memory consumption, along with a dynamic shared memory assignment(DSA) strategy to fully utilize the shared memory of the GPU.
Jiahao Tang, Jinxing Tu, Wei Xue 0003, Jianqiang Huang 0002
IEEE Trans. Parallel Distributed Syst.2
2025 Enhancing Visual Understanding in Multimodal Large Language Models with Efficient Feature Alignment and State Space Models
abstract
Multimodal Large Language Models (MLLMs) excel at processing complex tasks involving visual and textual data. However, existing Mamba-based MLLMs face significant challenges in visual feature extraction, resulting in poor cross-modal alignment and compromised performance. To tackle these issues, we present ML-Mamba, a novel architecture built on the Mamba framework to enhance multimodal learning. ML-Mamba features a robust visual encoder, a Mamba-Transformer Projector for optimizing feature alignment and interaction between modalities, and the advanced Mamba large language model. By utilizing techniques such as cluster-based scanning for improved visual feature extraction and a Shared-Specialized feed-forward mechanism, ML-Mamba enhances visual representation quality and overall model efficiency. Extensive benchmarking shows that ML-Mamba outperforms existing models in multimodal tasks, significantly improving inference speed and cross-modal alignment. This work underscores the potential of integrating structured state-space models with advanced transformer components to develop scalable and resource-efficient multimodal models.
Jiakai Pan, Jiahao Tang, Yifei Xing 0001, Zhengzhuo Wang, Shengzhi Shen, Jianguo Hu
ECAI3
2025 VisCompConText: Scaling Multi-Modal Contexts via Visual Token Compression and Language Model Guidance
abstract
Training multimodal models with extended text contexts faces significant challenges due to prohibitive GPU memory consumption and computational costs from traditional tokenization methods, compounded by limited semantic alignment between visual and linguistic representations. Our approach introduces two key innovations: (1) Visual tagging for processing long contexts: This component adaptively renders long text segments into spatially efficient visual tokens, significantly reducing GPU memory usage and floating-point operations (FLOPs) during both training and inference. (2) LLM-Guided Visual Encoder: By leveraging large language models (LLMs), we enhance the visual encoder’s ability to comprehend long-form text, and overcome the limited long-context semantic awareness of visual encoders. Experimental results demonstrate that VisCompConText complements existing methods for extending text context length, improving document understanding, and showing strong potential in long text Q&A tasks. This work presents the first systematic solution for efficient long-context multimodal learning without sacrificing semantic granularity.
Zhengzhuo Wang, Jiakai Pan, Shengzhi Shen, Jiahao Tang, Chaoxing Zhou, Jianguo Hu
ECAI6
2025 A U-net and Transformer Paralleled Network for Robust Cuffless Blood Pressure Estimation Based on CWT-transformed PPG Images
abstract
This paper proposes a parallel U-Net and Transformer network for blood pressure (BP) estimation using Photoplethysmography (PPG) signals transformed into timefrequency images using CWT. It combines the U-Net’s excellent capability of extracting local features with the Transformer’s advantage in capturing long-range dependencies when processing sequential data. Specifically, the PPG signal is first segmented into 5-second windows. These segments are then transformed into 2-D images using continuous wavelet transform (CWT), incorporating both the real and imaginary components of the signal fragments. These images are used as inputs to the deep-learning network. For continuous BP signals, the peaks of the waveform are extracted, and the 5-second average of systolic blood pressure (SBP) and diastolic blood pressure (DBP) are computed to serve as the labels and references. The effectiveness and robustness of the proposed model are evaluated on a dataset constructed by us consisting of 16 subjects under 3 different conditions (Sit, Ex and Db), and the results are compared with traditional PPG-based BP estimation methods. Experimental results demonstrate an average mean absolute error (MAE) of 6.99 mmHg for SBP and 7.27 mmHg for DBP. Specifically, our approach achieved a BHS Grade A on the sit dataset and a BHS Grade C on both the Ex and Db datasets. Cross dataset validation is further conducted to prove the effectiveness of our model.
Boyuan Gu, Jiahao Tang, Yunhan Tang, Changting Xie
SMC2
2025 UDA-DDA: Unsupervised domain adaptation with dynamic distribution alignment network for emotion recognition using EEG signals
Jiahao Tang, Youjun Li, Chun-Wang Su, Xiangting Fan, Yangxuan Zheng, Hadia Naeem, Nan Yao, Zi-Gang Huang
Neurocomputing1
2025 A Lensless Multifocal Fusion Imaging Method Focused on Microfluidic Chips for IoMT-Based Cell Analysis System
abstract
The combination of lensless diffraction imaging and microfluidics provides a relevant solution for the lab-on-a-chip analysis of cells in the context of the Internet of Medical Things (IoMT). However, the cells are on multiple planes in microfluidics, and they only can be sensed by means of complex algorithms. Taking into account the limited resources available on edge computing, a lightweight multifocal fusion reconstruction algorithm is proposed in this work, which still keeps a high-quality reconstruction. The reconstruction, autofocus, and multifocal fusion are integrated, and the reference light normalization and the object light extraction are achieved to adapt their combination. The integrated process is divided into two stages: the cell diffraction image segmentation and the multifocal fusion, based on mono-focal object light reconstruction and masking, respectively. The single autofocus with the golden section search algorithm is merged into the two stages. Finally, a validation system is built based on an FPGA, with a size of only 45 mm × 45 mm × 75 mm. The average error for 8 μm diameter microspheres is 4.67%. For the white blood cells, the evaluation functions of the criteria, the energy of gradient (EOG), and the variance are improved by 12.2%, 92.3%, and 57.3%, respectively. Under the same accuracy, the global number of searches for focus is reduced by 47%-59%, compared to the two-step golden section method, whereas the global speed is also improved by 53%.
Dian Tian, Ningmei Yu, Liangchen Lv, Jiahao Tang, Jihui Yu, Álvaro Hernández, Jesús Ureña
IEEE Internet Things J.4
2024 StepTC: Stepwise Triangle Counting on GPU with Two Efficient Set Intersection Methods
Jiahao Tang
DASFAA (4)1
2024 A Low-Latency Power Series Approximate Computing and Architecture for Co-Calculation of Division and Square Root
abstract
The calculation of division and square root is widely used in edge computing related to image processing, clustering, recognition, and reconstruction, among others. Their common multi-stage serial calculation takes longer, and requires a higher latency, and more redundant hardware resources for multi-dimensional parallel computations. This work proposes a low-latency power series approximate digital computing (PSADIC) paradigm and architecture, which achieves a fast low-latency calculation of multi-dimensional mathematical expressions, focusing on the co-calculation of division and square root. This approach allows to compute not only the division and the square root, but also the inverse and the inverse square root. Moreover, serial and pipelined architectures have been designed here to achieve smaller areas and lower latencies, respectively. Compared to the multi-stage calculation with CORDIC (COordinate Rotation DIgital Computer), the mean relative error distance (MRED) is reduced by 67% in PSADIC. Under TSMC (Taiwan Semiconductor Manufacturing Company) 40nm CMOS technology, the proposed serial and pipelined architectures achieve a 66.67% overall latency reduction compared to the CORDIC when using both its own maximum clock frequency, and a 75% cycle latency reduction. Meanwhile, both the ADP (Area Delay Product) and PDP (Power Delay Product) are also optimized.
Dian Tian, Ningmei Yu, Minghui Xie, Jiahao Tang, Zhuang Feng, Álvaro Hernández, Jesús Ureña
IEEE Trans. Circuits Syst. I Regul. Pap.4
2023 PAS: A new powerful and simple quantum computing simulator
abstract
Abstract In recent years, many researchers have been using CPU for quantum computing simulation. However, in reality, the simulation efficiency of the large‐scale simulator is low on a single node. Therefore, striving to improve the simulator efficiency on a single node has become a serious challenge that many researchers need to solve. After many experiments, we found that much computational redundancy and frequent memory access are important factors that hinder the efficient operation of the CPU. This paper proposes a new powerful and simple quantum computing simulator: PAS (power and simple). Compared with existing simulators, PAS introduces four novel optimization methods: efficient hybrid vectorization, fast bitwise operation, memory access filtering, and quantum tracking. In the experiment, we tested the QFT (quantum Fourier transform) and RQC (random quantum circuits) of 21 to 30 qubits and selected the state‐of‐the‐art simulator QuEST (quantum exact simulation toolkit) as the benchmark. After experiments, we have concluded that PAS compared with QuEST can achieve a mean speedup of (QFT), (RQC) (up to , ) on the Intel Xeon E5‐2670 v3 CPU.
Haodong Bian, Jianqiang Huang 0002, Jiahao Tang, Runting Dong, Xiaoying Wang 0002
Softw. Pract. Exp.3
2022 A Two-stage Algorithm Based on Prediction and Search for Maxk-Truss Decomposition
abstract
The cohesive subgraph K-truss is often used to detect communities in large-scale social networks. The truss decomposition algorithm calculates the number of triangles formed by each edge, then peels off the edges with the number of triangles less than k-2 iteratively. The incremental maxk-truss algorithm must detect k-truss one by one making it inefficient. We propose a two-stage maxk-ktruss decomposition algorithm PSKT. The innovation points of PSKT are as follows: 1) The k-core structure is used to predict the maxk value, and the structure helps us avoid unnecessary mass calculations. 2) The characteristic that the k-truss of graph G must be included in the (k-l)-core of graph G is utilized to ensure that the proposed algorithm can safely prune in the graph. 3) A suitable triangle counting algorithm and the dynamic stream compression method are used to improve the algorithm efficiency. Comprehensive experiments demonstrate that PSKT is 15.26 times faster than the incremental algorithm FMT-dec and 3 times faster than the state-of-the-art maxk-truss algorithm FMT-max.
Jiahao Tang, Lingbin Liu, Jinfang Jia, Xiaoying Wang 0002, Jianqiang Huang 0002
ICPADS1
2022 Explore-Bench: Data Sets, Metrics and Evaluations for Frontier-based and Deep-reinforcement-learning-based Autonomous Exploration
abstract
Autonomous exploration and mapping of unknown terrains employing single or multiple robots is an essential task in mobile robotics and has therefore been widely investigated. Nevertheless, given the lack of unified data sets, metrics, and platforms to evaluate the exploration approaches, we develop an autonomous robot exploration benchmark en-titled Explore-Bench. The benchmark involves various explo-ration scenarios and presents two types of quantitative metrics to evaluate exploration efficiency and multi-robot cooperation. Explore-Bench is extremely useful as, recently, deep rein-forcement learning (DRL) has been widely used for robot exploration tasks and achieved promising results. However, training DRL-based approaches requires large data sets, and additionally, current benchmarks rely on realistic simulators with a slow simulation speed, which is not appropriate for training exploration strategies. Hence, to support efficient DRL training and comprehensive evaluation, the suggested Explore-Bench designs a 3-level platform with a unified data flow and 12 × speed-up that includes a grid-based simulator for fast evaluation and efficient training, a realistic Gazebo simulator, and a remotely accessible robot testbed for high-accuracy tests in physical environments. The practicality of the proposed benchmark is highlighted with the application of one DRL-based and three frontier-based exploration approaches. Fur-thermore, we analyze the performance differences and provide some insights about the selection and design of exploration methods. Our benchmark is available at https://github.com/efc-robot/Explore-Bench.
Yuanfan Xu, Jiahao Tang, Jiantao Qiu, Jian Wang 0030, Yuan Shen 0001, Yu Wang 0002, Huazhong Yang
ICRA3