EDBT 2026 Demo / reviewers in the wild / expert
Yuhua Tang
dblp:41/556
· DBLP profile ↗
57ranked-venue papers
2as first author
31since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 5 since 2021Systems, architecture and hardware · 11 · 7 since 2021Databases, data management, data science and information retrieval · 8 · 6 since 2021Human-computer interaction and ubiquitous computing · 3Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fast Discovery of Functional Dependencies via Bayesian Network Learning
Shenglin Chen, Yuhua Tang, Ruochun Jin |
ICDE | 4 |
| 2026 | $L^{3}$ C: Leaf-Centric Continuous Codes for Natural Language-Driven Table Discovery
Ruochun Jin, Jixin Zhang, Yuhua Tang, Xiangyu Zhao 0001 |
ICDE | 4 |
| 2026 | Efficient Table Embeddings via Self-Supervised Structural-Semantic Graph Autoencoder
Jinlong Tian, Ruochun Jin, Yanfang Zhou, Xinhai Xu, Yuhua Tang |
Inf. Process. Manag. | 7 |
| 2025 | MagicNaming: Consistent Identity Generation by Finding a "Name Space" in T2I Diffusion ModelsabstractLarge-scale text-to-image diffusion models, (e.g., DALL-E, SDXL) are capable of generating famous persons by simply referring to their names. Is it possible to make such models generate generic identities as simple as the famous ones, e.g., just use a name? In this paper, we explore the existence of a ``Name Space'', where any point in the space corresponds to a specific identity. Fortunately, we find some clues in the feature space spanned by text embedding of celebrities' names. Specifically, we first extract the embeddings of celebrities' names in the Laion5B dataset with the text encoder of diffusion models. Such embeddings are used as supervision to learn an encoder that can predict the name (actually an embedding) of a given face image. We experimentally find that such name embeddings work well in promising the generated image with good identity consistency. Note that like the names of celebrities, our predicted name embeddings are disentangled from the semantics of text inputs, making the original generation capability of text-to-image models well-preserved. Moreover, by simply plugging such name embeddings, all variants (e.g., from Civitai) derived from the same base model (i.e., SDXL) readily become identity-aware text-to-image models. Heliang Zheng, Long Lan, Wanrong Huang, Yuhua Tang |
AAAI | 6 |
| 2025 | M2TQA: A Metacognitive Framework for Multi-Table Question Answering
Jinlong Tian, Yuhua Tang, Kejia Wan, Yanfang Zhou, Xinhai Xu |
CogSci | 2 |
| 2025 | Annotating Table Metadata with Knowledge-Enhanced Pre-Trained Language Model
Yuhua Tang, Jinlong Tian, Xudong Fang |
DASFAA (6) | 2 |
| 2025 | A Cost-effective Solution for Remote Sensing Image Segmentation via Train/Test-Time AdaptationabstractRemote Sensing Image (RSI) segmentation has made significant strides, emerging as a leading solution for interpreting remote sensing data. However, due to the substantial domain gap between different remote sensors and limited computational resources, existing RSI segmentation methods often suffer from poor generalization performance. To address these challenges, we propose a COst-effective and Whole-process Domain Adaptation solution, namely COWDA, which adapts models at both the train and test time through three key phases: 1) Source-data Domain Alignment: We employ traditional image stylization techniques to translate source images into target-style alternatives, which avoids the computationally intensive need for auxiliary neural network models. 2) Target-data Train-time Fine-tuning: We propose a joint positive and negative learning (JPNL) algorithm that adds both positive and negative samples to effectively learn domain-invariant knowledge from noisy pseudo-labeled target data. 3) Test-time Adaptation: We propose an entropy-weighted test-time adaptation strategy to update the trained model with online test samples, further enhancing its performance. Extensive experiments on two widely-used domain adaptation benchmarks for remote sensing show that COWDA improves state-of-the-art counterparts by 1.4% and 2.6% F1 scores, respectively. Wei Chen 0009, Xin Luo 0009, Yu-Lin He, Tianhang Guo, Yuhua Tang |
ICASSP | 8 |
| 2025 | Exploiting Foundation Models for Label-Efficient Few-Shot Learning via Feature Coupling: A Case Study of cardiac CT SegmentationabstractThe scarcity of labeled data poses a significant challenge for deep learning-based medical image segmentation. To address this, this study introduces the novel Foundation Model-based Few-Shot Segmentation (FM-FSS) paradigm. FM-FSS capitalizes on the knowledge distilled from pre-trained foundation models, such as the Segment Anything Model, to enhance segmentation performance in few-shot scenarios. The paradigm designs a feature coupling module that synergizes SAM’s powerful feature extraction capabilities with nnU-Net’s self-configuration strategy, enabling accurate segmentation with minimal labeled data and optional manual prompt inputs. Extensive experiments on a publicly available cardiac CT dataset demonstrate that FM-FSS outperforms state-of-the-art segmentation models. With only 20 labeled images, our method achieves an average Dice score of 94.33% and an ASD of 1.10 mm. Moreover, FM-FSS maintains its label-efficient performance in a one-shot setup, reducing the annotation requirements by at least fourfold. The code and pre-trained models will be released upon acceptance. Wei Chen 0009, Wenjuan Zhou, Tianhang Guo, Yuhua Tang |
ICASSP | 6 |
| 2025 | UniIVFT: Towards a Unified Framework for Infrared-Visible Fusion and TranslationabstractInfrared-visible image fusion (IVF) and infrared-to-visible image translation (I2V) are two closely related tasks in multimodal image processing, both aimed at combining or transforming infrared and visible modalities to enhance image information content. Existing methods typically focus on either fusion or translation, often requiring redundant construction of similar components for each task, which limits the effective utilization of cross-modal interactions and feature encoding capabilities. Furthermore, these approaches are often hindered by their reliance on complex feature extract models, limiting their overall effectiveness and adaptability. In this paper, we introduce the Unified Multimodal Infrared-Visible Image Fusion and Translation (UniIVFT) framework, which integrates both fusion and translation tasks within a single architecture. We employ a vision transformer (ViT) encoder-decoder structure augmented with task-specific tokens and introduce a contrastive loss to effectively align infrared and visible image features before multimodal encoding. This alignment enhances the encoder’s ability to capture cross-modal interactions. In UniIVFT, both IVF and I2V tasks share a unified encoder architecture and use task-specific tokens to control model outputs, reducing redundant model construction and training. Extensive experiments demonstrate that UniIVFT achieves performance on par with that of SOTAs across multiple tasks while maintaining a lightweight architecture with fewer model parameters. Xueqiong Li, Shaowu Yang, Huibin Tan, Yuhua Tang |
ICASSP | 5 |
| 2025 | Federated Dynamic Aggregation Learning Based on Parameter Decomposition to Combat Noisy Data
Xuyan Zhang, Da Huang 0002, Zhencheng Fan, Yuhua Tang, Xiyao Liu 0001 |
ICIC (15) | 4 |
| 2025 | AGFT-Tracker: Adaptive Game-Based PEFT for Object Tracking with PLMsabstractThe rise of pre-trained large models (PLMs) has sparked interest in vision tasks like object tracking. However, as PLMs scale, fully fine-tuning all parameters becomes impractical, highlighting the need for parameter-efficient fine-tuning (PEFT). While adapter tuning, which adds tunable parameters to Multi-Head Attention (MHA) or Feed-Forward Networks (FFN), is common, critical parameters like Layer Normalization (LN), vital for stability and convergence, are often overlooked. Furthermore, traditional fine-tuning strategies fail to differentiate module importance, limiting performance improvements. To solve these issues, we propose a new PEFT method for unlocking large model potential in object tracking: Adaptive Game-Based Fine-tuning Tracker (AGFT-Tracker). AGFT-Tracker combines adapter tuning with direct LN fine-tuning and adaptively allocates parameter budgets based on tracking attention losses. Important sensitive modules use higher-rank LoRA and frozen LN, while stable modules undergo lower-rank LoRA and LN adjustments. This approach improves effectiveness and efficiency, achieving state-of-the-art results on challenging benchmarks. Mingyu Cao, Xihuai He, Xueqiong Li, Kedi Zhang, Yuhua Tang, Wanrong Huang, Huibin Tan |
ICME | 5 |
| 2025 | Adaptive Distribution-Aware Modeling for Transformer TrackingabstractAdapting to changes in data distribution is a major challenge in visual object tracking. In Transformer-based tracking, Layer Normalization (LN) is often applied uniformly to both template and search features, limiting feature diversity. Additionally, models tend to converge to trivial solutions, and tracking samples are sensitive to distribution shifts, affecting robustness. To address these issues, we propose the Adaptive Distribution-Aware Transformer Tracker (ADAT), incorporating three key components: the Target-Aware Module (TAM), the Region-Aware Module (RAM), and the Self-Feedback-Aware Module (SFAM). TAM normalizes template and search features separately, preserving flexibility and enhancing target learning. RAM refines target perception by distinguishing between near and far target regions. SFAM filters out noisy samples and fine-tunes normalization parameters through self-feedback. While TAM and RAM regulate feature-level distribution, SFAM adjusts at the sample level. Extensive experiments show that ADAT outperforms existing methods, achieving superior performance on challenging benchmarks. Mingyu Cao, Huibin Tan, Xueqiong Li, Wanrong Huang, Kedi Zhang, Yuhua Tang, Shaowu Yang |
ICME | 6 |
| 2025 | Multi-Resolution Infrared-Visible Image Fusion using Multi-Scale Residual QuantizationabstractInfrared-visible image fusion (IVF) is an essential task in multimodal image processing that integrates infrared and visible modalities to enhance the overall image information content. However, existing methods often suffer from limited precision and efficiency. Furthermore, they fail to address practical requirements such as multi-resolution fusion and mutual translation. In this paper, we propose the Multi-Scale Residual Quantized Infrared-Visible Image Fusion (M-RQIVF) framework to efficiently generate high-quality fusion images. M-RQIVF trains multi-scale residual quantized infrared and visible autoencoders that convert images into multi-scale discrete token maps. This approach approximates the residuals from the features on a scale-by-scale basis, allowing for coarse-to-fine fused image generation that aligns well with human visual perception. Furthermore, by leveraging these discrete token maps, we train Visual Auto-Regressive (VAR) transformers using next-scale prediction. The VAR transformer ensures that features of corresponding sizes can be generated, even when the input infrared and visible images have different resolutions, facilitating fine-grained fusion. Additionally, the autoregressive structure enables image translation to be treated as a conditional generation task, thereby enabling mutual translation between infrared and visible images. Extensive experiments demonstrate that M-RQIVF outperforms the SOTAs while maintaining a much faster inference speed. Huibin Tan, Wanrong Huang, Yuhua Tang, Xueqiong Li |
ICME | 5 |
| 2025 | DFDUN: Deep Infrared and Visible Image Fusion with Diffusion Prior Unfolding NetworkabstractInfrared and Visible Image Fusion (IVF) intends to aggregate information from infrared and visible modalities, generating comprehensive images. While deep-unfolding-based methods and diffusion-based methods show satisfactory performances, the former suffers from insufficiently deterministic priors, and latter is limited by the inaccurate prior generation or opaque working mechanisms. In this paper, we propose Deep infrared and visible image Fusion with Diffusion prior Unfolding Network (DFDUN), aiming for effective fusion with the generative diffusion prior in a transparent mechanism. DFDUN starts with a model-based fusion optimization formulation, which is unfolded into a Denoising Diffusion Module (DDM) for generating informative diffusion prior, and a Data Consistent Module (DCM) that transparently and effectively aggregates complementary information from diffusion prior and source modalities. Moreover, DFDUN employs a hypernetwork for adaptive dictionary parameter generation in DCM, enhancing fusion flexibility. Experimental results indicate that DFDUN outperforms existing methods, providing superior IVF performance with efficient and transparent fusion. Code is available at https://github.com/XiongMaoyi2001/DFDUN. Maoyi Xiong, Tianrui Liu 0001, Xueqiong Li, Yuhua Tang |
ICME | 8 |
| 2025 | Intra- and Inter-Layer Scheduling Exploration and Optimization for ReRAM-Based DNN AcceleratorsabstractResistive Random Access Memory (ReRAM) based architectures have shown great potential for realizing energy-efficient Deep Neural Network (DNN) acceleration. When deploying a DNN, the ReRAM-based designs need a scheduling scheme to translate massive hardware resources into actual performance. Different scheduling schemes would lead to different levels of data reuse and computational parallelism, resulting in different energy efficiency and performance. However, the ReRAM-based scheduling scheme faces the following limitations. First, current studies mainly focus on intra-layer scheduling scheme optimizations by using the Weight Stationary (WS) data flow. These studies ignore the difference between layers and limit optimization opportunities. Second, there is no systematic definition and analysis for inter-layer scheduling schemes. Third, there is no co-optimization study on intra- and inter-layer scheduling schemes. Fourth, the complex network structure leads to intricate inter-layer data dependency, making the optimization of the scheduling scheme more challenging. These limitations restrict the comprehensive understanding of the scheduling schemes.Inspired by these observations, we identify the fundamental impact of intra-layer scheduling schemes on ReRAM-based designs, including the WS and Input Stationary (IS) data flows. We also systematically define and analyze inter-layer scheduling schemes according to different combinations of data flows, including the WS-WS, IS-IS, WS-IS, and IS-WS data flows. We analyze and explore different resource allocation strategies for these schemes. We also propose the intra- and inter-layer co-optimization to further improve performance and energy efficiency. Then, we propose a method for building a hybrid scheduling scheme by flexibly combining these inter-layer scheduling schemes for complex networks. Finally, we seek the potential to improve performance and energy efficiency for hybrid scheduling schemes. For deploying the MobileNet-V1, ResNet-18, VGG-16, and AlexNet, the hybrid scheduling scheme improves performance by 10.2×~130.7×, 1.9×~16.5×, 7.0×~56×, and 1×~153.1× than the WS-WS, IS-IS, IS-WS, and WS-IS based scheduling schemes, respectively. Similarly, the power efficiency can also be increased by 15× and 26× than the WS-WS and IS-IS based scheduling schemes, respectively. Yunping Zhao, Sheng Ma, Yuhua Tang |
IEEE Trans. Computers | 5 |
| 2025 | Tradeoff Performance and Energy Efficiency by Optimizing the Data Flow for PIM ArchitecturesabstractThe processing-in-memory (PIM) architecture becomes a promising candidate for deep learning accelerators by integrating computation and memory. Most PIM-based studies improve the performance and energy efficiency by using the weight stationary (WS) data flow due to its high parallelism. However, the WS data flow has some fundamental limitations. First, the WS data flow has huge activation movements between on-chip memory and off-chip memory due to the limited memory space of the resistive random-access memory (ReRAM) array. Second, the WS data flow needs to read the input activation repeatedly according to the convolution window. These data movements decrease the energy efficiency and performance of the PIM architecture. To address these issues, the input stationary (IS) data flow stores activations instead of weights to reduce data movements. But the IS data flow faces some challenges. First, the data dependency between adjacent layers limits the performance. Second, there are huge across-array computations due to the special mapping method. Third, the previous IS data flow cannot realize the high parallelism. Fourth, the IS data flow depends on the 3-D ReRAM structure. To address these issues, we propose a novel data flow for PIM architectures. We optimize the IS data flow to decrease the activation movement and propose a parallel computing method to realize high parallelism and reduce the across-array computations. We identify and analyze the fundamental limitations and impact of different interlayer data flows, including the WS-WS, IS-IS, WS-IS, and IS-WS. We also propose a method to build a hybrid data flow by combining these interlayer data flows to tradeoff performance and energy consumption. Our experimental results and analysis demonstrate the potential of our design. The performance and energy efficiency of our design reach 0.13–1.77 TFLOPS and 61–85 TOPS/J, respectively. Compared to the state-of-the-art design, the NEBULA, our design can improve performance by$1.4\times $,$2.3\times $, and$3.5\times $for deploying the MobileNet-V1, ResNet-18, and VGG-16, and also can improve energy efficiency by$3.3\times $,$2\times $, and$2\times $, respectively. Yunping Zhao, Sheng Ma, Yuhua Tang, Hengzhu Liu, Dongsheng Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | A Graph Embedded Feature Decoupling Model for Clustering Single Cell RNA-seq DataabstractThe rapid development in single-cell RNA sequencing (scRNA-seq) has dramatically enhanced our insight into cellular heterogeneity and disease mechanisms. In the analysis of scRNA-seq data, cell clustering plays a vital role in downstream tasks. The advent of deep learning has revolutionized the analysis method for cell clustering. However, due to the high dimensionality and sparsity of scRNA-seq data, neural network-based methods often capture a multitude of spurious correlations. This oversight results in redundancy within the low-dimensional feature space, making it difficult to distinguish cell subpopulations. To address these limitations, we introduce the Graph Embedded Feature Decoupling model (scGEFD), a two-phase approach for cell clustering. During the first phase, we capture cellular structural information by graph neural network. In the second phase, we propose a feature decoupling method inspired by the Barlow Twins. Specifically, we refine feature representations by minimizing discrepancies between two perturbed sample versions, which reduces redundancy in the latent space and further accentuates biologically pertinent features. Consequently, scGEFD outperforms state-of-the-art clustering methods on ten real-world datasets, providing more accurate and meaningful biological insights from scRNA-seq data. Haoang Chi, Huihui Yang, Yuhua Tang |
BIBM | 6 |
| 2024 | LLM-Based Processor Verification: A Case Study for Neuronnorphic ProcessorabstractWith the increasing complexity of the hardware design, conducting verification before the tapeout is of utmost importance. Simulation-based verification remains the primary method owing to its scalability and flexibility. A comprehensive verification of modern processors usually requires numerous effective tests to cover all possible conditions and use cases, leading to significant time, resource, and manual effort even with the EDA. Moreover, novel domain specific architecture (DSA), such as neuromorphic processors, will exacerbate the challenge of verification. Fortunately, emerging large language models (LLMs) have been demonstrating a powerful ability to complete specific tasks assigned by human instructions. In this paper, we explore the challenges and opportunities encountered when using the LLMs to accelerate the DSA verification using the proposed LLM-based workflow consisting of test generation, compilation&simulation, and result collection&processing. By verifying a RISC-V core and a neuromorphic processor, we examine the capabilities and limitations of the LLMs when using them for the function verification of traditional processors and emerging DSA. In the experiment, 36$C$programs and 128 assembly snippets for the RISC-V core and the neuromorphic processor are generated using an advanced LLM to demonstrate our claim. The experimental results show that the code coverage based on the LLM test generation can reach 89% and 91% for the above two architectures respectively, showing a promising research direction for the future processor verification in the new golden age for computer architecture. Yifei Deng, Renzhi Chen, Jingyue Zhao, Huadong Dai, Yuhua Tang |
DATE | 9 |
| 2024 | Modality Re-Balance for Visual Question Answering: A Causal FrameworkabstractVisual Question Answering (VQA) models often prioritize language cues over visual knowledge, leading to the "language prior" phenomenon. To address this, researchers have proposed methods to balance language and image information during training and inference. However, these approaches often struggle to capture important linguistic components due to the excessive exclusion of language information. Inspired by causal inference, we introduce a novel approach called the SyMmetrically Balanced Causal framework (SMBC) that rebalances visual and textual information in VQA tasks. This framework allows for an equal contribution of knowledge from both modalities to inference results. Experimental evaluation shows that SMBC: 1) applies to prevalent VQA models, including those with data augmentation, and 2) consistently improves performance on established benchmarks. Xinpeng Lv, Wanrong Huang, Haotian Wang 0001, Ruochun Jin, Xueqiong Li, Shuman Li, Yongquan Feng, Yuhua Tang |
ICASSP | 9 |
| 2024 | Continuous Review and Timely Correction: Enhancing the Resistance to Noisy Labels via Self-Not-True DistillationabstractDeep neural networks possess substantial learning capacities and robust expressive power, making them prone to overfitting mislabeled data. Fortunately, the memorization effect shows that the networks tend to memorize the clean data first, and then gradually memorize the mislabeled data. Correspondingly, early stopping is proposed and has proven to be effective in mitigating overfitting. However, the networks can still overfit some mislabeled data in the early training stage, resulting in forgotten knowledge of clean data. In addition, early stopping lacks correction of errors caused by mislabeled data. In this paper, we propose that the network should continuously review the knowledge it learned earlier to enhance clean data memorization while timely correcting the incorrect knowledge learned from the mislabeled data. To implement these two ideas, we first introduce self-distillation into training, which employs a teacher network from the previous stage to guide the current network, enhancing clean data memorization. Based on this, we further propose the not-true distillation. Before distilling knowledge from the teacher network, we mask the true class (i.e. label class) in the logits, focusing only on not-true classes to correct the accumulated incorrect knowledge. Extensive experiments on simulated and realworld benchmarks adequately validate the superior performance of our method. Xinghao Wu, Yuhua Tang, Long Lan |
ICASSP | 4 |
| 2024 | MOTPE/D: Hardware and Algorithm Co-design for Reconfigurable Neuromorphic ProcessorabstractRecent advances in hardware/algorithm co-design for spiking neural networks have demonstrated its potential for jointly optimizing algorithmic performance while minimizing hardware overhead. However, the gigantic mixed-variable hard-ware/algorithm co-design space and time-consuming hardware verification still pose an intractable challenge for solutions exploration. To tackle these problems, 1) we propose a generic three-phase hardware/algorithm co-design framework. In this framework, 2) we target a reconfigurable neuromorphic processor, and parameterize the hardware and network architecture in a unified design space. 3) We propose a generic analytical model to estimate the parameter size and power consumption, which can support fast candidate evaluation during the exploration. 4) We extend vanilla TPE (a single-objective optimization algorithm) to MOTPE/D, a generic Multi-objective optimization (MOO) algorithm, by introducing a decomposition strategy. Renzhi Chen, Xun Xiao, Jingyue Zhao, Zhenhua Zhu 0002, Huadong Dai, Yuhua Tang |
ICCD | 8 |
| 2024 | HPA: A Hybrid Data Flow for PIM ArchitecturesabstractThe Processing- In- Memory (PIM) architecture becomes a promising candidate for deep learning acceleration by integrating computation and memory. Due to the simple mapping method and high parallelism, the Weight Stationary (WS) data flow is widely used in PIM-based studies to improve performance and energy efficiency. However, the WS data flow leads to huge activation movements, becoming the bottleneck for reducing latency and energy consumption. To address this issue, the Input Stationary (IS) data flow stores activations instead of weights in the PIM architecture to reduce data movements. However, the traditional IS data flow also faces several challenges. First, the inter-layer data dependence and imbalance workload decrease pipeline efficiency. Second, the across-array computation reduces energy efficiency and performance. Third, the traditional IS data flow relies on the 3D ReRAM structure. Inspired by these observations, we propose a Hybrid data flow for PIM Architectures, named HPA. The HPA contains novel intra-layer and inter-layer data flows, named the PP-IS data flow and the IS- WS hybrid data flow, respectively. The PP- IS data flow optimizes the data mapping strategy and computing method to reduce activation movements. In addition, the PP- IS data flow uses the parallel computing method to decrease across-array computations. Based on the novel intra-layer data flow, we propose the IS- WS hybrid data flow to trade off performance and energy efficiency. Finally, we optimize the pipeline for the hybrid data flow to mitigate data depen-dence and balance inter-layer workloads, improving pipeline efficiency. Our experimental results and analysis demonstrate the potential of the HPA. The performance and power efficiency of the HPA reaches 1.64$GFLOPS\sim 63$G F LO P Sand 2.1$TOPS/W\sim 151\ TOPS/W$, respectively. Compared to the state-of-the-art design, the NEBULA, the HPA can significantly improve power efficiency and performance by$22.1\times$and$7.8\times$, respectively, when deploying the MobileNet VI. Sheng Ma, Yunping Zhao, Yuhua Tang |
ICCD | 3 |
| 2024 | Boosting Meaningful Dependency Mining with Clustering and Covariance AnalysisabstractFunctional dependencies (FDs) form a valuable ingredient for various data management tasks. However, existing methods can hardly discover practical and interpretable FDs, especially in large noisy real-life datasets. This paper studies the problem of discovering meaningful functional dependencies (FDms) that utilize support and error parameters to capture interesting dependencies in such datasets and proposes an efficient discovery algorithm called FDMε. In order to scale with large datasets, FDM ε employs an efficient sampling method with accuracy guarantees to capture the differences between tuple pairs and to quantify the connection between support/error of dependencies on samples and those on the entire dataset. Moreover, it adopts a clustering-based correlated attributes extraction to divide the exponentially large search space into multiple small sub-spaces and proposes an easy-first traversal strategy with covariance-based guidance that quickly detects candidate dependencies and validates them. Additionally, we prove a covariance lower bound as an additional pruning criterion to reduce the search space. Extensive experiments on real-life and synthetic datasets demonstrate that FDM ε is 14 times faster than existing discovery algorithms on average, up to 31 times, and scales to larger datasets with the least memory cost. Ruochun Jin, Wanrong Huang, Yuhua Tang |
ICDE | 4 |
| 2024 | Bioinspired sensing-memory-computing integrated vision systems: biomimetic mechanisms, design principles, and applications
Yinlong Tan, Yabo Chen, Yuhua Tang |
Sci. China Inf. Sci. | 5 |
| 2024 | Highly Efficient Active Learning With Tracklet-Aware Co-Cooperative Annotators for Person Re-IdentificationabstractSupervised person re-identification (ReID) has attracted widespread attentions in the computer vision community due to its great potential in real-world applications. However, the demand of human annotation heavily limits the application as it is costly to annotate identical pedestrians appearing from different cameras. Thus, how to reduce the annotation cost while preserving the performance remains challenging and has been studied extensively. In this article, we propose a tracklet-aware co-cooperative annotators' framework to reduce the demand of human annotation. Specifically, we partition the training samples into different clusters and associate adjacent images in each cluster to produce the robust tracklet which decreases the annotation requirements significantly. Besides, to further reduce the cost, we introduce a powerful teacher model in our framework to implement the active learning strategy and select the most informative tracklets for human annotator, the teacher model itself, in our setting, also acts as an annotator to label the relatively certain tracklets. Thus, our final model could be well-trained with both confident pseudo-labels and human-given annotations. Extensive experiments on three popular person ReID datasets demonstrate that our approach could achieve competitive performance compared with state-of-the-art methods in both active learning and unsupervised learning (USL) settings. Xiao Teng, Long Lan, Xueqiong Li, Yuhua Tang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | NUBA: Non-Uniform Bandwidth GPUsabstractThe parallel execution model of GPUs enables scaling to hundreds of thousands of threads, which is a key capability that many modern high-performance applications exploit. GPU vendors are hence increasing the compute and memory resources with every GPU generation — resulting in the need to efficiently stitch together a plethora of Symmetric Multiprocessors (SMs), Last-Level Cache (LLC) slices and memory controllers while maximizing bandwidth and keeping power consumption and design complexity in check. Conventional GPUs are Uniform Bandwidth Architectures (UBAs) as they provide equal bandwidth between all SMs and all LLC slices. UBA GPUs require a uniform high-bandwidth Network-on-Chip (NoC), and our key observation is that provisioning a NoC to match the LLC slice bandwidth incurs a hefty power and complexity overhead. We propose the Non-Uniform Bandwidth Architecture (NUBA), a GPU system architecture aimed at fully utilizing LLC slice bandwidth. A NUBA GPU consists of partitions that each feature a few SMs and LLC slices as well as a memory controller — hence exposing the complete LLC bandwidth to the SMs within a partition since they can be connected with point-to-point links — and a NoC between partitions — to enable access to remote data.Exploiting the potential of NUBA GPUs however requires carefully co-designing system software, the compiler and architectural policies. The critical system software component is our Local-And-Balanced (LAB) page placement policy which enables the GPU driver to place data in local partitions while avoiding load imbalance. Moreover, we propose Model-Driven Replication (MDR) which identifies read-only shared data with data-flow analysis at compile time. At run time, MDR leverages an architectural mechanism that replicates read-only shared data across LLC slices when this can be done without pressuring cache capacity. With LAB and MDR, our NUBA GPU improves average performance by 23.1% and 22.2% (and up to 183.9% and 182.4%) compared to iso-resource memory-side and SM-side UBA GPUs, respectively. When the NUBA concept is leveraged to reduce overhead while maintaining similar performance, NUBA reduces NoC power consumption by 12.1× and 9.4×, respectively. Xia Zhao 0004, Magnus Jahre, Yuhua Tang, Guangda Zhang, Lieven Eeckhout |
ASPLOS (2) | 3 |
| 2023 | Self-aware circular response-guided attention for robust siamese tracking
Huibin Tan, Mengzhu Wang, Tianyi Liang 0001, Yuhua Tang, Long Lan, Wenjing Yang 0002 |
Appl. Intell. | 5 |
| 2023 | Back to Homogeneous Computing: A Tightly-Coupled Neuromorphic Processor With Neuromorphic ISAabstractIn recent years, neuromorphic processors are widely used in many scenarios, showing extreme energy efficiency over traditional architectures. However, almost all existing neuromorphic hardware are following the heterogeneous computing methodology without Instruction Set Architecture (ISA), leading to inflexibility in programming. In this paper, we first propose a RISC-V Neuromorphic Extension (RVNE) to enable fine-grained and flexible homogeneous programming for neuromorphic algorithms while utilizing SNN sparsity from different levels of granularity and computing flows. Based on RVNE, we next implement a neuromorphic micro-architecture that is tightly coupled to the CPU pipeline to accelerate neuromorphic computing. To demonstrate the proposed homogeneous neuromorphic architecture, we implement a prototype processor called NeuroRVcore based on RISC-V ISA and an open-source RISC-V core. The evaluation results show that RVNE achieves a 2.8 × −4.3 × reduction in code density compared with the general-purpose ISAs. Compared with the state-of-the-art neuromorphic processor, the proposed homogeneous computing reduces energy consumption by 3.4%−22.5% while enabling fine-grained and flexible homogeneous programming. Lei Wang 0011, Yao Wang 0002, Junbo Tie, Feng Wang 0050, LingHui Peng, Xun Xiao, Gan Zhou, Xuhu Yu, Xia Zhao 0004, Yuhua Tang, Weixia Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 16 |
| 2021 | Identity-Based Data Augmentation via Progressive Sampling for One-Shot Person Re-identification
Runxuan Si, Shaowu Yang, Haoang Chi, Yuhua Tang |
ICONIP (4) | 5 |
| 2021 | Model Compression for a Plasticity Neural Network in a Maze Exploration Scenario
Baolun Yu, Wanrong Huang, Long Lan, Yuhua Tang |
ICONIP (5) | 4 |
| 2021 | Scalable Auto-weighted Discrete Multi-view ClusteringabstractMulti-view clustering has been widely studied in machine learning, which uses complementary information to improve clustering performance. However, challenges remain when handling large-scale multi-view data due to the traditional approaches’ high time complexity. Besides, the existing approaches suffer from parameter selection. Due to the lack of labeled data, parameter selection in practical clustering applications is difficult, especially in big data. In this paper, we propose a novel approach for large-scale multi-view clustering to overcome the above challenges. Our approach focuses on learning the low-dimensional binary embedding of multi-view data, preserving the samples’ local structure during binary embedding, and optimizing the embedding and clustering in a unified framework. Furthermore, we proposed to learn the parameters using a combination of data-driven and heuristic approaches. Experiments on five large-scale multi-view datasets show that the proposed method is superior to the state-of-the-art in terms of clustering quality and running time. Yuhua Tang |
WWW | 3 |
| 2020 | Multi-UAV Adaptive Path Planning in Complex Environment Based on Behavior Tree
Wendi Wu, Xiaoguang Ren, Yuhua Tang |
CollaborateCom (2) | 5 |
| 2020 | Adversarial Mixup Synthesis Training for Unsupervised Domain AdaptationabstractDomain adversarial training is a popular approach for Unsupervised Domain Adaptation (DA). However, the transferability of adversarial training framework may drop greatly on the adaptation tasks with a large distribution divergence between source and target domains. In this paper, we propose a new approach termed Adversarial Mixup Synthesis Training (AMST) to alleviate the issue. The AMST augments the training with synthesis samples by linearly interpolating between pairs of hidden representations and their domain labels. By this means, AMST encourages the model to make consistency domain prediction less confidently on interpolations points, which learn domain-specific representations with fewer directions of variance. Based on the previous work, we conduct a theoretical analysis on this phenomenon under ideal conditions and show that AMST could improve generalization ability. Finally, experiments on benchmark dataset demonstrate the effectiveness and practicability of AMST. We will publicly release our code on github soon. Yuhua Tang, Haotian Wang 0001 |
ICASSP | 1 |
| 2020 | Adaptive Inner-reward Shaping in Sparse Reward GamesabstractReinforcement learning focuses on goal-directed learning from interaction and the success of its applications strongly depends on how well the reward signal frames the problem and how well it assesses progress in solving it. But in many real-world scenarios, the agent is supplied with extremely sparse or even no rewards which makes learning fail and fall into ineffective exploration. In psychology, shaping is a method of animal training by reinforcing successive approximations of rewards to finally achieve the desired complex behavior. Inspired by this phenomenon of animal learning and reward as a signal in neuroscience, in this paper we solve the sparse reward problem by constructing a reward generator to generate inner-rewards and guide the agent learning control policies with deep neural networks. The proposed learning-based reward shaping does not require specific domain knowledge, but rather enable the agent to learn how to generate inner rewards to guide itself in any scenarios online jointly with the actual reinforcement learning process. To validate the performance in complex sparse reward problems, the proposed approach is evaluated in a challenging scenario, Football Academy in Google Research Football Environment, a newly released reinforcement learning environment with physics-based 3D simulator, instead of maze environments or grid world that are commonly used in research which are not sufficiently challenging. We compare the performance of our inner-rewards approach with two reinforcement algorithms (PPO and ICM + PPO). Experimental results show that our method improves the learning performance in terms of speed and quality, and also enables the agent to learn generalized skills applied to novel scenarios. Dong Yang 0010, Yuhua Tang |
IJCNN | 2 |
| 2020 | Online Binary Incomplete Multi-view Clustering
Yuhua Tang |
ECML/PKDD (1) | 3 |
| 2020 | Technological breakthroughs and scientific progress of the Chang'e-4 mission
Dengyun Yu, Jizhong Liu, Yuhua Tang |
Sci. China Inf. Sci. | 5 |
| 2020 | An effective few-shot learning approach via location-dependent partial differential equation
Haotian Wang 0001, Yuhua Tang |
Knowl. Inf. Syst. | 3 |
| 2019 | Non-Convex Transfer Subspace Learning for Unsupervised Domain AdaptationabstractTransfer subspace learning aims to learn robust subspace for the target domain by leveraging knowledge from the source domain. The traditional methods often adopt the convex norm to approximate the original sparse and low-rank constraints, which make the optimization problem be easily solved. However, such relax approximation leads to the performance deviation of the original non-convex model. In this paper, we propose a novel Non-convex Transfer Subspace Learning~(NTSL) method to provide a tighter approximation to the original sparse and low-rank constraints. Specifically, we design an objective function that leverages the Schatten p-norm and ℓ_2, p-norm to preserve the structure between the source and target domains. With Schatten p-norm, the objective function better approximates the rank minimization problem than the nuclear norm and preserves the structure of domains. Besides, the ℓ_2, p-norm can reduce the effect of noise and improve the robustness to outliers. Meanwhile, we develop an efficient algorithm to solve the non-convex minimization problem. Extensive experimental results on cross-domain tasks show the effectiveness of our proposed method. Tingjin Luo, Wenjing Yang 0002, Yongjun Zhang 0006, Yuhua Tang |
ICME | 6 |
| 2019 | Analyzing time-dimension communication characterizations for representative scientific applications on supercomputer systems
Juan Chen 0001, Yong Dong, Feihao Wu, Enqiang Zhou, Yuhua Tang |
Frontiers Comput. Sci. | 8 |
| 2019 | Rademacher dropout: An adaptive dropout for deep neural network via optimizing generalization gap
Haotian Wang 0001, Wenjing Yang 0002, Tingjin Luo, Ji Wang 0001, Yuhua Tang |
Neurocomputing | 6 |
| 2018 | MulAttenRec: A Multi-level Attention-Based Model for Recommendation
Wenjing Yang 0002, Yongjun Zhang 0006, Haotian Wang 0001, Yuhua Tang |
ICONIP (2) | 5 |
| 2018 | Multi-feature Fusion for Deep Reinforcement Learning: Sequential Control of Mobile Robots
Haotian Wang 0001, Wenjing Yang 0002, Wanrong Huang, Yuhua Tang |
ICONIP (7) | 5 |
| 2018 | Deep CNN-based Visual Target Tracking System Relying on Monocular Image SensingabstractThe one-on-one target tracking problem is important in robot vision. Previous studies mainly focused on locating, depth information and control mechanism. In this study, we construct an autonomously visual tracking system called learn-to-track (LtT) by using a novel approach. This system only depends on a monocular camera. The main component is a deep convolutional neural network called the LtT, which trains a supervised image classifier by using images captured by the monocular camera in the follower robot. By operating merely on two adjacent frames, the network can predict the estimated velocity of the target, i.e., the velocity control for the follower. To verify the effectiveness of the LtT system, we construct a large-scale dataset that supports download l in the simulator, in which the LtT network is trained and the LtT system performance is evaluated. Furthermore, a remarkable tracking performance is achieved. Yawen Cui, Bo Zhang 0007, Wenjing Yang 0002, Xiaodong Yi 0002, Yuhua Tang |
IJCNN | 5 |
| 2018 | Design of communication relay mission for supporting lunar-farside soft landing
Yuhua Tang, Dong Qiao |
Sci. China Inf. Sci. | 2 |
| 2017 | The Curve Boundary Design and Performance Analysis for DGM Based on OpenFOAM
Yongquan Feng, Xinhai Xu, Yuhua Tang, Yongjun Zhang 0006 |
ICA3PP | 3 |
| 2017 | Effect of orbital shadow at an Earth-Moon Lagrange point on relay communication mission
Yuhua Tang, Dong Qiao |
Sci. China Inf. Sci. | 1 |
| 2016 | Collaborative Communication in Multi-robot Surveillance Based on Indoor Radio Mapping
Yunlong Wu 0002, Bo Zhang 0007, Xiaodong Yi 0002, Yuhua Tang |
CollaborateCom | 4 |
| 2016 | Delay-reliability tradeoff for wireless-connected indoor robot surveillance based on radio environment mapabstractThis paper considers a surveillance scenario where a mobile robot monitors an indoor environment and transmits the monitored data to a base station. Considering the indoor radio environment is complex, we first build the radio environment map (REM) with two different interpolation methods. Then, we combine REM with the structural blueprint of the building to build an integrated map called radio-structural map (RSM). Based on RSM, we propose an optimal surveillance path search (OSPS) method which minimizes the data transmission delay of the patrol robot under a communication reliability constraint. In OSPS, two optimization methods are adopted, which may sharply reduce the computation cost. Besides the numerical simulations, we further discuss the relationship between the communication reliability and data transmission delay. Finally, we test the applicability of OSPS in the stage simulator of ROS. Yunlong Wu 0002, Bo Zhang 0007, Xuefeng Chang, Xiaodong Yi 0002, Yuhua Tang |
PIMRC | 6 |
| 2016 | Distributed graph regularized non-negative matrix factorization with greedy coordinate descentabstractGraph regularized non-negative matrix factorization (GNMF) decomposes a high-dimensional non-negative data matrix into two low-dimensional matrices with the non-negativity property kept and the geometric structure preserved. Due to its effectiveness, GNMF has been widely used in many fields such as computer vision and data mining. However, GNMF cannot process large-scale datasets on distributed system because the gradient of the graph regularization term costs huge amount of communication overheads among computing nodes. In this paper, we proposed a distributed GNMF (DGNMF) algorithm to overcome this deficiency. Particularly, DGNMF reformulates the graph regularization term to avoid multiplying graph Laplacian by factor matrix through introducing an auxiliary variable and incorporating an equality constraint over it. We optimize DGNMF by using greedy coordinate descent method in the frame of augmented Lagrange method and implement this algorithm on a distributed system. Since DGNMF requires quite few communication overheads among computing nodes, it can be applied to large scale dataset. The preliminary results illustrate efficiency, scalability, and effectiveness of DGNMF. Ziheng Gao, Naiyang Guan, Xuhui Huang, Xuefeng Peng, Zhigang Luo, Yuhua Tang |
SMC | 6 |
| 2016 | Detailed and clock-driven simulation for HPC interconnection network
Juan Chen 0001, Dezun Dong, Yuhua Tang |
Frontiers Comput. Sci. | 6 |
| 2016 | Reducing Static Energy in Supercomputer Interconnection Networks Using Topology-Aware PartitioningabstractThe key to reducing static energy in supercomputers is switching off their unused components. Routers are the major components of a supercomputer. Whether routers can be effectively switched off or not has become the key to static energy management for supercomputers. For many typical applications, the routers in a supercomputer exhibit low utilization. However, there is no effective method to switch the routers off when they are idle. By analyzing the router occupancy in time and space, for the first time, we present a routing-policy guided topology partitioning methodology to solve this problem. We propose topology partitioning methods for three kinds of commonly used topologies (mesh, torus and fat-tree) equipped with the three most popular routing policies (deterministic routing, directionally adaptive routing and fully adaptive routing). Based on the above methods, we propose the key techniques required in this topology partitioning based static energy management in supercomputer interconnection networks to switch off unused routers in both time and space dimensions. Three topology-aware resource allocation algorithms have been developed to handle effectively different job-mixes running on a supercomputer. We validate the effectiveness of our methodology by using Tianhe-2 and a simulator for the aforementioned topologies and routing policies. The energy savings achieved on a subsystem of Tianhe-2 range from 3.8 to 79.7 percent. This translates into a yearly energy cost reduction of up to half a million US dollars for Tianhe-2. Juan Chen 0001, Yuhua Tang, Yong Dong, Jingling Xue |
IEEE Trans. Computers | 2 |
| 2013 | A Message Logging Protocol Based on User Level Failure Mitigation
Xunyun Liu, Xinhai Xu, Xiaoguang Ren, Yuhua Tang, Ziqing Dai |
ICA3PP (1) | 4 |
| 2011 | Parallization of Adaboost Algorithm through Hybrid MPI/OpenMP and Transactional MemoryabstractThis paper proposes a parallelization of the Adaboost algorithm through hybrid usage of MPI, OpenMP, and transactional memory. After detailed analysis of the Adaboost algorithm, we show that multiple levels of parallelism exists in the algorithm. We develop the lower level of parallelism through OpenMP and higher level parallelism through MPI. Software transactional memory are used to facilitate the management of shared data among different threads. We evaluated the Hybrid parallelized Adaboost algorithm on a heterogeneous PC cluster. And the result shows that nearly linear speedup can be achieved given a good load balancing scheme. Moreover, the hybrid parallelized Adaboost algorithm outperforms Purely MPI based approach by about 14% to 26%. Yuhua Tang, Fudong Liu |
PDP | 2 |
| 2010 | Sim-spm: A SimpleScalar-Based Simulator for Multi-level SPM Memory Hierarchy ArchitectureabstractAs a fast on-chip SRAM managed by software (the application and/or compiler), Scratchpad Memory (SPM) is widely used in many fields. This paper presents a Simple Scalar-based multi-level SPM memory hierarchy architecture simulator Sim-spm. We simulate the hardware of the multi-level SPM memory hierarchy successfully by extending Sim-outorder, which is an out-of-order simulator from Simple Scalar. Through the simulating memory method, the simulation framework of the multi-level SPM memory hierarchy has been built under the existing ISA (Instruction Set Architecture), which largely reduces the requirement to modify the existing compiler. The experimental results show that Sim-spm can accurately simulate the running state of the processor with a multi-level SPM memory hierarchy architecture, and it has a good prospect for the research of multi-level SPM memory hierarchy architecture. Xiaoguang Ren, Yuhua Tang, Tao Tang 0006, Sen Ye, Huiquan Wang |
HPCC | 2 |
| 2010 | Managing Data-Objects in Dynamically Reconfigurable Caches
Xue-Jun Yang, Yuhua Tang |
J. Comput. Sci. Technol. | 4 |
| 2004 | Verify Memory Integrity Basing on Hash Tree and MAC Combined Approach
Fangyong Hou, Yuhua Tang, Jifeng Liu |
EUC | 3 |
| 2004 | Protecting integrity and confidentiality for data communicationabstractThis work presents a scheme to build data communication system that can effectively protect data integrity and confidentiality. Firstly, This work briefly introduces the situation of integrity and confidentiality protection. Then, This work brings forward a new cipher, which uses a keystream generator to produce infinite number of frame secret keys basing on an infinite root key space, and use a unique one-off frame secret key for each data encryption/decryption. Basing on this cipher, we construct a data communication system. This work illustrates how to build such a system and analyze its protections of data integrity and confidentiality. With the character of cipher, it offers high resistance against cryptanalysis to prevent data disclosure, and it gives little opportunities to those intractable attacks that can compromise data integrity. Fangyong Hou, Yuhua Tang |
ISCC | 3 |