Tianyou Chang

dblp:226/9249 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0002-1857-9482ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 MMM-RAG: A Multi-Agent Multi-Feature Method for Multimodal Retrieval-Augmented Generation
abstract
Retrieval-augmented generation (RAG) enhances large language models via external knowledge, but struggles with text-image multimodal heterogeneous data. Existing multimodal RAG methods extract only single-type image features (e.g., visual or semantic), failing to capture both structural and semantic information, leading to partial retrieval and degraded generation quality. This paper proposes MMM-RAG: A Multi-Agent MultiFeature Method for Multimodal Retrieval-Augmented Generation, a multi-agent collaborative framework for multimodal question answering. Its core innovations are multi-feature fusion retrieval and multi-agent coordination, with four modules: 1) constructing a multi-feature knowledge base (text embedded directly; images processed for ViT-extracted visual features and multimodal model-generated semantic descriptions, forming a dual-feature database); 2) decomposing complex queries into subtasks for different agents; 3) enabling each agent to retrieve from its corresponding feature database (visual or semantic); 4) fusing results via consensus voting to generate answers. The framework deeply integrates image visual-structural and semantic information, improving retrieval comprehensiveness and response quality. Experiments on ScienceQA and CrisisMM-D show MMM-RAG outperforms traditional unimodal/multimodal RAG in accuracy, especially in cross-modal reasoning. Specifically, it surpasses GPT-4o by 2.66 % on ScienceQA and 6.29 % on CrisisMMD, confirming robustness.
Tianyou Chang
ICPADS1
2025 Test Case Enhanced Self-Iterative Code Generation Framework
abstract
In the field of code generation, large language models(LLMs) have made significant advancements. In addition to straightforward code generation, the latest research integrates unit testing and debuggers into the code generation process of LLMs. This approach enables the large models to self-improve the generated code by incorporating feedback from test results. However, the current method relies only on a small number of test cases for code debugging, making it difficult to cover various scenarios and edge conditions of the program, especially in cases involving complex logic or extensive data operations. To address this issue, the paper proposes a Test Case Enhanced Self-Iterative code generation framework (TEI). This is a novel solution comprising analysts, designers, developers, and testers. Analysts and designers are responsible for refining code requirements and designing structured solutions, respectively. Developers encode based on the refined requirements and design solutions. Testers are required, firstly, to design multiple comprehensive, detailed, and precise test cases based on a given test case, and secondly, to systematically verify the generated code block by block and report its correctness and any potential errors, in order to provide feedback to the analysts in the next iteration. The experimental results demonstrate that our proposed framework (TEI), when using GPT-3.5 as the agent, outperforms existing code generation models and methods significantly. For instance, our approach achieves an 88.4% pass@1 rate in the HumanEval benchmarks, surpassing the previously state-of-the-art GPT-4, which only achieved 85.4%. We also investigated the accuracy of augmented test cases and their impact on the overall quality of code generation within the framework. “Failure is simply the opportunity to begin again, this time more intelligently.” Henry Ford
Tianyou Chang
ICPADS1
2024 A Vision-language Model Based on Prompt Learner for Few-shot Medical Images Diagnosis
abstract
In the real world, it can be challenging to annotate a large-scale dataset for all medical images, making few-shot medical image classification an important task. The latest advancements in pre-trained vision-language models as CLIP have demonstrated excellent performance in zero-shot natural image recognition and show advantages in medical applications. However, we have found that deploying such models in practical applications faces challenges in terms of engineering effort. It requires specialized medical domain knowledge and is time-consuming, as even slight variations in wording can have a significant impact on performance. Inspired by recent research on prompt learning in the field of Natural Language Processing (NLP), we propose a simple approach called Prompt Learner (PoLe) for automating the design of prompts in pre-trained vision-language models. This is a simple method specifically designed to fine-tune vision-language models, similar to CLIP, for downstream image recognition. In particular, PoLe models context tokens using continuous vectors that can automatically learn from medical images, thus avoiding the tedious process of handcrafting prompt engineering. Additionally, this approach maintains the frozen state of the large-scale pre-trained parameters, saving computational resources. Through extensive experiments on 5 medical image datasets, we have demonstrated that PoLe surpasses manually designed prompts with just one or two shots, and further training with more shots significantly improves the performance of image classification. For instance, when trained with 16 shots, the average improvement is approximately 20% (with a maximum improvement of over 31%). PoLe effectively transforms CLIP into a powerful few-shot learner. In terms of recognition performance, adjusting the CLIP model using PoLe yields better results than manually designed prompts for CLIP. When enhancing CLIP, PoLe demonstrates stronger learning capabilities compared to other few-shot learners such as linear probes. Furthermore, it outperforms CLIP models assisted by ChatGPT on most datasets. This indicates that PoLe possesses significant adaptability in the field of medical image analysis.
Tianyou Chang, Shizhan Chen, Zhiyong Feng 0002
CSCWD1
2023 A Self-Iteration Code Generation Method Based on Large Language Models
abstract
Although large language models (LLMs) have demonstrated impressive performance in code generation, they still face challenges when dealing with complex code generation tasks. In the software development process, humans often refine complex tasks iteratively and continuously modify and improve them. Inspired by this, we propose a self-iteration code generation framework based on large language models like ChatGPT. To realize this idea, we incorporated software development methodologies into the self-iteration framework. We introduced four roles into each cycle, including analyst, designer, developer, and tester. Each role performs different tasks during the self-iteration cycle, with analyst and designer continuously improving requirements analysis and task design based on the testing feedback provided by tester, while developer are responsible for refactoring or modifying code until it passes testing, concluding the entire self-iteration process. We conducted extensive experiments on multiple benchmarks. The experimental results indicate: (1) The code generated by the self-iteration framework achieves up to a 21.3% relative improvement in Pass@1 compared to direct code generation. (2) The self-iteration framework also exhibits strong generalization performance, enhancing code generation quality for different large language models."Failure is simply the opportunity to begin again, this time more intelligently."- Henry Ford
Tianyou Chang, Shizhan Chen, Zhiyong Feng 0002
ICPADS1
2023 RTCoder: An Approach based on Retrieve-template for Automatic Code Generation
abstract
Regarding code generation, researchers have recently proposed a retrieve-template-generation approach. This method involves retrieving similar code snippets through a retriever and providing them to a generator along with input descriptions. However, since the retrieved similar code can be influenced by various data types, it may lead the model to reference unrelated content, resulting in some discrepancies between the generated code and the target code. To mitigate this bias, we introduce a code generation method based on retrieve-template-generation called RTCoder. Specifically, RTCoder completes code generation through three steps: Retrieve. Using a natural language description, the retriever retrieves several similar code snippets from a corpus. Template. By comparing these similar code snippets, it employs the Rabin-Karp algorithm to extract their common substrings and represents different substrings with spaces, forming a code template. Generator. The generator, based on a specific natural language description and the corresponding code template, automatically generates the concrete target code. We conducted extensive comparative experiments on three datasets and used three widely used evaluation metrics. The experimental results demonstrate that: (1) Compared to mainstream code generation models, RTCoder shows improvements in all three metrics across different datasets. For instance, compared to the state-of-the-art CodeT5 base, the EM value is 5.98%, 3.34%, and 1.67% higher on the three datasets, respectively. (2) Our approach is effective for other models as well. Taking the CodeBLEU score on the Concode dataset as an example, the retrieval-template-based generation method improved by 3.73% and 1.80% compared to direct generation and retrieval-generation methods on the RNN model, respectively.
Tianyou Chang, Shizhan Chen, Zhiyong Feng 0002
ICPADS1