EDBT 2026 Demo / reviewers in the wild / expert
Tianyu Zheng
dblp:202/3981
· DBLP profile ↗
18ranked-venue papers
4as first author
17since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Security and privacy · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | MIGDG: A mutual information guided dual-space graph-embedding model for entity alignment
Yichao Zhang 0001, Fengfeng Yi, Jihong Guan, Shuigeng Zhou, Wengen Li, Tianyu Zheng |
Expert Syst. Appl. | 7 |
| 2026 | Two-Factor Authentication Can Harden Servers Against Offline Password Search
Xavier Boyen, Stanislaw Jarecki, Phillip Nazarian, Jiayu Xu 0001, Tianyu Zheng |
EUROCRYPT (2) | 5 |
| 2025 | Compressed Sigma Protocols: New Model and Aggregation Techniques
Yuxi Xue, Tianyu Zheng, Shang Gao 0006, Bin Xiao 0001, Man Ho Au |
ACISP (1) | 2 |
| 2025 | MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at ScaleabstractJiawei Guo, Tianyu Zheng, Yizhi Li, Yuelin Bai, Bo Li, Yubo Wang, King Zhu, Graham Neubig, Wenhu Chen, Xiang Yue. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tianyu Zheng, Yuelin Bai, Bo Li 0080, Yubo Wang 0019, King Zhu, Graham Neubig, Wenhu Chen, Xiang Yue |
ACL (1) | 2 |
| 2025 | MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding BenchmarkabstractXiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, Kai Zhang, Shengbang Tong, Yuxuan Sun, Botao Yu, Ge Zhang, Huan Sun, Yu Su, Wenhu Chen, Graham Neubig. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang 0019, Kai Zhang 0033, Shengbang Tong, Yuxuan Sun 0002, Botao Yu, Ge Zhang 0009, Huan Sun 0001, Yu Su 0001, Wenhu Chen, Graham Neubig |
ACL (1) | 2 |
| 2025 | MuPT: A Generative Symbolic Music Pretrained TransformerabstractIn this paper, we explore the application of Large Language Models (LLMs) to the pre-training of music. While the prevalent use of MIDI in music modeling is well-established, our findings suggest that LLMs are inherently more compatible with ABC Notation, which aligns more closely with their design and strengths, thereby enhancing the model's performance in musical composition.
To address the challenges associated with misaligned measures from different tracks during generation, we propose the development of a $\underline{S}$ynchronized $\underline{M}$ulti-$\underline{T}$rack ABC Notation ($\textbf{SMT-ABC Notation}$), which aims to preserve coherence across multiple musical tracks.
Our contributions include a series of models capable of handling up to 8192 tokens, covering 90\% of the symbolic music data in our training set. Furthermore, we explore the implications of the $\underline{S}$ymbolic $\underline{M}$usic $\underline{S}$caling Law ($\textbf{SMS Law}$) on model performance. The results indicate a promising research direction in music generation, offering extensive resources for further research through our open-source contributions. Xingwei Qu, Yuelin Bai, Yinghao Ma, Ziya Zhou, Ka Man Lo, Ruibin Yuan, Lejun Min, Xueling Liu 0001, Xeron Du, Shuyue Guo, Yiming Liang, Shangda Wu, Junting Zhou, Tianyu Zheng, Ziyang Ma 0001, Fengze Han, Wei Xue 0002, Gus Xia, Emmanouil Benetos, Xiang Yue, Chenghua Lin 0002, Xu Tan 0003, Wenhao Huang 0001, Jie Fu 0001, Ge Zhang 0009 |
ICLR | 17 |
| 2025 | Lattice-Based Zero-Knowledge Proofs for Blockchain Confidential Transactions
Shang Gao 0006, Tianyu Zheng, Yu Guo 0003, Zhe Peng, Bin Xiao 0001 |
PKC (5) | 2 |
| 2024 | MORE-3S: Multimodal-based Offline Reinforcement Learning with Shared Semantic SpacesabstractDrawing upon the intuition that aligning different modalities to the same semantic embedding space would allow models to understand states and actions more easily, we propose a new perspective to the offline reinforcement learning (RL) challenge. More concretely, we transform it into a supervised learning task by integrating multimodal and pre-trained language models. Our approach incorporates state information derived from images and action-related data obtained from text, thereby bolstering RL training performance and promoting long-term strategic thinking. We emphasize the contextual understanding of language and demonstrate how decision-making in RL can benefit from aligning states’ and actions’ representation with languages’ representation. Our method significantly outperforms current baselines as evidenced by evaluations conducted on Atari and OpenAI Gym environments. This contributes to advancing offline RL performance and efficiency while providing a novel perspective on offline RL. Tianyu Zheng, Ge Zhang 0009, Xingwei Qu, Ming Kuang, Wenhao Huang 0001, Zhaofeng He 0001 |
LREC/COLING | 1 |
| 2024 | MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIabstractWe introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. MMMU includes 11.5K meticulously collected multimodal questions from college exams, quizzes, and text-books, covering six core disciplines: Art & Design, Busi-ness, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering. These questions span 30 subjects and 183 subfields, comprising 30 highly het-erogeneous image types, such as charts, diagrams, maps, tables, music sheets, and chemical structures. Unlike existing benchmarks, MMMU focuses on advanced perception and reasoning with domain-specific knowledge, challenging models to perform tasks akin to those faced by experts. The evaluation of 28 open-source LMMs as well as the propri-etary GPT-4V(ision) and Gemini highlights the substantial challenges posed by MMMU. Even the advanced GPT-4V and Gemini Ultra only achieve accuracies of 56% and 59% respectively, indicating significant room for improvement. We believe MMMU will stimulate the community to build next-generation multimodal foundation models towards expert artificial general intelligence. Xiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 0033, Ruoqi Liu, Ge Zhang 0009, Samuel Stevens 0001, Dongfu Jiang, Weiming Ren, Yuxuan Sun 0002, Cong Wei 0001, Botao Yu, Ruibin Yuan, Renliang Sun, Boyuan Zheng 0001, Zhenzhu Yang, Wenhao Huang 0001, Huan Sun 0001, Yu Su 0001, Wenhu Chen |
CVPR | 3 |
| 2024 | MAmmoTH2: Scaling Instructions from the WebabstractInstruction tuning improves the reasoning abilities of large language models (LLMs), with data quality and scalability being the crucial factors. Most instruction tuning data come from human crowd-sourcing or GPT-4 distillation. We propose a paradigm to efficiently harvest 10 million naturally existing instruction data from the pre-training web corpus to enhance LLM reasoning. Our approach involves (1) recalling relevant documents, (2) extracting instruction-response pairs, and (3) refining the extracted pairs using open-source LLMs. Fine-tuning base LLMs on this dataset, we build MAmmoTH2 models, which significantly boost performance on reasoning benchmarks. Notably, MAmmoTH2-7B’s (Mistral) performance increases from 11% to 36.7% on MATH and from 36% to 68.4% on GSM8K without training on any in-domain data. Further training MAmmoTH2 on public instruction tuning datasets yields MAmmoTH2-Plus, achieving state-of-the-art performance on several reasoning and chatbot benchmarks. Our work demonstrates how to harvest large-scale, high-quality instruction data without costly human annotation or GPT-4 distillation, providing a new paradigm for building better instruction tuning data. Xiang Yue, Tianyu Zheng, Ge Zhang 0009, Wenhu Chen |
NeurIPS | 2 |
| 2024 | BFT-Net: A transformer-based boundary feedback network for kidney tumour segmentationabstractAbstract Kidney tumours are among the top ten most common tumours, the automatic segmentation of medical images can help locate tumour locations. However, the segmentation of kidney tumour images still faces several challenges: firstly, there is a lack of renal tumour endoscopic datasets and no segmentation techniques for renal tumour endoscopic images; secondly, the intra‐class inconsistency of tumours caused by variations in size, location, and shape of renal tumours; thirdly, difficulty in semantic fusion during decoding; and finally, the issue of boundary blurring in the localization of lesions. To address the aforementioned issues, a new dataset called Re‐TMRS is proposed, and for this dataset, the transformer‐based boundary feedback network for kidney tumour segmentation (BFT‐Net) is proposed. This network incorporates an adaptive context extract module (ACE) to emphasize local contextual information, reduces the semantic gap through the mixed feature capture module (MFC), and ultimately improves boundary extraction capability through end‐to‐end optimization learning in the boundary assist module (BA). Through numerous experiments, it is demonstrated that the proposed model exhibits excellent segmentation ability and generalization performance. The mDice and mIoU on the Re‐TMRS dataset reach 91.1% and 91.8%, respectively. Tianyu Zheng, Zhengping Li, Chao Nie, Rubin Xu, Minpeng Jiang, LeiLei Li |
IET Commun. | 1 |
| 2024 | MDSK-Net: Multi-scale dynamic segmentation kernel network for renal tumour endoscopic image segmentationabstractAbstract Automatic segmentation of renal tumours during renal cell carcinoma surgery can help doctors accurately locate the tumour region, protect the tissues and organs around the kidneys, enhance surgical efficiency, and reduce the possibility of leakage and misdiagnosis. However, since general polyp endoscopic image segmentation models have many problems when facing the task of renal tumour segmentation, there needs to be more research on the segmentation of endoscopic images of renal tumours. This paper proposes a multi‐scale dynamic segmentation kernel network for endoscopic image segmentation of kidney tumours. First, a spatial receptive field module is proposed to augment the feature information and improve the performance of the whole network. Second, an enhanced cross‐attention module is offered to attenuate the effect of a high‐similarity segmentation background. Finally, a multi‐scale dynamic segmentation kernel module is introduced to gradually refine the segmentation results from small to large sizes to obtain more accurate tumour boundaries. Extensive experiments on the established kidney tumour endoscopic dataset and publicly available endoscopic datasets show that this method exhibits enhanced performance and generalization capabilities compared to existing techniques. On this renal tumour dataset, MDSK‐Net achieved excellent results of 94.1% and 90.1% on mDice and mIoU. Minpeng Jiang, LeiLei Li, Zhengping Li, Chao Nie, Tianyu Zheng, Longyu Li |
IET Image Process. | 6 |
| 2024 | SWattention: designing fast and memory-efficient attention for a new Sunway SupercomputerabstractAbstract In the past few years, Transformer-based large language models (LLM) have become the dominant technology in a series of applications. To scale up the sequence length of the Transformer, FlashAttention is proposed to compute exact attention with reduced memory requirements and faster execution. However, implementing the FlashAttention algorithm on the new generation Sunway Supercomputer faces many constraints such as the unique heterogeneous architecture and the limited memory bandwidth. This work proposes SWattention, a highly efficient method for computing the exact attention on the SW26010pro processor. To fully utilize the 6 core groups (CG) and 64 cores per CG on the processor, we design a two-level parallel task partition strategy. Asynchronous memory access is employed to ensure that memory access overlaps with computation. Additionally, a tiling strategy is introduced to determine optimal SRAM block sizes. Compared with the standard attention, SWattention achieves around 2.0x speedup for FP32 training and 2.5x speedup for mixed-precision training. The sequence lengths range from 1k to 8k and scale up to 16k without being out of memory. As for the end-to-end performance, SWattention achieves up to 1.26x speedup for training GPT-style models, which demonstrates that SWattention enables longer sequence length for LLM training. Ruohan Wu, Xianyu Zhu, Junshi Chen 0003, Tianyu Zheng, Xin Liu 0081, Hong An |
J. Supercomput. | 5 |
| 2023 | Leaking Arbitrarily Many Secrets: Any-out-of-Many Proofs and Applications to RingCT ProtocolsabstractRing Confidential Transaction (RingCT) protocol is an effective cryptographic component for preserving the privacy of cryptocurrencies. However, existing RingCT protocols are instantiated from one-out-of-many proofs with only one secret, leading to low efficiency and weak anonymity when handling transactions with multiple inputs. Additionally, current partial knowledge proofs with multiple secrets are neither secure nor efficient to be applied in a RingCT protocol.In this paper, we propose a novel any-out-of-many proof, a logarithmic-sized zero-knowledge proof scheme for showing the knowledge of arbitrarily many secrets out of a public list. Unlike other partial knowledge proofs that have to reveal the number of secrets [ACF21], our approach proves the knowledge of multiple secrets without leaking the exact number of them. Furthermore, we improve the efficiency of our method with a generic inner-product transformation to adopt the Bulletproofs compression [BBB+18], which reduces the proof size to 2⌈log2(N)⌉+9.Based on our proposed proof scheme, we further construct a compact RingCT protocol for privacy cryptocurrencies, which can provide a logarithmic-sized communication complexity for transactions with multiple inputs. More importantly, as the only known RingCT protocol instantiated from the partial knowledge proofs, our protocol can achieve the highest anonymity level compared with other approaches like Omniring [LRR+19]. For other applications, such as multiple ring signatures, our protocol can also be applied with some modifications. We believe our techniques are also applicable in other privacy-preserving scenarios, such as multiple ring signatures and coin-mixing in the blockchain. Tianyu Zheng, Shang Gao 0006, Yubo Song, Bin Xiao 0001 |
SP | 1 |
| 2022 | BaGuaLu: targeting brain scale pretrained models with over 37 million coresabstractLarge-scale pretrained AI models have shown state-of-the-art accuracy in a series of important applications. As the size of pretrained AI models grows dramatically each year in an effort to achieve higher accuracy, training such models requires massive computing and memory capabilities, which accelerates the convergence of AI and HPC. However, there are still gaps in deploying AI applications on HPC systems, which need application and system co-design based on specific hardware features. Zixuan Ma, Jiaao He, Jiezhong Qiu, Huanqi Cao, Yuanwei Wang, Zhenbo Sun, Liyan Zheng 0001, Haojie Wang 0004, Shizhi Tang, Tianyu Zheng, Junyang Lin, Guanyu Feng, Zeqiang Huang, Aohan Zeng, Jianwei Zhang 0012, Runxin Zhong, Tianhui Shi, Jie Tang 0001, Hongxia Yang, Xin Liu 0086, Jidong Zhai |
PPoPP | 10 |
| 2021 | Establishing high performance AI ecosystem on Sunway platform
Xin Liu 0081, Zeqiang Huang, Tianyu Zheng |
CCF Trans. High Perform. Comput. | 5 |
| 2021 | Characterization of base station deployment distribution and coverage in heterogeneous networksabstractAbstract Considering different types of base stations (BSs) in future cellular networks are overlapping deployment with the status of dense, multi‐tier and heterogeneous in general, how to optimize the real BS deployment becomes a complicated problem. Based on it, repulsive BS dataset and clustering BS dataset are statically characterized with various types of spatial point processes. It shows that the improvement of coverage probability between 1‐tier and 2‐tier network fluctuates as the signal‐to‐interference ratio (SIR) threshold increases for different datasets. The authors' proposed hybrid model fits well with repulsive BS dataset, while the log‐Gaussian Cox point process (LGCP) and Cauchy model are reasonable models for clustering BS dataset. Besides, in order to dynamically analyze the coverage problem affected by adding new BSs, a cell boundary constructed by an irregular circle is introduced under an equal SIR constraint, and a BS placement scheme is proposed to place new BSs at the points of minimum interference. Numerical results show that the coverage probability may increase after adding BSs in the target area of heterogeneous network by using the proposed scheme. However, as the density of femto BS reaches a certain value, its coverage may remain unchanged even after adding more femto BSs. Tianyu Zheng, Dan Keun Sung |
IET Commun. | 3 |
| 2020 | Active switching multiple model method for tracking a noncooperative gliding flight vehicle
Tianyu Zheng, Yu Yao 0004, Fenghua He 0001, Denggao Ji |
Sci. China Inf. Sci. | 1 |