Tianshuo Yang

dblp:298/0830 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
0009-0006-9961-090XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Why Imperfect Team Management is Tolerable: An E-CARGO's Perspective
abstract
Team management, especially assignment, is important and challenging in governance, administration, and organizations. Managers try to use their wisdom, experience, and knowledge to make the managed team better. However, without effective tools, managers always make mistakes or make decisions with their bias, making the management imperfect. Many cases happen, i.e., that good intentions lead to bad consequences, and that official corruption cannot be extinguished. This paper explores the question of why people often tolerate imperfect management, drawing on the (Environments, Classes, Agents, Roles, Groups, and Objects) E-CARGO model to provide insights from a Group Role Assignment (GRA) perspective. In many organizational and social systems, individuals display resilience and adaptability to imperfect management. Using E-CARGO, we examine the factor of team performance that contributes to tolerance for management imperfections, mainly on role assignment. We use simulations to present that by simply accumulating the individual performance values to express team performance, imperfect assignment does not damage the optimized result much. This study uses a few management styles to simulate imperfect assignment, like, (WUWEIERZHI, i.e., rule by doing nothing), (ZEYOUSHANGGANG, i.e., select the best), and (SHIKEERZHI, i.e., know when to stop).
Haibin Zhu 0001, Tianshuo Yang
CSCWD2
2025 Lumina-T2X: Scalable Flow-based Large Diffusion Transformer for Flexible Resolution Generation
abstract
Sora unveils the potential of scaling Diffusion Transformer (DiT) for generating photorealistic images and videos at arbitrary resolutions, aspect ratios, and durations, yet it still lacks sufficient implementation details. In this paper, we introduce the Lumina-T2X family -- a series of Flow-based Large Diffusion Transformers (Flag-DiT) equipped with zero-initialized attention, as a simple and scalable generative framework that can be adapted to various modalities, e.g., transforming noise into images, videos, multi-view 3D objects, or audio clips conditioned on text instructions. By tokenizing the latent spatial-temporal space and incorporating learnable placeholders such as |[nextline]| and |[nextframe]| tokens, Lumina-T2X seamlessly unifies the representations of different modalities across various spatial-temporal resolutions. Advanced techniques like RoPE, KQ-Norm, and flow matching enhance the stability, flexibility, and scalability of Flag-DiT, enabling models of Lumina-T2X to scale up to 7 billion parameters and extend the context window to 128K tokens. This is particularly beneficial for creating ultra-high-definition images with our Lumina-T2I model and long 720p videos with our Lumina-T2V model. Remarkably, Lumina-T2I, powered by a 5-billion-parameter Flag-DiT, requires only 35% of the training computational costs of a 600-million-parameter naive DiT (PixArt-alpha), indicating that increasing the number of parameters significantly accelerates convergence of generative models without compromising visual quality. Our further comprehensive analysis underscores Lumina-T2X's preliminary capability in resolution extrapolation, high-resolution editing, generating consistent 3D views, and synthesizing videos with seamless transitions. All code and checkpoints of Lumina-T2X are released at https://github.com/Alpha-VLLM/Lumina-T2X to further foster creativity, transparency, and diversity in the generative AI community.
Peng Gao 0007, Le Zhuo, Ruoyi Du, Longtian Qiu, Rongjie Huang 0001, Shijie Geng, Renrui Zhang, Junlin Xie, Wenqi Shao, Zhengkai Jiang 0001, Tianshuo Yang, Weicai Ye, Tong He 0001, Jingwen He, Junjun He, Yu Qiao 0001, Hongsheng Li 0001
ICLR14
2025 MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
abstract
The capability to process multiple images is crucial for Large Vision-Language Models (LVLMs) to develop a more thorough and nuanced understanding of a scene. Recent multi-image LVLMs have begun to address this need. However, their evaluation has not kept pace with their development. To fill this gap, we introduce the Multimodal Multi-image Understanding (MMIU) benchmark, a comprehensive evaluation suite designed to assess LVLMs across a wide range of multi-image tasks. MMIU encompasses 7 types of multi-image relationships, 52 tasks, 77K images, and 11K meticulously curated multiple-choice questions, making it the most extensive benchmark of its kind. Our evaluation of nearly 30 popular LVLMs, including both open-source and proprietary models, reveals significant challenges in multi-image comprehension, particularly in tasks involving spatial understanding. Even the most advanced models, such as GPT-4o, achieve only 55.7\% accuracy on MMIU. Through multi-faceted analytical experiments, we identify key performance gaps and limitations, providing valuable insights for future model and data improvements. We aim for MMIU to advance the frontier of LVLM research and development. We release the data and code at https://github.com/MMIUBenchmark/MMIU.
Fanqing Meng, Chuanhao Li 0001, Quanfeng Lu, Hao Tian 0006, Tianshuo Yang, Jiaqi Liao, Xizhou Zhu, Jifeng Dai, Yu Qiao 0001, Ping Luo 0002, Kaipeng Zhang, Wenqi Shao
ICLR6
2025 4DSloMo: 4D Reconstruction for High Speed Scene with Asynchronous Capture
abstract
Reconstructing fast-dynamic scenes from multi-view videos is crucial for high-speed motion analysis and realistic 4D reconstruction. However, the majority of 4D capture systems are limited to frame rates below 30 FPS (frames per second), and a direct 4D reconstruction of high-speed motion from low FPS input may lead to undesirable results. In this work, we propose a high-speed 4D capturing system only using low FPS cameras, through novel capturing and processing modules. On the capturing side, we propose an asynchronous capture scheme that increases the effective frame rate by staggering the start times of cameras. By grouping cameras and leveraging a base frame rate of 25 FPS, our method achieves an equivalent frame rate of 100–200 FPS without requiring specialized high-speed cameras. On processing side, we also propose a novel generative model to fix artifacts caused by 4D sparse-view reconstruction, as asynchrony reduces the number of viewpoints at each timestamp. Specifically, we propose to train a video-diffusion-based artifact-fix model for sparse 4D reconstruction, which refines missing details, maintains temporal consistency, and improves overall reconstruction quality. Experimental results demonstrate that our method significantly enhances high-speed 4D reconstruction compared to synchronous capture. Project page: https://openimaginglab.github.io/4DSloMo/
Shi Guo, Tianshuo Yang, Lihe Ding, Xiuyuan Yu, Jinwei Gu, Tianfan Xue
SIGGRAPH Asia3
2024 Adaptive Group Multi-Role Assignment in Garbage Management System
abstract
In addition to prompt garbage collection and strategic routing, effective waste management also involves optimizing resource allocation and ensuring the sustainability of disposal practices. The E-CARGO (Environments – Classes, Agents, Roles, Groups, and Objects) model, coupled with the Role-Based Collaboration (RBC) methodology, offers a sophisticated approach to address these multifaceted challenges. By simulating the scheduling of garbage trucks for the transportation of bins within urban environments, the model provides valuable insights into the dynamics of waste collection operations. Furthermore, by integrating RBC, which dynamically assigns roles to agents involved in the waste management process, the model promotes coordination and cooperation among various entities, including municipal authorities, waste collection agencies, and residents. Through this adaptive framework, the article proposes a holistic approach to urban waste management that not only enhances operational efficiency but also fosters environmental sustainability and community engagement.
Tianshuo Yang, Haibin Zhu 0001, Phil Xing Yang
SMC1
2023 New Employee Training Scheduling Using the E-CARGO Model
abstract
New employee training scheduling is one of the most common events in many enterprises. Solving this problem has its significance and is useful in daily administrations and operations. Group Role Assignment (GRA) model is widely applied in the assignment problem. However, there are still many challenges to applying the GRA model. For example, when we need to assign different jobs for the same person at different times, GRA needs more structures to specify constraints. If we use the strategy that combines the time factor with the agents or roles to formalize new agents or roles, the problem can be converted to a solvable GRA problem with constraints. The focus of this article is to give a practical solution to this kind of problem by using the GRA formulations in expressing constraints. The formalization makes us resolve the problem easily through integer programming (IP) with the PuLP package of Python. Large-scale simulation experiments demonstrate the practicability and robustness of our method.
Tianshuo Yang, Haibin Zhu 0001
CSCWD1
2023 iVS-Net: Learning Human View Synthesis from Internet Videos
abstract
Recent advances in implicit neural representations make it possible to generate free-viewpoint videos of the human from sparse view images. To avoid the expensive training for each person, previous methods adopt the generalizable human model and demonstrate impressive results. However, these methods usually rely on limited multi-view images typically collected in the studio or commercial high-quality 3D scans for training, which heavily prohibits their generalization capability for in-the-wild images. To solve this problem, we propose a new approach to learn a generalizable human model from a new source of data, i.e., Internet videos. These videos capture various human appearances and poses and record the performers from abundant viewpoints. To exploit the Internet data, we present a video self-supervised pipeline to enforce the local appearance consistency of each body part over different frames of the same video. Once learned, the human model enables realistic novel view synthesis from a single input image. Experiments show that our method can generate high-quality view synthesis on in-the-wild images while only training on monocular videos.
Junting Dong, Tianshuo Yang, Qing Shuai, Chengyu Qiao, Sida Peng
ICCV3