VLDB 2026 Research / reviewers in the wild / expert
Mingyu Yang 0002
dblp:197/8171-2
· DBLP profile ↗
10ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0003-1301-6493ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SAM-Guided Pseudo Label Enhancement for Multi-Modal 3D Semantic SegmentationabstractMulti-modal 3D semantic segmentation is vital for applications such as autonomous driving and virtual reality (VR). To effectively deploy these models in real-world scenarios, it is essential to employ cross-domain adaptation techniques that bridge the gap between training data and real-world data. Recently, self-training with pseudo-labels has emerged as a predominant method for cross-domain adaptation in multi-modal 3D semantic segmentation. However, generating reliable pseudo-labels necessitates stringent constraints, which often result in sparse pseudo-labels after pruning. This sparsity can potentially hinder performance improvement during the adaptation process. We propose an image-guided pseudo-label enhancement approach that leverages the complementary 2D prior knowledge from the Segment Anything Model (SAM) to introduce more reliable pseudo-labels, thereby boosting domain adaptation performance. Specifically, given a 3D point cloud and the SAM masks from its paired image data, we collect all 3D points covered by each SAM mask that potentially belong to the same object. Then our method refines the pseudo-labels within each SAM mask in two steps. First, we determine the class label for each mask using majority voting and employ various constraints to filter out unreliable mask labels. Next, we introduce Geometry-Aware Progressive Propagation (GAPP) which propagates the mask label to all 3D points within the SAM mask while avoiding outliers caused by 2D-3D misalignment. Experiments conducted across multiple datasets and domain adaptation scenarios demonstrate that our proposed method significantly increases the quantity of high-quality pseudo-labels and enhances the adaptation performance over baseline methods. Mingyu Yang 0002, Jitong Lu, Hun-Seok Kim |
ICRA | 1 |
| 2025 | NBLoc: A Narrowband RF Localization System for Wide-Area Indoor ApplicationsabstractWe introduce NBLoc, a narrowband frequency hopping, long-range localization system designed for low-power Internet of Things (IoT) devices. Traditional high-accuracy localization systems typically require wide-bandwidth and power-demanding radio frequency (RF) circuits, leading to limitations such as short operational range and high power consumption to achieve decimeter-level accuracy. NBLoc overcomes these challenges by using narrowband symbols with a frequency-hopping mechanism across a wide bandwidth, enabling the localization of low-power tags over large areas. NBLoc features a novel custom-designed RF analog frontend (AFE) integrated circuit (IC), that eliminates the need for a conventional phase-locked loop, significantly reducing the cost and power consumption of the receiving tag. This advancement is enabled by NBLoc's thoughtful waveform design and specialized signal processing algorithms, which mitigate phase noise and uncertainty. Compared to previous solutions, NBLoc achieves lower power consumption and extended operational range due to its narrowband symbols while maintaining high localization accuracy by leveraging a 100 MHz localization bandwidth through frequency hopping. In NBLoc, system anchors transmit narrowband orthogonal symbols, hopping across the localization bandwidth in a predetermined pattern known to the tag. The tag, equipped with the custom low-power RF AFE IC, dynamically tunes its local oscillator (LO) frequency to match the hop pattern and capture these symbols, which are then used to estimate the channel impulse response (CIR). The tag calculates the time difference of arrival (TDOA) for each anchor pair from the CIRs, and determines its 2D location via multilateration. The system was implemented and tested using the low-power RF AFE IC in both line-of-sight (LOS) and non-line-of-sight (NLOS) environments, achieving decimeter-level accuracy across areas as large as 269 × 125 m$^{2}$. Demba Komma, Chien-Wei Tseng, Andrea Bejarano-Carbo, Mingyu Yang 0002, David T. Blaauw, Hun-Seok Kim |
IEEE Trans. Mob. Comput. | 4 |
| 2023 | Search for Efficient Deep Visual-Inertial Odometry Through Neural Architecture SearchabstractRecent deep learning based visual-inertial odometry (VIO) systems achieve impressive performance in various applications and challenging scenarios. However, it is difficult to deploy such VIO models directly on energy-constrained mobile platforms in real-time due to the extensive complexity of existing deep neural network (DNN) models. To address this issue, we propose to adopt the neural architecture search (NAS) technique to search for the most efficient VIO network architecture. Targeting the lowest number of operations and inference latency, our searched models achieve up to 97.4% complexity reduction with no performance degradation. The searched efficient visual encoder allows our VIO model to run at 83.3 frames per second on a single laptop CPU core. Moreover, the model complexity can be reduced by 99.1% when combined with a dynamic modality selection technique. Our searched efficient VIO models are available at https://github.com/unchenyu/NASVIO. Yu Chen 0070, Mingyu Yang 0002, Hun-Seok Kim |
ICASSP | 2 |
| 2023 | Efficient Computation Sharing for Multi-Task Visual Scene UnderstandingabstractSolving multiple visual tasks using individual models can be resource-intensive, while multi-task learning can conserve resources by sharing knowledge across different tasks. Despite the benefits of multi-task learning, such techniques can struggle with balancing the loss for each task, leading to potential performance degradation. We present a novel computation- and parameter-sharing framework that balances efficiency and accuracy to perform multiple visual tasks utilizing individually-trained single-task transformers. Our method is motivated by transfer learning schemes to reduce computational and parameter storage costs while maintaining the desired performance. Our approach involves splitting the tasks into a base task and the other sub-tasks, and sharing a significant portion of activations and parameters/weights between the base and sub-tasks to decrease inter-task redundancies and enhance knowledge sharing. The evaluation conducted on NYUD-v2 and PASCAL-context datasets shows that our method is superior to the state-of-the-art transformer-based multi-task learning techniques with higher accuracy and reduced computational resources. Moreover, our method is extended to video stream inputs, further reducing computational costs by efficiently sharing information across the temporal domain as well as the task domain. Our codes are available at https://github.com/sarashoouri/EfficientMTL. Sara Shoouri, Mingyu Yang 0002, Zichen Fan, Hun-Seok Kim |
ICCV | 2 |
| 2022 | Efficient Deep Visual and Inertial Odometry with Adaptive Visual Modality Selection
Mingyu Yang 0002, Yu Chen 0070, Hun-Seok Kim |
ECCV (38) | 1 |
| 2022 | Deep Joint Source-Channel Coding for Wireless Image Transmission with Adaptive Rate ControlabstractWe present a novel adaptive deep joint source-channel coding (JSCC) scheme for wireless image transmission. The proposed scheme supports multiple rates using a single deep neural network (DNN) model and learns to dynamically control the rate based on the channel condition and image contents. Specifically, a policy network is introduced to exploit the tradeoff space between the rate and signal quality. To train the policy network, the Gumbel-Softmax trick is adopted to make the policy network differentiable and hence the whole JSCC scheme can be trained end-to-end. To the best of our knowledge, this is the first deep JSCC scheme that can automatically adjust its rate using a single network model. Experiments show that our scheme successfully learns a reasonable policy that decreases channel bandwidth utilization for high SNR scenarios or simple image contents. For an arbitrary target rate, our rate-adaptive scheme using a single model achieves similar performance compared to an optimized model specifically trained for that fixed target rate. To reproduce our results, we make the source code publicly available at https://github.com/mingyuyng/Dynamic_JSCC. Mingyu Yang 0002, Hun-Seok Kim |
ICASSP | 1 |
| 2022 | Deep Learning Based Near-Orthogonal Superposition Code for Short Message TransmissionabstractMassive machine type communication (mMTC) has attracted new coding schemes optimized for reliable short message transmission. In this paper, a novel deep learning based near-orthogonal superposition (NOS) coding scheme is proposed for reliable transmission of short messages in the additive white Gaussian noise (AWGN) channel for mMTC applications. Similar to recent hyper-dimensional modulation (HDM), the NOS encoder spreads the information bits to multiple near-orthogonal high dimensional vectors to be combined (superimposed) into a single vector for transmission. The NOS decoder first estimates the information vectors and then performs a cyclic redundancy check (CRC)-assisted K-best tree-search algorithm to further reduce the packet error rate. The proposed NOS encoder and decoder are deep neural networks (DNNs) jointly trained as an auto-encoder and decoder pair to learn a new NOS coding scheme with near-orthogonal codewords. Simulation results show the proposed deep learning-based NOS scheme outperforms HDM and Polar code with CRC-aided list decoding for short (32-bit) message transmission. Chenghong Bian, Mingyu Yang 0002, Chin-Wei Hsu, Hun-Seok Kim |
ICC | 2 |
| 2021 | Deep Joint Source Channel Coding for Wireless Image Transmission with OFDMabstractWe present a deep learning based joint source channel coding (JSCC) scheme for wireless image transmission over multipath fading channels with non-linear signal clipping. The proposed encoder and decoder use convolutional neural networks (CNN) and directly map the source images to complex-valued baseband samples for orthogonal frequency division multiplexing (OFDM) transmission. The proposed model-driven machine learning approach eliminates the need for separate source and channel coding while integrating an OFDM datapath to cope with multipath fading channels. The end-to-end JSCC communication system combines trainable CNN layers with non-trainable but differentiable layers representing the multipath channel model and OFDM signal processing blocks. Our results show that injecting domain expert knowledge by incorporating OFDM baseband processing blocks into the machine learning framework significantly enhances the overall performance compared to an unstructured CNN. Our method outperforms conventional schemes that employ state-of-the-art but separate source and channel coding such as BPG and LDPC with OFDM. Moreover, our method is shown to be robust against non-linear signal clipping in OFDM for various channel conditions that do not match the model parameter used during the training. Mingyu Yang 0002, Chenghong Bian, Hun-Seok Kim |
ICC | 1 |
| 2021 | mSAIL: milligram-scale multi-modal sensor platform for monarch butterfly migration trackingabstractEach fall, millions of monarch butterflies across the northern US and Canada migrate up to 4,000 km to overwinter in the exact same cluster of mountain peaks in central Mexico. To track monarchs precisely and study their navigation, a monarch tracker must obtain daily localization of the butterfly as it progresses on its 3-month journey. And, the tracker must perform this task while having a weight in the tens of milligram (mg) and measuring a few millimeters (mm) in size to avoid interfering with monarch's flight. This paper proposes mSAIL, 8 × 8 × 2.6 mm and 62 mg embedded system for monarch migration tracking, constructed using 8 prior custom-designed ICs providing solar energy harvesting, an ultra-low power processor, light/temperature sensors, power management, and a wireless transceiver, all integrated and 3D stacked on a micro PCB with an 8 × 8 mm printed antenna. The proposed system is designed to record and compress light and temperature data during the migration path while harvesting solar energy for energy autonomy, and wirelessly transmit the data at the overwintering site in Mexico, from which the daily location of the butterfly can be estimated using a deep learning-based localization algorithm. A 2-day trial experiment of mSAIL attached on a live butterfly in an outdoor botanical garden demonstrates the feasibility of individual butterfly localization and tracking. Inhee Lee 0001, Roger Hsiao, Gordy A. Carichner, Chin-Wei Hsu, Mingyu Yang 0002, Sara Shoouri, Katherine Ernst, Tess Carichner, Yuyang Li 0001, Jaechan Lim, Cole R. Julick, Eunseong Moon, Jamie Phillips, Kristi L. Montooth, Delbert A. Green II, Hun-Seok Kim, David T. Blaauw |
MobiCom | 5 |
| 2019 | iLPS: Local Positioning System with Simultaneous Localization and Wireless CommunicationabstractThis paper presents a novel RF local positioning system, iLPS, specifically designed for challenging indoor non-lineof-sight (NLOS) scenarios and/or urban canyons where global positioning systems (GPS) fail to reliably operate. iLPS enables decimeter-level localization of numerous tags concurrently with wireless communication using frequency-shifting active reflector anchors and orthogonal frequency division multiple access (OFDMA) waveforms. OFDMA signals are devised with carefully assigned subcarriers so that each tag can estimate the time-difference-of-arrival (TDoA) by analyzing the channel impulse response (CIR) from the main and reflector anchors without interfering each other. The proposed active reflection scheme efficiently eliminates the stringent time synchronization requirement while providing the diversity gain for enhanced information decoding reliability at the tag. Significant challenges from NLOS multipaths and the usage of relatively narrow bandwidth of 80MHz in the ISM-band are successfully mitigated by machine learning assisted algorithms. Field trials with the prototype system on Universal Software Radio Peripheral (USRP) confirm that iLPS can achieve decimeter-level accuracy localization and concurrent wireless communication over up to 100m distances. Mingyu Yang 0002, Li-Xuan Chuo, Karan Suri, Hun-Seok Kim |
INFOCOM | 1 |