Yunjia Li

dblp:39/7451 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-5728-9795ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Teaching Audio-Language Models to Reason over Time
abstract
Large Audio-Language Models (LALMs) have achieved remarkable progress in general audio understanding. Nevertheless, they still exhibit significant limitations when reasoning about the complex temporal relationships between sound events. This bottleneck arises from two key problems: the lack of a systematic benchmark for evaluation and the absence of a reasoning-oriented training paradigm for audio. To resolve this, we introduce TASC-Bench (Temporal Audio Scene Comprehension Benchmark), a large-scale question-answering (QA) benchmark focused on the temporal understanding of audio scenes, built upon AudioSet. TASC-Bench comprises over 320 hours of audio and 9.5 million QA pairs, covering seven dimensions, including temporal boundaries, event durations, and inter-event relations. To teach LALMs to reason over time and overcome the limitations of Supervised Fine-Tuning (SFT) in tackling complex temporal reasoning, we propose a cross-modal Chain-of-Thought (CoT) distillation strategy. This progressive approach begins by bootstrapping the model’s reasoning capabilities using a large language model (LLM) as a teacher to generate CoT, then transitions to self-driven reasoning, and culminates in aligning its behavior with preference data. For this final alignment stage, we introduce our improved Equilibrated Chain Direct Preference Optimization (EC-DPO) method alongside SFT to mitigate hallucinations and enhance stability. Our experiments demonstrate that our dataset reveals the current shortcomings of LALMs in temporal reasoning, while our method significantly enhances these capabilities. We will make the TASC-Bench dataset publicly available to foster further research in audio temporal reasoning.
Yunjia Li, Wei Li 0012
ICMR2
2025 A Multifaceted Multi-Agent Framework for Zero-Shot Emotion Analysis and Recognition of Symbolic Music
abstract
ICMI '25: International Conference on Multimodal Interaction; Canberra, Australia; October 13 - 17, 2025
Yunjia Li, Kazuyoshi Yoshii
ICMI2
2024 A Self-powered Wireless Sensor System for Monitoring the Deformation of Gas Insulated Switchgear Expansion Joint
abstract
This work reports on the design and implementation of a self-powered wireless GIS expansion joint deformation monitoring system based on extremely low amplitude stray magnetic field energy harvester. A prototype of the wireless deformation monitoring system that can be powered by energy harvesting technology is designed, fabricated, and characterized. A cantilever beam-type stray magnetic field energy harvester (MEH) based on the magneto-mechano-electric coupling effect is designed for harvesting stray magnetic field energy from the environment around the GIS expansion joint. The proposed MEH device is capable of generating an open-circuit output peak-to-peak voltage of 4.24 V and a power of 113.1 μW in a magnetic field environment of 50 μT. The system is of great significance to improve the safe operation of the power system and meets the technical requirements of building passive wireless Internet of Things (IoT).
Xitong Sun, Yunjia Li
IECON6
2023 Electromagnetic Vibration Energy Harvester for Low Frequency and High Amplitude Applications
abstract
An electromagnetic vibrational energy harvester (EVEH) capable of harvesting low-frequency and high-amplitude vibration energy is designed, fabricated and characterized in this paper. The high impact/shock-resistance of the proposed EVEH is enabled by a pair of specially designed dampers. The dampers, combined with a pair of nonlinear magnetic springs, also expands the effective bandwidth of the EVEH device. The fabricated EVEH, with a size of Φ2 cm X 10 cm, is capable of generating an open-circuit peak-to-peak voltage of 4.35 V and an average output power of 55.3 mW under a sinusoidal excitation of ±10g at 48 Hz. The designed shock/impact-resistance of the EVEH is verified by a sinusoidal excitation of ±20g at 18 Hz, and the fabricated EVEH is capable of generating an output voltage of 4.21V steadily.
Qinghong Zhang, Siyu Deng, Weitao Dou, Yunjia Li
IECON6
2021 Torsional Electromagnetic Vibrational Energy Harvester Based on Stacked Flexible Coils
abstract
In this work, we report on an electromagnetic vibrational energy harvester (EVEH), based on the torsional movement of a permanent magnet over a stack of flexible planar coils. The disc magnet is glued to a silicon plate suspended by straight torsional springs. With a size of 1 cm×1 cm×1.08 cm, the proposed EVEH is capable of generating an open-circuit peak-to-peak voltage of 138.9 mV and a power of 4.6 μW, under a sinusoidal excitation of ±0.5g and frequency of 104 Hz. At elevated acceleration levels, the maximum peak-to-peak output voltage is 185.2 mV under the acceleration of 7g (±3.5g).
Jiaxing Li 0014, Chenyuan Zhou, Kai Tao, Dayong Qiao, Yunjia Li
IECON6
2019 Quantitative Detection of Mixed Gases by Sensor Array Using C-Means Clustering and Artificial Neural Network
abstract
Due to the cross-sensitivity to kinds of gases, traditional single sensor is impossible to selectively detect gases, which limits its application. In this paper, a sensor array with supervised and unsupervised algorithms was employed to selectively detect NO2, CO and their mixtures. To improve the recognition accuracy, average resistance over a period of time was introduced to acquire features, including response value, response time and recovery time. Firstly, principal component analysis (PCA) was utilized to reduce the dimensionality of samples. With the help of unsupervised fuzzy C-means clustering algorithm, samples have exhibited obvious clustering. Furthermore, artificial neural network (ANN) has reached a high accuracy of quantitative identification on six different detected gases varying from 0 to 50 ppm.
Jifeng Chu, Mingzhe Rong, Weijuan Li, Dawei Wang 0011, Chengyu Fan, Aijun Yang, Yunjia Li, Xiaohua Wang 0001
IECON9
2019 A miniaturized electromagnetic energy harvester with off-axis magnet and stacked flexible coils
abstract
In this work, a miniaturized Electromagnetic Vibration Energy Harvester (EVEH) device based on stacked flexible coils and flexible springs is designed and characterized. The energy harvesting is realized by the combined piston and torsional movement of the magnet. At an acceleration of ±1g peak-to-peak amplitude (±9.8 m/s2), a maximum output peak-to-peak voltage of 704.28 mV and an output power of 20.75 μW are measured, at a frequency of 58 Hz.
Yunjia Li, Aijun Yang, Kai Tao, Dayong Qiao
IECON1
2019 Spherical electret generator for water wave energy harvesting by folded structure
abstract
With the increasing demand for energy resource, the development of marine energy harvesting technology has attracted widespread attention. Herein, a spherical electret wave power generator is designed and constructed for ocean wave energy harvesting. Based on electrostatic induction from electret material, the generator converts wave energy into electric energy through the bias voltage of the folded electret power generation unit. With ultra-low excitation frequency at 2Hz, the peak-to-peak output voltage and output power can be reached up to around 300 V and 20 mW, respectively. Compared with other traditional wave energy generators, the utilization of space in the sphere is improved by the folded structure, which also has the advantages in terms of simple structure, high-energy conversion efficiency, light weight and small size. The generator is also capable of collecting low-frequency vibration energy in random directions, as well as realizing large-scale power supply after deploying an array of structures.
Shishi Li, Haiping Yi, Yunjia Li, Honglong Chang, Kai Tao
IECON7
2019 A Carbon Nanotube Based Ionization Sensor Array to Gas Mixture
abstract
The three kinds of gases such as NO2, NO, and SO2produced by industrial combustion process are the main sources of atmospheric pollution. Almost all the available technologies of measuring three components in mixed gases are limited by detection range, operating temperature, and integration, etc. Here we report an array with three triple-electrode carbon nanotube sensors and its sensitivity to the three flue gas components. The three sensors operate at the same gas ionization mechanism and are comprised of a common long cathode, three extracting electrodes and three collecting electrodes, and are set by different electrode separations with$75\ \boldsymbol{\mu} \mathbf{m}, 100\ \boldsymbol{\mu} \mathbf{m}$, and$120\ \boldsymbol{\mu} \mathbf{m}$, respectively. We explored and obtained distinct single-valued sensitive characteristics to the three-component gases at given voltages applied on electrodes, which should be higher than a critical voltage to obtain single-valued sensitivities of the three sensors. The standard deviation of repeatability experiment is less than 0.133 nA, which indicates that the sensor has a good repeatability. The array displays a high level of integration and small size, and also has the potential to detect the concentrations of different components of gas mixtures.
Yunjia Li, Wen Han, Pinghai Lyu, Zhenzhen Cheng, Xu Song, Longlong Ke
IECON2
2019 Recommendations from Cold Starts in Big Data
abstract
This paper examines the challenging problem of new user cold starts in subset labelled and extremely sparsely labelled big data. We introduce a new Isle of Wight Supply Chain (IWSC) dataset demonstrating these characteristics. We also introduce a new technique addressing these challenges, the Transitive Semantic Relationships (TSR) model, which infers potential relationships from user and item text content and few labelled examples. We perform both implicit and explicit evaluation of TSR as a recommender system and from new user cold starts we achieve a hit-rate@10 of 77% on a collection of 630 items with only 376 supply-chain consumer labels, and 67% with only 142 supply-chain supplier labels, demonstrating a high level of performance even with extremely few labels in challenging cold-start scenarios. TSR is suitable for any dataset featuring few labels and user and item content, where similarity of content indicates similar relationship forming capability. TSR can be used as a standalone recommender system or to complement existing high-performance recommender models that require more labels or do not support cold starts.
David Ralph, Yunjia Li, Gary B. Wills, Nicolas G. Green
IoTBDS2
2017 Research on Link Layer Topology Discovery Algorithm Based on Dynamic Programming
Yunjia Li
ICIC (3)1
2014 Synote Second Screening: Using Mobile Devices for Video Annotation and Control
Mike Wald, Yunjia Li, George Cockshull, David Hulme, Douglas Moore, Aidan Purdy-Say
ICCHP (1)2
2012 Synote: Important Enhancements to Learning with Recorded Lectures
abstract
This paper explains three new important enhancements to Synote, the freely available, award winning, open source, web based application that makes web hosted recordings easier to access, search, manage, and exploit for learners, teachers and other users. Synote uniquely achieves this through the creation of synchronized notes, bookmarks, tags, links, images and text captions, enabling users to easily find, or associate their notes or resources with, any part of a recording available on the web. Students surveyed would like to be able to access all their lectures through Synote. The facility to convert and import narrated PowerPoint PPTX files means that teachers can capture their lectures without requiring institution-wide expensive lecture capture systems. Crowdsourcing correction of speech recognition errors allows for sustainable captioning of the lecture while the development of an integrated mobile speech recognition application enables synchronized live verbal contributions from the class to also be captured through captions.
Mike Wald, Yunjia Li
ICALT2
2009 Synchronised Annotation of Multimedia
abstract
Multimedia has become technically easier to create (e.g. recording lectures) but while users can easily bookmark, search, link to, or tag the WHOLE of a podcast or video recording available on the web they cannot easily find, or associate their notes or resources with, PART of that recording. This paper describes the development of a web based application that makes multimedia web resources (e.g. podcasts) easier to access, search, manage, and exploit for learners, teachers and other users through the creation of notes, bookmarks, tags, links, images and text captions synchronized to any part of the recording.
Mike Wald, Gary B. Wills, David E. Millard, Lester Gilbert, Shakeel Ahmed Khoja, Jiri Kajaba, Yunjia Li
ICALT7