Da Yang 0001

dblp:64/6513-1 · DBLP profile ↗
← Back
5ranked-venue papers in the field
0as first author
4since 2021 · last 2026
0000-0001-5782-894XORCID · conflict

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 4Other / Interdisciplinary · 1
YearPublicationVenuePosition
2026 Bridging the Gap: More Powerful Residual Fusion for Deep BNNs
Chengshuo Bai, Shuai Wang 0027, Hao Sheng 0001, Da Yang 0001, Hailong Zhao, Guanqun Su
KSEM (4)4
2025 Expert Data - Assisted Diagnosis: An INFO - iTransformer - XGBoost Combined Discriminative System for Prenatal Diagnosis of Fetal Congenital Heart Disease
Hao Sheng 0001, Xiaoyan Gu 0005, Jiancheng Han, Da Yang 0001, Xuefei Huang, Yihua He, Haogang Zhu
KSEM (4)7
2025 Depth State Space Model for Light Field Depth Estimation via Text-Similar Representation
Zexin Sun, Tun Wang, Da Yang 0001, Zhenglong Cui, Rongshan Chen, Ying Li 0122, Guanqun Su, Hao Sheng 0001
KSEM (1)3
2025 InstructTrack: Language-Guided Multi-Object Tracking with Semantic-Aware Association
abstract
Locating and continuously tracking individuals in videos using natural-language descriptions is essential for human-AI collaboration, surveillance analytics, and video-based question answering. However, there are still three gaps: (i) although existing methods can reliably associate trajectories in most scenarios, they still fail to capture semantic understanding; (ii) large vision–language models (VLMs) grasp semantics but lack temporal identity stability; and (iii) person re-identification (ReID) excels at identity discrimination but ignores linguistic intent and often discards contextual cues. We present InstructTrack, an instruction-driven tracking agent that bridges these gaps. Using VLM backbone as a semantic hub, the video frame is parsed to localize the referred target, decide whether contextual cues are required, and extract initial semantic embeddings. The system then aligns VLM proposals with a lightweight detector via Hungarian matching to initialize or update track IDs. Subsequently, a context-gated ReID head learns identity and instruction relevant context embeddings and fuses them under language control; a tailored triplet objective jointly optimizes identity and context consistency. Integrated into an online MOT loop, InstructTrack delivers instruction-controllable, long-term person tracking, and single-video ReID. On MOT17 and MOT20, our method outperforms strong online baselines, achieving HOTA 68.4/68.4 and IDF1 86.1/81.6 while halving identity switches.
Zishun Zhou, Shuai Wang 0027, Hao Sheng 0001, Dazhi Yang 0003, Sentan Li, Da Yang 0001, Zhenglong Cui
MMAsia6
2018 W-Shaped Selection for Light Field Super-Resolution
Bing Su 0004, Hao Sheng 0001, Shuo Zhang 0003, Da Yang 0001, Nengcheng Chen, Wei Ke 0001
KSEM (1)4