Portrait of Ruilong Ren

Ruilong Ren 任瑞龙

Large Model Algorithm Engineer @ Autel Robotics · Foundation Model Team

M.Eng., Peking University B.Eng., Shandong University

Biography

I am a Large Model Algorithm Engineer in the Foundation Model team at Autel Robotics, working on UAV-oriented Vision-Language-Action (VLA) models, agentic systems, and world models. Previously, I was an AI Engineer at Huawei (Computing Product Line), focusing on spatial intelligence and autonomous-driving VLA on Ascend.

I received my M.Eng. in Electronic Information (AI) from Peking University (2022-2025; recommended admission, top of the department) and my B.Eng. in Electronic Information Science and Technology from Shandong University (2018-2022; Provincial Outstanding Graduate). During my studies I interned at TeleAI, Baidu Apollo, and DiDi.

My research interests include:

  • 2D / 3D visual understanding
  • Multimodal large language models
  • Embodied AI & Vision-Language-Action models
  • World models for robotics / UAVs

News

  • 2026.07: CosFly-VLA preprint released; submitting to AAAI 2027.
  • 2026.03: Joined Autel Robotics as a Large Model Algorithm Engineer.
  • 2026.01: One paper accepted to AAAI 2026.
  • 2025.11: Released TeleEgo - egocentric AI assistant benchmark.
  • 2025.07: Joined Huawei Computing Product Line as an AI Engineer.
  • 2025.06: Graduated from Peking University (M.Eng.). Survey paper accepted to KBS 2025.
  • 2024.10: One paper accepted to IROS 2024.
  • 2024.11: Started research internship at TeleAI (China Telecom).

Experience

Autel Robotics · Foundation Model Team
Large Model Algorithm Engineer
2026.03 - Present
  • Proposed CosFly-VLA for UAV dynamic target tracking under occlusion (Spatial-grounded CPT + Curriculum SFT + Expert-guided GRPO on Qwen3.5), achieving 30%+ gains over prior methods; submitting to AAAI 2027.
  • Led architecture design of Autel-Agentic-Brain for security area-search: unified cloud Agent, onboard VLA, 2.5D spatial memory, and aircraft capabilities into a closed-loop workflow; delivered demo in simulation.
  • Pre-research on UAV world models; designed Autel-WM / Autel-WAM based on Cosmos-3 toward an infinite data engine and physics foundation model for UAVs.
Huawei · Computing Product Line
AI Engineer
2025.07 - 2026.03
  • Contributed to spatial intelligence / world model / VLA technical roadmap and autonomous-driving VLA team formation on Ascend.
  • Built Ascend-native driving VLA (Qwen3-VL + DiT / Flow Matching) on Bench2Drive; open-loop trajectory prediction Top-3 and closed-loop Top-5 among public baselines.
TeleAI · Multimodal LLM
Algorithm Intern
2024.11 - 2025.06
  • Led construction of TeleEgo: a real-world egocentric long-video / omni-modal / streaming benchmark for AI assistants; evaluated SOTA MLLMs and highlighted gaps in long-term memory, omni-modality, and real-time reasoning.
Baidu Apollo · Autonomous Driving Perception
Multimodal LLM Intern
2024.05 - 2024.11
  • Built end-to-end driving MLLM training with high-resolution multi-view inputs and visual token compression; intention prediction success rate reached 92% and shipped in internal pre-annotation tooling.
DiDi · Perception
Algorithm Intern
2023.03 - 2023.08
  • Lane localization via image classification; active learning over tens of millions of road images with large-scale A100 pretraining, yielding 50%+ gains reused by peer teams.

Selected Publications

* equal contribution | Full list on Google Scholar.

CosFly-VLA task overview figure

CosFly-VLA: A Spatially-Aware Vision-Language-Action Model for UAV Tracking

Ruilong Ren, Songsheng Cheng, Yunpeng Zhou, et al.

AAAI 2027 · Submitting

Reframes UAV tracking as visibility prediction + closed-loop re-acquisition; Spatial-grounded CPT, curriculum SFT, and expert-guided GRPO improve occlusion recovery by 30%+.

3D-OCR framework overview

Boosting 3D Visual Grounding by Object-Centric Referring Network

Ruilong Ren, Jian Cao, Weichen Xu, et al.

IROS 2024

Fine-grained 2D semantics (SAM+CLIP) projected to 3D with explicit/implicit relation modeling; SOTA on ScanRefer / Nr3D.

Typical examples of language-grounded multimodal 3D scene understanding

A Survey of Language-grounded Multimodal 3D Scene Understanding

Ruilong Ren, et al.

KBS 2025

First comprehensive survey of language-grounded multimodal 3D scene understanding since 2020 (100+ papers), with taxonomy, benchmarks, and future outlook.

TeleEgo benchmark overview

TeleEgo: Benchmarking Egocentric AI Assistants in the Wild

Jiaqi Yan*, Ruilong Ren*, Jingren Liu*, et al.

arXiv 2025

Long-duration streaming omni-modal benchmark for egocentric AI assistants with memory, understanding, and cross-memory reasoning diagnostics.

CosFly-Track data generation pipeline

CosFly-Track: A Large-Scale Multi-Modal Dataset for UAV Visual Tracking

Xiangyue Wang, Hanxuan Chen, Songsheng Cheng, Ruilong Ren, et al.

NeurIPS 2026 · Under Review
CosFly teaser with UAV tracking and multi-modal outputs

CosFly: Plan in the Matrix, Fly in the World

Hanxuan Chen, Xiangyue Wang, Songsheng Cheng, Ruilong Ren, et al.

arXiv 2026
UAV-VLN survey contrasting paradigms figure

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap

Hanxuan Chen, Jie Zheng, Siqi Yang, et al., Ruilong Ren, et al.

arXiv 2026
Two Heads distillation framework overview

Two Heads Are Better than One: Distilling LLM Features into Small Models with Feature Decomposition and Mixture

Tianhao Fu, Xinxin Xu, Weichen Xu, Jue Chen, Ruilong Ren, et al.

AAAI 2026
3DMIT architecture and tasks

3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding

Zeju Li, Chao Zhang, Xiaoyan Wang, Ruilong Ren, et al.

ICMEW 2024

Infinite Video Understanding

Dell Zhang, Xiangyu Chen, Jixiang Luo, et al., Ruilong Ren, et al.

arXiv 2025

Education

Peking University - M.Eng. in Electronic Information (AI), 2022.09 - 2025.06
GPA 3.72 / 4.0 (Top 10%); recommended admission (1st in department)
Shandong University - B.Eng. in Electronic Information Science and Technology, 2018.09 - 2022.06
GPA 93.94 / 100 (1 / 76); Provincial Outstanding Graduate

Honors & Awards

National Scholarship (2020, 2021) - Top 1%, 1st in department
Shandong Provincial Outstanding Student - Top 0.1%
Shandong University Presidential Award - Top 0.1% undergraduate honor
CUMCM Shandong First Prize · MCM/ICM Meritorious Winner

Skills

Stack
Python, PyTorch, Linux; multi-hundred-GPU / NPU cluster training
Models
Transformer, ViT, CLIP; LLaVA, Qwen-VL, InternVL and related VLMs / VLAs
Data
Text, image, video, 3D point cloud; preprocessing & cleaning pipelines
Tools
Manus, Cursor, Claude Code for engineering / docs / presentation acceleration