- Proposed CosFly-VLA for UAV dynamic target tracking under occlusion (Spatial-grounded CPT + Curriculum SFT + Expert-guided GRPO on Qwen3.5), achieving 30%+ gains over prior methods; submitting to AAAI 2027.
- Led architecture design of Autel-Agentic-Brain for security area-search: unified cloud Agent, onboard VLA, 2.5D spatial memory, and aircraft capabilities into a closed-loop workflow; delivered demo in simulation.
- Pre-research on UAV world models; designed Autel-WM / Autel-WAM based on Cosmos-3 toward an infinite data engine and physics foundation model for UAVs.
Ruilong Ren 任瑞龙
Large Model Algorithm Engineer @ Autel Robotics · Foundation Model Team
Biography
I am a Large Model Algorithm Engineer in the Foundation Model team at Autel Robotics, working on UAV-oriented Vision-Language-Action (VLA) models, agentic systems, and world models. Previously, I was an AI Engineer at Huawei (Computing Product Line), focusing on spatial intelligence and autonomous-driving VLA on Ascend.
I received my M.Eng. in Electronic Information (AI) from Peking University (2022-2025; recommended admission, top of the department) and my B.Eng. in Electronic Information Science and Technology from Shandong University (2018-2022; Provincial Outstanding Graduate). During my studies I interned at TeleAI, Baidu Apollo, and DiDi.
My research interests include:
- 2D / 3D visual understanding
- Multimodal large language models
- Embodied AI & Vision-Language-Action models
- World models for robotics / UAVs
News
- 2026.07: CosFly-VLA preprint released; submitting to AAAI 2027.
- 2026.03: Joined Autel Robotics as a Large Model Algorithm Engineer.
- 2026.01: One paper accepted to AAAI 2026.
- 2025.11: Released TeleEgo - egocentric AI assistant benchmark.
- 2025.07: Joined Huawei Computing Product Line as an AI Engineer.
- 2025.06: Graduated from Peking University (M.Eng.). Survey paper accepted to KBS 2025.
- 2024.10: One paper accepted to IROS 2024.
- 2024.11: Started research internship at TeleAI (China Telecom).
Experience
- Contributed to spatial intelligence / world model / VLA technical roadmap and autonomous-driving VLA team formation on Ascend.
- Built Ascend-native driving VLA (Qwen3-VL + DiT / Flow Matching) on Bench2Drive; open-loop trajectory prediction Top-3 and closed-loop Top-5 among public baselines.
- Led construction of TeleEgo: a real-world egocentric long-video / omni-modal / streaming benchmark for AI assistants; evaluated SOTA MLLMs and highlighted gaps in long-term memory, omni-modality, and real-time reasoning.
- Built end-to-end driving MLLM training with high-resolution multi-view inputs and visual token compression; intention prediction success rate reached 92% and shipped in internal pre-annotation tooling.
- Lane localization via image classification; active learning over tens of millions of road images with large-scale A100 pretraining, yielding 50%+ gains reused by peer teams.
Selected Publications
* equal contribution | Full list on Google Scholar.
CosFly-VLA: A Spatially-Aware Vision-Language-Action Model for UAV Tracking
Reframes UAV tracking as visibility prediction + closed-loop re-acquisition; Spatial-grounded CPT, curriculum SFT, and expert-guided GRPO improve occlusion recovery by 30%+.
Boosting 3D Visual Grounding by Object-Centric Referring Network
Fine-grained 2D semantics (SAM+CLIP) projected to 3D with explicit/implicit relation modeling; SOTA on ScanRefer / Nr3D.
A Survey of Language-grounded Multimodal 3D Scene Understanding
First comprehensive survey of language-grounded multimodal 3D scene understanding since 2020 (100+ papers), with taxonomy, benchmarks, and future outlook.
Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap
Two Heads Are Better than One: Distilling LLM Features into Small Models with Feature Decomposition and Mixture
Education
GPA 3.72 / 4.0 (Top 10%); recommended admission (1st in department)
GPA 93.94 / 100 (1 / 76); Provincial Outstanding Graduate