| name | d3vl-understanding-driving-scenes-from-3d-time-ser |
| description | Skill generated from arXiv paper 2607.19528: D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models |
| metadata | {"arxiv":{"id":"2607.19528","title":"D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models","authors":["Heesang Han","A. Lynn Abbott","Abhijit Sarkar"],"published":"2026-07-21","categories":["cs.CV","cs.AI"],"url":"https://arxiv.org/abs/2607.19528","utility":0.93}} |
D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models
arXiv: 2607.19528
Published: 2026-07-21
Authors: Heesang Han, A. Lynn Abbott, Abhijit Sarkar
Categories: cs.CV, cs.AI
Utility: 0.93
Key Innovation
Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs for autonomous driving. However, the main emphasis to date has been for MLLMs using 2D images and videos. In contrast, this paper considers MLLM effectiveness using 3D sensors, particularly LiDAR and stereo cameras. LiDAR presents unique challenges to integration within an MLLM, largely because of data sparsity and lack of a grid structure for the data. For similar reasons, fusion of ca...
Potential Application
This paper presents advancements that could be applied to enhance agent capabilities in the areas of cs.CV, cs.AI.
References