Transformers, a ground-breaking neural network architecture born from natural language processing (NLP), have since transformed various domains, including the automotive sector. Their evolution into vision transformers has revolutionized computer vision tasks – significantly boosting perception and scene understanding in Advanced Driver Assistance Systems (ADAS) and Autonomous Driving (AD).

Recent advancements in Large Language Models (LLMs) have pushed the boundaries of reasoning capabilities. For instance, OpenAI’s o3 recently surpassed the ARC’s AGI benchmark, demonstrating robust chain-of-thought reasoning and problem-solving. Innovations such as Large Concept Models and Titan models are further extending these capabilities, enhancing not only natural language understanding but also enabling complex decision-making and planning. This makes them highly relevant for next-gen ADAS and AD applications.

What Are Multimodal Large Language Models?

Traditional LLMs vs. Multimodal LLMs

Traditional LLMs are designed primarily for text-based tasks – understanding and generating human language. In contrast, Multimodal Large Language Models (MMLLMs) are capable of processing and integrating multiple sensor modalities such as cameras, microphones, LiDAR, and RADAR. By fusing this sensor data, MMLLMs can generate meaningful outputs like predictions, actions, or commands, resulting in a more holistic understanding of the environment.

Vision-Language Models (VLMs)

A prominent subset of MMLLMs, Vision-Language Models (VLMs) combine visual data with textual information. Originating from early image captioning systems around 2015, VLMs have matured alongside LLMs to tackle nuanced challenges in ADAS and AD – particularly long-tail scenarios, where rare but critical events may occur.

Key Applications of Multimodal LLMs in ADAS and Autonomous Driving

MMLLMs are currently being employed across a spectrum of tasks extending well beyond offline simulations to real-time applications.

center

arXiv:2310.14414v2

Industry Adoption

While many MMLLMs remain in the research or offline stages, Cyient uses MLLMs for auto annotation, dataset generation and test scenario generation for robust training and faster time to market. Several players are pushing toward real-time integration such as:

NVIDIA:Demonstrated a two-staged anomaly detection system using MMLLMs on its DRIVE platform. A faster module detects anomalies and queries a slower LLM for context-based control decisions —a blend of speed and depth in reasoning.

center

https://www.youtube.com/watch?v=TSC_mVH5abI

Nuro:Employs MMLLMs within its LAMBDA system for enhanced scene understanding, rider interaction and explainable AI.

center

https://medium.com/nuro/lambda-the-nuro-drivers-real-time-language-reasoning-model-7c3567b2d7b4

Waymo:Introduced EMMA, a multimodal, end-to-end Vision-Language-Action model that consolidates perception, localization, planning and control within a single neural network.

center

arXiv:2410.23262v1

Li Auto:In collaboration with Tsinghua University, Li Auto developedDriveVLM—a VLM, then integrated a VLM into their ADMax platform and deployed via OTA (Over the Air). This system supports real-time decision-making based on multimodal inputs.

center

https://kr-asia.com/transcript-why-li-auto-is-revving-up-its-smart-driving-efforts-chasing-tesla

Current Challenges in Real-Time Deployment

Despite their promise MMLLMs face hurdles in live automotive environments:

Conclusion

From their roots in NLP to advanced multimodal architectures, transformers have significantly enhanced the perception and scene understanding capabilities in ADAS and autonomous driving. With cutting-edge reasoning capabilities and sensor fusion, multimodal LLMs are poised to play a critical role in the next wave of ADAS and autonomous driving systems. While challenges remain in scaling these models for real-time use, ongoing advancements signal a future where perception and action converge seamlessly through intelligent, unified models.

Table of Contents

You may also like

Explore All Insights