AV Causal Scenario Retrieval Challenge banner

AV Causal Scenario Retrieval Challenge

Retrieve complex driving scenarios from video sequences with focus on spatio-temporal and causal reasoning

01 Overview

Understanding complex driving scenarios requires more than object detection. It demands reasoning about the causal chains that connect agents, actions, and environmental context over time. The AV Causal Scenario Retrieval Challenge invites the research community to develop models that can retrieve driving scenarios from dense, structured annotations capturing spatial, temporal, and causal relationships.

Dataset identity, download links, task splits, and development instructions are provided only after a participant applies and the organizers accept the registration. Sign in on the participation page to apply or check your status.

02 Challenge Description

This is a query-based video retrieval challenge focused on autonomous-driving scenarios. Given a natural-language query describing events, situations, interactions, or spatial, temporal, and causal relationships in a driving scene, a system must search a corpus of videos and return a ranked list of the clips that best match the query.

The held-out test set spans approximately 2,000 queries evaluated against a corpus of 600+ videos. A query can match more than one clip, so systems must retrieve all relevant scenarios and rank them highly.

Evaluation metrics

For every test query, a submission produces an ordered list of video IDs. Results are measured with the following retrieval metrics, macro-averaged so that every query contributes equally:

Average R-Precision Primary metric
If a query has R relevant videos, R-Precision is the fraction of those videos found in the first R retrieved results. The official leaderboard score is the mean R-Precision across test queries.
Mean Average Precision (MAP)
Rewards relevant videos for appearing early throughout the ranked results.
Precision@k
The fraction of the first k retrieved videos that are relevant.
Recall@k
The fraction of all relevant videos retrieved within the first k results.
Hit@k
Whether at least one relevant video appears within the first k results.

Precision, Recall, and Hit are reported at k = 1, 3, 5, and 10. Submissions are ranked by Average R-Precision. Ties are broken by MAP, Recall@10, Precision@10, and then earliest completion time.

Participating in the Challenge

Who can participate?

  • You may participate individually or as part of a team with multiple members.
  • Every participant must be at least the age of majority in their jurisdiction of residence when they register. By registering, participants confirm that they meet this requirement.
  • Participants must be affiliated with a recognized university, research institute, or commercial organization, declare that affiliation, and register with an official institutional email address. Personal email addresses require explicit organizer approval. Organizers may verify these details and reject or remove participants whose affiliation or email cannot be verified.
  • Each team must designate one submitter during registration. Only that person may submit for the team, and the designation cannot change after the challenge begins unless the organizers approve an exception.
  • Each person may belong to only one team. Multiple accounts may not be created or used to bypass submission limits or other challenge rules.
  • Each individual entrant or team may make up to five submissions. Organizers may adjust this limit when necessary for the fair and efficient operation of the challenge.

How do I participate?

  1. Sign in and apply.

    Open the participation page, sign in with Hugging Face, and submit the team name, contact email, and institution that the organizers should review.

  2. Wait for organizer review.

    The admin dashboard records an ACCEPTED, PENDING, or REJECTED decision. Challenge data and Official Terms and Conditions remain unavailable until the registration is accepted.

  3. Review the protected participant materials.

    Once accepted, return to the participation page to open the dataset and development guide and read the current Official Terms and Conditions.

  4. Develop and validate your system.

    Use the approved-participant guide to obtain the data, task-owned train and validation split, devkit, and local evaluation workflow.

  5. Submit for evaluation.

    Create or reconnect the private artifact repository from the participation page, publish your self-contained container artifact with the devkit, review and explicitly accept the current Agreement, and queue the evaluation. Acceptance is required at submission time.

More about submissions

  • Each individual entrant or team may submit up to five entries during the challenge.
  • Approved participants must read and accept the current Official Terms and Conditions when submitting.
  • Submissions run without network access. All code, dependencies, model weights, and other components needed for inference must be packaged in the submission artifact.
  • Each submission must complete within six hours on one NVIDIA A100 80 GB GPU (compute capability 8.0, sm_80). The devkit does not require a specific CUDA version; your packaged framework runtime and compiled extensions must be compatible with the evaluation host and include sm_80 support.

03 Timeline

  1. 2026-06-03 Challenge announcement
  2. Competition goes live Registration opens; accepted participants receive data, terms, tools, and submission instructions.
  3. 2026-11-20 Public leaderboard closes Final submissions accepted for the leaderboard.
  4. 2026-11-29 Final results released Final results and winners are released to participants.
  5. NeurIPS 2026 Results announced at the World Models in Physical AI Workshop Challenge results and winners will be announced at the World Models in Physical AI Workshop at NeurIPS 2026 in Sydney, Australia.

Dates are tentative and subject to update.

04 Communications

Official challenge announcements and updates will be published on the challenge Hugging Face site or sent to participants by email.

05 Organizers

This challenge is organized by NVIDIA's Spatial Intelligence Lab.

Host

NVIDIA Spatial Intelligence Lab

NVIDIA Research lab working on spatial intelligence for autonomous systems, 3D understanding, and scene reasoning.