No episode media available here
This is a representative still, not playable footage. Episode video for this dataset lives with its original source.
CALVIN
Long-horizon language instruction chains
Freiburg Robot Learning Lab
RobotData indexes this dataset's metadata. Attribution and licensing stay with the original publisher.
Composing Actions from Language and Vision: four simulated environments (A/B/C/D) with 24 hours of teleoperated play and 20,000 language annotations. The standard benchmark for chaining five instructions in a row with zero-shot environment transfer.
Reported structure
Values below come from the dataset's own documentation. Missing fields are shown as not provided rather than estimated.
2 cameras
7 DoF
30 Hz
Not provided
Data modalities
Episode explorer
No episode media has been added for this dataset, so there is nothing to play. The explorer appears as soon as real clips exist — either uploaded here or extracted from the source.
Contribute episodesLinked models, robots & papers
Resolving related models, datasets and papers…
Related datasets
BridgeData V2
Source: UC Berkeley RAIL
- 60K
- 333.9
- 4
LeRobot Community SO-100 Datasets
Source: Hugging Face LeRobot
- 23K
- 481
- 2
DROID
Source: Stanford IRIS & REAL Labs
- 76K
- 350
- 3
RT-1 Robot Action Dataset
Source: Google DeepMind Robotics
- 130K
- 17.4K
- 1