#
ENJA

calvin

mees/calvin

CALVIN is an open-source simulated benchmark designed for training and evaluating language-conditioned policies in long-horizon robotic manipulation tasks. It enables researchers to develop agents capable of executing continuous control based on unconstrained natural language instructions and visual sensor inputs.

967 123MITUpdated 2025-09-08

Overview

CALVIN (Composing Actions from Language and Vision) addresses the challenge of long-horizon robotics by offering a comprehensive simulation platform complete with varied sensor observation spaces (static and gripper cameras, depth, tactile, and proprioception) and flexible action spaces. The benchmark evaluates agents on executing a sequence of diverse robotic manipulation tasks specified through human language, promoting research in vision-language-action learning and generalist robot policies.

Capabilities

  • Simulated robotic environment with multi-step instruction sequences
  • Support for absolute cartesian pose, relative cartesian displacement, and joint action spaces
  • Multimodal sensory support (static RGB-D, gripper RGB-D, tactile, proprioception)
  • Built-in baseline training scripts using PyTorch Lightning and Hydra
  • Evaluation suites for long-horizon multi-task language control (LH-MTLC)

Best for

Training and benchmarking long-horizon language-conditioned robotic manipulation models, Experimenting with different observation spaces including RGB, depth, tactile, and proprioceptive data, Evaluating continuous control policies under multi-task language control challenges, Relabeling raw language annotations with custom sentence-transformer language models

Evoa Score breakdown= Σ (score × weight)
71
Task usefulness20%88 → +17.6

実タスクにどれだけ役立つか(機能の豊富さ・用途の明確さ)。 AIによるcapabilities/use-cases解析

Code quality15%85 → +12.8

実装・指示の品質。 AIによるSKILL.md/README解析

Maintenance15%50 → +7.5

リポジトリがどれだけ活発に保守されているか。 GitHub 最終push日時の新しさ

Documentation12%52 → +6.2

ドキュメントの充実度・分かりやすさ。 README/独自要約の情報量

Security15%80 → +12.0

危険・不審な挙動が無いか。 AIによるセキュリティレビュー

Originality10%90 → +9.0

ありふれたラッパーではない独自性。 AIによる独自性判定

Popularity8%77 → +6.2

コミュニティの採用度。 GitHub Stars/Forks(対数スケール)

Compatibility5%0 → +0.0

対応AIエージェントの広さ。 AIによる対応エージェント判定

Weighted total71.3 / 100

ライセンス不明/制限あり(Red)のSkillは総合スコアに0.85倍の補正を適用します。 ランキングはこのScoreのみで決まり、広告で変わりません。 算出方法の詳細 →

Security considerations

The skill executes standard training and simulation loops in Python. Ensure all downloaded datasets and model checkpoints are sourced from trusted repositories.

Categories

Summary and analysis are original content generated by AI Skills Rank. The skill's source text is not reproduced here — view it on the linked repository.