StonyBrookNLP/appworld
AppWorld is an interactive benchmark and simulation environment featuring 9 apps and 457 APIs for evaluating function calling and coding agents.
AppWorld provides a comprehensive, controllable simulated digital world populated by ~100 people and day-to-day applications. It serves as a benchmark and execution environment where autonomous agents can perform complex tasks via interactive Python coding, API calls, and MCP (Model Context Protocol) integration.
Benchmarking and evaluating interactive coding and function-calling AI agents, Running reinforcement learning rollouts and parallel world simulations, Testing agent performance across diverse, realistic digital workflows (e.g. Spotify, banking, utilities)
実タスクにどれだけ役立つか(機能の豊富さ・用途の明確さ)。 — AIによるcapabilities/use-cases解析
実装・指示の品質。 — AIによるSKILL.md/README解析
リポジトリがどれだけ活発に保守されているか。 — GitHub 最終push日時の新しさ
ドキュメントの充実度・分かりやすさ。 — README/独自要約の情報量
危険・不審な挙動が無いか。 — AIによるセキュリティレビュー
ありふれたラッパーではない独自性。 — AIによる独自性判定
コミュニティの採用度。 — GitHub Stars/Forks(対数スケール)
対応AIエージェントの広さ。 — AIによる対応エージェント判定
ライセンス不明/制限あり(Red)のSkillは総合スコアに0.85倍の補正を適用します。 ランキングはこのScoreのみで決まり、広告で変わりません。 算出方法の詳細 →
Executes dynamic agent-generated code inside the simulated world environment, requiring appropriate containment and safety precautions.
Summary and analysis are original content generated by AI Skills Rank. The skill's source text is not reproduced here — view it on the linked repository.