FlagScale-Agent
flagos-ai/FlagScale-Agent
FlagScale-Agent is an autonomous AI agent designed for large-scale distributed training, inference, and serving infrastructure management.
Overview
Built on the ReAct framework, FlagScale-Agent automates complex infrastructure workflows ranging from environment setup and data preparation to training execution, performance monitoring, and systematic debugging. It incorporates domain-specific knowledge domains, safety guardrails, persistent memory, and structured execution plans tailored for deep learning operations like Megatron-LM, TransformerEngine, and NCCL.
Capabilities
- ▸Distributed training management
- ▸Automated error classification and debugging
- ▸Hardware topology detection
- ▸Model porting and adaptation
- ▸Safety guardrail enforcement
Best for
Setting up and configuring distributed training environments, Launching, monitoring, and debugging large-scale training jobs, Porting models between HuggingFace and Megatron-LM, Diagnosing infrastructure failures like NCCL timeouts, CUDA OOMs, and shape mismatches
Works with
実タスクにどれだけ役立つか(機能の豊富さ・用途の明確さ)。 — AIによるcapabilities/use-cases解析
実装・指示の品質。 — AIによるSKILL.md/README解析
リポジトリがどれだけ活発に保守されているか。 — GitHub 最終push日時の新しさ
ドキュメントの充実度・分かりやすさ。 — README/独自要約の情報量
危険・不審な挙動が無いか。 — AIによるセキュリティレビュー
ありふれたラッパーではない独自性。 — AIによる独自性判定
コミュニティの採用度。 — GitHub Stars/Forks(対数スケール)
対応AIエージェントの広さ。 — AIによる対応エージェント判定
ライセンス不明/制限あり(Red)のSkillは総合スコアに0.85倍の補正を適用します。 ランキングはこのScoreのみで決まり、広告で変わりません。 算出方法の詳細 →
Security considerations
Executes shell commands and manages infrastructure nodes, requiring strict access controls and credential isolation to prevent unintended system modifications.
Categories
Summary and analysis are original content generated by AI Skills Rank. The skill's source text is not reproduced here — view it on the linked repository.