این اسکیل چه میکند؟
برای "run an evaluation"، "evaluate my ADK agent"، "write an eval dataset"، "analyze eval failures"، "compare eval results"، "optimize agent" یا راهنمای Agent Platform eval و Quality Flywheel؛ metric، schema دیتاست، LLM-as-judge و علت رایج شکست.
This skill should be used when the user wants to "run an evaluation", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures", "compare eval results", "optimize agent", or needs guidance on the Agent Platform eval methodology and the Quality Flywheel. Covers eval metrics, dataset schema, LLM-as-judge scoring, and common failure causes.
چه زمانی به کار میآید؟
در دادهٔ اصلی، کاربرد جداگانهای ثبت نشده است. توضیح بالا و صفحهٔ منبع را ببینید.