DevOps-Eval

The DevOps-Eval is an industrial-first evaluation benchmark specifically designed for Large Language Models (LLMs) in the DevOps/AIOps domain¹. It was released by Ant Group in collaboration with Peking University³.

The goal of DevOps-Eval is to help developers, especially those in the DevOps field, track the progress and analyze the strengths and shortcomings of their models¹. The repository contains questions and exercises related to DevOps, including the AIOps and ToolLearning¹.

Here are some key features of DevOps-Eval¹:

  • It currently contains 7486 multiple-choice questions spanning 8 diverse general categories.
  • There are a total of 2840 samples in the AIOps subcategory, covering scenarios such as log parsing, time series anomaly detection, time series classification, time series forecasting, and root cause analysis.
  • There are a total of 1509 samples in the ToolLearning subcategory, covering 239 tool scenes across 59 fields.

The benchmark also includes a leaderboard that presents the zero-shot and five-shot accuracies from the models evaluated in the initial release¹. This allows for a comparison of different models' performance in the DevOps domain.

(1) GitHub - codefuse-ai/codefuse-devops-eval: Industrial-first evaluation .... https://github.com/codefuse-ai/codefuse-devops-eval. (2) DevOps-Eval:蚂蚁集团联合北京大学发布首个面向DevOps .... https://developer.aliyun.com/article/1365893. (3) codefuse-devops-eval: A DevOps Domain Knowledge .... https://gitee.com/codefuse-ai/codefuse-devops-eval. (4) codefuse-devops-eval/resources/tutorial_zh.md at main .... https://github.com/codefuse-ai/codefuse-devops-eval/blob/main/resources/tutorial_zh.md. (5) DevOps-Eval:蚂蚁集团联合北京大学发布首个面向DevOps .... https://juejin.cn/post/7296513628332359692.