SuperBench
SuperBench is a comprehensive evaluation system for large language models that includes five benchmark datasets: ExtremeGLUE for semantics, CodeBench for code, AlignBench for alignment, AgentBench for intelligent agents, and SafetyBench for safety. These benchmarks cover a wide range of tasks and dimensions to assess the overall capabilities of large language models. The SuperBench team aims to provide objective and scientific evaluation standards for large models to promote their healthy development in terms of technology, applications, and ecosystem.