PlatinumBench
CC-BY-SA-4.0Introduced 2025-02-05
Platinum Benchmarks are benchmarks that are are carefully curated to minimize label errors and ambiguity, allowing us to measure reliability of models.
This dataset contains fifteen platinum benchmarks created by manually revising questions from existing datasets (see the github repo for details on accessing our revised subset of VQA). To revise each benchmark, we ran a variety of frontier models on individual examples and manually re-annotated any example for which at least one model made an error. See the paper for further details on the revision process.