Introducing the MLPerf End-to-End RAG Inference Benchmark
MLCommons' MLPerf Inference working group has launched the first end-to-end retrieval-augmented generation (RAG) inference benchmark, measuring the full pipeline from document ingestion to multi-hop question answering across two workloads.