| Mode | Score | Pass rate | Cases | Latency | Date |
|---|---|---|---|---|---|
| No runs yet. Trigger an evaluation above. | |||||
RAG mode
Evaluates the hybrid retrieval pipeline — dense embeddings, BM25, and Cohere reranking — against 10 hand-built cases.
Single agent mode
Evaluates the LangGraph 9-tool agent's ability to select the right tool and complete the task correctly.
Multi-agent mode
Evaluates the Phase 4 supervisor system — whether it delegates correctly and synthesises specialist findings coherently.