AgentOven

RAGAS Evaluation

Measure RAG pipeline quality with faithfulness, relevancy, and precision metrics.


AgentOven includes a RAGAS evaluation sidecar — a Python FastAPI service that measures the quality of your RAG pipeline using three core metrics: faithfulness, answer relevancy, and context precision.


What is RAGAS?

RAGAS (Retrieval-Augmented Generation Assessment) is an evaluation framework for RAG pipelines. It provides metrics that measure how well your retrieval and generation components work together, without requiring human-labeled ground truth.

Evaluation Metrics
Faithfulness

Score: 0.0 – 1.0

Measures whether the generated answer is faithful to the retrieved context. A score of 1.0 means every claim in the answer is supported by the provided context. Detects hallucination.

Answer Relevancy

Score: 0.0 – 1.0

Measures how relevant the generated answer is to the original question. A score of 1.0 means the answer directly addresses the question without unnecessary information.

Context Precision

Score: 0.0 – 1.0

Measures whether the retrieved context is relevant and precise. High precision means the retriever is returning only relevant documents, not noise.


Architecture

The RAGAS evaluator runs as a separate Python sidecar service (port 8400). The Go control plane communicates with it via HTTP, and the evaluation is also exposed as an MCP tool for agent-driven quality monitoring.

text

┌─────────────────┐     HTTP      ┌──────────────────────┐
│  Go Control      │ ──────────→ │  RAGAS Sidecar        │
│  Plane (:8080)   │             │  Python FastAPI (:8400)│
│                  │ ←────────── │                        │
│  MCP Tool:       │   JSON      │  Metrics:              │
│  ragas_evaluate  │   response  │  - faithfulness        │
│                  │             │  - answer_relevancy    │
└─────────────────┘             │  - context_precision   │
                                └──────────────────────┘
Running the Sidecar

terminal

# Using Docker (recommended)
$ docker build -t agentoven-ragas \
  -f control-plane/internal/integrations/ragas/Dockerfile .
$ docker run -p 8400:8400 agentoven-ragas

# Or run directly with Python
$ cd control-plane/internal/integrations/ragas
$ pip install fastapi uvicorn ragas langchain-openai
$ uvicorn server:app --host 0.0.0.0 --port 8400
Usage via API

terminal

$ curl -X POST http://localhost:8080/api/v1/rag/evaluate \
  -H "Content-Type: application/json" \
  -H "X-Kitchen: default" \
  -d '{
    "question": "What is AgentOven?",
    "answer": "AgentOven is an open-source agent control plane.",
    "contexts": [
      "AgentOven is an open-source, framework-agnostic enterprise agent control plane that standardizes how AI agents are built, deployed, observed, and orchestrated."
    ]
  }'
MCP Tool Integration

The RAGAS evaluator is also available as an MCP tool, allowing agents to self-evaluate their RAG responses:

json

{
  "method": "tools/call",
  "params": {
    "name": "ragas_evaluate",
    "arguments": {
      "question": "What protocols does AgentOven support?",
      "answer": "AgentOven supports A2A and MCP protocols.",
      "contexts": ["AgentOven is built on two open protocols: A2A and MCP."]
    }
  }
}
Pro: RAG Quality Monitor

AgentOven Pro includes a RAG Quality Monitor — a standalone A2A agent that continuously evaluates RAG pipeline quality and sends alerts when metrics drop below configured thresholds.