RAG Pipelines
Retrieval-Augmented Generation with 5 built-in retrieval strategies.
AgentOven includes a built-in Retrieval-Augmented Generation (RAG) pipeline with 5 retrieval strategies. Ingest documents, chunk them, generate embeddings, store vectors, and retrieve context for LLM generation — all through a unified API.
Retrieval Strategies
Choose the right strategy based on your use case. Each strategy optimizes for different trade-offs between speed, accuracy, and cost.
Naive
Simple embed → search → generate pipeline. Best for straightforward Q&A where the query closely matches document content.
Query → Embed → Vector Search → Top-K → LLM Generate
Sentence Window
Retrieves matching sentences then expands context by including surrounding sentences. Provides focused yet contextualized results.
Query → Embed → Search Sentences → Expand Window → LLM Generate
Parent Document
Searches at the chunk level but returns full parent documents. Great when small chunks match but the full document context is needed.
Query → Embed → Search Chunks → Return Parents → LLM Generate
HyDE
Hypothetical Document Embeddings — first generates a hypothetical answer using the LLM, then embeds and searches for real matches. Excels when queries are abstract.
Query → LLM Hypothesize → Embed Hypothesis → Search → LLM Generate
Agentic
LLM decides whether to search, refine the query, or synthesize from existing context. Most intelligent but highest latency and cost.
Query → LLM Decide → [Search | Refine | Synthesize] → Iterate → LLM Generate
Document Ingestion
The ingestion engine handles the full pipeline from raw documents to searchable vectors:
- Upload documents via API
- Text is split into chunks (configurable size, overlap, separator)
- Chunks are embedded in batches using the configured embedding driver
- Vectors are upserted into the configured vector store
- Original text and metadata preserved for retrieval
Usage Example
Ingest documents
terminal
$ curl -X POST http://localhost:8080/api/v1/rag/ingest \ -H "Content-Type: application/json" \ -H "X-Kitchen: default" \ -d '{ "documents": [ { "id": "doc-1", "content": "AgentOven is an open-source agent control plane...", "metadata": { "source": "readme", "version": "0.2" } } ], "chunk_size": 512, "chunk_overlap": 50 }'
Query with RAG
terminal
$ curl -X POST http://localhost:8080/api/v1/rag/query \ -H "Content-Type: application/json" \ -H "X-Kitchen: default" \ -d '{ "question": "What is AgentOven?", "strategy": "naive", "top_k": 5, "model": "gpt-4o-mini" }'
Chunker Configuration
The text chunker is configurable to optimize for your document structure:
json
{
"chunk_size": 512,
"chunk_overlap": 50,
"separator": "\n\n"
}- chunk_size — Maximum characters per chunk
- chunk_overlap — Overlap between adjacent chunks to preserve context
- separator — Primary split boundary (falls back to sentence/word splitting)
API Endpoints
| Method | Path | Description |
|---|---|---|
POST | /api/v1/rag/ingest | Ingest and chunk documents |
POST | /api/v1/rag/query | RAG query with selected strategy |
GET | /api/v1/rag/status | Pipeline status and configuration |