ArchitectureAWS Bedrock
Designing retrieval-augmented generation on Amazon Bedrock Knowledge Bases
Bedrock Knowledge Bases manages the ingest, retrieve and generate steps of retrieval-augmented generation, but the choices made at each layer still decide answer quality. This guide covers managed and customer-managed knowledge bases, source preparation, chunking, embeddings and vector stores, retrieval settings, citations, grounding checks and evaluation, and ends with the signs that a custom pipeline would serve you better.
On this page
- The layers of a Knowledge Bases application
- Managed or customer-managed knowledge base
- Preparing sources so retrieval has something good to find
- Picking a chunking strategy for your documents
- Choosing an embedding model and vector store together
- Retrieval settings to tune before changing the model
- Citations and the no-answer path
- What contextual grounding checks can and cannot catch
- Evaluating retrieval and answers before release
- When a custom RAG pipeline is the better choice
- Questions and answers
- Sources
The layers of a Knowledge Bases application
- Application and citations
Presents answers with the sources they came from and handles the no-answer case.
- Guardrails
Content policies plus contextual grounding and relevance checks on generated answers.
- Generation
RetrieveAndGenerate with a prompt template, or your own prompt over Retrieve results.
- Retrieval configuration
Number of results, search type, metadata filters, reranking and query decomposition.
- Vector store
Holds embeddings, chunk text and metadata, run by Bedrock or provisioned by you.
- Chunking and embeddings
Splits parsed documents and converts each chunk into a vector.
- Data sources and sync
Connectors, document formats, metadata files and ingestion jobs.
Managed or customer-managed knowledge base
Bedrock now offers two kinds of knowledge base, and the choice constrains most later decisions1.
| Aspect | Bedrock Managed Knowledge Base | Customer-managed knowledge base |
|---|---|---|
| Data store | An auto-scaling store for embeddings, text, metadata and files, run by Bedrock | A vector store you choose, provision and scale |
| Connectors | Native connectors including S3, SharePoint, Confluence, web crawler, Google Drive and OneDrive, plus custom | S3 and custom sources for new connections |
| Embedding and reranking | Service-managed models by default, or your own Bedrock models within stated limits | Any Bedrock embedding model, and reranking models you select |
| Search | Managed hybrid and agentic retrieval | Your own search strategy, within what the store supports |
| Best suited to | End-to-end managed RAG with native connectors | Teams that need a particular vector database or direct access to it |
Since September 30, 2026, new Confluence, SharePoint, Salesforce and web crawler connectors cannot be created on customer-managed knowledge bases, although existing ones keep working2.
Preparing sources so retrieval has something good to find
Retrieval quality is capped by the documents. Before connecting anything, remove superseded versions, duplicates and drafts, because the retriever cannot tell an obsolete policy from a current one unless metadata says so. Files read from S3 must use supported formats such as plain text, Markdown, HTML, Word, CSV, Excel or PDF, within a per-file size quota3.
Attach metadata early. For S3 sources, a companion metadata file per document can carry fields such as department, document type, effective date or audience, which later drive filters and access rules4. Then decide how sync runs, whether on a schedule, on upload events or after an approval step, and what should happen to vectors when a source document is deleted.
Picking a chunking strategy for your documents
The right strategy depends on how your documents are structured and how questions are asked5.
| Strategy | How it splits | Suits | Watch for |
|---|---|---|---|
| Default | Chunks of a set approximate token length that keep sentences whole | Mixed, unstructured text with no better signal | Answers that span chunk boundaries |
| Fixed-size | Your chosen maximum tokens per chunk, with a percentage overlap | Consistent prose where you want control over size | Too small loses context; too large dilutes relevance |
| Hierarchical | Small child chunks for matching, replaced by larger parent chunks at retrieval | Long structured documents such as manuals and policies | Fewer results than requested; not recommended with S3 Vectors |
| Semantic | Breaks where meaning shifts between sentences, using a model | Documents whose topics change without clear headings | Extra model cost during ingestion |
| No chunking | Each file becomes a single chunk | Content already split into small, self-contained files | No page numbers in citations and no page filters |
Choosing an embedding model and vector store together
The embedding model's vector dimensions must match the index, so these choices are made as a pair6.
- If
Users search for part numbers, codes or names as well as concepts.
ThenUse a store that supports hybrid search with a filterable text field: Aurora PostgreSQL, OpenSearch Serverless or MongoDB Atlas4.
Other stores fall back to semantic-only search.
- If
Query volume is low and storage cost matters more than latency.
ThenConsider Amazon S3 Vectors, avoiding hierarchical chunking with large token counts6.
AWS positions S3 Vectors for infrequent queries, and hierarchical metadata can exceed its per-vector limits.
- If
You want binary embeddings to reduce storage.
ThenUse OpenSearch Serverless or an OpenSearch managed cluster6.
They are the only supported stores for binary vectors.
- If
Relationships between entities matter as much as passage similarity.
ThenEvaluate Neptune Analytics with GraphRAG6.
A graph can connect facts spread across documents that plain vector search returns separately.
- If
Documents or questions span several languages.
ThenChoose a multilingual embedding model and test retrieval with queries written by native speakers.
An English-oriented embedding can rank other-language queries poorly even when the answer is present.
Retrieval settings to tune before changing the model
Citations and the no-answer path
Citations make a RAG answer checkable. With RetrieveAndGenerate, a custom generation prompt template must keep the $output_format_instructions$ placeholder, or responses come back without citations4. If you call Retrieve and write your own prompt, carry chunk identifiers and source locations through to the response yourself.
Design the no-answer path deliberately. Instruct the model to say when the retrieved passages do not answer the question, show what was searched and offer a route to a person. An application that always answers will eventually answer from the model's general knowledge while appearing to cite your documents.
What contextual grounding checks can and cannot catch
Evaluating retrieval and answers before release
Build a question set with ground truth
Collect real questions with the passages that should be retrieved and a reference answer, including questions the documents cannot answer.
Test retrieval on its own
Check whether the right passages appear for each question before judging generated answers. Bedrock RAG evaluations can score retrieval alone, or retrieval with generation, against your ground truth8.
Test answers with retrieval held fixed
Score answers for correctness, completeness and faithfulness to the passages, and confirm that unanswerable questions are declined.
Test permissions
Query as users with different access rights and confirm that restricted documents never appear in results or answers.
Re-run after every change
Repeat the set after re-chunking, a new embedding model, a prompt change or a large sync, and compare with the previous run.
When a custom RAG pipeline is the better choice
Some requirements point to building the pipeline yourself, often still on Bedrock models and AWS storage: document parsing that needs domain-specific logic, retrieval that must blend vector search with structured queries in one ranking, access rules too complex for metadata filters, or full control of every intermediate step for audit. The cost is ownership of ingestion, indexing, retrieval and evaluation code. Prototype on Knowledge Bases first and measure where it falls short on your question set; those gaps become the specification for the custom parts.
Questions and answers
How do Bedrock Knowledge Bases handle document permissions?
Retrieval does not know who is asking unless you tell it. In a customer-managed knowledge base, the usual approach is to store access attributes as metadata and apply filters derived from the authenticated user on the server, never from the prompt. Managed knowledge bases also document an access-control-list awareness option1. Either way, test with real user roles that restricted documents never surface.
How do we keep a Bedrock knowledge base up to date?
Run ingestion jobs whenever sources change, on a schedule or triggered by upload events through the StartIngestionJob API. Decide how deletions are handled, because the data source's deletion policy controls whether vectors for removed documents are deleted or retained. After large syncs, rerun your evaluation set to confirm answer quality has not shifted.
Can a knowledge base serve content in several languages?
Yes, but language support comes from the embedding and generation models you choose rather than from the knowledge base itself. Pick a multilingual embedding model when documents or questions span languages, then test retrieval with queries written by native speakers in each one. Store language as a metadata field so you can filter, and report quality, per language.
Do we need Bedrock Agents to use a knowledge base?
No. An application can call the Retrieve or RetrieveAndGenerate APIs directly, which is the simplest pattern for question answering over documents. Agents become relevant when the application must choose between several tools or take actions, which is the territory of our Bedrock and AgentCore practice.
Sources
- Build a managed knowledge base (Amazon Bedrock User Guide) — Amazon Web Services · checked 10 October 2026
- Connect a data source to your knowledge base — Amazon Web Services · checked 10 October 2026
- Prerequisites for your Amazon Bedrock knowledge base data — Amazon Web Services · checked 10 October 2026
- Configure and customize queries and response generation — Amazon Web Services · checked 10 October 2026
- How content chunking works for knowledge bases — Amazon Web Services · checked 10 October 2026
- Prerequisites for using a vector store you created for a knowledge base — Amazon Web Services · checked 10 October 2026
- Use contextual grounding check to filter hallucinations in responses — Amazon Web Services · checked 10 October 2026
- Evaluate the performance of Amazon Bedrock resources — Amazon Web Services · checked 10 October 2026