ArchitectureAWS Bedrock

Designing retrieval-augmented generation on Amazon Bedrock Knowledge Bases

Bedrock Knowledge Bases manages the ingest, retrieve and generate steps of retrieval-augmented generation, but the choices made at each layer still decide answer quality. This guide covers managed and customer-managed knowledge bases, source preparation, chunking, embeddings and vector stores, retrieval settings, citations, grounding checks and evaluation, and ends with the signs that a custom pipeline would serve you better.

Reviewed 7 min read

On this page
  1. The layers of a Knowledge Bases application
  2. Managed or customer-managed knowledge base
  3. Preparing sources so retrieval has something good to find
  4. Picking a chunking strategy for your documents
  5. Choosing an embedding model and vector store together
  6. Retrieval settings to tune before changing the model
  7. Citations and the no-answer path
  8. What contextual grounding checks can and cannot catch
  9. Evaluating retrieval and answers before release
  10. When a custom RAG pipeline is the better choice
  11. Questions and answers
  12. Sources

The layers of a Knowledge Bases application

Application and citations01Guardrails02Generation03Retrieval configuration04Vector store05Chunking and embeddings06Data sources and sync07
  1. Application and citations

    Presents answers with the sources they came from and handles the no-answer case.

  2. Guardrails

    Content policies plus contextual grounding and relevance checks on generated answers.

  3. Generation

    RetrieveAndGenerate with a prompt template, or your own prompt over Retrieve results.

  4. Retrieval configuration

    Number of results, search type, metadata filters, reranking and query decomposition.

  5. Vector store

    Holds embeddings, chunk text and metadata, run by Bedrock or provisioned by you.

  6. Chunking and embeddings

    Splits parsed documents and converts each chunk into a vector.

  7. Data sources and sync

    Connectors, document formats, metadata files and ingestion jobs.

Conceptual layer view from user (top) to data (bottom); a design aid, not an official AWS diagram.

Managed or customer-managed knowledge base

Bedrock now offers two kinds of knowledge base, and the choice constrains most later decisions1.

AspectBedrock Managed Knowledge BaseCustomer-managed knowledge base
Data storeAn auto-scaling store for embeddings, text, metadata and files, run by BedrockA vector store you choose, provision and scale
ConnectorsNative connectors including S3, SharePoint, Confluence, web crawler, Google Drive and OneDrive, plus customS3 and custom sources for new connections
Embedding and rerankingService-managed models by default, or your own Bedrock models within stated limitsAny Bedrock embedding model, and reranking models you select
SearchManaged hybrid and agentic retrievalYour own search strategy, within what the store supports
Best suited toEnd-to-end managed RAG with native connectorsTeams that need a particular vector database or direct access to it

Since September 30, 2026, new Confluence, SharePoint, Salesforce and web crawler connectors cannot be created on customer-managed knowledge bases, although existing ones keep working2.

Preparing sources so retrieval has something good to find

Retrieval quality is capped by the documents. Before connecting anything, remove superseded versions, duplicates and drafts, because the retriever cannot tell an obsolete policy from a current one unless metadata says so. Files read from S3 must use supported formats such as plain text, Markdown, HTML, Word, CSV, Excel or PDF, within a per-file size quota3.

Attach metadata early. For S3 sources, a companion metadata file per document can carry fields such as department, document type, effective date or audience, which later drive filters and access rules4. Then decide how sync runs, whether on a schedule, on upload events or after an approval step, and what should happen to vectors when a source document is deleted.

Picking a chunking strategy for your documents

The right strategy depends on how your documents are structured and how questions are asked5.

StrategyHow it splitsSuitsWatch for
DefaultChunks of a set approximate token length that keep sentences wholeMixed, unstructured text with no better signalAnswers that span chunk boundaries
Fixed-sizeYour chosen maximum tokens per chunk, with a percentage overlapConsistent prose where you want control over sizeToo small loses context; too large dilutes relevance
HierarchicalSmall child chunks for matching, replaced by larger parent chunks at retrievalLong structured documents such as manuals and policiesFewer results than requested; not recommended with S3 Vectors
SemanticBreaks where meaning shifts between sentences, using a modelDocuments whose topics change without clear headingsExtra model cost during ingestion
No chunkingEach file becomes a single chunkContent already split into small, self-contained filesNo page numbers in citations and no page filters

Choosing an embedding model and vector store together

The embedding model's vector dimensions must match the index, so these choices are made as a pair6.

  • If

    Users search for part numbers, codes or names as well as concepts.

    Then

    Use a store that supports hybrid search with a filterable text field: Aurora PostgreSQL, OpenSearch Serverless or MongoDB Atlas4.

    Other stores fall back to semantic-only search.

  • If

    Query volume is low and storage cost matters more than latency.

    Then

    Consider Amazon S3 Vectors, avoiding hierarchical chunking with large token counts6.

    AWS positions S3 Vectors for infrequent queries, and hierarchical metadata can exceed its per-vector limits.

  • If

    You want binary embeddings to reduce storage.

    Then

    Use OpenSearch Serverless or an OpenSearch managed cluster6.

    They are the only supported stores for binary vectors.

  • If

    Relationships between entities matter as much as passage similarity.

    Then

    Evaluate Neptune Analytics with GraphRAG6.

    A graph can connect facts spread across documents that plain vector search returns separately.

  • If

    Documents or questions span several languages.

    Then

    Choose a multilingual embedding model and test retrieval with queries written by native speakers.

    An English-oriented embedding can rank other-language queries poorly even when the answer is present.

Retrieval settings to tune before changing the model

0 of 7 checked

Citations and the no-answer path

Citations make a RAG answer checkable. With RetrieveAndGenerate, a custom generation prompt template must keep the $output_format_instructions$ placeholder, or responses come back without citations4. If you call Retrieve and write your own prompt, carry chunk identifiers and source locations through to the response yourself.

Design the no-answer path deliberately. Instruct the model to say when the retrieved passages do not answer the question, show what was searched and offer a route to a person. An application that always answers will eventually answer from the model's general knowledge while appearing to cite your documents.

What contextual grounding checks can and cannot catch

Evaluating retrieval and answers before release

  1. Build a question set with ground truth

    Collect real questions with the passages that should be retrieved and a reference answer, including questions the documents cannot answer.

  2. Test retrieval on its own

    Check whether the right passages appear for each question before judging generated answers. Bedrock RAG evaluations can score retrieval alone, or retrieval with generation, against your ground truth8.

  3. Test answers with retrieval held fixed

    Score answers for correctness, completeness and faithfulness to the passages, and confirm that unanswerable questions are declined.

  4. Test permissions

    Query as users with different access rights and confirm that restricted documents never appear in results or answers.

  5. Re-run after every change

    Repeat the set after re-chunking, a new embedding model, a prompt change or a large sync, and compare with the previous run.

When a custom RAG pipeline is the better choice

Some requirements point to building the pipeline yourself, often still on Bedrock models and AWS storage: document parsing that needs domain-specific logic, retrieval that must blend vector search with structured queries in one ranking, access rules too complex for metadata filters, or full control of every intermediate step for audit. The cost is ownership of ingestion, indexing, retrieval and evaluation code. Prototype on Knowledge Bases first and measure where it falls short on your question set; those gaps become the specification for the custom parts.

Questions and answers

How do Bedrock Knowledge Bases handle document permissions?

Retrieval does not know who is asking unless you tell it. In a customer-managed knowledge base, the usual approach is to store access attributes as metadata and apply filters derived from the authenticated user on the server, never from the prompt. Managed knowledge bases also document an access-control-list awareness option1. Either way, test with real user roles that restricted documents never surface.

How do we keep a Bedrock knowledge base up to date?

Run ingestion jobs whenever sources change, on a schedule or triggered by upload events through the StartIngestionJob API. Decide how deletions are handled, because the data source's deletion policy controls whether vectors for removed documents are deleted or retained. After large syncs, rerun your evaluation set to confirm answer quality has not shifted.

Can a knowledge base serve content in several languages?

Yes, but language support comes from the embedding and generation models you choose rather than from the knowledge base itself. Pick a multilingual embedding model when documents or questions span languages, then test retrieval with queries written by native speakers in each one. Store language as a metadata field so you can filter, and report quality, per language.

Do we need Bedrock Agents to use a knowledge base?

No. An application can call the Retrieve or RetrieveAndGenerate APIs directly, which is the simplest pattern for question answering over documents. Agents become relevant when the application must choose between several tools or take actions, which is the territory of our Bedrock and AgentCore practice.

Sources

  1. Build a managed knowledge base (Amazon Bedrock User Guide) — Amazon Web Services · checked 10 October 2026
  2. Connect a data source to your knowledge base — Amazon Web Services · checked 10 October 2026
  3. Prerequisites for your Amazon Bedrock knowledge base data — Amazon Web Services · checked 10 October 2026
  4. Configure and customize queries and response generation — Amazon Web Services · checked 10 October 2026
  5. How content chunking works for knowledge bases — Amazon Web Services · checked 10 October 2026
  6. Prerequisites for using a vector store you created for a knowledge base — Amazon Web Services · checked 10 October 2026
  7. Use contextual grounding check to filter hallucinations in responses — Amazon Web Services · checked 10 October 2026
  8. Evaluate the performance of Amazon Bedrock resources — Amazon Web Services · checked 10 October 2026

More in AWS Bedrock

Back to AWS Bedrock

Next step

Test your Bedrock knowledge base design against real questions

Share the document types, rough volumes and a handful of questions users ask. We will suggest a knowledge base type, chunking strategy and store, and how to evaluate the design before release.

Discuss a Bedrock RAG design