RAG in Microsoft Agent Framework – Overview

Microsoft’s Agent Framework treats retrieval as a first‑class capability so agents can fetch only what they need (or always fetch), attach source metadata, and call search as a tool during reasoning. The result: more efficient, auditable, and controllable Retrieval‑Augmented Generation (RAG) for production assistants.

tl;dr

  • Microsoft Agent Framework implements RAG via TextSearchProvider (an AIContextProvider) and a Semantic Kernel bridge to many vector stores.
  • Two retrieval modes: BeforeAIInvoke (automatic injection) and OnDemandFunctionCalling (agent calls search as a tool).
  • Recommended starting defaults: top_k = 3–5, chunk size ≈ 500–1,000 characters with 10–20% overlap, and prefer OnDemandFunctionCalling for cost/latency control.
  • Key production concerns: chunking, metadata for citations, latency and cost management, freshness, security and telemetry.

Key terms

  • AIContextProvider: a component that supplies contextual data to an agent before the model is invoked. TextSearchProvider implements retrieval as an AIContextProvider.
  • Tool / Function calling: exposes actions (search, API calls) as callable functions the agent can invoke on demand during reasoning. The TextSearchProvider can be advertised as such (OnDemandFunctionCalling).
  • VectorStore / TextSearchStore: VectorStore holds embeddings and indexes; TextSearchStore is a convenience schema for text chunks + metadata built on a VectorStore.

Core components

  • VectorStore: stores embeddings + metadata. Backends supported via Semantic Kernel include InMemory, Qdrant, Pinecone, Redis, Weaviate, Azure AI Search, etc.
  • TextSearchStore: wraps a VectorStore with a text‑centric schema (collectionName, namespace, vector dimensions, chunk metadata).
  • TextSearchProvider: the AIContextProvider that performs searches and either injects results or exposes search as a callable tool.
  • Kernel bridge: converts Semantic Kernel search functions into Agent Framework tools so the same agent logic works across backends.
  • Agent / AgentThread: the runtime that combines user messages, context providers, tools, and the LLM to produce grounded responses.

How RAG is implemented — simple flow
1) Index your docs into a VectorStore:

  • Generate embeddings (Azure OpenAI, OpenAI, etc.) and store vectors with metadata (source URL, chunk id, section).
    2) Wrap the VectorStore in a TextSearchStore (choose collectionName, namespaces).
    3) Create a TextSearchProvider backed by the TextSearchStore and add it to the agent thread’s AIContextProviders.
    4) Choose SearchTime:
  • BeforeAIInvoke (default): run searches automatically before each model call and inject the top results into the prompt.
  • OnDemandFunctionCalling: advertise search as a callable tool and let the agent call it while reasoning.
    5) Run the agent: retrieved text is combined with the prompt and sent to the LLM; results can include source metadata for inline citations.

Injection mechanics — what actually gets passed to the model

  • In BeforeAIInvoke mode, the provider runs a vector search (by default top_k hits) and concatenates the retrieved chunks into the agent’s context. That context is typically appended as extra system/assistant content and is subject to truncation/prioritization to respect the model’s context window.
  • In OnDemandFunctionCalling mode, the search appears as a callable tool; the LLM receives the tool’s output (chunks + metadata) only when the agent invokes the tool.
  • Retrieved results include metadata (source URL, document id, chunk id, score). Use that metadata for citations and audit trails.
  • You control ranking limits and filtering via TextSearchProviderOptions (top_k, namespaces, recency filters, message memory limits).

Defaults and practical parameter guidance

  • top_k (number of chunks returned): start with 3–5. More adds context but increases token use and noise.
  • Chunk size: aim for 500–1,000 characters per chunk (roughly 75–200 tokens). This balances retrieval granularity and coherent passages. If you prefer token‑based chunks, 200–500 tokens is a reasonable upper bound for longer passages.
  • Overlap: 10–20% overlap between adjacent chunks helps prevent losing relevant sentence boundaries.
  • Relevance filtering: use namespace/collectionName to scope queries (multi‑tenant or multi‑corpus setups).
  • Embedding model: choose a semantic embedding suitable for your domain; embedding quality directly affects retrieval relevance.

BeforeAIInvoke vs OnDemandFunctionCalling — choose by use case

  • BeforeAIInvoke (automatic):
  • Best when almost every user query must be grounded (e.g., compliance answers).
  • Simpler to reason about: search runs, results are always available to the model.
  • Downsides: higher cost and possible token bloat.
  • Trace: user query -> provider runs search -> top_k chunks injected -> model call -> response.
  • OnDemandFunctionCalling (agentic/tool-based):
  • Best when many queries are casual or do not need grounding, and you want the agent to decide when to fetch data.
  • Enables multi‑step reasoning (agent thinks, calls search, examines results, calls other tools, returns final).
  • Lower baseline cost and conditional latency.
  • Trace: user query -> agent begins reasoning -> decides to call Search tool -> search returns chunks -> agent may call another tool or ask follow-up -> final model call -> response.

Example: a multi-step agentic sequence (conceptual)
1) User: “How do I roll back build 1.2.3?”
2) Agent (thinking): Not sure. Calls Search tool with query “roll back build 1.2.3 runbook”.
3) Search returns runbook chunks A, B (with source URLs).
4) Agent inspects chunks, calls a “Validate-Runbook” tool to confirm commands are safe.
5) Agent composes final answer quoting steps and adds “[source: Runbook / sectionX | URL]” inline for each step.

Prompt and output formatting — keep answers auditable

  • When injecting retrieved content, format chunks with clear attribution. Example snippet used in the prompt:
    [Retrieved 1/3] Title: “Rollback Procedure” — Source: https://contoso/docs/runbook#sectionX
    “Step 1: … Step 2: …” (chunk id: abc123)
  • When returning a final answer, include inline citations:
    “To roll back build 1.2.3, follow steps 1–3 (see Runbook: https://contoso/docs/runbook#sectionX).”
  • If using OnDemandFunctionCalling, have the tool return structured metadata (title, url, chunk_id, score) so the agent can produce precise citations.

Code outline (C#) — on‑demand search example (conceptual)

// 1) Create embedding generator (IEmbeddingGenerator), vector store and TextSearchStore
var embeddingGenerator = /* AzureOpenAI embedding client */;
var vectorStore = new InMemoryVectorStore(new() { EmbeddingGenerator = embeddingGenerator });
using var textSearchStore = new TextSearchStore(vectorStore, collectionName: “Docs”, vectorDimensions: 1536);

// 2) Create TextSearchProvider with on‑demand behavior
var options = new TextSearchProviderOptions { SearchTime = TextSearchProviderOptions.TextSearchBehavior.OnDemandFunctionCalling, TopK = 4 };
var textSearchProvider = new TextSearchProvider(textSearchStore, options);

// 3) Attach to agent thread
var agentThread = new ChatHistoryAgentThread();
agentThread.AIContextProviders.Add(textSearchProvider);

// 4) Invoke the agent — the agent may call the search tool during reasoning
var response = await agent.InvokeAsync(“How do I roll back build 1.2.3?”, agentThread).FirstAsync();
// response includes final answer; the framework supplies tool outputs when the agent invoked the search tool

Production checklist and operational considerations

  • Latency: measure search latency and embedding latency; cache frequent queries; prefer on‑demand to avoid unnecessary embedding at request time.
  • Cost: monitor embedding and LLM call spend; use top_k and chunk limits strategically; cache and reuse embeddings where possible.
  • Freshness: plan an ingestion cadence and a strategy for reindexing changed documents.
  • Chunking/metadata: store rich metadata (source URL, section, timestamp) to make citations reliable.
  • Hallucination & prompt injection: sanitize retrieved text, require provenance for critical facts, and apply verification steps for high‑risk actions.
  • Scaling: choose a production VectorStore that supports the throughput and replication you need (Qdrant, Pinecone, Redis, Azure AI Search).
  • Security & permissions: treat vector stores and connectors as sensitive; enforce least privilege and secure credentials for connectors (Oracle, SQL, etc.).
  • Telemetry & observability: capture search latency, top_k, cache hit rate, tool call counts, and a hallucination/error metric (mismatch between cited source and assertion). Log search queries and returned metadata for auditing.
  • Limitations: retrieval quality depends on embeddings and chunk strategy; RAG does not replace the need for verification for time‑sensitive facts unless you keep the index fresh.

Where to learn more

Note: The initial draft of this post was written by BlogWriter and then edited by Jesse Liberty
Illustrations by Copilot. Caution: LLMs make mistakes; this post is offered as is.

Unknown's avatar

About Jesse Liberty

Jesse Liberty has three decades of experience writing and delivering software projects and is the author of 2 dozen books and a couple dozen online courses. Liberty is a Senior AI Engineer at the University of Pittsburgh Medical Center, and was a Team Lead and Senior Software Engineer for various corporations, a Senior Technical Evangelist for Microsoft, a Distinguished Software Engineer for AT&T, a VP for Information Services for Citibank and a Software Architect for PBS. He is a 21 year Microsoft MVP.
This entry was posted in AI. Bookmark the permalink.