Skip to main content

RAG (Retrieval-Augmented Generation)

Overview

RAG combines the power of retrieval-based systems with generative AI models to provide accurate, context-aware responses. Instead of relying solely on the LLM's training data, RAG fetches relevant information from your document repository before generating responses.

The existing rag agent type has two execution modes; clients do not create a separate agent type:

  • Agentic RAG is selected for models whose ML Platform metadata explicitly reports function-calling support. It uses a constrained agent harness with Content Lake corpus search, document-scoped search, and adjacent-chunk context tools.
  • Legacy RAG is selected for models whose metadata does not report function-calling support or when capability lookup fails. It retains the fixed workflow.

Agentic RAG always performs retrieval and uses maxRetrievalCalls; it ignores enableMultihopQueryRefinement. Legacy RAG ignores maxRetrievalCalls; its optional multi-hop workflow uses a fixed two-hop refinement limit. Deep Search remains a legacy-only mode and is rejected when agentic RAG is selected.

Configuration

Here's how to configure a RAG agent with our enhanced architecture:

{
"name": "document-helper",
"displayName": "Document Helper",
"description": "Document assistant",
"agentType": "rag",
"notes": "Agent for document queries",
"config": {
"hxqlQuery": "SELECT * FROM SysContent",
"limit": 5,
"hybridSearch": true,
"adjacentChunkRange": 1,
"adjacentChunkMerge": true,
"rerankerTopN": 5,
"enableHallucinationCheck": true,
"enableMultihopQueryRefinement": false,
"systemPrompt": "Context information is below.\n---------------------\n{context_str}\n---------------------\nGiven the context information and not prior knowledge, answer the query.\nQuery: {query_str}\nAnswer: ",
"llmModelId": "amazon.nova-micro-v1:0",
"inferenceConfig": {
"maxTokens": 4000,
"temperature": 0.7
},
"guardrails": ["HAIP-Profanity", "HAIP-Insults-High"]
}
}

Configuration Parameters

ParameterTypeDescriptionRequiredDefault
hxqlQuerystringHXQL query for Content Lake retrieval filteringNonull
limitinteger (≥ 1)Maximum number of chunks to retrieve from Content LakeNo50
hybridSearchbooleanEnable/disable hybrid search (embeddings + full-text)Notrue
adjacentChunkRangeinteger (≥ 0)Number of adjacent chunks to fetch around each retrieved chunk. For range=N, fetches N chunks before and N chunks after each result (0 = disabled)No0
adjacentChunkMergebooleanWhen true, adjacent chunk text is merged into the parent chunk in document order. When false, adjacent chunks are returned as separate nodesNofalse
rerankerTopNinteger (≥ 1)Number of top results to keep after rerankingNo13
enableHallucinationCheckbooleanEnable hallucination detection. When enabled, the agent validates that the generated response is supported by the retrieved chunks and retries if hallucination is detectedNofalse
enableMultihopQueryRefinementbooleanEnable the fixed two-hop query-refinement workflow for legacy RAG. Ignored by agentic RAGNofalse
maxRetrievalCallsinteger (3–10)Maximum combined corpus-search, document-search, and document-context tool calls for agentic RAG. Always enforced by agentic RAG and ignored by legacy RAGNo3
systemPromptstringLegacy RAG synthesis template or supplemental answer/formatting instructions for agentic RAG. Agentic RAG preserves mandatory retrieval and grounding rulesNonull
llmModelIdstringID of the language model to useYesnull
inferenceConfigobjectLLM parameters (see Inference Config)NoSee defaults
guardrailsarrayList of guardrail names or parameterized guardrails (see Guardrails guide)Nonull
tip

Parameters like enableHallucinationCheck, enableMultihopQueryRefinement, adjacentChunkRange, adjacentChunkMerge, rerankerTopN, and hybridSearch can also be overridden per invocation request. maxRetrievalCalls is agent configuration and cannot be overridden per invocation.

System Prompt Compatibility

Recommendation: omit systemPrompt unless the agent needs domain-specific answer or formatting instructions. The built-in agentic prompt already requires retrieval for every question, rejects unsupported answers, responds in the user's language, and formats grounded answers in Markdown.

Legacy RAG uses systemPrompt as its synthesis template, where {context_str} and {query_str} refer to retrieved context and the current query. Agentic RAG appends the configured prompt as constrained custom instructions; it maps those legacy placeholders to semantic references to retrieval evidence and the current question, so existing templates can transition safely. Custom instructions cannot disable retrieval, allow general knowledge, or replace the exact empty-result fallback.

Invocation Parameters

These optional fields can be passed in the request body when invoking a RAG agent (/invoke or /invoke-stream):

ParameterTypeDescriptionDefault
hxqlQuerystringOverride the agent-level HXQL query for this invocationAgent config value
hybridSearchbooleanEnable/disable hybrid search (embeddings + full-text). Uses config default if not specified.Agent config value
guardrailsstring[]Additional guardrails for this invocation[]

Max Retries

When invoking agents that error for any reason (E.g., input is too large, network issue, etc.), the client you're using to make the request could time out. To help with this, you could set maxRetries in the inferenceConfig to a low number (E.g., 2).

Best Practices

  1. Document Preparation

    • Ensure documents are properly chunked
    • Maintain consistent formatting
    • Include metadata for filtering
  2. Query Formation

    • Be specific in queries
    • Use natural language
    • Include relevant context
  3. Filter Usage

    • The hxqlQuery parameter follows the HxQL filtering structure
    • Common filters include document type, department, and date ranges
    • Combine with semantic search for better results
tip

The filtering capabilities are provided by the Semantic API. Check their documentation for the complete list of supported filters and proper filter syntax.

  1. Guardrails
    • For comprehensive information on discovering, configuring, and merging guardrails, see the Guardrails guide.

Invoking a RAG Agent

Standard Invocation

POST /v1/agents/{agent_id}/versions/{version_id}/invoke

Example Request:

{
"messages": [
{
"role": "user",
"content": "What's in our HR policy about vacation days?"
}
],
"hxqlQuery": "SELECT * FROM SysContent"
}

Example Response:

{
"object": "response",
"createdAt": 1741705447,
"model": "amazon.nova-micro-v1:0",
"output": [
{
"type": "message",
"status": "completed",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": "According to the HR policy, full-time employees receive 15 vacation days per year..."
}
]
}
],
"customOutputs": {
"sourceNodes": [
{
"docId": "6e4a7f58-13f1-4d3f-83b9-ec86c5b7df60",
"chunkId": "b9f622c9-7a62-46fc-8cf1-0d1f4d3c35d7",
"score": 0.95,
"text": "VACATION POLICY: Full-time employees are entitled to 15 vacation days per calendar year..."
}
],
"ragMode": "normal"
}
}

Response Fields

FieldDescription
outputList of response messages from the agent
customOutputs.sourceNodesDocuments retrieved from Content Lake that were used as context
customOutputs.sourceNodes[].scoreRelevance score (0–1) of the retrieved chunk
customOutputs.sourceNodes[].textText content of the retrieved chunk
customOutputs.ragModeRAG mode used (normal or deepResearch)
tip

Use latest as the version_id to invoke the most recent version of the agent.

Multi-Turn Conversation

RAG Agents support multi-turn conversations by passing previous messages in the messages array. The last message must always have role: "user".

Session-Based Memory

RAG agents support optional session-based short-term memory. Add an X-Session-ID header with a consistent UUID to maintain conversation context across multiple invocations. The platform does not auto-generate session IDs — you must generate and manage them yourself (any valid UUID).

{
"messages": [
{
"role": "user",
"content": "What's our vacation policy?"
},
{
"role": "assistant",
"content": "Full-time employees receive 15 vacation days per year..."
},
{
"role": "user",
"content": "How do I request time off?"
}
],
"hxqlQuery": "SELECT * FROM SysContent"
}

Streaming Support

RAG agents support streaming responses through the unified /v1/agents/{agent_id}/versions/{version_id-or-latest}/invoke-stream endpoint:

POST /v1/agents/{agent_id}/versions/{version_id-or-latest}/invoke-stream

Example Request:

{
"messages": [
{
"role": "user",
"content": "What's in our HR policy about vacation days?"
}
],
"hxqlQuery": "SELECT * FROM SysContent",
"enableHallucinationCheck": true,
"enableMultihopQueryRefinement": false
}
tip

When using streaming, responses come in chunks. Each chunk is a valid JSON object containing a portion of the complete response. Agentic RAG suppresses internal tool events and buffers answer text until the final grounded graph state, so the answer may arrive as one terminal text delta rather than token-by-token. The completed chunk contains the final output and retrieved source nodes.

tip

Due to the nature of streaming, returning data chunk by chunk, failures that occur during runtime are difficult to return to the client. Therefore, the returned status code might be a 200. Please check the logs for errors that occur.

Error Responses

When a request fails, the API returns a JSON error response in the following format:

{
"status": 400,
"error": "HTTPException",
"message": "Description of what went wrong"
}

Common Errors

StatusErrorDescription
400Bad RequestInvalid agent configuration (e.g., missing hxqlQuery at both config and invocation level) or invalid invocation parameters.
401UnauthorizedMissing or expired access token. Re-authenticate to obtain a new token.
403ForbiddenInsufficient permissions for the requested operation. Verify your IAM user group permissions.
404Not FoundAgent ID or version does not exist. Use GET /v1/agents to verify available agents.