22.3 Context Grounding: Retrieval-Augmented Generation with Company Data
Key Takeaways
Context grounding in the AI Trust Layer indexes company documents for retrieval-augmented generation without separate vector databases or model subscriptions.
Indexes use storage buckets or Confluence Cloud, Dropbox, Google Drive, and OneDrive & SharePoint via connections with Files.Read scope.
Supported file types are CSV, DOCX, JPG, JSON, PDF, PNG, TXT, and XLSX, with ten indexes per tenant by default.
Indexes inherit folder permissions, so organize one index per audience and folder path.
Refresh indexes with Update Context Grounding Index, search with Context Grounding Search, and use DeepRAG for PDF analysis.
22.3 Context Grounding: Retrieval-Augmented Generation with Company Data
Core Concept: Context grounding is part of the AI Trust Layer. It makes business documents usable by UiPath's generative AI features through retrieval-augmented generation (RAG): documents are indexed as embeddings, the most relevant passages are retrieved for each question, and the model answers using them. No separate embedding model, vector database, or LLM subscription is needed.
Where It Is Used
- GenAI Activities, such as Content Generation grounded on an index and Context Grounding Search.
- Autopilot for Everyone, to answer questions from company knowledge.
- Agents, as a knowledge source.
Indexes and Data Sources
An index holds the embeddings of documents from one data source. Indexes are created and managed in Orchestrator (the folder-level Indexes permission governs them) and are then available across the platform.
| Data source | Access |
|---|---|
| Orchestrator storage buckets | Upload documents directly, through Studio activities, or through the API |
| Confluence Cloud, Dropbox, Google Drive, Microsoft OneDrive & SharePoint | Through Integration Service connections authorized with at least the Files.Read scope |
Limits and rules
- Supported file types: CSV, DOCX, JPG, JSON, PDF, PNG, TXT, and XLSX.
- Index limit: ten indexes per tenant in Automation Cloud, expandable on request.
- UiPath recommends a 1:1 relationship between indexes and folder paths in the data source, which keeps data for different audiences separate.
- Indexes inherit folder permissions: users without access to the folder cannot view, update, delete, or use its indexes.
Keeping Indexes Current
When documents change, the index must be refreshed:
- Use Update Context Grounding Index in a workflow, for example after a process uploads new policy documents.
- Schedule regular refreshes for sources that change often.
- An Index Completed event lets automations continue when indexing has finished.
Using Context Grounding in Workflows
Search, then generate
- Context Grounding Search takes a query and returns the most relevant passages from an index.
- Pass the passages to Content Generation with instructions such as "Answer only from the provided passages; say you don't know otherwise."
- Validate the answer, and log which sources were used.
Grounded content generation
Content Generation can reference an index directly, retrieving context behind the scenes.
DeepRAG
For deeper analysis across documents, a DeepRAG analysis can be run and its result retrieved with Get DeepRAG Analysis by ID. DeepRAG works with PDF documents.
Designing a Good Knowledge Base
- Curate the documents. Remove outdated versions, because the index cannot tell which policy is current.
- One index per audience and topic, for example "HR policies" and "IT runbooks", matching folder paths.
- Prefer text-based documents. Scanned documents with poor quality produce weak passages.
- Write clear queries in automations, for example "vacation carry-over rules for employees in Germany" rather than "vacation".
- Test with real questions and compare answers to the source documents.
Security and Governance
- Folder permissions protect indexes, so place indexes in folders whose members may see the documents.
- Data source connections should use accounts that can read only the intended folders.
- AI Trust Layer policies govern which AI features may use company data.
Worked Example: HR Question Answering
- HR uploads the handbook PDFs and policy documents to a storage bucket in the HR folder.
- An administrator creates an index HR-Policies on that bucket.
- An attended automation for HR assistants asks for a question, runs Context Grounding Search on HR-Policies, and sends the top passages to Content Generation with instructions to answer only from them and cite the document names.
- When HR publishes a new policy, a small process uploads it and runs Update Context Grounding Index.
- Employees outside HR cannot use the index, because they lack access to the HR folder.
Exam-Style Scenarios
Scenario 1: Answers quote a policy that was replaced last week. Cause: the index was not refreshed. Fix: run Update Context Grounding Index after publishing new documents, or schedule refreshes.
Scenario 2: A company wants to ground answers in documents stored in SharePoint. Use a Microsoft OneDrive & SharePoint connection with at least the Files.Read scope as the index's data source.
Scenario 3: Legal and HR documents must not be visible to each other's users. Use separate indexes on separate folder paths in separate Orchestrator folders; indexes inherit folder permissions.
Scenario 4: The tenant already has ten indexes and needs another. Consolidate indexes that serve the same audience, or request an increase, because ten per tenant is the default limit.
Common Traps
- Forgetting to refresh the index after documents change.
- Creating one large index for all departments, which mixes audiences and permissions.
- Assuming grounding guarantees correctness; answers still need validation for important decisions.
- Using unsupported file types or connections without the Files.Read scope.
Which data sources can a context grounding index use?
Only local folders on the robot machine.
Only Data Fabric entities.
Orchestrator queues and assets.
Orchestrator storage buckets, and Confluence Cloud, Dropbox, Google Drive, or OneDrive & SharePoint through Integration Service connections with at least Files.Read scope.
A user in the Sales folder tries to use an index created on the HR folder's bucket. What happens?
They cannot view or use it, because indexes inherit folder permissions.
They can use it, because indexes are tenant-wide.
They can use it read-only.
They get a copy of the index.
How many context grounding indexes does an Automation Cloud tenant get by default?
One
Ten, expandable on request
Unlimited
One per storage bucket, with no maximum
Sections you finish are checked off in the contents.