13 KiB
MCP Best Practices, Market Reality, and Cached Knowledge Rationale
Executive Summary
MCP (Model Context Protocol) has rapidly evolved from Anthropic's November 2024 release to become the de-facto standard for AI-to-tool integration, with adoption by OpenAI in March 2025 and donation to the Linux Foundation's Agentic AI Foundation in December 2025. However, its effectiveness for document-centric use cases depends heavily on implementation patterns. This analysis addresses whether MCP with cached/pre-computed knowledge layers makes sense versus alternatives.
1. Market Best Practices and Real Examples
1.1 Production MCP Document Implementations
| Implementation | Approach | Key Features |
|---|---|---|
| Chroma MCP Server | Vector-native semantic search | Sets the standard for semantic document management using vector search, supporting both ephemeral and persistent storage |
| Knowledge-Base-MCP (Puran) | Production-grade RAG | Agent-directed hybrid retrieval with auto mode choosing among dense, hybrid, sparse, and rerank routes; when scores are low, returns abstain so client can decide whether to retry |
| AWS Bedrock KB + MCP | Enterprise knowledge bases | Connects to Amazon Bedrock Knowledge Bases for semantic search capabilities, providing unified access pattern regardless of underlying AWS service |
| Basic Memory | Local-first knowledge graphs | Local-first knowledge management system that builds a semantic graph from Markdown files, enabling persistent memory across conversations |
| AtScale MCP | Semantic layer for BI | Exposes semantic models to any MCP-compatible AI agent with real-time model discovery where new models become queryable instantly after deployment |
1.2 Dominant Architecture Patterns
Pattern A: MCP + Vector Store (Most Common)
Documents → Chunking → Embeddings → Vector DB → MCP Server → LLM Client
Prepare the knowledge base by collecting and preprocessing data, chunk into reasonably sized pieces, embed using an embedding model, and load into a vector database like FAISS, Weaviate, or Pinecone
Pattern B: MCP + Knowledge Graph
Documents → Entity Extraction → Graph DB → MCP Server → LLM Client
The MCP server exposes Graphiti's core capabilities including episode management, entity management, search capabilities with semantic and hybrid search for facts and node summaries
Pattern C: MCP + Hybrid (Best Practice)
Documents → [Vectors + Graph + BM25] → Unified MCP Interface → LLM Client
Combines vector search plus BM25 lexical search using RRF, then reranks. Best for complex queries with both conceptual and specific keyword requirements
2. When MCP Beats Alternatives
2.1 MCP vs. Direct Document Attachment
| Scenario | Winner | Rationale |
|---|---|---|
| One-off Q&A on small docs | Direct Attachment | Zero setup, full context visible |
| Repeated queries on same corpus | MCP | Avoid re-processing, selective retrieval |
| Multi-user access | MCP | AI applications can seamlessly access up-to-date information and context as needed through unified protocol |
| Agentic workflows | MCP | Enables "agentic" AI systems that can autonomously interact with multiple systems, retrieve the latest information, and even take actions |
| Version-controlled knowledge | MCP | Git-based updates, deterministic ingestion |
2.2 MCP vs. Traditional RAG API
| Aspect | Traditional RAG | MCP-wrapped RAG | Advantage |
|---|---|---|---|
| Standardization | Custom per-service | Universal protocol | MCP |
| Tool Discovery | Manual documentation | Auto-discovery | MCP |
| Multi-source | N×M integrations | M+N integrations | MCP flips the N×M problem to an M+N model: tool providers implement one standard MCP server, and AI app developers implement MCP client support once |
| Agent autonomy | Fixed pipeline | LLM itself makes contextual decisions about how to interact with the data, determining query strategy and prompt formulation | MCP |
2.3 When NOT to Use MCP
- Simple, one-time document analysis - Direct attachment wins
- Highly dynamic real-time data - Direct API calls may be simpler
- No multi-client requirement - Overhead not justified
- Prototype/exploratory phase - Start simple, add MCP later
3. Cached/Pre-computed Knowledge: The Rationale
3.1 What is "Cached Knowledge" in MCP Context?
Three tiers of pre-computation that MCP servers can provide:
| Tier | What's Cached | When Generated | Example |
|---|---|---|---|
| 1. Embeddings | Vector representations | At ingestion | Semantic search index |
| 2. Summaries | Condensed content | At ingestion or scheduled | "This document describes X" |
| 3. Knowledge Graph | Entity-relationship extractions | At ingestion | "Goal M1 → implemented by → System X" |
3.2 Rationale FOR Cached Knowledge Layers
A. Token Economics
Simply stuffing all potentially relevant data into the prompt is inefficient and sometimes impossible. MCP enables dynamically retrieving just-in-time context from external sources as needed instead of front-loading everything
Cost comparison (200-page document):
- Direct attachment: ~100K tokens every query = $0.30-3.00/query
- MCP with embeddings: ~2K tokens retrieved = $0.006-0.06/query
- MCP with summaries: ~500 tokens = $0.0015-0.015/query
B. Response Quality
Tune the number of retrieved documents included in the prompt - often 3-5 good snippets are better than 10 - sometimes using too many can overwhelm the model
Pre-computed summaries ensure the LLM gets:
- Condensed, relevant context
- Pre-extracted key facts
- Relationship context from knowledge graphs
C. Caching Benefits
Implement caching at multiple levels. Cache the results of common retrieval queries — for example, if many users ask "What is the refund policy?", you can cache the answer or at least the retrieved document so the agent doesn't vector-search the same question repeatedly
Implementing advanced caching (exact, semantic, task-aware) to avoid redundant API calls, tracking and optimizing costs across providers
D. Offline/Latency Benefits
Pre-computed knowledge enables:
- Faster response times (no embedding at query time)
- Offline capability (no API calls for embeddings)
- Deterministic behavior (same query = same retrieval)
3.3 Implementation: Hierarchical Memory
Provides hierarchical memory storage with three-tier compression (chunks, micro-summaries, meta-summaries)
Example architecture:
# Pre-computed knowledge layers
raw_chunks:
- content: "Full text chunk"
- embedding: [0.1, 0.2, ...]
micro_summaries:
- chunk_ids: [1, 2, 3]
- summary: "These chunks describe the cultural heritage preservation goals"
- keywords: ["heritage", "preservation", "M1"]
meta_summaries:
- scope: "02-goals subdomain"
- summary: "Six strategic goals (M1-M6) covering preservation through cybersecurity"
- entity_count: 6
3.4 Rationale AGAINST Over-caching
| Risk | Description | Mitigation |
|---|---|---|
| Staleness | Summaries out of sync with source | Git-triggered regeneration |
| Loss of nuance | Summarization loses detail | Keep raw chunks accessible |
| Hallucination risk | LLM-generated summaries may be wrong | Human review for critical content |
| Storage cost | Multiple representations | Tiered storage, compress cold data |
4. Recommended Architecture for KISC Use Case
4.1 Current State Analysis
Your Architecture-as-Code repo has:
- ✅ Structured YAML entities (good for knowledge graph)
- ✅ Explicit relationships in
edges.yaml - ❌ No vector embeddings
- ❌ No pre-computed summaries
- ❌ Primitive substring search (not semantic)
4.2 Proposed Enhanced Architecture
┌─────────────────────────────────────────────────────────────────┐
│ KISC Architecture MCP Server │
├─────────────────────────────────────────────────────────────────┤
│ Layer 1: Raw Data │
│ ├── YAML entities (goals, systems, services, etc.) │
│ └── Markdown documentation │
├─────────────────────────────────────────────────────────────────┤
│ Layer 2: Pre-computed Knowledge (NEW) │
│ ├── embeddings.index (vector search via FAISS/Qdrant) │
│ ├── summaries.yaml (entity-level summaries in Latvian) │
│ ├── knowledge_graph.json (Neo4j-style graph export) │
│ └── glossary.yaml (term definitions for natural language) │
├─────────────────────────────────────────────────────────────────┤
│ Layer 3: MCP Tools (ENHANCED) │
│ ├── semantic_search(query) → vector similarity │
│ ├── get_entity(id) → full entity + related context │
│ ├── explain_concept(term) → natural language explanation │
│ ├── find_relationships(entity) → graph traversal │
│ ├── summarize_domain(domain) → pre-computed summary │
│ └── natural_query(latvian_question) → LLM-friendly response │
└─────────────────────────────────────────────────────────────────┘
4.3 Pre-computation Pipeline
# On every git commit to Architecture-as-Code:
1. Load all YAML entities
2. Generate embeddings for each entity (title + description)
3. Generate micro-summaries for each subdomain
4. Build/update knowledge graph from edges.yaml
5. Create glossary from all entity titles/codes
6. Store in /mcp/cache/ directory
4.4 Query Flow (Enhanced)
User asks: "Kādas sistēmas realizē valodas tehnoloģiju mērķi?" (What systems implement the language technology goal?)
Current behavior: Returns [] (no match for Latvian natural language)
Enhanced behavior:
natural_querytool receives Latvian question- Extracts intent: "systems implementing language technology goal"
- Maps to goal.m4 (Latviešu valoda digitālajā laikmetā)
- Traverses knowledge graph: goal.m4 → implemented_by → [sys.valoda.01, sys.valoda.02, ...]
- Returns pre-computed summary + entity list
5. Decision Framework: When to Add Cached Knowledge
| Question | If YES | If NO |
|---|---|---|
| Will multiple users query the same corpus? | Add caching | Skip |
| Is query latency critical (<1s)? | Add embeddings | Direct retrieval OK |
| Do users ask in natural language (not IDs)? | Add semantic search | ID-based lookup OK |
| Is the corpus >100 entities? | Add summaries | Full scan OK |
| Do queries require cross-entity reasoning? | Add knowledge graph | Flat search OK |
| Is the corpus updated less than daily? | Pre-compute aggressively | Real-time generation |
6. Conclusion
Is MCP with Cached Knowledge Worth It?
YES, when:
- You have a stable, structured knowledge base (like Architecture-as-Code)
- Multiple consumers need consistent access
- Natural language queries are required
- Token cost optimization matters
- Cross-entity reasoning is needed
NO, when:
- One-time document analysis
- Rapidly changing data (real-time feeds)
- Simple keyword lookup suffices
- No multi-client requirement
For KISC Specifically:
Your Architecture-as-Code approach is fundamentally correct but needs:
- Semantic search layer (embeddings for Latvian content)
- Pre-computed summaries (domain/subdomain level)
- Natural language interface (Latvian query handling)
- Enhanced MCP tools (beyond primitive substring search)
The investment in these layers will pay off as the architecture grows and more stakeholders (internal teams, external auditors, automated agents) need to query it.
References
- Model Context Protocol Specification: https://modelcontextprotocol.io/specification
- MCP Server Registry: https://github.com/modelcontextprotocol/servers
- Agentic RAG + MCP Integration Guide: https://becomingahacker.org/integrating-agentic-rag-with-mcp-servers
- AWS MCP Implementation: https://aws.amazon.com/blogs/machine-learning/unlocking-the-power-of-model-context-protocol-mcp-on-aws/