# MCP Best Practices, Market Reality, and Cached Knowledge Rationale ## Executive Summary MCP (Model Context Protocol) has rapidly evolved from Anthropic's November 2024 release to become the de-facto standard for AI-to-tool integration, with adoption by OpenAI in March 2025 and donation to the Linux Foundation's Agentic AI Foundation in December 2025. However, its effectiveness for document-centric use cases depends heavily on implementation patterns. This analysis addresses whether MCP with cached/pre-computed knowledge layers makes sense versus alternatives. --- ## 1. Market Best Practices and Real Examples ### 1.1 Production MCP Document Implementations | Implementation | Approach | Key Features | |----------------|----------|--------------| | **Chroma MCP Server** | Vector-native semantic search | Sets the standard for semantic document management using vector search, supporting both ephemeral and persistent storage | | **Knowledge-Base-MCP (Puran)** | Production-grade RAG | Agent-directed hybrid retrieval with auto mode choosing among dense, hybrid, sparse, and rerank routes; when scores are low, returns abstain so client can decide whether to retry | | **AWS Bedrock KB + MCP** | Enterprise knowledge bases | Connects to Amazon Bedrock Knowledge Bases for semantic search capabilities, providing unified access pattern regardless of underlying AWS service | | **Basic Memory** | Local-first knowledge graphs | Local-first knowledge management system that builds a semantic graph from Markdown files, enabling persistent memory across conversations | | **AtScale MCP** | Semantic layer for BI | Exposes semantic models to any MCP-compatible AI agent with real-time model discovery where new models become queryable instantly after deployment | ### 1.2 Dominant Architecture Patterns **Pattern A: MCP + Vector Store (Most Common)** ``` Documents → Chunking → Embeddings → Vector DB → MCP Server → LLM Client ``` Prepare the knowledge base by collecting and preprocessing data, chunk into reasonably sized pieces, embed using an embedding model, and load into a vector database like FAISS, Weaviate, or Pinecone **Pattern B: MCP + Knowledge Graph** ``` Documents → Entity Extraction → Graph DB → MCP Server → LLM Client ``` The MCP server exposes Graphiti's core capabilities including episode management, entity management, search capabilities with semantic and hybrid search for facts and node summaries **Pattern C: MCP + Hybrid (Best Practice)** ``` Documents → [Vectors + Graph + BM25] → Unified MCP Interface → LLM Client ``` Combines vector search plus BM25 lexical search using RRF, then reranks. Best for complex queries with both conceptual and specific keyword requirements --- ## 2. When MCP Beats Alternatives ### 2.1 MCP vs. Direct Document Attachment | Scenario | Winner | Rationale | |----------|--------|-----------| | **One-off Q&A on small docs** | Direct Attachment | Zero setup, full context visible | | **Repeated queries on same corpus** | MCP | Avoid re-processing, selective retrieval | | **Multi-user access** | MCP | AI applications can seamlessly access up-to-date information and context as needed through unified protocol | | **Agentic workflows** | MCP | Enables "agentic" AI systems that can autonomously interact with multiple systems, retrieve the latest information, and even take actions | | **Version-controlled knowledge** | MCP | Git-based updates, deterministic ingestion | ### 2.2 MCP vs. Traditional RAG API | Aspect | Traditional RAG | MCP-wrapped RAG | Advantage | |--------|-----------------|-----------------|-----------| | **Standardization** | Custom per-service | Universal protocol | MCP | | **Tool Discovery** | Manual documentation | Auto-discovery | MCP | | **Multi-source** | N×M integrations | M+N integrations | MCP flips the N×M problem to an M+N model: tool providers implement one standard MCP server, and AI app developers implement MCP client support once | | **Agent autonomy** | Fixed pipeline | LLM itself makes contextual decisions about how to interact with the data, determining query strategy and prompt formulation | MCP | ### 2.3 When NOT to Use MCP 1. **Simple, one-time document analysis** - Direct attachment wins 2. **Highly dynamic real-time data** - Direct API calls may be simpler 3. **No multi-client requirement** - Overhead not justified 4. **Prototype/exploratory phase** - Start simple, add MCP later --- ## 3. Cached/Pre-computed Knowledge: The Rationale ### 3.1 What is "Cached Knowledge" in MCP Context? Three tiers of pre-computation that MCP servers can provide: | Tier | What's Cached | When Generated | Example | |------|---------------|----------------|---------| | **1. Embeddings** | Vector representations | At ingestion | Semantic search index | | **2. Summaries** | Condensed content | At ingestion or scheduled | "This document describes X" | | **3. Knowledge Graph** | Entity-relationship extractions | At ingestion | "Goal M1 → implemented by → System X" | ### 3.2 Rationale FOR Cached Knowledge Layers **A. Token Economics** Simply stuffing all potentially relevant data into the prompt is inefficient and sometimes impossible. MCP enables dynamically retrieving just-in-time context from external sources as needed instead of front-loading everything Cost comparison (200-page document): - Direct attachment: ~100K tokens every query = $0.30-3.00/query - MCP with embeddings: ~2K tokens retrieved = $0.006-0.06/query - MCP with summaries: ~500 tokens = $0.0015-0.015/query **B. Response Quality** Tune the number of retrieved documents included in the prompt - often 3-5 good snippets are better than 10 - sometimes using too many can overwhelm the model Pre-computed summaries ensure the LLM gets: - Condensed, relevant context - Pre-extracted key facts - Relationship context from knowledge graphs **C. Caching Benefits** Implement caching at multiple levels. Cache the results of common retrieval queries — for example, if many users ask "What is the refund policy?", you can cache the answer or at least the retrieved document so the agent doesn't vector-search the same question repeatedly Implementing advanced caching (exact, semantic, task-aware) to avoid redundant API calls, tracking and optimizing costs across providers **D. Offline/Latency Benefits** Pre-computed knowledge enables: - Faster response times (no embedding at query time) - Offline capability (no API calls for embeddings) - Deterministic behavior (same query = same retrieval) ### 3.3 Implementation: Hierarchical Memory Provides hierarchical memory storage with three-tier compression (chunks, micro-summaries, meta-summaries) Example architecture: ```yaml # Pre-computed knowledge layers raw_chunks: - content: "Full text chunk" - embedding: [0.1, 0.2, ...] micro_summaries: - chunk_ids: [1, 2, 3] - summary: "These chunks describe the cultural heritage preservation goals" - keywords: ["heritage", "preservation", "M1"] meta_summaries: - scope: "02-goals subdomain" - summary: "Six strategic goals (M1-M6) covering preservation through cybersecurity" - entity_count: 6 ``` ### 3.4 Rationale AGAINST Over-caching | Risk | Description | Mitigation | |------|-------------|------------| | **Staleness** | Summaries out of sync with source | Git-triggered regeneration | | **Loss of nuance** | Summarization loses detail | Keep raw chunks accessible | | **Hallucination risk** | LLM-generated summaries may be wrong | Human review for critical content | | **Storage cost** | Multiple representations | Tiered storage, compress cold data | --- ## 4. Recommended Architecture for KISC Use Case ### 4.1 Current State Analysis Your Architecture-as-Code repo has: - ✅ Structured YAML entities (good for knowledge graph) - ✅ Explicit relationships in `edges.yaml` - ❌ No vector embeddings - ❌ No pre-computed summaries - ❌ Primitive substring search (not semantic) ### 4.2 Proposed Enhanced Architecture ``` ┌─────────────────────────────────────────────────────────────────┐ │ KISC Architecture MCP Server │ ├─────────────────────────────────────────────────────────────────┤ │ Layer 1: Raw Data │ │ ├── YAML entities (goals, systems, services, etc.) │ │ └── Markdown documentation │ ├─────────────────────────────────────────────────────────────────┤ │ Layer 2: Pre-computed Knowledge (NEW) │ │ ├── embeddings.index (vector search via FAISS/Qdrant) │ │ ├── summaries.yaml (entity-level summaries in Latvian) │ │ ├── knowledge_graph.json (Neo4j-style graph export) │ │ └── glossary.yaml (term definitions for natural language) │ ├─────────────────────────────────────────────────────────────────┤ │ Layer 3: MCP Tools (ENHANCED) │ │ ├── semantic_search(query) → vector similarity │ │ ├── get_entity(id) → full entity + related context │ │ ├── explain_concept(term) → natural language explanation │ │ ├── find_relationships(entity) → graph traversal │ │ ├── summarize_domain(domain) → pre-computed summary │ │ └── natural_query(latvian_question) → LLM-friendly response │ └─────────────────────────────────────────────────────────────────┘ ``` ### 4.3 Pre-computation Pipeline ```bash # On every git commit to Architecture-as-Code: 1. Load all YAML entities 2. Generate embeddings for each entity (title + description) 3. Generate micro-summaries for each subdomain 4. Build/update knowledge graph from edges.yaml 5. Create glossary from all entity titles/codes 6. Store in /mcp/cache/ directory ``` ### 4.4 Query Flow (Enhanced) **User asks:** "Kādas sistēmas realizē valodas tehnoloģiju mērķi?" (What systems implement the language technology goal?) **Current behavior:** Returns `[]` (no match for Latvian natural language) **Enhanced behavior:** 1. `natural_query` tool receives Latvian question 2. Extracts intent: "systems implementing language technology goal" 3. Maps to goal.m4 (Latviešu valoda digitālajā laikmetā) 4. Traverses knowledge graph: goal.m4 → implemented_by → [sys.valoda.01, sys.valoda.02, ...] 5. Returns pre-computed summary + entity list --- ## 5. Decision Framework: When to Add Cached Knowledge | Question | If YES | If NO | |----------|--------|-------| | Will multiple users query the same corpus? | Add caching | Skip | | Is query latency critical (<1s)? | Add embeddings | Direct retrieval OK | | Do users ask in natural language (not IDs)? | Add semantic search | ID-based lookup OK | | Is the corpus >100 entities? | Add summaries | Full scan OK | | Do queries require cross-entity reasoning? | Add knowledge graph | Flat search OK | | Is the corpus updated less than daily? | Pre-compute aggressively | Real-time generation | --- ## 6. Conclusion ### Is MCP with Cached Knowledge Worth It? **YES, when:** - You have a stable, structured knowledge base (like Architecture-as-Code) - Multiple consumers need consistent access - Natural language queries are required - Token cost optimization matters - Cross-entity reasoning is needed **NO, when:** - One-time document analysis - Rapidly changing data (real-time feeds) - Simple keyword lookup suffices - No multi-client requirement ### For KISC Specifically: Your Architecture-as-Code approach is **fundamentally correct** but needs: 1. **Semantic search layer** (embeddings for Latvian content) 2. **Pre-computed summaries** (domain/subdomain level) 3. **Natural language interface** (Latvian query handling) 4. **Enhanced MCP tools** (beyond primitive substring search) The investment in these layers will pay off as the architecture grows and more stakeholders (internal teams, external auditors, automated agents) need to query it. --- ## References - Model Context Protocol Specification: https://modelcontextprotocol.io/specification - MCP Server Registry: https://github.com/modelcontextprotocol/servers - Agentic RAG + MCP Integration Guide: https://becomingahacker.org/integrating-agentic-rag-with-mcp-servers - AWS MCP Implementation: https://aws.amazon.com/blogs/machine-learning/unlocking-the-power-of-model-context-protocol-mcp-on-aws/