1
0
Files
Architecture-as-Code/ikt-arh-kultura-valodu-tehnologijas/docs/mcp_best_practices_and_rationale.md
2026-10-11 14:36:51 +03:00

13 KiB

MCP Best Practices, Market Reality, and Cached Knowledge Rationale

Executive Summary

MCP (Model Context Protocol) has rapidly evolved from Anthropic's November 2024 release to become the de-facto standard for AI-to-tool integration, with adoption by OpenAI in March 2025 and donation to the Linux Foundation's Agentic AI Foundation in December 2025. However, its effectiveness for document-centric use cases depends heavily on implementation patterns. This analysis addresses whether MCP with cached/pre-computed knowledge layers makes sense versus alternatives.


1. Market Best Practices and Real Examples

1.1 Production MCP Document Implementations

Implementation Approach Key Features
Chroma MCP Server Vector-native semantic search Sets the standard for semantic document management using vector search, supporting both ephemeral and persistent storage
Knowledge-Base-MCP (Puran) Production-grade RAG Agent-directed hybrid retrieval with auto mode choosing among dense, hybrid, sparse, and rerank routes; when scores are low, returns abstain so client can decide whether to retry
AWS Bedrock KB + MCP Enterprise knowledge bases Connects to Amazon Bedrock Knowledge Bases for semantic search capabilities, providing unified access pattern regardless of underlying AWS service
Basic Memory Local-first knowledge graphs Local-first knowledge management system that builds a semantic graph from Markdown files, enabling persistent memory across conversations
AtScale MCP Semantic layer for BI Exposes semantic models to any MCP-compatible AI agent with real-time model discovery where new models become queryable instantly after deployment

1.2 Dominant Architecture Patterns

Pattern A: MCP + Vector Store (Most Common)

Documents → Chunking → Embeddings → Vector DB → MCP Server → LLM Client

Prepare the knowledge base by collecting and preprocessing data, chunk into reasonably sized pieces, embed using an embedding model, and load into a vector database like FAISS, Weaviate, or Pinecone

Pattern B: MCP + Knowledge Graph

Documents → Entity Extraction → Graph DB → MCP Server → LLM Client

The MCP server exposes Graphiti's core capabilities including episode management, entity management, search capabilities with semantic and hybrid search for facts and node summaries

Pattern C: MCP + Hybrid (Best Practice)

Documents → [Vectors + Graph + BM25] → Unified MCP Interface → LLM Client

Combines vector search plus BM25 lexical search using RRF, then reranks. Best for complex queries with both conceptual and specific keyword requirements


2. When MCP Beats Alternatives

2.1 MCP vs. Direct Document Attachment

Scenario Winner Rationale
One-off Q&A on small docs Direct Attachment Zero setup, full context visible
Repeated queries on same corpus MCP Avoid re-processing, selective retrieval
Multi-user access MCP AI applications can seamlessly access up-to-date information and context as needed through unified protocol
Agentic workflows MCP Enables "agentic" AI systems that can autonomously interact with multiple systems, retrieve the latest information, and even take actions
Version-controlled knowledge MCP Git-based updates, deterministic ingestion

2.2 MCP vs. Traditional RAG API

Aspect Traditional RAG MCP-wrapped RAG Advantage
Standardization Custom per-service Universal protocol MCP
Tool Discovery Manual documentation Auto-discovery MCP
Multi-source N×M integrations M+N integrations MCP flips the N×M problem to an M+N model: tool providers implement one standard MCP server, and AI app developers implement MCP client support once
Agent autonomy Fixed pipeline LLM itself makes contextual decisions about how to interact with the data, determining query strategy and prompt formulation MCP

2.3 When NOT to Use MCP

  1. Simple, one-time document analysis - Direct attachment wins
  2. Highly dynamic real-time data - Direct API calls may be simpler
  3. No multi-client requirement - Overhead not justified
  4. Prototype/exploratory phase - Start simple, add MCP later

3. Cached/Pre-computed Knowledge: The Rationale

3.1 What is "Cached Knowledge" in MCP Context?

Three tiers of pre-computation that MCP servers can provide:

Tier What's Cached When Generated Example
1. Embeddings Vector representations At ingestion Semantic search index
2. Summaries Condensed content At ingestion or scheduled "This document describes X"
3. Knowledge Graph Entity-relationship extractions At ingestion "Goal M1 → implemented by → System X"

3.2 Rationale FOR Cached Knowledge Layers

A. Token Economics

Simply stuffing all potentially relevant data into the prompt is inefficient and sometimes impossible. MCP enables dynamically retrieving just-in-time context from external sources as needed instead of front-loading everything

Cost comparison (200-page document):

  • Direct attachment: ~100K tokens every query = $0.30-3.00/query
  • MCP with embeddings: ~2K tokens retrieved = $0.006-0.06/query
  • MCP with summaries: ~500 tokens = $0.0015-0.015/query

B. Response Quality

Tune the number of retrieved documents included in the prompt - often 3-5 good snippets are better than 10 - sometimes using too many can overwhelm the model

Pre-computed summaries ensure the LLM gets:

  • Condensed, relevant context
  • Pre-extracted key facts
  • Relationship context from knowledge graphs

C. Caching Benefits

Implement caching at multiple levels. Cache the results of common retrieval queries — for example, if many users ask "What is the refund policy?", you can cache the answer or at least the retrieved document so the agent doesn't vector-search the same question repeatedly

Implementing advanced caching (exact, semantic, task-aware) to avoid redundant API calls, tracking and optimizing costs across providers

D. Offline/Latency Benefits

Pre-computed knowledge enables:

  • Faster response times (no embedding at query time)
  • Offline capability (no API calls for embeddings)
  • Deterministic behavior (same query = same retrieval)

3.3 Implementation: Hierarchical Memory

Provides hierarchical memory storage with three-tier compression (chunks, micro-summaries, meta-summaries)

Example architecture:

# Pre-computed knowledge layers
raw_chunks:
  - content: "Full text chunk"
  - embedding: [0.1, 0.2, ...]
  
micro_summaries:
  - chunk_ids: [1, 2, 3]
  - summary: "These chunks describe the cultural heritage preservation goals"
  - keywords: ["heritage", "preservation", "M1"]

meta_summaries:
  - scope: "02-goals subdomain"
  - summary: "Six strategic goals (M1-M6) covering preservation through cybersecurity"
  - entity_count: 6

3.4 Rationale AGAINST Over-caching

Risk Description Mitigation
Staleness Summaries out of sync with source Git-triggered regeneration
Loss of nuance Summarization loses detail Keep raw chunks accessible
Hallucination risk LLM-generated summaries may be wrong Human review for critical content
Storage cost Multiple representations Tiered storage, compress cold data

4.1 Current State Analysis

Your Architecture-as-Code repo has:

  • ✅ Structured YAML entities (good for knowledge graph)
  • ✅ Explicit relationships in edges.yaml
  • ❌ No vector embeddings
  • ❌ No pre-computed summaries
  • ❌ Primitive substring search (not semantic)

4.2 Proposed Enhanced Architecture

┌─────────────────────────────────────────────────────────────────┐
│                    KISC Architecture MCP Server                  │
├─────────────────────────────────────────────────────────────────┤
│  Layer 1: Raw Data                                              │
│  ├── YAML entities (goals, systems, services, etc.)             │
│  └── Markdown documentation                                      │
├─────────────────────────────────────────────────────────────────┤
│  Layer 2: Pre-computed Knowledge (NEW)                          │
│  ├── embeddings.index (vector search via FAISS/Qdrant)          │
│  ├── summaries.yaml (entity-level summaries in Latvian)         │
│  ├── knowledge_graph.json (Neo4j-style graph export)            │
│  └── glossary.yaml (term definitions for natural language)      │
├─────────────────────────────────────────────────────────────────┤
│  Layer 3: MCP Tools (ENHANCED)                                  │
│  ├── semantic_search(query) → vector similarity                 │
│  ├── get_entity(id) → full entity + related context             │
│  ├── explain_concept(term) → natural language explanation       │
│  ├── find_relationships(entity) → graph traversal               │
│  ├── summarize_domain(domain) → pre-computed summary            │
│  └── natural_query(latvian_question) → LLM-friendly response    │
└─────────────────────────────────────────────────────────────────┘

4.3 Pre-computation Pipeline

# On every git commit to Architecture-as-Code:
1. Load all YAML entities
2. Generate embeddings for each entity (title + description)
3. Generate micro-summaries for each subdomain
4. Build/update knowledge graph from edges.yaml
5. Create glossary from all entity titles/codes
6. Store in /mcp/cache/ directory

4.4 Query Flow (Enhanced)

User asks: "Kādas sistēmas realizē valodas tehnoloģiju mērķi?" (What systems implement the language technology goal?)

Current behavior: Returns [] (no match for Latvian natural language)

Enhanced behavior:

  1. natural_query tool receives Latvian question
  2. Extracts intent: "systems implementing language technology goal"
  3. Maps to goal.m4 (Latviešu valoda digitālajā laikmetā)
  4. Traverses knowledge graph: goal.m4 → implemented_by → [sys.valoda.01, sys.valoda.02, ...]
  5. Returns pre-computed summary + entity list

5. Decision Framework: When to Add Cached Knowledge

Question If YES If NO
Will multiple users query the same corpus? Add caching Skip
Is query latency critical (<1s)? Add embeddings Direct retrieval OK
Do users ask in natural language (not IDs)? Add semantic search ID-based lookup OK
Is the corpus >100 entities? Add summaries Full scan OK
Do queries require cross-entity reasoning? Add knowledge graph Flat search OK
Is the corpus updated less than daily? Pre-compute aggressively Real-time generation

6. Conclusion

Is MCP with Cached Knowledge Worth It?

YES, when:

  • You have a stable, structured knowledge base (like Architecture-as-Code)
  • Multiple consumers need consistent access
  • Natural language queries are required
  • Token cost optimization matters
  • Cross-entity reasoning is needed

NO, when:

  • One-time document analysis
  • Rapidly changing data (real-time feeds)
  • Simple keyword lookup suffices
  • No multi-client requirement

For KISC Specifically:

Your Architecture-as-Code approach is fundamentally correct but needs:

  1. Semantic search layer (embeddings for Latvian content)
  2. Pre-computed summaries (domain/subdomain level)
  3. Natural language interface (Latvian query handling)
  4. Enhanced MCP tools (beyond primitive substring search)

The investment in these layers will pay off as the architecture grows and more stakeholders (internal teams, external auditors, automated agents) need to query it.


References