Towards AIblog

ContextFusion: The Context Brain Your LLM Apps Are Missing

Tuesday, September 1, 2026Rohan RView original
Author(s): Rohan R Originally published on Towards AI. A deep dive for users who want results and developers who want control TL;DR (For the Impatient) Normal users: Install context-portfolio-optimizer, run cpo compile ./your-docs --budget 4000, and stop overpaying for tokens. Developers: Middleware pipeline that ingests heterogeneous sources → normalizes → precomputes → optimizes via multi-objective knapsack → compiles provider-specific payloads with delta fusion for agents. Both groups get 60–99% token reduction with identical answer quality. Part 1: For Normal Users — “Just Make My LLM Cheaper and Faster” The Problem You Actually Face You’re building with LLMs. Maybe it’s a chatbot over your company docs. Maybe it’s a coding assistant. Maybe it’s an agent that needs to remember context across 20 turns. You keep hitting the same frustrations: “Why is this API call so expensive?” — You’re sending 8,000 tokens when 800 would suffice “Why does it take 10 seconds to respond?” — Latency scales with prompt size “Why does my agent forget everything?” — You’re not managing context deltas across turns “Why do I have to rewrite everything when I switch from GPT-4 to Claude?” — Hardcoded prompt formats You’ve tried RAG. You’ve tried chunking. But you’re still blindly stuffing retrieved chunks into prompts without knowing which ones actually matter. What ContextFusion Does (No Jargon) Think of it like a smart travel packer for your LLM trips. You have a weight limit (token budget). You have dozens of items (documents, code, images). Some items are essential. Some are nice-to-have. Some are duplicates. Some are risky (outdated, untrusted). ContextFusion: Unpacks everything — PDFs, Word docs, spreadsheets, images, code files Weighs and labels each item — How useful? How risky? How heavy? Packs the optimal suitcase — Maximum value within your weight limit Formats it for your destination — OpenAI’s preferred style, Anthropic’s format, or local Ollama And for return trips (agent conversations), it remembers what you already packed and only adds what’s new. Real Results Benchmarks run with Claude Sonnet 4.6 on production-like workloads. Full methodology at github.com/rotsl/context-fusion/benchmarks Getting Started (Three Options) Option A: NPM Wrapper (Easiest — No Python Required) # One-time setupnpm install -g @rotsl/contextfusionnpx @rotsl/contextfusion setup# Create API keys filenpx @rotsl/contextfusion env# Edit .env with your OPENAI_API_KEY or ANTHROPIC_API_KEY# Run optimizationnpx @rotsl/contextfusion run ./my-documents \ --query "Summarize key findings" \ --provider anthropic \ --model claude-sonnet-4-6 \ --budget 4000# Launch Web UInpx @rotsl/contextfusion ui --port 8080 Option B: Python Package (More Control) pip install context-portfolio-optimizer# Set up environmentcat > .env << 'EOF'ANTHROPIC_API_KEY=your_key_hereOPENAI_API_KEY=your_key_hereEOF# Run CLIcpo run ./my-documents --budget 4000 --query "What are the main points?"# Or compile for specific task typecpo compile ./my-codebase \ --task "Explain this function" \ --provider openai \ --model gpt-5-mini \ --mode code \ --budget 3000 Option C: Docker (Isolated, Reproducible) docker build -t context-fusion:latest .docker run --rm -it -v "$(pwd)":/app context-fusion:latest run ./data --budget 3000 The Web UI: See What Your LLM Actually Receives Run cpo ui --port 8080 and open your browser. You'll see: Run stats: Files ingested, blocks selected, total tokens Representation usage: Which compact variants were chosen Selected blocks: Source, representation type, utility score, token estimate Context preview: Exactly what gets sent to the LLM Model answer: Optional direct comparison This transparency is rare. Most RAG tools are black boxes. ContextFusion shows its work. Common Use Cases When ContextFusion Helps Most ✅ Multi-provider setups — Same pipeline, different output formats✅ Cost-sensitive production — 60–99% token reduction✅ Agent conversations — Delta fusion prevents token churn✅ Complex ingestion — PDFs, images, code, spreadsheets unified✅ Latency requirements — Precomputation + caching When You Might Not Need It ❌ Simple single-turn Q&A with tiny documents❌ You’re already heavily invested in a specific RAG framework and happy with costs❌ You need real-time streaming with sub-100ms latency (ContextFusion adds 50–200ms optimization overhead) Part 2: For Developers — “How This Actually Works” Architecture Overview ┌─────────────────────────────────────────────────────────────────┐│ INGESTION LAYER ││ PDF │ DOCX │ CSV │ JSON │ Images (OCR) │ Code │ Markdown │└─────────────────────────────────────────────────────────────────┘ ↓┌─────────────────────────────────────────────────────────────────┐│ NORMALIZATION LAYER ││ Convert all sources to uniform ContextBlock objects ││ - source_type, content_hash, created_at, metadata │└─────────────────────────────────────────────────────────────────┘ ↓┌─────────────────────────────────────────────────────────────────┐│ REPRESENTATION LAYER ││ Precompute compact variants per block: ││ - universal_summary (general purpose) ││ - qa_extractive (question-answering focused) ││ - code_signature (functions, classes, dependencies) ││ - agent_condensed (working memory format) │└─────────────────────────────────────────────────────────────────┘ ↓┌─────────────────────────────────────────────────────────────────┐│ PRECOMPUTE PIPELINE ││ Store: fingerprints, summaries, token stats, ││ retrieval features, compact variants in .cpo_cache/ │└─────────────────────────────────────────────────────────────────┘ ↓┌─────────────────────────────────────────────────────────────────┐│ RETRIEVAL LAYER ││ Query classification → Lexical retrieval (top-100) ││ → Fast rerank (top-20/25) → Candidate set │└─────────────────────────────────────────────────────────────────┘ ↓┌─────────────────────────────────────────────────────────────────┐│ MULTI-OBJECTIVE PLANNER (Core) ││ ││ maximize Σ( w_u·utility - w_r·risk - w_t·token_cost ││ - w_l·latency + w_c·cacheability + w_d·diversity ) ││ ││ subject to: Σ(token_i) ≤ budget ││ ││ Selects optimal representation variant per block │└─────────────────────────────────────────────────────────────────┘ ↓┌─────────────────────────────────────────────────────────────────┐│ COMPRESSION LAYER ││ - JSON minification ││ - Citation compaction (Source URI → [id]) ││ - Schema field pruning ││ Levels: none │ light │ medium │ aggressive │└─────────────────────────────────────────────────────────────────┘ ↓┌─────────────────────────────────────────────────────────────────┐│ DELTA FUSION (Agent Mode) ││ Compute ContextDelta: ││ - added_blocks: new since last turn ││ - updated_blocks: changed content ││ - removed_blocks: no longer relevant ││ - unchanged_block_ids: reuse from cache │└─────────────────────────────────────────────────────────────────┘ ↓┌─────────────────────────────────────────────────────────────────┐│ PROVIDER ADAPTER LAYER ││ Compile provider-specific payloads: ││ - openai: chat.completions format ││ - anthropic: messages with XML citations ││ - ollama: local API structure ││ - openai_compatible: generic wrapper │└─────────────────────────────────────────────────────────────────┘ ↓┌─────────────────────────────────────────────────────────────────┐│ CACHE-AWARE ASSEMBLY ││ Segment into: ││ - stable: system instructions, citation maps, cacheable blocks ││ - dynamic: volatile content, real-time data │└─────────────────────────────────────────────────────────────────┘ The Knapsack Formulation: Why This Isn’t Just “Smart Chunking” Most RAG tools use semantic similarity: embed query, embed chunks, return top-k. This fails when: Your budget is 4,000 tokens and you have 50 relevant chunks of 500 tokens each Some chunks are high-utility but high-risk (outdated documentation) Some chunks are cacheable, others must be fresh You need diversity (don’t send 5 versions of the same information) ContextFusion’s planner treats this as a constrained optimization problem: # Pseudocode of the core algorithmdef select_context_blocks(candidates, budget, weights): """ candidates: List[ContextBlock with multiple representation variants] budget: int (token limit) weights: dict[str, float] (utility, risk, latency, cacheability, diversity) […]