Enterprise AI Knowledge Mesh & RAG Pipeline
Semantic Search Engine over 10M+ Enterprise Documents with Granular RBAC
The Business & Technical Problem
Corporate legal and consulting teams spent over 3.5 hours every day manually searching across disjointed file shares, cloud drives, and internal databases to verify contract obligations and precedent clauses. Standard keyword search failed to understand legal synonyms and semantic intent, while off-the-shelf public AI models posed severe risks of confidential data leakage and hallucinated citations.
Delivered Scope & Responsibilities
Designed hybrid dense-sparse RAG architecture with sub-second streaming feedback and deterministic source citation verification.
Engineered asynchronous FastAPI microservices and Node.js ingestion background workers for parsing and embedding generation.
Built Next.js 15 App Router web interface with Server-Sent Events (SSE) streaming token parser and citation drawer.
Architected PostgreSQL schema with pgvector extension for 1536-dimensional embeddings with HNSW indexing.
Inlined Role-Based Access Control (RBAC) user group IDs directly into vector retrieval SQL queries for zero-leakage security.
Configured containerized Docker deployment on AWS with Redis query caching and automated health checks.
Implemented automated evaluation harness checking context retrieval precision and answer faithfulness.
Architecture Decisions, Rejected Alternatives & Accepted Trade-Offs
PostgreSQL with pgvector instead of third-party hosted vector databases
Eliminated cross-network hops, preserved existing ACID relational transaction guarantees, and avoided recurring per-vector SaaS licensing fees.
Pinecone and Weaviate were rejected due to data residency compliance hurdles and high ongoing monthly costs.
Required careful tuning of PostgreSQL shared_buffers, work_mem, and HNSW maintenance parameters during bulk vector ingestion.
In-Query Vector Filtering at Retrieval Level instead of Post-Query Application Filtering
Post-filtering application results discarded unauthorized hits after retrieval, resulting in pagination gaps and leaking aggregate counts.
Application-layer role filtering after vector retrieval.
Increased query planner complexity requiring composite indexes on vector embeddings and tenant/role JSON metadata.
Server-Sent Events (SSE) Streaming over Full-Duplex WebSockets
Unidirectional HTTP streaming is lightweight, traversable through corporate enterprise firewalls, and naturally benefits from HTTP/2 multiplexing.
Full-duplex WebSockets (unnecessary connection state overhead for query-response patterns).
Required separate REST endpoints for user follow-up prompt submissions.
Measured Project Outcomes
Every metric specifies the measurement environment, method, and Abin's contribution. Unverified benchmarks are not presented as contractual SLAs.
| Metric Name | Baseline | Final Value | Environment | Measurement Method | Abin's Contribution |
|---|---|---|---|---|---|
| Time-To-First-Token (TTFT) | 3,800ms (buffered completion) | 320ms | Benchmark | Browser Performance API measuring initial SSE chunk arrival(Q3 2024) | Engineered FastAPI token generator and Next.js client-side streaming hook. |
| Search Relevance Precision (mAP) | 68.4% (standard keyword search) | 94.2% | Lab | Mean Average Precision evaluated across 250 curated domain query test suites(Q4 2024) | Configured reciprocal rank fusion (RRF) combining dense OpenAI embeddings with sparse BM25 reranking. |
| Document Chunks Vectorized | 0 chunks | 10,000,000+ chunks | Benchmark | PostgreSQL pgvector table row count telemetry(2024) | Designed recursive paragraph-aware text chunker and batch ingestion worker pipeline. |
| Permission Enforcement Latency | N/A | < 2ms overhead | Benchmark | PostgreSQL EXPLAIN ANALYZE comparison of filtered vs unfiltered vector queries(Q4 2024) | Designed compound indexing on tenant_id and role bitmasks. |
Evidence & Confidentiality Notice
Client details and production screenshots are omitted under confidentiality. The technical description and benchmarks have been generalized with client permission.
- Reported retrieval metrics reflect tested legal and compliance corpus benchmarks; accuracy on highly unformatted scanned OCR documents requires specialized preprocessing.
- Benchmarks do not constitute a contractual SLA for arbitrary document types without custom chunking and prompt engineering tuning.
Plan an Enterprise AI Knowledge System
Need private document intelligence, RAG pipelines, or streaming AI workflows for your company?
Technology Stack & Tools Used
Need Similar Architecture or Development for Your Team?
AI Integration & RAG Development Services
Production-ready AI integrations, Retrieval-Augmented Generation (RAG) search, document intelligence, and secure LLM workflows for modern web apps and SaaS.
Custom Web Application & Website Development
High-performance custom web application and website development with Next.js 15, React, Node.js, and PostgreSQL for modern businesses.
Let's Build Your Custom Web Application
Whether you need a full-stack web app, a Node.js REST API, a Next.js SaaS MVP, an admin dashboard, or performance optimization for an existing application — I am available for freelance projects in Kerala, India, and worldwide.