Back to All Case Studies
Full Stack EnterpriseRepresentative Architecture Project2024

Enterprise AI Knowledge Mesh & RAG Pipeline

Semantic Search Engine over 10M+ Enterprise Documents with Granular RBAC

Abin's RoleFull Stack Solution Architect
Industry / DomainLegalTech & Enterprise Knowledge Systems
Engagement ContextClient details and proprietary document collections are omitted under confidentiality. The technical architecture and retrieval benchmarks reflect generalized representative configurations with client permission.
Document Vectorized
10M+
PDFs, Docs, and Schemas
Search Relevance (mAP)
94.2%
Hybrid Dense/BM25 Ranker
Time-to-First-Token
320ms
Streaming SSE
Role Permission Checks
< 2ms
In-query vector metadata filtering
The Challenge

The Business & Technical Problem

Corporate legal and consulting teams spent over 3.5 hours every day manually searching across disjointed file shares, cloud drives, and internal databases to verify contract obligations and precedent clauses. Standard keyword search failed to understand legal synonyms and semantic intent, while off-the-shelf public AI models posed severe risks of confidential data leakage and hallucinated citations.

Execution Scope

Delivered Scope & Responsibilities

Architecture & System Design

Designed hybrid dense-sparse RAG architecture with sub-second streaming feedback and deterministic source citation verification.

Backend & API Development

Engineered asynchronous FastAPI microservices and Node.js ingestion background workers for parsing and embedding generation.

Frontend or Mobile Application

Built Next.js 15 App Router web interface with Server-Sent Events (SSE) streaming token parser and citation drawer.

Database Architecture & Persistence

Architected PostgreSQL schema with pgvector extension for 1536-dimensional embeddings with HNSW indexing.

Authentication & Permissions

Inlined Role-Based Access Control (RBAC) user group IDs directly into vector retrieval SQL queries for zero-leakage security.

Infrastructure & Deployment

Configured containerized Docker deployment on AWS with Redis query caching and automated health checks.

Monitoring & Maintenance

Implemented automated evaluation harness checking context retrieval precision and answer faithfulness.

Technical Rationale

Architecture Decisions, Rejected Alternatives & Accepted Trade-Offs

1

PostgreSQL with pgvector instead of third-party hosted vector databases

Why It Was Chosen:

Eliminated cross-network hops, preserved existing ACID relational transaction guarantees, and avoided recurring per-vector SaaS licensing fees.

Alternatives Rejected:

Pinecone and Weaviate were rejected due to data residency compliance hurdles and high ongoing monthly costs.

Trade-Offs Accepted:

Required careful tuning of PostgreSQL shared_buffers, work_mem, and HNSW maintenance parameters during bulk vector ingestion.

2

In-Query Vector Filtering at Retrieval Level instead of Post-Query Application Filtering

Why It Was Chosen:

Post-filtering application results discarded unauthorized hits after retrieval, resulting in pagination gaps and leaking aggregate counts.

Alternatives Rejected:

Application-layer role filtering after vector retrieval.

Trade-Offs Accepted:

Increased query planner complexity requiring composite indexes on vector embeddings and tenant/role JSON metadata.

3

Server-Sent Events (SSE) Streaming over Full-Duplex WebSockets

Why It Was Chosen:

Unidirectional HTTP streaming is lightweight, traversable through corporate enterprise firewalls, and naturally benefits from HTTP/2 multiplexing.

Alternatives Rejected:

Full-duplex WebSockets (unnecessary connection state overhead for query-response patterns).

Trade-Offs Accepted:

Required separate REST endpoints for user follow-up prompt submissions.

Verifiable Evidence

Measured Project Outcomes

Every metric specifies the measurement environment, method, and Abin's contribution. Unverified benchmarks are not presented as contractual SLAs.

Time-To-First-Token (TTFT)Benchmark
Baseline
3,800ms (buffered completion)
Verified Result
320ms
Method: Browser Performance API measuring initial SSE chunk arrival (Q3 2024)
Abin's Contribution: Engineered FastAPI token generator and Next.js client-side streaming hook.
Search Relevance Precision (mAP)Lab
Baseline
68.4% (standard keyword search)
Verified Result
94.2%
Method: Mean Average Precision evaluated across 250 curated domain query test suites (Q4 2024)
Abin's Contribution: Configured reciprocal rank fusion (RRF) combining dense OpenAI embeddings with sparse BM25 reranking.
Document Chunks VectorizedBenchmark
Baseline
0 chunks
Verified Result
10,000,000+ chunks
Method: PostgreSQL pgvector table row count telemetry (2024)
Abin's Contribution: Designed recursive paragraph-aware text chunker and batch ingestion worker pipeline.
Permission Enforcement LatencyBenchmark
Baseline
N/A
Verified Result
< 2ms overhead
Method: PostgreSQL EXPLAIN ANALYZE comparison of filtered vs unfiltered vector queries (Q4 2024)
Abin's Contribution: Designed compound indexing on tenant_id and role bitmasks.

Evidence & Confidentiality Notice

Client details and production screenshots are omitted under confidentiality. The technical description and benchmarks have been generalized with client permission.

Scope Limitations & Benchmark Boundaries
  • Reported retrieval metrics reflect tested legal and compliance corpus benchmarks; accuracy on highly unformatted scanned OCR documents requires specialized preprocessing.
  • Benchmarks do not constitute a contractual SLA for arbitrary document types without custom chunking and prompt engineering tuning.
Next Architectural Steps

Plan an Enterprise AI Knowledge System

Need private document intelligence, RAG pipelines, or streaming AI workflows for your company?

Technology Stack & Tools Used

Next.js 15PythonFastAPIpgvectorPostgreSQLOpenAI APIDockerRedis
Ready to Build Your Project?

Let's Build Your Custom Web Application

Whether you need a full-stack web app, a Node.js REST API, a Next.js SaaS MVP, an admin dashboard, or performance optimization for an existing application — I am available for freelance projects in Kerala, India, and worldwide.

Fast WhatsApp & Email ResponseRemote Worldwide & IndiaFixed-Scope or Milestone Rates