Skip to content
ansezz.

▸ Problem · Production RAG

RAG on Laravel that still works after the demo.

Demo RAG answers look fine on three happy PDFs. Production RAG fails on chunking, retrieval, cost, and silent regressions. This page is for teams who need Laravel + pgvector RAG that holds up under evals.

01

The problem

Most RAG failures are not "the model is dumb." They are bad chunks, weak retrieval, missing keyword fallback, and no eval gate before release.

A notebook pipeline does not become a product feature. You need ingest jobs, tenant isolation, caching, token budgets, and a way to catch quality drops when someone changes a prompt.

Laravel and PostgreSQL with pgvector are enough for many SaaS RAG paths if the architecture is honest about hybrid search and measurement.

02

Who this is for

  • Product teams adding search-and-answer to an existing Laravel SaaS.
  • Engineering leads who refuse to ship RAG without evals and cost limits.
  • Founders burned by a vendor demo who want an owned pgvector pipeline instead.

What you get

  • Ingest and retrieval pipeline on Laravel + PostgreSQL / pgvector
  • Hybrid search when pure vectors miss obvious keyword hits
  • Eval harness and cost guardrails your team can keep running
  • Tenant-aware document isolation when the SaaS is multi-tenant
  • Handoff docs and a path into the AI & MCP package for tool use next
03

How I approach it

▸ Problem-specific, not a generic playbook

  1. 01

    Fix retrieval before prompts

    Chunking, metadata filters, and hybrid (vector + keyword) search get attention first. Fancy prompts on bad retrieval just fail more confidently.

  2. 02

    Keep it in your stack

    pgvector on PostgreSQL, Laravel jobs for ingest, and Redis caching where it earns its keep. Fewer mystery services to debug at 2am.

  3. 03

    Eval before users

    A small golden set and regression checks run in CI or on a schedule. Quality drops block release the same way a failing test would.

  4. 04

    Guard the bill

    Token budgets, cache hits, and rate limits are part of the design so a viral query pattern cannot empty the API wallet overnight.

Related reading

04

Questions, answered.

Why Laravel and pgvector instead of a vector SaaS?

If your product already runs on Laravel and PostgreSQL, pgvector keeps documents, tenants, and retrieval next to the rest of the app. Fewer sync bugs. A hosted vector DB is still fine when the constraints say so.

Will this include an eval harness?

Yes. Shipping without evals is how demo quality dies quietly. You get a starter golden set and a way to run it on a schedule or in CI.

How long is the engagement?

A focused RAG path usually fits the 2-4 week AI integration sprint. Larger corpora or multi-collection setups get a longer proposal after discovery.

Can this sit beside MCP or Claude tool use?

Yes. Many teams start with retrieval, then expose tools over MCP. Same package family; we sequence so each piece earns its place.

Ready for RAG that survives evals?

Free discovery call. We pick one corpus and one answer surface, then measure before we scale.