Skip to content
ansezz.
← Back to blog
Architecture May 22, 2026 8 min read 1,450 words

Message queues for heavy-duty document processing

Stop running embeddings in the request cycle. Build a document pipeline on message queues with staged workers, retries, dead-letter queues, and autoscaling.

Anass Ez-zouaine

Backend · Architect · AI

▸ Share

Document processing pipeline fed by a message queue with multiple worker stages
▸ On this page (6)

If your embeddings run inside the request cycle, the first 50MB PDF will take your app down. Put that work behind a queue.

If you run document embeddings inside your request-response cycle instead of behind a message queue, you are playing with fire.

I have seen many developers build a beautiful RAG application that falls over the second a user uploads a 50MB PDF. The browser spins, the proxy timeout hits and the database locks up while your worker tries to chunk 500 pages of legal jargon in real time.

The heavy lifting problem

This is the classic “heavy lifting” problem in AI engineering. Document processing (OCR, text extraction, semantic chunking and embedding) is slow, unpredictable and resource-heavy.

Forcing it into a synchronous web request gives you a bad user experience and a fragile system.

Decouple it

The fix is decoupling with message queues. Here is why async work belongs in a queue and how to build an ingestion pipeline that doesn’t melt your server.

If you want the broker-level foundations first, I cover them in scaling with RabbitMQ.

The synchronous trap

Imagine a user uploads a document to your SaaS. Your code receives the file, sends it to an extraction API and waits.

Then it loops through the text to create chunks, sends each chunk to an embedding model and finally saves it all to pgvector.

Where it breaks

If the whole chain runs long, you hit a timeout somewhere: a proxy gateway timeout (Nginx defaults to 60s), a PHP max_execution_time or a platform request cap. The connection drops.

If the embedding API has a momentary blip, the whole process fails and the user has to start over. While your server is busy with this heavy work, it isn’t responding to other users.

The 100ms rule

A good rule of thumb: if it takes more than 100ms, consider making it async.

Moving this work to a message queue gives your users immediate feedback (“we’re processing your file!”) while the heavy lifting happens safely in the background.

In the request

Upload, extract, chunk, embed and save in one request. A timeout or API blip throws away all the work, and other users wait.

Behind a queue

Upload, save to S3, push a small message and return right away. Workers handle each stage, retry failures and report progress.

The anatomy of a document pipeline

A solid RAG pipeline is a series of decoupled stages, not one big function. I break it into modular steps, each triggered by a message in a queue, so you can scale each part on its own.

Four stages

Here is how I usually structure it:

  1. Ingestion and discovery: a user uploads a file. You save it to S3 and push a small message to the queue with the file_path and tenant_id.
  2. Parsing and normalization: a worker picks up the message, downloads the file and runs it through a parser like pdfplumber or an OCR service. It sends the raw text to the next queue.
  3. Chunking: this worker splits the text into semantic sections. Keeping it in its own stage means you can swap chunking strategies (for example recursive character vs semantic) without re-running the heavy parsing step.
  4. Embedding and indexing: the final stage batches the chunks, calls your embedding API (like OpenAI or a local model) and pushes the vectors into your vector DB.

Backpressure for free

This stage-based approach is what I discuss in my post on 7 RAG mistakes to avoid in production.

It gives you backpressure control. If your vector DB slows down, the “index” queue grows, but the “parsing” workers keep going.

When one stage slows down, its queue grows. Every other stage keeps working.
Comic panel of four robots on a conveyor line unpacking papers, reading pages, cutting text into strips and turning strips into glowing blue cubes
Each pipeline stage works at its own pace with a queue in between.

Retries and dead letters

In the real world, things break. APIs time out. PDFs are malformed. Workers crash.

Opt in to retries

A message queue like Redis (with BullMQ or Laravel Queues) or SQS gives you retries with almost no code, but you do have to opt in:

  • BullMQ defaults to a single attempt, so set attempts and a backoff strategy explicitly.
  • Laravel queues take --tries and backoff.
  • SQS retries up to maxReceiveCount.

When a worker fails, the message goes back onto the queue to be tried again after a delay. Use exponential backoff instead of hammering a failing API every 5 seconds: wait 10, then 60, then 300.

Quarantine what can’t succeed

Some documents simply cannot be processed. Maybe it’s a password-protected PDF or a corrupted file, and you don’t want it retrying forever and clogging your workers.

This is where a dead-letter queue (DLQ) comes in. After a set number of failed attempts, the message moves to the DLQ, a “quarantine” zone.

SQS gives you this natively through the queue’s redrive policy. With BullMQ or Laravel, you route exhausted jobs there yourself (listen for the failed event in BullMQ, or use failed_jobs plus a handler in Laravel). You can then inspect the failed jobs, fix the issue and re-queue them. My PHP unserialize tool decodes a failed_jobs payload so you can see exactly what the job received.

Batching for efficiency

If you process 10,000 chunks, you don’t want 10,000 separate API calls to your embedding provider. That is slow and expensive.

Most embedding APIs and vector databases work much better with batches. A good worker pulls several messages from the queue (or gathers them in memory) and sends them as one bulk request.

Show progress

In Laravel, I often use job batching to track the progress of a large document. I can see exactly when 95% of a PDF is processed and update a progress bar for the user.

For how this fits a larger architecture, see my thoughts on event-driven pub/sub systems.

Event-driven prefetching

Queues help beyond ingestion. You can also use them for prefetching.

If a user is chatting with an AI agent and the conversation is heading toward a topic, you can fire a background job to fetch related documents and warm the cache before they ask the next question. The AI feels very fast because the context is already “ready” when retrieval runs.

Keep the chat app simple

An event bus decouples the chat interface from these optimization tasks. The chat app just emits a user_asked_question event, and a background worker decides whether to prefetch more data or update the semantic cache.

Monitoring your message queue

Once you move to a queue-based system, request latency stops being the only number that matters. You need to watch your queue depth.

If the queue grows faster than your workers can clear it, you have a bottleneck.

Add workers when it grows

Tools like Docker and Coolify make this easy. I can spin up five more worker containers to handle a sudden surge in uploads.

You can read more about how I manage this infra in my Coolify and Docker guide.

Comic panel of a robot pointing at a pressure gauge over a tall stack of envelopes while a squad of robots charges in to help
When queue depth climbs, add workers.

Key takeaways

  • Never store large files in the queue. Pass references (like an S3 key) and keep messages small.
  • Make tasks idempotent. Assume a message might be processed twice, and use upsert instead of insert in your vector DB.
  • Use structured logging. Every worker log should include the doc_id and tenant_id, or “why did this file fail?” has no answer.
  • Scale on queue depth. Add workers based on how many messages are waiting, not CPU alone.
  • Separate worker pools. Keep “fast” tasks (like metadata updates) apart from “slow” ones (like OCR and embedding), so a huge PDF never blocks a simple name change.

A document pipeline is about respecting the time it takes to process data. If big uploads are timing out your RAG app, here’s how I help teams build queue-based ingestion pipelines with retries and autoscaling, or drop a note via contact.

How are you handling long-running AI tasks today: still fighting request timeouts, or already using a queue?

▸ Made it to the end? Send it around.

▸ Share

▸ Comments