Standard PHP rebuilds your whole Laravel app for every request. Octane builds it once and keeps it warm.
Standard PHP is a bit like a restaurant that fires its entire staff and rebuilds the kitchen for every customer order. You walk in, they hire a chef, buy a stove, cook your meal and then demolish the building as soon as you leave.
It’s consistent and safe, but it wastes a lot of energy when you’re trying to serve thousands of people at once.
The PHP-FPM lifecycle
Every request boots the entire Laravel framework, loads your service providers, parses your config and instantiates your objects. For low-traffic sites, that’s fine.
At real scale, say a busy multi-tenant SaaS, those milliseconds of “boot time” become a wall you can’t climb without throwing expensive hardware at it. Horizontal scaling buys you room, but it doesn’t fix the underlying latency.
Stop rebuilding the kitchen
Laravel Octane boots your application once, keeps it in memory and feeds requests to it through a fast worker pool. PHP goes from “short-lived script” to “long-lived process.”
Why your app feels heavy without Laravel Octane
The overhead of traditional PHP costs you efficiency as well as speed. In a standard request, your CPU spends a big chunk of time just getting the application ready to work.
Once it finally starts, it runs the database query, renders the view and then dies.
Paying for setup again and again
If you run a complex Laravel monolith with dozens of packages and custom service providers, your boot time might be 30ms to 50ms before a single line of business logic runs.
Under high traffic, this leads to CPU thrashing. You’re paying for the “setup” over and over again.
Warm workers
Octane removes this boot cycle. With application servers like Swoole or RoadRunner, your app stays resident in memory.
The first request boots the framework, and later requests hit a “warm” application. You can move from 50ms responses to sub-10ms responses just by changing how the process is managed.
PHP-FPM
Octane
Choosing your application server
Octane runs on top of an application server, and it supports four: FrankenPHP, Swoole, Open Swoole and RoadRunner.
FrankenPHP (a Go-based server built on Caddy) is the first option in the Octane docs, and it’s the easiest path to HTTP/2, HTTP/3 and automatic HTTPS. In production, the two I reach for most are Swoole and RoadRunner, and they solve the “persistent state” problem differently.
Swoole
Swoole is a C extension for PHP. It’s a fast networking engine that lets PHP handle asynchronous tasks, coroutines and long-lived connections.
- Pros: it is very fast. Because it lives as an extension, it has deep access to PHP’s internals. It also enables the Octane cache, an in-memory store (backed by Swoole tables) that the docs clock at up to 2 million reads/writes per second, plus concurrent tasks, ticks and intervals.
- Cons: it can be a bit of a nightmare to install and debug. As a binary extension, you have to compile it or find the right package for your OS. Xdebug doesn’t always play nice with it, and it can be picky about your environment.
RoadRunner
RoadRunner is written in Go. It acts as a load balancer and process manager that talks to your PHP workers over a fast binary protocol (Goridge).
- Pros: no extensions required. It’s a single binary you drop into your project. It’s much easier to set up in a Docker container and feels more “cloud-native.”
- Cons: it’s slightly slower than Swoole because of the communication overhead between the Go binary and the PHP processes, though for most web apps, where time goes to the database and other I/O, you won’t notice the difference.
Which one to pick
| Server | Install effort | Pick it when |
|---|---|---|
| FrankenPHP | Smoothest setup | You’re starting out and want HTTP/3 for free |
| RoadRunner | Single Go binary | You want no PHP extension to manage |
| Swoole | C extension, fiddly | You need every millisecond, task workers or Octane cache |
For most developers starting out, I’d pick FrankenPHP at the install prompt and not overthink it. Reach for Swoole only when you’re chasing every last millisecond or need its concurrency features like async task workers, ticks and the Octane cache.
Tuning your worker pool
The secret sauce of Octane is the “worker pool.” Instead of one process per request, you have a fixed number of workers waiting for incoming traffic.
If you misconfigure it, you’ll either leave performance on the table or crash your server. My server capacity calculator turns requests per second and latency into the number of workers you need.
CPU-bound vs I/O-bound
The right pool size depends on whether your app is CPU-bound or I/O-bound:
- CPU-bound apps: if you do heavy data processing or image manipulation, set your worker count to your number of CPU cores. More workers just cause context switching and slow things down.
- I/O-bound apps: most web apps spend 90% of their time waiting for a database, Redis or an external API. Here you can scale workers to 2x or even 4x your core count, so one worker can wait on the DB while another handles a new request.
# Starting Octane with 16 workers for an I/O-heavy app
php artisan octane:start --server=swoole --workers=16
Task workers
Don’t forget task workers if you’re using Swoole. Sized separately with the --task-workers flag, they run in their own processes.
They’re a good fit for fanning out concurrent queries without blocking the main request cycle.
The danger of persistent state
The main hurdle when moving to Octane is the shift in mindset. In traditional PHP, “leaky” code barely matters because the process dies after 100ms.
What to watch
If a service provider has a static array that you append to on every request, that array grows until your server runs out of RAM. Be careful with:
- Static properties: don’t use them for request-specific data.
- Singletons: a singleton you register in the app container stays alive. If it caches data, make sure that data is cleared or managed between requests.
- Global state: avoid
globalvariables at all costs (which you should be doing anyway).
Laravel “resets” some core services between requests, but it can’t catch everything. If you’re migrating an old codebase from monolith to microservices, audit your service providers for any long-lived state.
Protecting yourself with max-requests
The best insurance against memory leaks is the --max-requests flag. It tells Octane to gracefully restart a worker after it has handled a set number of requests.
Octane already does this every 500 requests by default. Tune the number to match how fast your workers accumulate memory.
# Restart workers every 1000 requests to prevent memory bloat
php artisan octane:start --max-requests=1000
This keeps memory use predictable while still giving you the speed of a warm application.
Real-world optimization tips
Once Octane is running, a few technical levers can squeeze out more performance.
- Database connections: in Octane your workers stay alive, so each one holds its own DB connection open between requests instead of reconnecting every time. That’s a win, but watch the math: more workers means more open connections. Check your
database.phpconfig and make sure the worker count doesn’t blow past your DB server’s max connection limit. This matters even more once you add read replicas. - Octane cache: if you’re using Swoole, reach for
Cache::store('octane'). It’s an in-memory Swoole table that’s very fast and shared across every worker on the box. Use it for frequently accessed configuration or small datasets that rarely change. Remember it’s wiped when the server restarts. - Bytecode caching: make sure OPcache is enabled and tuned. Set
opcache.validate_timestamps=0in production since your code won’t change while the server is running. - Graceful reloads: when you deploy new code, run
php artisan octane:reload. It gracefully restarts the workers without dropping current connections, which you need for zero-downtime deployments.
Key takeaways
- Identify the bottleneck: only use Octane if boot time is the problem. If your database queries take 2 seconds, Octane won’t help.
- Test for leaks: watch how your app behaves under sustained load and profile memory growth over time.
- Monitor workers: keep an eye on CPU and RAM to find the right worker count.
- Use concurrency: on Swoole,
Octane::concurrently()runs independent operations in parallel via task workers and returns their results together, cutting total response time for complex pages.
Octane lets PHP compete with Node.js and Go for high-concurrency apps while keeping the Laravel developer experience. If you’re moving a busy Laravel app to Octane, here’s how I help teams tune workers, connections and deploys for high traffic, or drop a note via contact.
Are you running Octane in production yet, or is the fear of memory leaks keeping you on PHP-FPM?