Skip to content
ansezz.
← Back to blog
AI Oct 4, 2026 9 min read 1,778 words

AI writes the code. Your job is the gates.

AI made code cheap. The engineer's job is the gate stack bad code cannot pass: tests, mutation testing, PHPStan, required CI checks, monitoring and rollback.

Anass Ez-zouaine

Backend · Architect · AI

▸ Share

Comic pop-art panel of a code cannon firing crates down a conveyor through checkpoint gates while a robot at a control panel bins a cracked crate
▸ On this page (8)

An AI can write 10,000 lines of code in a minute. I don’t care. I care whether those 10,000 lines survive my constraints.

Code generation stopped being the slow part of my week. An agent can scaffold a Laravel module, its migrations and its tests before my coffee cools.

That speed is real, and I use it every day. It also changed the question. “Can we write this?” is cheap now. “Should this reach production?” is the expensive one.

So I stopped judging my work by how much code I write. I judge the system around the code: what it refuses, how fast it refuses it, and how fast it recovers when something slips through.

Cheap code moves the job

When code was slow to write, the writing itself filtered out bad ideas. You thought twice before typing 400 lines by hand.

That filter is gone. Volume goes up, and every weak spot in your process gets hit more often.

I covered the trade-off in AI vs traditional development and the loop itself in agentic workflows. This post is the next step: the concrete gates I put between generated code and real users.

Think of them as a row of doors. Each one is cheap to run, catches one kind of failure, and says no without caring who wrote the code. An agent, a teammate, or me at 1am: same doors.

Engineer as typist

Writes most lines by hand. Quality depends on how careful you were that day. Review catches what it can.

Engineer as gatekeeper

Designs the checks every change must pass. Quality depends on the gates, and they run the same way on every push.

Gate 1: static analysis and types

The cheapest gate runs before any test. PHPStan reads your code and finds whole classes of bugs without executing it. Larastan teaches it about Laravel: Eloquent models, facades and container bindings.

PHPStan has levels from 0 to 10. Level 10 arrived with PHPStan 2.0 and treats every mixed type strictly, including the implicit ones you get from missing types.

On an existing app you won’t jump to level 10 on day one, and you don’t need to.

Generate a baseline first, then add it to includes. Old errors are recorded and new code has to meet the higher bar.

vendor/bin/phpstan analyse --generate-baseline
# phpstan.neon
includes:
    - vendor/larastan/larastan/extension.neon
    - phpstan-baseline.neon

parameters:
    paths:
        - app/
    level: 8

Then ratchet. Every few weeks, clear part of the baseline, and raise the level by one when it gets small. That file is a to-do list that can only get shorter.

Agents are happy here. Point one at a failing PHPStan run and it will fix types all afternoon without complaint. That’s exactly the work I want to hand over.

Gate 2: tests that describe behavior

Unit tests pin down the rules: pricing, proration, state changes, the logic that costs money when it’s wrong. Integration tests (Laravel calls them feature tests) prove the pieces work together: the route, the policy, the database and the queued job.

I lean on feature tests for AI-written code. An agent can make every unit test pass and still wire the controller to the wrong policy. A feature test that hits the real route as a real user catches that.

Pest also gives you architecture tests, which turn “we don’t do that here” into a failing build:

// tests/ArchTest.php
arch()->preset()->laravel();
arch()->preset()->security();

arch('no debug calls ship')
    ->expect('App')
    ->not->toUse(['dd', 'dump']);

On an older app, start with ->ignoring() for the folders you haven’t cleaned up yet.

When I write the spec first, I write it as tests. The agent gets a red suite and a clear finish line. If it can’t turn the suite green honestly, I usually learn something about my spec.

Gate 3: mutation testing

Coverage tells you a line ran. It doesn’t tell you a test would notice if that line were wrong. That gap is where generated tests like to hide, with assertions like “the result is not null”.

Mutation testing closes it. The tool changes your code on purpose (flips a > to >=, drops a method call, returns an empty array) and reruns your tests. If the tests still pass, that mutant escaped, and you just found a test that checks nothing.

Comic pop-art panel of a sneaky robot tampering with a lever inside a machine while a test robot with a red siren points and catches it
Mutation testing breaks your code on purpose to see if a test notices.

Pick the tool for your stack:

StackToolFail the build with
PHP on PestPest’s built-in --mutate--min=70
PHP on PHPUnitInfection--min-msi=70
JavaScript or TypeScriptStrykerJSthresholds.break

One detail to know before you wire this up: Infection’s Pest support covered Pest 1 only. From Pest 2 on, Pest runs its own mutation engine, so a Pest suite uses pest --mutate, and it mutates the classes your tests name with covers() or mutates().

On a PHPUnit suite, Infection is the usual choice. Its --git-diff-lines flag mutates only the lines a pull request touched, which keeps CI fast on a big codebase.

For JavaScript, Stryker exits with code 1 when the score drops below break:

{
  "thresholds": { "high": 80, "low": 60, "break": 60 }
}

The 70 and 60 in these examples are starting points. Pick the score your suite hits today, gate on it, and raise it as the tests improve. Why this matters so much for agent-written tests is in trust is not a QA strategy.

Gate 4: CI that nobody can skip

All of this is a suggestion until the merge button respects it. Here is the workflow I start with on a Laravel app that uses Pest:

name: quality-gates
on: [pull_request]

jobs:
  static:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
      - uses: shivammathur/setup-php@v2
        with:
          php-version: "8.4"
      - run: composer install --no-interaction --prefer-dist
      - run: vendor/bin/phpstan analyse --no-progress

  tests:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
      - uses: shivammathur/setup-php@v2
        with:
          php-version: "8.4"
      - run: composer install --no-interaction --prefer-dist
      - run: cp .env.example .env && php artisan key:generate
      - run: vendor/bin/pest --parallel

  mutation:
    needs: tests
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
        with:
          fetch-depth: 0
      - uses: shivammathur/setup-php@v2
        with:
          php-version: "8.4"
          coverage: pcov
      - run: composer install --no-interaction --prefer-dist
      - run: cp .env.example .env && php artisan key:generate
      - run: vendor/bin/pest --mutate --parallel --min=70
      # PHPUnit suite? The same gate with Infection:
      # vendor/bin/infection --git-diff-lines --git-diff-base=origin/main --min-msi=70 --threads=max

The mutation job only has work to do once your tests name the classes they cover with covers() (or mutates()), so add those as you write tests.

Then open branch protection (or a ruleset) for main and turn on “Require status checks to pass before merging”. Add all three jobs. Also tick “Require branches to be up to date before merging”, so each pull request is tested against the code it will actually land on.

Now nobody merges red. No agent, no teammate in a hurry, and no version of me on a deadline. The wider picture of the pipeline is in CI vs CD.

Gate 5: production is a gate too

Some bugs only show up with real data and real traffic. Your tests won’t catch the query that is fine with 200 rows and slow with two million. So the last gate sits after the deploy.

I want three things there:

  • Error tracking. Sentry’s Laravel SDK reports unhandled exceptions with their stack trace as they happen. I hear about the bug before the support ticket arrives.
  • A smoke check. Laravel 11 and later ship a /up health route that returns 200 when the app boots. Hit it right after deploy and fail the pipeline if it doesn’t answer.
  • A rollback you have practiced. Keep the previous build ready to switch back to. If rolling back needs a meeting, it’s too slow.
curl --fail --silent --max-time 10 https://app.example.com/up || ./rollback.sh

Monitoring is also how the gates improve. When something escapes to production, I fix it, then I ask which earlier gate should have caught it, and I add the test or the rule there. Every incident should leave the doors tighter. The difference between having logs and actually being told is in logging vs monitoring.

Comic pop-art panel of a robot pulling a big rollback lever as a cracked rocket gets winched back onto the launch pad
The last gate runs in production: spot the failure fast and roll back faster.

Code review, when it matters

With the machines doing the mechanical checks, review changes shape. I don’t read every line hunting for missing types anymore. The pipeline already did that, on every file, with the same attention.

My review time goes to the questions no tool answers. Should this feature exist? Does this new queue or table fit the system? What happens to the money if this webhook arrives twice?

GitHub can route that attention for you. A CODEOWNERS file plus “Require review from Code Owners” means risky paths wait for a human, and everything else merges on green:

# .github/CODEOWNERS
/database/migrations/  @your-org/leads
/app/Policies/         @your-org/leads
/app/Billing/          @your-org/leads

My full read list, and how I handle agent-sized diffs, is in stop reading every line of AI-generated code.

Generating code is the easy part now. The craft is in the system that decides what ships.

What moving up the stack looks like on a Tuesday

“Moving engineers up the stack” sounds like a conference slide. Here is what it means in my actual week:

  • Less boilerplate. Agents write the CRUD, the DTOs, the form requests and the first draft of the tests. I don’t miss typing them.
  • More specs. I write acceptance tests and invariants before the agent starts. A refund never exceeds the captured amount. A discount never drops a total below zero.
  • More pipeline work. I treat the CI config like product code. It gets reviewed, versioned and improved whenever an incident shows a gap.
  • More design. Schema choices, module boundaries, what goes on a queue. These decisions outlive any single file, and no gate can make them for you.
  • Fewer heroics. When the gates are good, a Friday deploy is boring. Boring is the goal.

The engineers who do well with AI can say exactly what “correct” means, then make a machine check it on every push.

Key takeaways

  • Cheap code moves the hard work to deciding what may merge and ship.
  • Order your gates from cheap to expensive: static analysis, tests, mutation testing, required CI checks, then monitoring and rollback.
  • Use a baseline and a ratchet so standards rise without a big rewrite.
  • Spend human review on design and risk, and route it with CODEOWNERS.
  • Every bug that escapes should add a rule to an earlier gate.

If you want an outside look at which of these gates your Laravel pipeline is missing, my architecture audit covers the code, the CI and the deploy path.

Which gate is missing from your pipeline right now?

▸ Made it to the end? Send it around.

▸ Share

▸ Comments