Skip to content
Quality Engineering•12 min read•Published September 13, 2026

Testing Strategy for Laravel SaaS: PHPUnit, Dusk, Load Tests

A passing CI pipeline and a production incident can coexist. Here's the Laravel testing architecture that closes that gap for real SaaS teams.

Aqib Javaid
Aqib Javaid
Senior Full-Stack Engineer
Dark tech blog cover for a Laravel SaaS testing strategy article, showing a terminal panel with PHPUnit, Dusk, and k6 results

Introduction: A Green Pipeline Is Not the Same Thing as a Stable Product#

Every Laravel project eventually gets a test suite. Fewer of them get a testing strategy. The difference shows up the first time a fully green CI pipeline ships a regression straight into production: a webhook handler that works fine in isolation but deadlocks under concurrent load, a real-time dashboard that renders correctly in a unit test but drops updates the moment two browser tabs are open at once, a tier-gated feature that passes every assertOk() because the test data happened to belong to the one tenant with unlimited seats. None of these are testing failures in the sense of "nobody wrote a test." They're architecture failures - the test pyramid was shaped wrong for what the product actually does.

I built out the testing infrastructure for On The Dot Global, a community platform, using PHPUnit for backend logic and Laravel Dusk for full browser-driven flows, backed by a structured approach to test data management specifically so that test accuracy didn't depend on whoever last ran the seeder. That work sits alongside testing problems I've had to solve differently on other products: SignageFlow, where the thing that breaks in production is timing - a screen not receiving a content push - which no assertDatabaseHas() will ever catch; MindWrite AI, where billing-gated feature access has to be tested without either mocking away the tier logic or burning real OpenAI credits on every CI run; and Expreco, where the riskiest code in the app is the code that talks to FedEx and UPS's APIs, neither of which you want your test suite calling live on every push.

This is a breakdown of how to structure a Laravel testing strategy for a real SaaS product - one with tenants, subscriptions, real-time updates, and third-party integrations - so that "tests pass" and "safe to deploy" actually mean the same thing.

Architecture: Match the Test Layer to the Way the Bug Actually Ships#

The mistake most Laravel test suites make is treating every code path the same way: a feature test that hits a route and asserts a status code and a database row. That's correct for a form submission. It tells you nothing about a websocket broadcast, nothing about what happens when two requests race for the same resource, and nothing about how the app behaves at 200 concurrent users instead of one. A SaaS product needs at least four distinct layers, each aimed at a different class of bug.

1. Unit and feature tests for business logic and access rules#

This is the layer most teams already have, and it should stay the largest by volume - fast, isolated, run on every commit. The part teams get wrong is testing access rules against convenient fixture data instead of data shaped like the actual failure mode. A tier-gating test that only ever uses a tenant on the top plan will never catch a bug in the downgrade path.

php
class ScreenLimitTest extends TestCase
{
    use RefreshDatabase;

    public function test_tenant_at_screen_limit_cannot_add_another_screen(): void
    {
        $tier = SubscriptionTier::factory()->create(['limits' => ['max_screens' => 5]]);
        $tenant = Tenant::factory()
            ->for(Subscription::factory()->for($tier, 'tier'))
            ->has(Screen::factory()->count(5))
            ->create();

        $response = $this->actingAs($tenant->owner)
            ->postJson("/tenants/{$tenant->id}/screens", ['name' => 'Lobby Display']);

        $response->assertStatus(422);
        $this->assertEquals(5, $tenant->screens()->count());
    }

    public function test_tenant_mid_downgrade_keeps_old_limit_until_period_end(): void
    {
        $tenant = Tenant::factory()
            ->for(Subscription::factory()->state([
                'tier_id' => SubscriptionTier::factory()->create(['limits' => ['max_screens' => 25]]),
                'pending_tier_id' => SubscriptionTier::factory()->create(['limits' => ['max_screens' => 10]]),
                'pending_tier_effective_at' => now()->addDays(12),
            ]), 'subscription')
            ->has(Screen::factory()->count(18))
            ->create();

        $response = $this->actingAs($tenant->owner)
            ->postJson("/tenants/{$tenant->id}/screens", ['name' => 'New Display']);

        $response->assertStatus(201); // still under the OLD limit of 25, downgrade hasn't landed yet
    } 
}

That second test only exists because a downgrade-in-progress is a real, recurring tenant state - not an edge case someone thought up in the abstract. The test data has to model the actual states the product produces, not just the happy-path ones.

2. Contract tests for every third-party integration, with the real API mocked at the boundary#

Expreco's core value is FedEx and UPS rate calculation. If the test suite calls those APIs live, it's slow, it's flaky on their infrastructure's schedule instead of yours, and it silently stops testing anything the day a sandbox credential expires. The fix isn't skipping the test - it's mocking exactly at the HTTP boundary, using response fixtures captured from the real API, so the test still exercises your actual parsing and error-handling code:

php
class FedExRateQuoteTest extends TestCase
{
    public function test_parses_fedex_rate_response_into_quote_object(): void
    {
        Http::fake([
            'apis.fedex.com/rate/v1/rates/quotes' => Http::response(
                json_decode(file_get_contents(base_path('tests/fixtures/fedex_rate_success.json')), true),
                200
            ),
        ]);

        $quote = app(FedExRateService::class)->getQuote($this->shipmentPayload());

        $this->assertEquals('12.50', $quote->amount);
        $this->assertEquals('USD', $quote->currency);
    }

    public function test_gracefully_handles_fedex_service_unavailable(): void
    {
        Http::fake(['apis.fedex.com/*' => Http::response(null, 503)]);

        $quote = app(FedExRateService::class)->getQuote($this->shipmentPayload());

        $this->assertNull($quote);
        Log::assertLogged('warning', fn ($message) => str_contains($message, 'FedEx rate service unavailable'));
    }
}

The same pattern covers MindWrite AI's OpenAI calls: Http::fake() against captured completion responses means the AI generation pipeline, credit deduction, and error handling are all tested on every push, with zero API spend and zero flakiness from a third-party outage.

3. Browser tests for the flows where timing and JavaScript state actually matter#

This is where PHPUnit structurally can't help you, and it's also the layer teams most often skip because it's slower to write and run. On SignageFlow, the feature that matters most - a screen picking up a content push in real time - can pass every backend test while being completely broken in the browser, because the bug is in how the frontend subscribes to the broadcast, not in whether the broadcast fired. Dusk is the layer that catches that:

php
class RealtimeContentPushTest extends DuskTestCase
{
    public function test_screen_receives_pushed_content_without_manual_refresh(): void
    {
        $screen = Screen::factory()->create();

        $this->browse(function (Browser $browser) use ($screen) {
            $browser->visit("/display/{$screen->uuid}")
                ->waitForText('Waiting for content');

            ContentPushService::push($screen, Content::factory()->create(['title' => 'Q3 Promo']));

            $browser->waitForTextIn('.display-title', 'Q3 Promo', 8)
                ->assertSee('Q3 Promo'); // proves the broadcast AND the subscription both work
        });
    } 
}

On On The Dot Global, this same layer covered the flows a form-submission test can't: multi-step signup, live-updating activity feeds, and the interactions between them - the reason a structured, reusable set of Dusk page objects and seeded test fixtures mattered more there than raw test count.

4. Load tests for the endpoints where correctness and capacity are the same question#

A webhook handler that's correct for one request and a webhook handler that's correct for fifty arriving in the same second are not automatically the same code. This is the layer a feature test suite has no way to express, and it's the one most SaaS teams add only after an incident proves they needed it.

javascript
// k6 load test - hits the webhook endpoint under realistic burst concurrency
import http from 'k6/http';
import { check } from 'k6';

export const options = {
  scenarios: {
    webhook_burst: { executor: 'ramping-vus', startVUs: 0,
      stages: [{ duration: '10s', target: 50 }, { duration: '30s', target: 50 }, { duration: '10s', target: 0 }] },
  },
};

export default function () { 
  const res = http.post('https://staging.example.com/webhooks/stripe', JSON.stringify(samplePayload()), { 
    headers: { 'Content-Type': 'application/json', 'Stripe-Signature': testSignature() },
  });
  check(res, {
    'status is 200': (r) => r.status === 200,
    'no duplicate processing': (r) => !r.body.includes('duplicate'),
  });
}

Run against a staging environment wired to the same idempotency table as production, this is the test that actually proves a webhook handler is safe under Stripe's real retry behavior, rather than just correct in the single-request case a feature test checks.

Step-by-Step: Wiring It Into a Pipeline That Blocks Bad Deploys#

  1. Split CI into stages by speed and cost, not by test type alone. Unit and feature tests (seconds, free) run on every push. Dusk (slower, needs a browser) runs on every PR before merge. Load tests (expensive, needs a real environment) run nightly against staging or as a manual gate before a major release - never on every commit, or the team will start ignoring the pipeline entirely because it's too slow to wait for.
yaml
jobs:
  fast-tests:
    steps:
      - run: php artisan test --parallel --testsuite=Unit,Feature
  browser-tests:
    needs: fast-tests
    steps:
      - run: php artisan dusk
  nightly-load-test:
    if: github.event.schedule
    steps:
      - run: k6 run load-tests/webhook-burst.js
  1. Seed test data that represents real tenant states, not just fresh accounts. A factory library that only ever produces a brand-new tenant on the default plan will never generate the downgrade-in-progress, near-limit, or past-due states where entitlement bugs actually live. Build factory states for every subscription and lifecycle status the product can actually be in, and use them deliberately in the tests that check access rules.
  1. Fake third-party HTTP calls at the transport layer, using fixtures captured from real responses. Http::fake() with a captured JSON fixture tests your parsing, retry, and error-handling logic on every run, with no dependency on FedEx, UPS, Stripe, or OpenAI actually being reachable, and no risk of API spend from a CI job.
  1. Give Dusk tests explicit waits tied to real signals, never a fixed sleep. waitForText() and waitForTextIn() wait for the actual DOM state the broadcast is supposed to produce; a hardcoded sleep(2) either wastes time when the update is fast or flakes out when it's briefly slow, and tells you nothing true either way.
  1. Run load tests against an environment wired to the same infrastructure as production - same queue driver, same database indexing, same idempotency store. A load test against SQLite or a locally-mocked queue proves nothing about how the real Redis-backed queue and MySQL indexes behave under concurrent writes; it has to point at a staging environment that's structurally identical to prod.
  1. Track flaky tests as a defect, not as background noise to re-run past. A Dusk test that fails one run in twenty because of unrelated timing is a signal that the underlying UI or broadcast has a real race condition - muting it with a retry annotation buries the exact bug this layer of testing exists to catch.

Pitfalls I've Seen Cost Real Time and Trust#

Testing access control against only the happiest-path tenant. A subscription-gating test suite that never constructs a downgrade-in-progress, past-due, or at-limit tenant will pass on every commit while that exact combination breaks in production the first week real customers hit it - because those states aren't hypothetical, they're the normal lifecycle of a paying account.

Mocking third-party APIs so loosely that the test stops testing anything. A blanket Http::fake() that returns ['success' => true] for every call proves the code runs, not that it parses a real FedEx or Stripe response correctly. Fixtures should be captured from actual API responses, error payloads included, or the contract test is theater.

Skipping browser tests because they're slower to write. The bugs that live in JavaScript subscription state, broadcast timing, and multi-tab interaction are invisible to PHPUnit by construction - no amount of backend test coverage substitutes for the one layer that actually renders a page and waits on it.

Adding load testing only after an incident, instead of before scale. SignageFlow and MindWrite AI both have endpoints - content push broadcasts and AI generation queues - where correctness under concurrency is the actual requirement, not a nice-to-have; treating load testing as a post-launch afterthought means the first real load test is a production outage.

Fixed sleep() calls in browser tests. They make a real-time feature's test suite either slower than it needs to be or randomly flaky depending on network jitter that day - waiting on the actual DOM signal is both faster and more honest about what's being verified.

No test-data isolation across tenants. In a multi-tenant product, a test suite that doesn't scope its factories and assertions per-tenant can pass while silently never testing the cross-tenant leakage bugs that are the single most damaging class of defect a SaaS product can ship.

Key Takeaways#

A test suite that's green on every commit and a product that's stable in production are related, not identical - the gap between them is exactly the set of bugs the wrong test layer can't see. The Laravel SaaS testing strategies that actually hold up share the same shape:

  • Unit and feature tests carry the volume, but only if their fixture data models real tenant lifecycle states, not just fresh accounts
  • Every third-party integration gets contract tests against captured real-response fixtures - never a live call, never a meaningless blanket mock
  • Browser tests cover the flows where JavaScript state and real-time timing are the actual behavior being verified, waited on by real signals, never fixed sleeps
  • Load tests run against production-equivalent infrastructure, targeting the specific endpoints where concurrency is part of correctness - webhooks, real-time push, queued AI work
  • CI is staged by cost - fast tests on every push, browser tests before merge, load tests nightly or pre-release - so speed doesn't force the team to start ignoring red pipelines

I've built this layered approach into a community platform's full regression suite, a real-time signage platform where timing bugs are the ones that matter, and AI and logistics products where the riskiest code is the code calling someone else's API. The stack changes; the discipline of matching the test layer to the way the bug actually ships doesn't.

If your test suite is green but your team still dreads deploy day, or you're building the testing architecture for a SaaS product from scratch, get in touch about your testing strategy or see the full case studies from platforms where these patterns are running today.

Share this technical insight with your network

Share to LinkedIn or Facebook with key takeaways, featured media, and direct links.

📁 Production Case Study

Case Study: Violerts - Enterprise NYC PropTech Compliance & Violation Monitoring SaaS

Violerts is a PropTech SaaS platform that consolidates fragmented NYC municipal property data into a single compliance intelligence platform. I led the modernization of the React frontend and Laravel backend, building multi-agency data ingestion, GIS mapping, asynchronous scraping, real-time alerts, team collaboration, and Stripe-powered SaaS billing.

Related Technical Articles

View all articles →
✦ Let's Build Together

Have a complex technical project in mind?

Available for full-stack engineering, performance audits, cloud deployments, and high-concurrency systems architecture.

Need a web or software development partner?

Tell me what you’re building, what’s getting in the way, and where you need help. Whether you need a custom web application, SaaS platform, API integration, or full-stack development, I’ll give you a clear answer on scope, cost, and timeline usually within one business day.

AqibJavaid

Senior Full-Stack Engineer building backend systems, cloud infrastructure and product platforms for teams that need them to stay up.

Available for new projects

Get in touch

© 2026 Aqib Javaid. All rights reserved.

Built and maintained by Aqib Javaid