Skip to content
Laravel & PHP•16 min read•Published September 23, 2026•Updated 9/27/2026

Building Observability into a Laravel SaaS Platform

A stack trace tells you what broke. It doesn't tell you which tenant, which request, or whether it's still happening right now.

Aqib Javaid
Aqib Javaid
Senior Full-Stack Engineer
Dark tech blog cover showing Laravel logs, metrics, and error traces converging into a unified observability dashboard

Introduction: Log::error($e->getMessage()) Is Not Observability#

Every Laravel app starts logging the same way: a try/catch block somewhere, a Log::error($e->getMessage()), and a storage/logs/laravel.log file nobody reads until something breaks. It feels like enough, because in local development it is — there's one request at a time, one tenant, one developer who already knows what they just changed. Then it ships, and the questions that actually matter in production are ones a bare error message can't answer. Which tenant hit this? Was it one request or four hundred? Did it start after last night's deploy, or has it been happening quietly for a week? Is the AI generation queue backed up because of one slow customer's prompt, or because the OpenAI API itself is degraded? A log line that says "Call to a member function on null" with no tenant ID, no request ID, and no context around it doesn't answer any of that — it just confirms that something, somewhere, broke.

That gap is invisible until the first real incident, and then it's the difference between a ten-minute fix and a three-hour investigation conducted half from memory and half from guessing. None of it is a logging feature problem — Laravel's logging, exception handling, and job system all give you plenty of hooks. It's an architecture problem: whether every log line, every metric, and every reported error carries enough structured context to answer "who, what, when, and how often" without someone having to reconstruct it by hand at 2am.

I've had to build this discipline into products where "we'll check the logs" isn't an acceptable incident response: SwapPad, the bulk SIM activation platform I built inside the CelleUp dealer back office, where a carrier API timeout mid-batch needs to reach an on-call Slack channel within seconds, tagged with exactly which batch and which carrier, not discovered when a dealer opens a support ticket the next morning; SafetySpace, the AI-driven safety platform I run as CTO, where a failed document generation or an authorization error has compliance implications, and "we're not sure which customer was affected" isn't a sentence I can say to an enterprise client; MindWrite AI, the subscription AI writing tool, where every OpenAI call has a real dollar cost and a latency that directly affects a paying customer's experience, and both need to be measurable per tenant, not just in aggregate; SignageFlow, the digital signage platform, where a screen going silent needs to be distinguishable from "no content was scheduled" the moment it happens, not the next time someone happens to glance at a dashboard; and ReplyVibe and Expreco, both built around third-party APIs (Google Business Profile, FedEx, UPS) where knowing which upstream call failed, with what payload, is the entire difference between a fast fix and reverse-engineering the bug from a vague user complaint.

This is a breakdown of how to build observability into a Laravel SaaS product that actually helps you find and fix problems fast — structured logging with correlation across requests and queued jobs, metrics that answer specific operational questions instead of feeding a dashboard nobody checks, error tracking that groups and tags instead of just collecting, and alerting that reaches a human without drowning them.

Architecture: What Observability Needs Beyond Log::error()#

Laravel gives you the primitives — the Log facade, an exception handler, queued jobs, a Context facade for structured data. It doesn't give you the discipline of using them so that a production incident is traceable instead of a guessing game. That discipline has four parts, and skipping any one of them is invisible until the exact moment you need it.

1. Every log line and every reported error carries structured context, propagated across the request and into queued jobs#

A Log::info("Batch started") and a Log::info("Batch started for batch {$batchId}") look similar. They're not. The first is unsearchable — there's no way to ask "show me every log line for batch 4821" without grepping raw text and hoping the format never changed. The second is slightly better but still requires remembering to manually interpolate the right ID into every single log call across a codebase, which nobody does consistently under deadline pressure.

Laravel's Context facade solves the actual problem: attach structured data once, near the entry point of a request, and every subsequent Log:: call and every reported exception picks it up automatically, without being passed explicitly through every function signature in between.

php
class AddRequestContext
{ 
    public function handle(Request $request, Closure $next)
    {
        $requestId = $request->header('X-Request-Id', (string) Str::uuid());

        Context::add('request_id', $requestId);
        Context::add('tenant_id', $request->user()?->tenant_id);
        Context::add('user_id', $request->user()?->id);
        Context::add('route', $request->route()?->getName());

        return $next($request)->header('X-Request-Id', $requestId);
    }
}
php
// From anywhere downstream — a controller, a service class, a listener —
// this line now automatically includes request_id, tenant_id, user_id,
// and route in the log output, with zero extra plumbing:
Log::warning('Carrier activation returned an unexpected status', [
    'carrier' => 'nexio-mobile',
    'activation_id' => $activation->id,
]);

The part that actually matters for a SaaS product built on queues — and this is where most logging setups quietly fall apart — is that context established during an HTTP request doesn't automatically survive into a job dispatched from inside it unless the job re-establishes its own. A job that runs three minutes later, on a different worker process, has no HTTP request to inherit context from, so it has to set its own correlation ID deliberately:

php
class ProcessSwapPadBatchJob implements ShouldQueue
{ 
    public function __construct(
        private int $batchId,
        private string $correlationId,
    ) {}

    public function handle(): void
    {
        Context::add('batch_id', $this->batchId);
        Context::add('correlation_id', $this->correlationId);

        // every Log:: call and every exception reported inside this job,
        // and inside anything it dispatches further, now carries both
    }
}

Passing $correlationId from the request that triggered the dispatch — instead of generating a fresh one inside the job — is what lets you trace one customer action end-to-end: the HTTP request that uploaded the CSV, the job that processed row 340, and the Slack alert that fired when row 340's carrier call timed out, all tagged with the same ID, searchable as one thread instead of three disconnected log lines.

2. Metrics answer a specific operational question, not "look, a dashboard"#

It's easy to end up with a metrics dashboard that's full of numbers and answers nothing — CPU percent, memory, request count. Those numbers move but rarely tell you what a customer is actually experiencing. The metrics worth collecting are the ones tied to a specific failure mode you already know matters: queue wait time (covered in Laravel queue architecture, because 30 jobs stuck for fifteen minutes is an incident and depth alone won't tell you that), cache hit rate (from Redis caching architecture, because a silently collapsing hit rate is often the earliest signal of a cache-key bug before the database even feels it), webhook delivery success rate (from webhook delivery architecture), and — specific to MindWrite AI — AI generation latency and token cost, broken down per tenant, because a single customer's prompt patterns can quietly become the majority of the OpenAI bill without a single error ever being thrown.

A lightweight metrics recorder, backed by Redis, is enough to start answering these questions without standing up a full metrics stack on day one:

php
class Metrics
{
    public static function timing(string $name, float $milliseconds, array $tags = []): void
    {
        Redis::connection('metrics')->pipeline(function ($pipe) use ($name, $milliseconds, $tags) {
            $key = self::key($name, $tags);
            $pipe->rpush($key, (int) $milliseconds);
            $pipe->ltrim($key, -1000, -1); // rolling window — bounded, not unbounded growth
            $pipe->expire($key, 86400);
        });
    }

    public static function increment(string $name, array $tags = [], int $by = 1): void
    {
        Redis::connection('metrics')->incrby(self::key($name, $tags), $by);
    }

    private static function key(string $name, array $tags): string
    {
        ksort($tags);
        $suffix = collect($tags)->map(fn ($v, $k) => "{$k}={$v}")->implode(',');

        return $suffix ? "metrics:{$name}:{$suffix}" : "metrics:{$name}";
    }
}
php
// MindWrite AI — cost and latency, tagged per tenant, so "which customer's
// usage is driving the OpenAI bill this month" is a query, not a guess
$start = microtime(true);
$result = $provider->generate($prompt);

Metrics::timing('ai.generation.duration_ms', (microtime(true) - $start) * 1000, [
    'tenant_id' => $tenant->id,
    'model' => $result->model,
]);

Metrics::increment('ai.generation.tokens_used', ['tenant_id' => $tenant->id], $result->totalTokens);

The Redis::connection('metrics') detail is deliberate, not decoration — metrics writes are high-frequency and disposable, and they have no business sharing a connection pool or memory budget with the cache or queue Redis instances that the rest of the app depends on for correctness.

3. Error tracking groups, tags, and separates expected failures from real bugs#

Log::error() scattered across a codebase gets you a file full of text. An error tracker like Sentry (or Bugsnag) gets you something categorically different: identical errors grouped into one issue instead of a thousand duplicate log lines, a stack trace with the actual variable state at the point of failure, and — critically for a multi-tenant product — the ability to tag every reported exception with the same context established above, so "how many tenants hit this, and which ones" is answered in the tool, not reconstructed from logs after the fact.

php
// bootstrap/app.php (Laravel 11+)
->withExceptions(function (Exceptions $exceptions) {
    $exceptions->reportable(function (Throwable $e) {
        \Sentry\configureScope(function (\Sentry\State\Scope $scope): void {
            $scope->setTag('tenant_id', (string) Context::get('tenant_id'));
            $scope->setTag('request_id', (string) Context::get('request_id'));
            $scope->setContext('correlation', [
                'batch_id' => Context::get('batch_id'),
                'correlation_id' => Context::get('correlation_id'),
            ]);
        });

        if ($e instanceof CarrierTimeoutException) {
            // Expected and transient — a carrier sandbox being slow isn't
            // an engineering bug. Count it, don't page anyone for it.
            Metrics::increment('carrier.timeout', ['carrier' => $e->carrier]);

            return false; // stop propagation — this never reaches Sentry
        }
    });

    $exceptions->dontReport([
        ValidationException::class,
        AuthorizationException::class,
    ]);
})

The dontReport list and the early return false matter more than they look — an error tracker that captures every 422 validation failure and every expected 403 alongside genuine 500s trains the whole team to skim past it, because the signal-to-noise ratio makes it unusable. The exceptions worth an engineer's attention are the ones that mean something's actually wrong, not the ones that mean the system is correctly rejecting bad input.

4. Alerts reach a human, with enough restraint that they still get read#

None of the above matters if the one alert that mattered arrives buried under forty identical ones from the same flapping dependency. A carrier API having a bad ten minutes shouldn't generate forty separate Slack messages — it should generate one, with a cooldown that suppresses the repeat until either the condition clears or enough time has passed that it's worth re-surfacing.

php
class AlertDispatcher
{
    public function critical(string $key, string $message, array $context = []): void
    {
        $cooldownKey = "alert-cooldown:{$key}";

        if (Cache::has($cooldownKey)) {
            return; // already paged for this exact condition recently
        }

        Cache::put($cooldownKey, true, now()->addMinutes(15));

        Notification::route('slack', config('services.slack.ops_webhook'))
            ->notify(new OpsAlert($message, array_merge($context, [
                'tenant_id' => Context::get('tenant_id'),
                'correlation_id' => Context::get('correlation_id'),
            ]), severity: 'critical'));
    }
}
php
// Used from the queue-health check described in the queue architecture
// piece — same underlying signal, now routed through a dispatcher that
// won't flood the channel if the backlog doesn't clear in one check cycle
app(AlertDispatcher::class)->critical(
    key: "queue-backlog:{$queue}",
    message: "Queue [{$queue}] has a job waiting {$waitSeconds}s",
    context: ['queue' => $queue, 'wait_seconds' => $waitSeconds],
);

The $key is what makes the cooldown precise instead of blunt — a backlog on the broadcasts queue and a separate backlog on batch-imports are different conditions and should both be able to alert, even while either one individually is being rate-limited against repeating itself every sixty seconds.

Step-by-Step: Building the Pipeline#

  1. Add a request-correlation middleware first. Everything else in this piece depends on request_id and tenant_id existing before a single log line is written — retrofitting correlation onto an app with years of bare Log::info() calls is a much bigger job than starting with it in place.
  1. Replace string-interpolated log messages with structured context arrays, one log call at a time as you touch that code — Log::info("Batch {$id} started") becomes Log::info('Batch started', ['batch_id' => $id]), which is queryable in any log aggregator instead of only greppable as raw text.
  1. Give each concern its own log channel where it deserves one. An audit trail for privileged actions (the discipline from RBAC architecture and secure file storage) and a search-index drift warning (from full-text search architecture) both deserve dedicated channels, separate from general application noise, so a compliance review or a drift investigation isn't wading through unrelated debug output to find what matters:
php
// config/logging.php
'channels' => [
    'audit' => [
        'driver' => 'daily',
        'path' => storage_path('logs/audit.log'),
        'days' => 365, // compliance retention, not the 14-day default
    ],
    'integrations' => [
        'driver' => 'daily',
        'path' => storage_path('logs/integrations.log'),
        'days' => 30,
    ],
],
  1. Integrate an error tracker and tag every event with tenant and request context, as shown above — an untagged error tracker answers "something broke" the same way a bare Log::error() does; a tagged one answers "this broke, for this tenant, on this request, N times this hour."
  1. Pick metrics tied to a failure mode you already care about, not everything you can measure. Queue wait time, cache hit rate, webhook and third-party API success rate, and — where it applies — per-tenant cost on metered dependencies like OpenAI. A dashboard with forty panels nobody checks is worse than five numbers everyone knows the healthy range for.
  1. Route alerts through severity and a cooldown, never a raw Notification::send() per incident. A critical alert (carrier down, queue backlog past a threshold, error rate spike) pages Slack or PagerDuty immediately with suppression; a warning-level signal can batch into a daily digest instead of interrupting anyone in real time.
  1. Test that your error handler actually captures what you think it captures. The same discipline from testing strategy in Laravel SaaS: a test that asserts a thrown exception is tagged with the right tenant ID catches the regression where someone refactors the exception handler and quietly drops the configureScope call.
php
public function test_reported_exception_is_tagged_with_tenant_context(): void
{
    $tenant = Tenant::factory()->create();
    $user = User::factory()->for($tenant)->create();

    $events = [];
    \Sentry\Client::class; // illustrative — capture via your test double of choice
    Event::listen(SentryEventCaptured::class, function ($event) use (&$events) {
        $events[] = $event;
    });

    $this->actingAs($user)->getJson('/api/reports/triggers-error');

    $this->assertNotEmpty($events);
    $this->assertEquals((string) $tenant->id, $events[0]->tags['tenant_id']);
}
  1. Run a quarterly "can we actually find this" drill, not just a passive review — pick a real incident from the last quarter and time how long it takes to answer "which tenants were affected, and for how long" using only the logging and metrics in place. If the answer takes longer than the original incident did to fix, the observability layer isn't doing its job.

Real-World Pitfalls to Avoid#

String-interpolated log messages with no structured context. Log::error("Failed for {$id}") is only ever searchable as raw text — the moment you need "every failure for tenant 41 in the last hour," a structured ['tenant_id' => 41] context array is the difference between a query and an afternoon of grepping.

A single generic exception handler that reports everything to the error tracker. Every validation failure, every expected 403, and every genuine 500 landing in the same feed as equally important trains the team to stop reading it — the exceptions that deserve attention need to be separable from the ones that are the system working correctly.

No correlation ID that survives into a queued job. A request that dispatches a job, and a job that generates its own fresh identity instead of carrying the one it was given, breaks the exact trace you need most — the ability to follow one customer action from the HTTP request that started it through to the background work it triggered.

Logging full payloads that include secrets or PII in cleartext. A Stripe token, an API key, or a SafetySpace incident report's personal details sitting unredacted in a log file that's readable by anyone with server access is a compliance problem waiting to be discovered during an audit, not an engineering shortcut with no cost.

Alerts with no cooldown or dedup. A flapping dependency generating one Slack message per failed check is how a channel that used to get read stops getting read — a suppression window per alert key keeps the signal legible instead of drowned.

Metrics that don't map to a customer-facing consequence. Server CPU and memory are useful for capacity planning; they don't tell you a customer is staring at a stuck upload. Queue wait time, delivery success rate, and cache hit rate are useful specifically because a regression in any of them is a regression someone downstream actually feels.

No retention or cost plan for logs at real volume. A single unindexed laravel.log file, or an error-tracker plan sized for a demo's traffic, becomes either unsearchable or unaffordable the moment usage grows — decide retention per channel deliberately (30 days for integration noise, a year for an audit trail) instead of defaulting everything to the same setting and discovering the cost or the gap later.

Key Takeaways#

Observability earns its place in a Laravel SaaS product the same way queueing, caching, and authorization do — as a system with a specific job, not a storage/logs folder that happens to exist.

  • Every log line and every reported exception carries structured, tenant-aware context, established once at the request boundary and deliberately re-established inside every queued job it touches
  • Metrics are chosen for the specific operational questions they answer — queue wait time, cache hit rate, delivery success rate, per-tenant cost — not collected because they're easy to collect
  • Error tracking groups and tags instead of just accumulating, with expected, transient failures kept out of the same feed as genuine bugs
  • Alerts reach a human with severity and a cooldown, so the one that matters isn't buried under forty duplicates from the same flapping dependency
  • The whole system is periodically tested against a real incident, not assumed to work because it was configured once and never revisited

I've built this discipline into a bulk SIM activation platform where a carrier failure needs to reach on-call within seconds, an AI safety platform where "which customer was affected" has to be a fast, confident answer, a subscription AI tool where cost and latency are tracked per tenant down to the token, a signage platform where a silent screen has to be distinguishable from an idle one immediately, and integration-heavy products where a third-party API's bad day needs to be diagnosable from the first alert, not the fifth support ticket. The stack changes; the requirement that a production incident be traceable in minutes, not reconstructed from memory, doesn't.

If your incident response still starts with "let me check the logs" and a lot of hoping — get in touch about your observability architecture or see the full case studies from platforms running structured, tenant-aware observability in production today.

Share this technical insight with your network

Share to LinkedIn or Facebook with key takeaways, featured media, and direct links.

📁 Production Case Study

Case Study: SafetySpace AI-Powered Safety Management Platform & Workflow Automation

SafetySpace is an AI-powered safety management platform designed to help organizations manage safety processes, access critical safety information, and streamline documentation through intelligent, configurable workflows. As CTO, I led the technical direction and development of the platform across backend, frontend, architecture, and AI integrations—turning complex safety workflows into a more intelligent and streamlined digital experience.

Related Technical Articles

View all articles →
✦ Let's Build Together

Have a complex technical project in mind?

Available for full-stack engineering, performance audits, cloud deployments, and high-concurrency systems architecture.

Need a web or software development partner?

Tell me what you’re building, what’s getting in the way, and where you need help. Whether you need a custom web application, SaaS platform, API integration, or full-stack development, I’ll give you a clear answer on scope, cost, and timeline usually within one business day.

AqibJavaid

Senior Full-Stack Engineer building backend systems, cloud infrastructure and product platforms for teams that need them to stay up.

Available for new projects

Get in touch

© 2026 Aqib Javaid. All rights reserved.

Built and maintained by Aqib Javaid