Introduction: ShouldQueue Is Not An Architecture#
Every Laravel job starts the same way: implement ShouldQueue, dispatch it, watch it disappear off the request thread and into a worker somewhere. It feels solved the moment php artisan queue:work picks it up in a local terminal. Then production traffic arrives, and the questions a local dev environment never asks all land at once. What happens when a worker gets killed mid-job by a deploy, halfway through updating three rows? What happens when the same webhook fires the same job twice because the queue connection blipped and Laravel's own retry-after mechanism kicked in? What happens when one tenant's nightly import job is queued behind 4,000 other jobs because everything shares one default queue, and a paying customer's screen content update is stuck waiting behind it?
None of that is a queue failure. It's a job design failure — the kind that a single worker processing jobs one at a time in a demo will never surface, because a demo never has two jobs racing, a worker restarting under a deploy, or a queue depth past a few dozen.
I've built and operated queue-driven job pipelines across products where background work isn't incidental, it's core to the product working at all: SafetySpace, the AI-powered safety platform I run as CTO, where AI-generated SWMS/SOP documents are queued because a 30-second model completion has no business holding an HTTP connection open; SwapPad, a bulk multi-carrier SIM activation platform I built inside the CelleUp dealer back office, where a single CSV upload can spawn a queued job that runs for up to two hours against rate-limited carrier sandboxes, activating hundreds of SIMs row by row without double-charging a dealer's wallet on retry; SignageFlow, where every real-time content-push broadcast to a customer's screen fleet is dispatched through a queue so a slow WebSocket server never becomes a slow API response; and Reply Vibe, where sentiment classification on incoming reviews runs as a queued pipeline so a burst of 200 new reviews doesn't block anything a user is actively waiting on.
This is a breakdown of how to design Laravel queue architecture so jobs survive worker restarts, retries, and real concurrency — not just the first successful run in local dev.
Architecture: A Job Is a Contract With Failure, Not a Deferred Function Call#
The mindset shift that matters most: dispatching a job doesn't mean "run this later." It means "this code has to produce the correct result whether it runs zero times, once, or three times, on a worker that might be killed at any line." Laravel's queue system gives you ShouldQueue, retries, and backoff for free. It does not give you idempotency, correct job sizing, or queue prioritization — those are architecture decisions you make, and getting them wrong is invisible until the exact conditions that expose them show up in production.
1. Pass IDs into a job, never hydrated models#
The most common queue bug I inherit is a job constructed with a full Eloquent model:
// Don't do this
class SendWelcomeEmail implements ShouldQueue
{
public function __construct(private User $user) {}
}Laravel serializes that model into the queue payload at dispatch time — a snapshot of the user's attributes right now. If the job sits in the queue for even a few seconds before a worker picks it up, and something else updates that user in the meantime, the job runs against stale data it never re-fetches. Worse, on a job that retries after a failure, it retries against that same stale snapshot, not the current row.
class SendWelcomeEmail implements ShouldQueue
{
public function __construct(private int $userId) {}
public function handle(): void
{
$user = User::findOrFail($this->userId);
// always the current row, on first attempt and on every retry
}
}This is a small change with an outsized effect: it makes every retry correct by construction, because the job re-reads reality instead of replaying a memory of it.
2. Make every job idempotent against its own side effects, not just its trigger#
Idempotent webhook handling (recording a processed event_id before acting) is the well-known half of this problem. The less obvious half is that the job itself needs to be safe to run twice, independent of whether its trigger was deduplicated — because Laravel's own queue retries a job that times out or throws, and a worker that's SIGKILLed mid-execution can leave a job's database "processed" flag unset even though most of its side effects already ran.
On SwapPad, ProcessSwapPadBatchJob activates SIMs row by row against carrier APIs that charge non-refundable wholesale cost per activation. A naive retry-from-scratch on a job failure would re-activate — and re-charge — every row that had already succeeded before the failure.
class ProcessSwapPadBatchJob implements ShouldQueue
{
public $timeout = 7200; // some carrier sandboxes are slow
public function __construct(private int $batchId) {}
public function handle(): void
{
$batch = ISPActivationBatch::findOrFail($this->batchId);
foreach ($batch->activations()->where('status', 'pending')->cursor() as $activation) {
// Only rows still pending get touched — a retry skips
// everything a previous attempt already completed or failed
$this->activateRow($activation);
}
}
}The job doesn't ask "have I run before?" — it asks "which rows still need work?" on every execution. That distinction is what makes a two-hour job survive a worker restart at the 90-minute mark without duplicating cost or activations.
3. Use WithoutOverlapping and unique jobs where concurrent runs would corrupt state#
Some jobs are dangerous specifically when two copies run at once — not because of external side effects, but because they read-modify-write the same row. A nightly reconciliation job that syncs a tenant's subscription state, or a job that recalculates a dealer's wallet balance, can produce a corrupted result if two instances interleave their reads and writes.
class ReconcileTenantSubscription implements ShouldQueue, ShouldBeUnique
{
public function __construct(private int $tenantId) {}
public function uniqueId(): string
{
return "reconcile-subscription-{$this->tenantId}";
}
public $uniqueFor = 300; // seconds — release the lock if something goes badly wrong
}ShouldBeUnique uses your cache driver's atomic lock to guarantee only one instance of that job, for that tenant, is on the queue at a time — a second dispatch for the same tenant while one is still running is silently dropped rather than queued to run concurrently.
Step-by-Step: Designing the Pipeline, Not Just the Job#
- Separate queues by latency sensitivity, not just by feature. A
defaultqueue that mixes a user-facing password-reset email with a four-hour batch import means the email waits behind the import the moment the import gets there first. On SignageFlow, real-time content-push broadcasts run on their ownbroadcastsqueue, entirely separate fromdefault, so a backlog in one never delays the other:
class ScreenContentUpdated implements ShouldBroadcast, ShouldQueue
{
public $connection = 'redis';
public $queue = 'broadcasts';
}// config/horizon.php
'environments' => [
'production' => [
'supervisor-broadcasts' => [
'connection' => 'redis',
'queue' => ['broadcasts'],
'balance' => 'auto',
'maxProcesses' => 10,
],
'supervisor-default' => [
'connection' => 'redis',
'queue' => ['default'],
'balance' => 'auto',
'maxProcesses' => 4,
],
'supervisor-batch' => [
'connection' => 'redis',
'queue' => ['batch-imports'],
'balance' => 'auto',
'maxProcesses' => 2,
'timeout' => 7200,
],
],
],A long-running SwapPad batch and a fast SignageFlow broadcast are structurally incapable of blocking each other, because they never share a worker pool.
- Size jobs by what a single failure should cost, not by what's convenient to write. A job that processes 500 rows in one
handle()method means a failure on row 481 either loses the first 480 rows' progress or requires exactly the kind of "re-check what's already done" idempotency logic shown above. Job batching splits the difference — each row is its own job, tracked under one batch, so a single row's failure doesn't threaten the other 499:
$batch = Bus::batch(
$activationRows->map(fn ($row) => new ActivateSingleSim($row->id))
)->then(function (Batch $batch) use ($batchId) {
ISPActivationBatch::find($batchId)->update(['status' => 'completed']);
})->catch(function (Batch $batch, Throwable $e) use ($batchId) {
Log::error('SwapPad batch had failures', ['batch_id' => $batchId, 'error' => $e->getMessage()]);
})->allowFailures()->dispatch();allowFailures() matters here specifically: without it, one carrier timeout on row 12 cancels the remaining 488 rows instead of letting them keep processing while row 12's failure is surfaced separately.
- Rate-limit jobs that call an external API, at the job layer, not just with
Http::retry(). SwapPad's carrier integrations and Reply Vibe's Google Business Profile sync both hit providers with real per-minute quotas. A burst of 300 queued jobs hitting a 60-requests-per-minute API doesn't fail gracefully on its own — it needs the queue itself to throttle:
class SyncReviewFromGoogle implements ShouldQueue
{
public function middleware(): array
{
return [(new RateLimited('google-business-api'))->releaseAfterMs(2000)];
}
}// AppServiceProvider
RateLimiter::for('google-business-api', function () {
return Limit::perMinute(60);
});Jobs that exceed the limit are released back onto the queue with a delay instead of firing anyway and eating a 429 — the same discipline that matters for any third-party integration, expressed at the job-scheduling layer instead of inside each HTTP call.
- Give every retryable job a real backoff curve, not a fixed delay. A flat
$backoff = 30retries a downstream outage every 30 seconds regardless of whether the outage is a five-second blip or a ten-minute one, which just adds load to an already-struggling dependency. An increasing backoff spreads retries out as failures persist:
class GenerateSwmsDocument implements ShouldQueue
{
public $tries = 5;
public function backoff(): array
{
return [10, 30, 90, 300, 900]; // seconds — widens as failures persist
}
public function handle(AiGenerationProvider $provider): void
{
// ...
}
}- Watch the queue in production the way you'd watch any critical infrastructure — not by tailing worker logs after a customer complains. Laravel Horizon's dashboard surfaces queue depth, throughput, and failed-job counts per queue in real time, but the number that actually matters operationally is wait time, not job count — 4,000 jobs on a queue that drains in ten seconds is healthy; 40 jobs on a queue that hasn't moved in twenty minutes is an incident. Alert on wait time, not depth:
class MonitorQueueHealth extends Command
{
public function handle(): void
{
foreach (['default', 'broadcasts', 'batch-imports'] as $queue) {
$oldestJob = Redis::connection('horizon')->zrange("queues:{$queue}", 0, 0, 'WITHSCORES');
if ($oldestJob && (now()->timestamp - $oldestJob[1]) > 300) {
Notification::route('slack', config('services.slack.ops_webhook'))
->notify(new QueueBacklogAlert($queue, now()->timestamp - $oldestJob[1]));
}
}
}
}Pitfalls I've Seen Cost Real Time and Money#
Hydrating models into job constructors. Covered above, but it's the single most common queue bug I fix in inherited codebases — a job holding a stale snapshot of a row means retries silently operate on data that's already wrong, and nobody notices until a support ticket doesn't match what's actually in the database.
One shared default queue for everything. A long-running batch import and a password-reset email competing for the same worker pool means the fast, user-facing job is only as fast as whatever got queued ahead of it. Separate queues by latency sensitivity before it becomes a support ticket about "why didn't I get my email," not after.
Treating job failure as "retry and hope." $tries = 3 with no backoff strategy just hits a struggling dependency three times in rapid succession instead of once — that's not resilience, it's amplification. A real backoff curve, and a failed() method that actually does something (alerts, compensating action, dead-letter queue) instead of silently landing in the failed_jobs table, is what makes retry logic meaningful instead of decorative.
No idempotency on jobs with real-money or real-inventory side effects. SwapPad activates physical SIM inventory and charges dealer wallets — a job retry that doesn't check what already succeeded doesn't just create a data inconsistency, it burns non-refundable carrier cost twice for one customer action. Any job with an external, costly side effect needs to answer "what's already done" before it does anything, every single time it runs.
Batch jobs with no allowFailures(). A job batch that cancels 499 healthy rows because row 12 hit a transient carrier timeout turns one recoverable failure into a total batch loss — allowFailures() plus a catch() callback that surfaces exactly which rows failed is what keeps a partial failure partial.
Monitoring queue depth instead of wait time. A queue with 5,000 jobs that's draining in under a minute is fine. A queue with 30 jobs that hasn't moved in fifteen minutes because every worker died on a bad deploy is not — and depth alone won't tell you which situation you're in. Wait time on the oldest job is the number that actually correlates with customer impact.
No $timeout set on long-running jobs. SwapPad's batch job runs up to two hours against slow carrier sandboxes; without an explicit $timeout matched to that reality, Laravel's default worker timeout kills it well before it finishes, and a job killed mid-execution without idempotent row-level tracking is exactly how partial batches happen in the first place.
Key Takeaways#
A queue isn't a place where code runs later — it's a contract that code has to hold up under retries, concurrent execution, and workers that can die mid-job at any line. The queue architectures that hold up in production share the same shape:
- Jobs constructed from IDs, re-fetching current state on every execution — never hydrated models carrying a stale snapshot
- Idempotency built at the row or unit-of-work level, so a retry after partial failure never repeats a costly or destructive side effect
- Separate queues by latency sensitivity, so a slow batch job structurally can't delay a fast, user-facing one
- Job batching for multi-step work, with
allowFailures()so one row's failure doesn't cancel hundreds of healthy ones - Rate limiting and real backoff curves at the job layer for anything calling a metered external dependency
- Monitoring wait time on the oldest queued job, not just job count, because that's the number that actually predicts customer impact
I've built this discipline into a bulk SIM activation engine where a job failure has real carrier cost attached, an AI safety platform queuing document generation as CTO, a signage platform broadcasting to thousands of screens without ever blocking an API response, and a review-sync pipeline throttled against a provider's real rate limits. The product changes; the discipline around what a job has to survive doesn't.
If your failed_jobs table is quietly growing, or you're designing background processing for a SaaS product that can't afford a job to run twice, get in touch about your queue architecture or see the full case studies from platforms where these patterns are running today.
Share this technical insight with your network
Share to LinkedIn or Facebook with key takeaways, featured media, and direct links.
Case Study: Multi-Carrier Bulk SIM Activation Platform
Wireless dealers and distributors needed to activate SIMs and eSIMs across a dozen-plus disparate carrier and MVNO backends one line at a time; I built a bulk CSV-driven activation engine inside the existing CelleUp dealer platform that validates multi-tier dealer funds, routes each row to the correct carrier API, processes activations asynchronously with retries, and streams live progress back to the dealer.
Related Technical Articles
View all articles →
Testing Strategy for Laravel SaaS: PHPUnit, Dusk, Load Tests
A passing CI pipeline and a production incident can coexist. Here's the Laravel testing architecture that closes that gap for real SaaS teams.

Subscription Billing Architecture in Laravel: Building Tier-Based SaaS Pricing That Survives Renewals, Downgrades, and Failed Cards
Charging a card once is easy. Charging the right amount, on the right day, for a plan that changed twice last month that's the real billing problem.
Have a complex technical project in mind?
Available for full-stack engineering, performance audits, cloud deployments, and high-concurrency systems architecture.

