Introduction: throttle:60,1 Is a Placeholder, Not a Policy#
Every Laravel API starts with the same line in routes/api.php: Route::middleware('throttle:60,1')->group(...). It ships, it works, nobody thinks about it again — until the day one integration partner writes a polling loop with no backoff and starts hammering an endpoint every 200ms, and every other tenant sharing that API suddenly sees their own, completely unrelated requests getting rejected. Or the opposite happens: an enterprise customer who's paying for a higher tier and a heavier integration keeps hitting the same flat limit as a customer on the free plan, because the limit was never actually tied to who's asking or what they're paying for. Both failures come from the same root cause — treating rate limiting as a single global number bolted onto a route, instead of a policy that has to answer "who is this," "what plan are they on," and "how expensive is this specific request" before it can answer "should this be allowed right now."
I've had to get this right on products where the API itself is the product, not an afterthought: MindWrite AI, the subscription-based AI writing tool, where every generation request costs real money in OpenAI tokens and a monthly credit limit has to be enforced request-by-request, not just checked at the dashboard; SignageFlow, the digital signage platform, where hundreds of physical screens poll the API on a schedule for content updates, and a naive per-IP or per-account limit either throttles a customer's entire screen fleet during a legitimate traffic spike or does nothing at all because every screen shares one tenant's quota; SafetySpace, the AI-driven safety platform I run as CTO, where enterprise customers integrate directly against our API and expect documented, predictable limits they can build their own retry logic around; and the third-party integrations I wrote about separately — FedEx, UPS, Stripe, OpenAI, Google Business Profile — where the rate limiting problem runs in the other direction: it's not just about protecting my API from abuse, it's about not getting my own outbound calls banned by someone else's.
This is a breakdown of how to build rate limiting in Laravel that's tenant-aware, tier-based, algorithmically correct under concurrency, and considerate enough to the client on the other end that a 429 response is something their integration can actually recover from — not a mystery error that shows up in a support ticket.
Architecture: What Rate Limiting Needs Beyond the throttle Middleware#
Laravel's built-in RateLimiter facade and throttle middleware are a solid foundation — they're not the whole architecture. The gap is everything between "count requests" and "count the right requests, against the right limit, for the right identity, without a race condition."
1. Limits are keyed by tenant and plan tier, not by IP or a single global bucket#
An IP-based limit punishes every user behind the same NAT or corporate proxy as a single client, and a flat global limit ignores the fact that a customer paying for a higher tier should get a proportionally higher ceiling — the same principle from subscription billing architecture: entitlements, including rate limits, are something a tier grants, not a constant hardcoded into a route definition.
class RateLimitPolicy
{
public function limitFor(Tenant $tenant, string $bucket): Limit
{
$config = $tenant->subscription->effectiveTier()->limits['rate_limits'][$bucket]
?? config("rate_limits.defaults.{$bucket}");
return Limit::perMinute($config['requests'])
->by("tenant:{$tenant->id}:{$bucket}");
}
}Registering this in AppServiceProvider means every named limiter resolves the same way, for every route, instead of a route-by-route guess at what number felt reasonable:
RateLimiter::for('api', function (Request $request) {
$tenant = $request->user()->tenant;
return app(RateLimitPolicy::class)->limitFor($tenant, 'api');
});
RateLimiter::for('generations', function (Request $request) {
// MindWrite AI: a separate, much tighter bucket for the endpoint that actually costs money
return app(RateLimitPolicy::class)->limitFor($request->user()->tenant, 'generations');
});On MindWrite AI, generations and api are deliberately separate buckets on the same tenant — a customer browsing their document list shouldn't burn down the same allowance as a customer generating content, because the two requests don't cost the same thing on the backend.
2. Expensive endpoints are weighted, not counted as one request each#
A rate limiter that treats GET /projects and POST /generate as equally "one request" is measuring the wrong thing. The generate endpoint on MindWrite AI calls OpenAI and costs real tokens; a plain read from Postgres costs almost nothing. Counting them identically means either the cheap endpoint is throttled too aggressively or the expensive one isn't throttled enough:
class WeightedRateLimiter
{
public function attempt(string $key, int $weight, int $maxPoints, int $decaySeconds): bool
{
return Redis::eval(<<<'LUA'
local current = tonumber(redis.call('GET', KEYS[1]) or '0')
if current + tonumber(ARGV[1]) > tonumber(ARGV[2]) then
return 0
end
redis.call('INCRBY', KEYS[1], ARGV[1])
redis.call('EXPIRE', KEYS[1], ARGV[3])
return 1
LUA, 1, $key, $weight, $maxPoints, $decaySeconds) === 1;
}
}
// A list endpoint costs 1 point; an AI generation call costs 20
$allowed = app(WeightedRateLimiter::class)->attempt(
"tenant:{$tenant->id}:points",
weight: $request->routeIs('generations.store') ? 20 : 1,
maxPoints: $tenant->subscription->effectiveTier()->limits['rate_limits']['points_per_minute'],
decaySeconds: 60
);The Lua script matters here for a reason that has nothing to do with style: GET then INCRBY as two separate PHP calls is a check-then-act race condition. Two concurrent requests can both read the count before either writes it back, and both pass a check that should have let only one through. Running the check-and-increment as a single atomic script on Redis is what makes the limit actually hold under real concurrent load, not just in a single-threaded test.
3. Distinguish burst limits from sustained limits#
A single per-minute ceiling can't express "a short burst is fine, sustained hammering isn't" — a dashboard load that fires eight requests in the same second is normal UI behavior, not abuse, but the same eight-per-second rate sustained for ten minutes is exactly the polling-loop-gone-wrong pattern that took down shared infrastructure on SignageFlow when one customer's screen-management script had a bug in its retry interval. Two limiters, layered:
RateLimiter::for('signage-sync', function (Request $request) {
$tenant = $request->user()->tenant;
return [
Limit::perSecond(5)->by("tenant:{$tenant->id}:burst"),
Limit::perMinute(120)->by("tenant:{$tenant->id}:sustained"),
];
});Laravel evaluates an array of limiters as an AND — a request has to pass both the burst check and the sustained check — which is exactly the shape needed to allow a legitimate spike while still catching the runaway loop the spike-only check would miss.
4. Outbound calls need the same discipline, in the opposite direction#
Rate limiting isn't only about protecting your own API — the third-party API integration work I did against FedEx, UPS, and Google's APIs means respecting their limits, or risking a temporary ban that takes down every tenant's integration at once, not just the one that caused it:
class ThrottledApiClient
{
public function __construct(private string $provider, private int $maxPerSecond) {}
public function call(Closure $request): mixed
{
$key = "outbound:{$this->provider}";
while (!RateLimiter::attempt($key, $this->maxPerSecond, fn () => true, 1)) {
usleep(50_000);
}
return $request();
}
}On Expreco, FedEx and UPS rate quotes are funneled through exactly this kind of self-imposed throttle — not because Laravel needs protecting, but because a burst of quote requests from an unusually large logistics order shouldn't be the thing that gets the account rate-limited by the carrier for every customer at once.
Step-by-Step: Building the Pipeline#
- Define named limiters per concern, not per route —
api,generations,webhooks,signage-sync— registered once inAppServiceProvider, so a limit change is a config or database update, not a search-and-replace across route files.
- Resolve every limiter's ceiling from the tenant's subscription tier, using the same
SubscriptionTier.limitsstructure from subscription billing architecture — a rate limit is an entitlement like any other, and it should change the moment a tier changes, not require a deploy.
- Return headers that let the client actually behave correctly, not just a bare 429:
RateLimiter::for('api', function (Request $request) {
return Limit::perMinute($limit)
->by("tenant:{$request->user()->tenant_id}")
->response(function (Request $request, array $headers) {
return response()->json([
'message' => 'Rate limit exceeded. Retry after the window resets.',
], 429, $headers);
});
});Laravel populates X-RateLimit-Limit, X-RateLimit-Remaining, and Retry-After automatically — the same headers documented in the webhook delivery system I built, because an integration that can read Retry-After can back off correctly on its own, while one that only sees a bare 429 usually just retries immediately and makes the problem worse.
- Use a shared Redis store for the rate limiter, never the default array/file cache, the instant the app runs on more than one server. A per-server in-memory limiter means a tenant's real ceiling is
limit × number of app servers, silently, and it drifts further out of sync every time the fleet scales up or down:
// config/cache.php — the limiter must resolve through a store that's shared across every app server
'default' => env('CACHE_DRIVER', 'redis'),- Give webhook-receiving and public unauthenticated endpoints their own, stricter limiter, keyed by IP or signature rather than tenant, since these are the endpoints most likely to be probed by something that isn't a legitimate integration at all.
- Log rate-limit rejections with enough context to diagnose them, not just increment a counter silently — when a customer emails asking why their integration is failing, "which limiter, which tenant, how far over" needs to be a query against a log, not a guess:
class LogRateLimitRejection
{
public function handle(RateLimitExceeded $event): void
{
Log::channel('rate-limits')->warning('Rate limit exceeded', [
'tenant_id' => $event->tenantId,
'limiter' => $event->limiterName,
'route' => $event->route,
]);
}
}- Test that a tenant on a lower tier is actually capped, and a tenant on a higher tier isn't — the same negative-case discipline from testing strategy in Laravel SaaS: a passing suite that never asserts a 429 will never catch a regression where every tenant quietly inherited the same default limit.
public function test_free_tier_tenant_is_throttled_past_its_limit(): void
{
$tenant = Tenant::factory()->withTier('free')->create();
$user = User::factory()->for($tenant)->create();
Http::fakeSequence()->pushResponse(status: 200)->times(60)->pushResponse(status: 429);
collect(range(1, 60))->each(fn () => $this->actingAs($user)->getJson('/api/projects'));
$this->actingAs($user)->getJson('/api/projects')->assertStatus(429);
}Real-World Pitfalls to Avoid#
A single global limit for every tenant, regardless of plan. It's the fastest thing to ship and the first thing an enterprise customer complains about — they're paying for a higher tier and getting throttled at the same ceiling as a free account.
Counting every endpoint as one request, regardless of cost. An endpoint that triggers an AI generation call or a third-party API round-trip isn't the same "one request" as a cached list read — weight them differently, or the cheap endpoints end up over-throttled while the expensive ones aren't protected enough.
Rate limiting on the app server's local memory or file cache. The moment there's more than one server behind a load balancer, the effective limit silently multiplies by server count, and it changes every time the fleet autoscales — always resolve limiters through a shared store like Redis.
A bare 429 with no Retry-After header. A client integration that can't tell how long to wait either retries immediately, making the overload worse, or gives up and files a support ticket that a header would have prevented.
Check-then-increment logic that isn't atomic. Reading a counter and writing it back as two separate operations is a race condition under real concurrency — two simultaneous requests can both pass a check that should have only let one through. Use an atomic Redis operation or Lua script, not sequential GET/SET calls.
No separate, stricter limiter for public and webhook-receiving endpoints. These are the routes most exposed to something that isn't a legitimate client at all, and they deserve tighter, IP- or signature-keyed limits distinct from the authenticated tenant traffic.
Ignoring your own outbound rate limits against third parties. Protecting your API is half the job — the other half, covered in the API integration patterns I wrote about separately, is not getting your own account rate-limited or banned by FedEx, Stripe, or OpenAI because one large customer's traffic burst went out uncapped.
Key Takeaways#
Rate limiting done well is the same story as tenant isolation, subscription billing, and authorization: a policy resolved as data, tied to who's actually asking, enforced in one consistent place — not a number picked once and pasted into a route file.
- Limits are keyed by tenant and resolved from the subscription tier, so a plan upgrade changes the ceiling automatically instead of requiring a deploy
- Expensive endpoints are weighted higher than cheap ones, because "one request" doesn't cost the same thing everywhere
- Burst and sustained limits are layered together, so a legitimate spike isn't punished the same as a runaway polling loop
- The rate limiter's store is shared across every app server, or the effective limit silently multiplies with the fleet
- 429 responses carry
Retry-AfterandX-RateLimit-*headers, so a client integration can actually recover instead of guessing - Outbound calls to third-party APIs get the same throttling discipline, so one tenant's traffic burst can't get the whole platform rate-limited by someone else's provider
I've built this into an AI product where every generation request has a real backend cost, a signage platform where hundreds of devices poll the same API on a schedule, a safety platform where enterprise customers integrate directly and expect documented limits, and logistics integrations where respecting someone else's rate limits matters as much as enforcing my own. The endpoints change; the requirement that a limit reflect who's asking, what it costs, and what they're paying for doesn't.
If your API's rate limiting is still one throttle:60,1 pasted across every route — get in touch about your API architecture or see the full case studies from platforms running tiered, tenant-aware rate limiting in production today.
Share this technical insight with your network
Share to LinkedIn or Facebook with key takeaways, featured media, and direct links.
Case Study: Violerts - Enterprise NYC PropTech Compliance & Violation Monitoring SaaS
Violerts is a PropTech SaaS platform that consolidates fragmented NYC municipal property data into a single compliance intelligence platform. I led the modernization of the React frontend and Laravel backend, building multi-agency data ingestion, GIS mapping, asynchronous scraping, real-time alerts, team collaboration, and Stripe-powered SaaS billing.
Related Technical Articles
View all articles →
Secure Multi-Tenant File Storage in Laravel SaaS
A public S3 bucket and a predictable file path is how one tenant's uploads end up in another tenant's browser tab.

Zero-Downtime Deployments: CI/CD for Laravel SaaS
git pull && composer install && php artisan migrate is not a deploy strategy — it's how live users get 500 errors mid-request.
Have a complex technical project in mind?
Available for full-stack engineering, performance audits, cloud deployments, and high-concurrency systems architecture.

