Profile    Mohammed Shiroz Status   Loading  
Logo
Share This
Back to blog
Filter by:
Tags
//Article title

Rate Limiting Your API: The Algorithms and a Laravel Example

About Post

Your API is fine for months. Then one day somebody writes a script with a loop and no sleep(). Or a mobile app release has a bug that retries a failed request instantly, forever. Or someone decides to try every password in a dictionary against your login endpoint.

Each of these looks different, but your server experiences them the same way: one client using far more than its share, and everyone else getting slower.

Rate limiting is the bouncer at the door. Not exciting until the night it really matters. Let's look at how the common algorithms work, where each one is a bit dishonest, and how to set up sensible limits in a Laravel 11 app.

What a rate limit actually needs

Every rate limiter answers one question per request: "has this client done too much recently?" That hides three decisions:

  • Who is "this client"? A user ID, an API key, an IP address, or a combination.
  • How much is too much? 60 per minute, 5 login attempts, 1,000 per day.
  • What does "recently" mean? That's where the algorithms differ.

Fixed window: simple, with a seam

Keep a counter per client per window, say per minute. Increment on every request; reject when it passes the limit; reset when the minute is over.

It's cheap and easy to understand. Its weakness is the boundary. With a limit of 60 per minute, a client can send 60 requests in the last second of one window and 60 more in the first second of the next. That's 120 requests in two seconds while technically never breaking the rule.

For most APIs that's an acceptable imperfection. For a login endpoint or an expensive report, it might not be.

Sliding window: smoothing the seam

A sliding window asks "how many requests in the last 60 seconds?" at every moment, instead of per calendar minute. There are two common versions:

  • Sliding log: store the timestamp of every request and count the ones newer than 60 seconds. Exact, but it stores a lot for busy clients.
  • Sliding window counter: keep only this window's and last window's counts, and estimate. If we're 25% into the current minute, count = current + 75% of the previous minute. Cheap, and close enough in practice.

Token bucket: allow bursts, cap the average

Picture a bucket that holds 10 tokens and gets refilled at one token per second. Each request takes a token. No token, no request.

This models real usage nicely. A mobile app that opens and fires eight requests at once is fine, because the bucket was full. A client that keeps hammering is limited to the refill rate. You get two dials: burst size (bucket capacity) and sustained rate (refill speed).

// Simplified token bucket. Not atomic: real versions do this in one
// Redis Lua script so two requests can't read the same state.
function allowRequest(string $key, int $capacity, float $perSecond): bool
{
    $now = microtime(true);
    $bucket = Cache::get($key, ['tokens' => $capacity, 'at' => $now]);

    $tokens = min($capacity, $bucket['tokens'] + ($now - $bucket['at']) * $perSecond);
    $allowed = $tokens >= 1;

    Cache::put($key, ['tokens' => $allowed ? $tokens - 1 : $tokens, 'at' => $now], 3600);

    return $allowed;
}

A close relative, the leaky bucket, processes requests at a steady rate from a queue. Same idea, but it smooths traffic out instead of allowing bursts.

AlgorithmStrengthWeakness
Fixed windowSimplest, cheapestDouble burst at window edges
Sliding logExactMemory grows with traffic
Sliding counterCheap and smoothAn estimate, not exact
Token bucketAllows natural burstsTwo parameters to tune

Rate limiting in Laravel

Laravel's built-in limiter is a counter with a decay window that starts at the client's first hit, so it behaves like a fixed window. That's plenty for most apps. You define named limiters in the boot method of AppServiceProvider:

use Illuminate\Cache\RateLimiting\Limit;
use Illuminate\Support\Facades\RateLimiter;

RateLimiter::for('api', function (Request $request) {
    return $request->user()
        ? Limit::perMinute(120)->by('user:'.$request->user()->id)
        : Limit::perMinute(30)->by('ip:'.$request->ip());
});

RateLimiter::for('login', function (Request $request) {
    return [
        Limit::perMinute(5)->by('login:'.$request->input('email').'|'.$request->ip()),
        Limit::perMinute(20)->by('login-ip:'.$request->ip()),
    ];
});

Then attach them to routes with the throttle middleware:

Route::post('/login', [AuthController::class, 'login'])->middleware('throttle:login');

Route::middleware(['auth:sanctum', 'throttle:api'])->group(function () {
    // ...
});

The login limiter returns two limits: a tight one per email and IP (to slow down guessing one account) and a looser one per IP (to slow down one machine trying many accounts). Prefixing the keys keeps the counters separate.

For limits inside your own code, like "three OTP texts per user per ten minutes", there's RateLimiter::attempt() and RateLimiter::tooManyAttempts(), which work with any key you choose.

Per user or per IP?

Both, for different reasons. Per user is fairer for logged-in traffic: an office full of people behind one IP shouldn't share one limit. Per IP is all you have before login, which is exactly where abuse tends to happen.

Two production gotchas here:

  • Behind a load balancer, $request->ip() may return the balancer's address for every request, so everyone shares one bucket. Configure trusted proxies (in Laravel 11, $middleware->trustProxies(...) in bootstrap/app.php) so the real client IP comes through.
  • With several app servers, the limiter's cache must be shared, such as Redis. A file cache gives each server its own counter, quietly multiplying your limit.

Say no politely: 429 and Retry-After

When a client is over the limit, respond with 429 Too Many Requests and a Retry-After header saying how many seconds to wait. Laravel's throttle middleware does this for you and adds X-RateLimit-Limit and X-RateLimit-Remaining headers too.

The other half is on the client. A well-behaved app reads Retry-After and waits, instead of retrying instantly and digging the hole deeper. If you build the mobile app too, that's worth a few lines of code.

Start here: a generous per-user limit on the whole API, strict limits on login, password reset and OTP endpoints, and a shared cache store. Tune from your logs, not from guesses.

The short version

  • Fixed window is simple and fine for most endpoints.
  • Sliding windows smooth out the edge bursts.
  • Token bucket allows natural bursts while capping the average.
  • Key logged-in traffic by user, anonymous traffic by IP, and trust your proxies correctly.
  • Answer with 429 and Retry-After, and make your clients respect it.

The Laravel rate limiting docs cover the remaining options, including custom responses.

Which endpoint in your app would you rate limit first if you had to pick just one?

Comments (0)
Leave your review

Thanks for your valuable comments. Your comments has been updated and appreciate your getting in touch...

01. About Shiroz

Mohammed Shiroz

Hi, I'm Mohammed Shiroz, a software engineer and AI enthusiast from Sri Lanka who turns ideas into intelligent, real-world solutions. With over 9 years of hands-on experience, I currently lead real estate ERP development at Kate Group, a...

03.My Projects

04. Categories

Ready To order Your Project ?

Get in Touch
Close