Profile    Mohammed Shiroz Status   Loading  
Logo
Share This
Back to blog
Filter by:
Tags
//Article title

Laravel Horizon in Production: Queue Monitoring That Actually Tells You Something

About Post

Queues fail quietly. That's their whole personality.

A web request that breaks shows an error page, and someone complains within minutes. A queued job that breaks just... doesn't happen. The welcome email never arrives. The PDF is never generated. The payment webhook sits unprocessed. And you find out days later from a confused user, not from your monitoring.

If your Laravel app runs on Redis queues, Horizon is the tool that turns that silence into something you can see. Installing it takes minutes. Setting it up so it actually protects you in production takes a little more thought, so here's the setup I'd use, organised around the four questions you need answered at any moment.

Question 1: are the workers actually running?

Horizon is a long-running process. It starts and supervises your queue workers, so if Horizon itself dies, everything stops. The first job is making sure something restarts it.

On a Linux server that's usually Supervisor (the process manager, not to be confused with Horizon's own "supervisors"):

[program:horizon]
process_name=%(program_name)s
command=php /var/www/app/artisan horizon
autostart=true
autorestart=true
user=www-data
redirect_stderr=true
stdout_logfile=/var/www/app/storage/logs/horizon.log
stopwaitsecs=3600

stopwaitsecs should be longer than your longest job, so a restart waits for running jobs to finish instead of killing them halfway.

Then, in your deploy script, after the new code is in place:

php artisan horizon:terminate

Workers keep your application in memory. Without this line, they keep running yesterday's code after a deploy, which produces some of the most confusing bugs you'll ever chase. horizon:terminate lets them finish their current job and exit, and Supervisor starts fresh ones.

For an external health check, php artisan horizon:status tells you whether Horizon is running, paused or inactive.

Question 2: are the queues keeping up?

This is what config/horizon.php is for. The part that matters most is the supervisor configuration per environment. A setup I like separates urgent work from slow work, so a batch of heavy reports can never delay a payment confirmation:

'environments' => [
    'production' => [
        'critical' => [
            'connection' => 'redis',
            'queue' => ['payments', 'notifications'],
            'balance' => 'auto',
            'minProcesses' => 2,
            'maxProcesses' => 8,
            'tries' => 3,
            'timeout' => 30,
        ],
        'background' => [
            'connection' => 'redis',
            'queue' => ['default', 'reports'],
            'balance' => 'auto',
            'maxProcesses' => 4,
            'tries' => 2,
            'timeout' => 300,
        ],
    ],
],

A few settings worth understanding rather than copying:

  • balance: auto moves worker processes towards the busiest queues, simple splits them evenly, and false works through queues in the order listed.
  • minProcesses / maxProcesses: the floor and ceiling for auto-balancing. The minimum keeps critical queues warm even when they're quiet.
  • timeout: must be shorter than the retry_after value of your Redis connection in config/queue.php. If it isn't, a slow job can be picked up a second time while the first attempt is still running. That's a classic source of duplicate emails and double processing.

Then tell Horizon what "falling behind" means for you:

'waits' => [
    'redis:payments' => 30,
    'redis:default' => 120,
],

If a job waits longer than that number of seconds, Horizon fires a LongWaitDetected event. On its own, nobody sees it. Which brings us to notifications.

Question 3: will someone hear about it?

A dashboard is only useful if someone is looking at it. Route Horizon's alerts somewhere people already look, in the boot method of HorizonServiceProvider:

Horizon::routeMailNotificationsTo('[email protected]');
Horizon::routeSlackNotificationsTo(config('services.slack.horizon_webhook'), '#alerts');

And turn on the metrics graphs, which need a regular snapshot. In routes/console.php:

Schedule::command('horizon:snapshot')->everyFiveMinutes();

Without this, the Metrics tab stays empty, and it's one of the most useful screens: throughput and runtime per job and per queue. A job whose runtime creeps up week after week is telling you something about a growing table or a slow external API long before it times out.

Question 4: what failed, and why?

The Failed Jobs screen shows the exception, the stack trace and the job's payload, and lets you retry. That retry button is wonderful and slightly dangerous: retrying a job that isn't idempotent can send the same notification twice. Know which of your jobs are safe to replay before you click it in a hurry.

Two features make this screen much more useful:

  • Tags. Horizon automatically tags jobs with the Eloquent models they receive, like App\Models\Invoice:4821. When a user reports a problem with one record, search the tag and see every job that touched it.
  • Retention. The trim section of the config controls how long recent, completed and failed jobs are kept (in minutes). Keep failed jobs long enough to survive a weekend.

The production rule: every queue needs an owner, a "too slow" threshold and an alert that reaches a human. A dashboard nobody opens is just a nicer way to fail silently.

Lock the door

Horizon's dashboard shows job payloads, which can include emails, names and IDs. Outside the local environment, access is controlled by the viewHorizon gate in HorizonServiceProvider:

protected function gate(): void
{
    Gate::define('viewHorizon', function ($user) {
        return $user->hasRole('developer');
    });
}

(Adapt hasRole to however your app handles roles.) Also think about what you put in job payloads in the first place. Pass IDs, not whole documents or secrets.

Your Horizon checklist

  • Supervisor restarts Horizon; deploys run horizon:terminate.
  • Urgent and slow queues run under separate supervisors.
  • Every timeout is below the connection's retry_after.
  • waits thresholds are set and alerts go to mail or Slack.
  • horizon:snapshot is scheduled, so metrics exist.
  • The dashboard is behind the viewHorizon gate.
  • You know which jobs are safe to retry.

Horizon only works with Redis queues. If you're on the database or SQS driver, the same four questions still apply; you'll just answer them with other tools.

What's the first metric you look at when a queue starts misbehaving: wait time, failures or job runtime?

Comments (0)
Leave your review

Thanks for your valuable comments. Your comments has been updated and appreciate your getting in touch...

01. About Shiroz

Mohammed Shiroz

Hi, I'm Mohammed Shiroz, a software engineer and AI enthusiast from Sri Lanka who turns ideas into intelligent, real-world solutions. With over 9 years of hands-on experience, I currently lead real estate ERP development at Kate Group, a...

03.My Projects

04. Categories

Ready To order Your Project ?

Get in Touch
Close