If your Flask app is deployed behind Gunicorn, you’ve probably seen this in your logs at some point: [CRITICAL] WORKER TIMEOUT, followed by the worker getting killed and restarted mid-request. It usually shows up under real production load, not in local testing, which makes it particularly annoying to track down.

Quick Answer: A Gunicorn worker timeout means a worker process didn’t respond to Gunicorn’s internal heartbeat within the configured timeout (default 30 seconds), so Gunicorn assumes it’s stuck and kills it. The most common causes are slow database queries, blocking network calls (external APIs, SMTP, etc.), and CPU-heavy work running on Gunicorn’s default sync worker class, which can only handle one request at a time per worker. Fix it by finding what’s actually blocking, then either speeding it up, moving it to a background task, or switching to an async-capable worker class like gevent.

How Gunicorn’s Worker Timeout Actually Works

Gunicorn runs a master process and a pool of worker processes. Each worker handles requests, and periodically it needs to check in with the master to say “I’m still alive.” That check-in happens between requests when you’re using the default sync worker — meaning if a single request takes longer than the timeout, the worker never gets a chance to report back, and the master assumes it has hung.

By default, that timeout is 30 seconds. You’ll set it explicitly in your Gunicorn command or config:

gunicorn app:app --workers 4 --timeout 30

When the master doesn’t hear from a worker in time, it sends SIGKILL to that worker and spins up a replacement. Any request that worker was handling gets dropped — the client sees a connection reset or a 502 from whatever’s in front of Gunicorn (nginx, an ALB, etc.). You’ll see something like this in your Gunicorn logs:

[2026-09-28 14:32:07 +0000] [1842] [CRITICAL] WORKER TIMEOUT (pid:1855)
[2026-09-28 14:32:07 +0000] [1855] [INFO] Worker exiting (pid: 1855)
[2026-09-28 14:32:08 +0000] [1901] [INFO] Booting worker with pid: 1901

This is the important part to understand: the timeout isn’t Flask’s timeout, it’s Gunicorn’s. Flask (via Werkzeug) has no idea any of this is happening — it just processes the request until it’s told to stop. The kill is entirely Gunicorn’s doing, based on whether the worker process itself is responsive at the OS level.

Diagnosing Which Request Is Actually Stuck

Before you touch any code, figure out what the worker was doing when it got killed. The WORKER TIMEOUT log line alone won’t tell you — you need a bit more context.

Turn on Gunicorn’s access logs with response times, if you haven’t already:

gunicorn app:app --workers 4 --timeout 30 \
  --access-logfile - \
  --access-logformat '%(t)s "%(r)s" %(s)s %(D)s'

%(D)s is the request duration in microseconds. Grep for the slowest entries and you’ll usually spot a pattern — one specific route, one specific customer, or a spike that lines up with a deploy or a traffic surge.

If the worker is already stuck and you need a live snapshot, py-spy can attach to a running process without restarting it:

py-spy dump --pid 1855

That gives you a Python stack trace of exactly where the worker is stuck — waiting on a socket read, a database cursor, a lock, whatever it is — without needing to reproduce the issue locally. It’s often the fastest way to go from “something is timing out” to “line 42 of reports.py is waiting on cursor.execute().”

Once you know which endpoint and which line are responsible, you’re usually looking at one of the three patterns below.

Common Pitfalls

A few patterns show up over and over when people hit this in production:

  • Assuming it’s a Flask bug. It’s almost never Flask itself — it’s something the request handler is waiting on.
  • Bumping the timeout instead of fixing the cause. Setting --timeout 120 makes the symptom go away for a while, but it also means a genuinely broken dependency (say, a downstream API that’s hanging) now ties up a worker for two minutes instead of thirty seconds, making the problem worse under load.
  • Not knowing which worker class is in use. With the default sync worker, one slow request blocks that entire worker — no other request can use it until it finishes. People are often surprised their “4 workers” setup falls over with only a handful of slow concurrent requests.
  • Running expensive work synchronously inside the request/response cycle — image processing, PDF generation, large CSV exports — instead of offloading it.
  • Ignoring --graceful-timeout and --keep-alive, which interact with --timeout in ways that aren’t obvious until you dig into Gunicorn’s docs.

Real-World Examples

Case 1: A Slow Database Query Under Load

This is the classic one. A query that returns in 200ms with an empty table takes 8 seconds once the table has a few million rows and the query is missing an index.

# routes/reports.py
@app.route("/api/reports/monthly")
def monthly_report():
    # No index on created_at — full table scan under load
    orders = Order.query.filter(
        Order.created_at >= start_of_month()
    ).all()
    return jsonify([o.to_dict() for o in orders])

Locally, with a small dataset, this responds instantly. In production, once the orders table grows and a few of these requests land concurrently, each one competes for database connections and CPU, and response times creep past the 30-second timeout.

The fix — add the missing index, and don’t fetch more than you need:

# migrations/versions/xxxx_add_orders_created_at_index.py
def upgrade():
    op.create_index(
        "ix_orders_created_at", "orders", ["created_at"]
    )
@app.route("/api/reports/monthly")
def monthly_report():
    orders = (
        Order.query.filter(Order.created_at >= start_of_month())
        .limit(500)
        .all()
    )
    return jsonify([o.to_dict() for o in orders])

If you’re not sure whether the database is the bottleneck, turn on SQLAlchemy’s query logging temporarily (SQLALCHEMY_ECHO=True or echo=True on create_engine) and watch for queries that take longer than a second or two under real traffic.

Case 2: A Blocking Call to a Third-Party API

Outbound HTTP calls without a timeout are a frequent cause, especially when the third party is having a bad day.

# ❌ Before: no timeout, worker hangs if the API stalls
@app.route("/api/verify-address")
def verify_address():
    response = requests.post(
        "https://address-validator.example.com/verify",
        json=request.get_json(),
    )
    return jsonify(response.json())

If address-validator.example.com starts hanging (not erroring — hanging, which is worse), requests will wait indefinitely by default. That request sits there until Gunicorn’s timeout kills the worker.

# ✅ After: explicit timeout + graceful handling
@app.route("/api/verify-address")
def verify_address():
    try:
        response = requests.post(
            "https://address-validator.example.com/verify",
            json=request.get_json(),
            timeout=(3, 5),  # (connect timeout, read timeout)
        )
        response.raise_for_status()
    except requests.exceptions.RequestException:
        return jsonify({"error": "address validation unavailable"}), 502
    return jsonify(response.json())

Setting timeout on every outbound call is one of the highest-leverage things you can do to prevent worker timeouts. Make it a habit — a request library call without a timeout is a liability waiting to happen.

Case 3: CPU-Heavy Work on a Sync Worker

Image resizing, PDF generation, and large data transforms don’t hang, exactly — they just take a while, and while they’re running, that worker can’t do anything else.

# This runs synchronously inside the request — fine at low traffic,
# a bottleneck the moment two of these land at once
@app.route("/api/thumbnails", methods=["POST"])
def generate_thumbnail():
    image = Image.open(request.files["photo"])
    image.thumbnail((200, 200))
    resized = save_to_storage(image)
    return jsonify({"url": resized.url})

If thumbnail generation takes 4 seconds and you only have 4 workers, five concurrent uploads already means someone’s waiting behind a worker that’s still busy with someone else’s image — and under enough concurrency, you’ll eventually cross the timeout threshold too.

The fix — move it to a background job (Celery, RQ, or similar) and respond immediately:

@app.route("/api/thumbnails", methods=["POST"])
def generate_thumbnail():
    file_path = save_upload(request.files["photo"])
    task = generate_thumbnail_task.delay(file_path)
    return jsonify({"task_id": task.id}), 202

The request now returns in milliseconds, and the actual work happens outside Gunicorn’s timeout window entirely. If you’re new to background tasks in Flask, our post on Flask Celery tasks stuck in pending covers the setup and common gotchas.

Advanced Tips

Switch worker classes if your app is I/O-bound. The sync worker is simple and safe, but if most of your slow time is spent waiting on network I/O (databases, APIs, external services), a gevent or eventlet worker can handle many concurrent connections per process instead of blocking one worker per request:

pip install gevent
gunicorn app:app --worker-class gevent --workers 4 --worker-connections 100 --timeout 30

This isn’t a silver bullet — CPU-bound work still blocks the event loop under gevent — but for apps that are mostly waiting on the database or downstream APIs, it can dramatically increase how much concurrent traffic the same number of workers can absorb.

Set --max-requests to guard against slow memory leaks. Long-running workers can accumulate memory over time (a common issue with certain C extensions or connection pools). Recycling workers periodically limits the blast radius:

gunicorn app:app --workers 4 --timeout 30 --max-requests 1000 --max-requests-jitter 50

Distinguish --timeout from --graceful-timeout. --timeout is the hard kill threshold discussed above. --graceful-timeout controls how long a worker gets to finish up during a graceful restart (e.g., during a deploy) before it’s forced to stop. Confusing the two leads to either deploys that hang or deploys that cut off in-flight requests unnecessarily.

Add timing middleware to catch slow endpoints before they become timeouts. A simple before_request/after_request pair logging duration gives you visibility without needing a full APM setup:

import time
from flask import g

@app.before_request
def start_timer():
    g.start_time = time.monotonic()

@app.after_request
def log_slow_requests(response):
    duration = time.monotonic() - g.start_time
    if duration > 5:
        app.logger.warning(
            f"Slow request: {request.path} took {duration:.2f}s"
        )
    return response

This won’t fix anything on its own, but it turns “the app randomly times out sometimes” into “endpoint X takes 6+ seconds when Y happens” — which is a much easier problem to solve.

Key Takeaways

  • A Gunicorn worker timeout means a worker didn’t check in within the timeout window (default 30s) and got killed — it’s Gunicorn’s doing, not a Flask error.
  • Don’t just raise the timeout value; find what’s actually slow first, or you’ll just make failures take longer to surface.
  • Always set explicit timeouts on outbound HTTP calls (requests, database drivers, third-party SDKs) — a hanging dependency is one of the most common root causes.
  • Move genuinely slow work (image processing, report generation, bulk exports) out of the request/response cycle and into a background task queue.
  • Consider gevent/eventlet workers if your bottleneck is I/O-bound waiting rather than raw CPU work.
  • Add lightweight request timing logging so you catch slow endpoints before they escalate into timeouts under load.

If your Gunicorn logs are full of long, noisy tracebacks alongside those WORKER TIMEOUT messages, use Debugly’s trace formatter to quickly parse and analyze Python tracebacks so you can pinpoint the exact line that’s stalling. If you’re running a similar setup with FastAPI and Uvicorn, our guide on Uvicorn worker timeouts walks through the equivalent issue for that stack.