fix(supervisor): report healthy during planned process restarts

The eager restart fired after --chromium-restart-after or
--libreoffice-restart-after conversions made Healthy() report false,
so a client probing /health between two conversions got a 503 from an
otherwise serving node. Tasks arriving during that window are requeued
by acquireSlot, not rejected.

Track whether the in-flight restart is planned and keep reporting
healthy for those. Unplanned restarts still report unhealthy so load
balancers get honest information.

Closes #1648
This commit is contained in:
Julien Neuhart
2026-09-03 16:04:20 +02:00
parent 8944db131c
commit 88ddaed09b
4 changed files with 295 additions and 21 deletions

View File

@@ -1,6 +1,5 @@
# TODO:
# 1. Check if down for each module.
# 2. Restarting modules do not make health check fail.
@health
Feature: /health
@@ -106,7 +105,18 @@ Feature: /health
When I make a "HEAD" request to Gotenberg at the "/foo/health" endpoint
Then the response status code should be 200
# A planned restart, the eager cycle after LIBREOFFICE_RESTART_AFTER
# conversions, must not fail the health check: requests arriving during it
# are requeued, not rejected. Setting the limit to 1 restarts LibreOffice
# after every conversion, so each probe lands right on a restart.
# See https://github.com/gotenberg/gotenberg/issues/1648.
Scenario: GET /health (Planned LibreOffice Restart)
Given I have a Gotenberg container with the following environment variable(s):
| LIBREOFFICE_RESTART_AFTER | 1 |
When I make 5 sequential "POST" requests to Gotenberg at the "/forms/libreoffice/convert" endpoint, probing "/health" after each, with the following form data and header(s):
| files | testdata/page_1.docx | file |
Then all probe response status codes should be 200
# TODO:
# 1. Check if down for each module.
# 2. Restarting modules do not make health check fail.
# 1. Check if down for each module.