Compare commits

..

79 Commits

Author SHA1 Message Date
Julien Neuhart
430ed38cc5 build(makefile): re-enable the embedlit go fix modernizer 2026-09-11 16:13:06 +02:00
Julien Neuhart
623e93bf3e chore(deps): update Go to 1.27.1 2026-09-11 16:13:06 +02:00
dependabot[bot]
fe4fb9416d chore(deps): bump golang.org/x/sync from 0.22.0 to 0.23.0 (#1658)
Bumps [golang.org/x/sync](https://github.com/golang/sync) from 0.22.0 to 0.23.0.
- [Commits](https://github.com/golang/sync/compare/v0.22.0...v0.23.0)

---
updated-dependencies:
- dependency-name: golang.org/x/sync
  dependency-version: 0.23.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-11 07:20:22 +02:00
Julien Neuhart
fb7536a526 fix(chromium): disable the WebUI omnibox popup preloaded at browser start 2026-09-09 19:38:40 +02:00
Julien Neuhart
b16ce08da7 perf(chromium): skip formatting a discarded debug message per response 2026-09-08 18:35:03 +02:00
Julien Neuhart
fcfd590169 perf(api): build the access log with a single record 2026-09-08 18:35:03 +02:00
Julien Neuhart
b39b8c76aa perf(gotenberg): match allow and deny patterns without recompiling them 2026-09-08 18:34:58 +02:00
Julien Neuhart
06ed58b6e7 ci(build): reuse the Docker build cache across runs 2026-09-08 18:34:54 +02:00
Julien Neuhart
34b7b4845e perf(pdfengines): process a request's files concurrently behind an opt-in ceiling 2026-09-08 18:34:51 +02:00
Julien Neuhart
ab3832d9e5 test(integration): drop the per-scenario Docker network 2026-09-08 18:34:43 +02:00
Julien Neuhart
cddaa0fa57 fix(chromium): ceiling the CONNECT tunnels in flight and bound tunnel writes 2026-09-07 18:13:29 +02:00
Julien Neuhart
ade6a327a4 fix(chromium): drop settled requests from the per-conversion network map 2026-09-07 18:13:29 +02:00
Julien Neuhart
35ddf81812 fix(libreoffice): stop the outbound proxy when the daemon fails to start 2026-09-07 18:13:29 +02:00
Julien Neuhart
1891a9ea68 fix(chromium): bound CONNECT tunnels so a silent upstream cannot pin them forever 2026-09-07 18:13:29 +02:00
Julien Neuhart
e1e0a80883 fix(webhook): close the response body when the callback returns an error status 2026-09-07 18:13:29 +02:00
Max Freedom Pollard
c21ceacd4b fix(libreoffice): report an encrypted .xlsb as password-protected (#1655)
Uploading a password-protected .xlsb workbook without its password
answered 500 with the unattributable-failure message instead of 400 with
the remedy.

DetectPasswordProtection in pkg/modules/libreoffice/api/protection.go
infers encryption from a compound-file header carried by an extension
whose unencrypted form is always a ZIP package. The ooxmlExtensions map
listed .xlsx, .xlsm, .xltx and .xltm but not .xlsb, so an encrypted
workbook under that extension fell through to PasswordProtectionUnknown.
The convert route in pkg/modules/libreoffice/routes.go then matched
neither password branch of its exit-code switch and returned the 500
default.

An Excel Binary Workbook is an Open Packaging Conventions ZIP holding
binary parts, so a compound file under that extension is encrypted for
the same reason .xlsx is. Adding .xlsb to the map restores the 400 that
names the 'password' form field.
2026-09-06 13:05:20 +02:00
Julien Neuhart
ac825a2c03 feat(api): warn at startup when the debug route has no authentication 2026-09-05 14:04:21 +02:00
Julien Neuhart
2c9fa6b6ed fix(exiftool): reject metadata keys that collide with ExifTool options 2026-09-05 14:04:21 +02:00
Julien Neuhart
ca8b45cd3a fix(api): keep the filter verdict generic when a redirect is blocked 2026-09-05 14:04:21 +02:00
Julien Neuhart
40cf48442f fix(outbound): treat CGNAT and benchmarking ranges as non-public 2026-09-05 11:34:03 +02:00
Julien Neuhart
df3bac99ed fix(api): keep upload order when de-duplicating repeated filenames 2026-09-05 11:26:10 +02:00
Julien Neuhart
17868b8c02 fix(api): stop upload filenames from failing or silently dropping a request 2026-09-05 10:07:47 +02:00
Julien Neuhart
78284df590 fix(api): snapshot the output filename before echo recycles its context 2026-09-05 10:05:12 +02:00
Julien Neuhart
8f415186d5 fix(api): bound downloadFrom decoding by the entry limit 2026-09-05 10:03:56 +02:00
Julien Neuhart
f675f78f77 fix(webhook): bound webhook delivery with its own retry budget 2026-09-05 10:01:29 +02:00
Julien Neuhart
83b01c2baa fix(api): bound downloadFrom requests by the request deadline 2026-09-05 09:59:29 +02:00
Julien Neuhart
9a46fdd681 style: escape backslashes in the Chromium allow-list scenario table 2026-09-04 20:00:22 +02:00
Julien Neuhart
4de9b0f68b docs(outbound): document the allow-list bypass on every allow-list flag and terminate the example patterns 2026-09-04 19:55:57 +02:00
Julien Neuhart
201e80b9d7 feat(outbound): warn at startup about allow-list patterns that grant more than intended 2026-09-04 19:52:41 +02:00
Julien Neuhart
86a013b664 fix(outbound): strip URL userinfo before allow and deny list matching 2026-09-04 19:45:58 +02:00
Julien Neuhart
9c5acd7418 chore(deps): update Go dependencies 2026-09-04 10:00:28 +02:00
Julien Neuhart
e29b7cb4f5 build(makefile): fix lint-todo for golangci-lint v2 2026-09-03 19:23:14 +02:00
Julien Neuhart
23d59f3133 refactor: adopt strings.Cut and strings.SplitSeq
Applies what go fix now proposes, so make fmt is a no-op on a clean tree
instead of dirtying these two files on every run. Both rewrites are
equivalent: Cut's first result matches SplitN(s, sep, 2)[0], and SplitSeq
iterates the same substrings without building the intermediate slice.
2026-09-03 19:22:23 +02:00
Julien Neuhart
8a0de7d5d7 build(makefile): disable two go fix modernizers that emit broken code
go fix runs the modernize suite since Go 1.26, and two of its analyzers
rewrite this codebase into code that does not compile, so make fmt broke
the build and then make lint. CONTRIBUTING tells every contributor to run
both before opening a PR.

embedlit folded telemetryCfg.LogLevel into a literal that already set that
key, and errorsastype rewrote errors.As to errors.AsType[HttpError] even
though HttpError does not embed error. Disable both, with a TODO covering
when each can come back.
2026-09-03 19:22:23 +02:00
Julien Neuhart
3c691cbebc test(health): assert conversions succeed across a planned restart 2026-09-03 17:55:35 +02:00
Julien Neuhart
0e83f737b4 docs(supervisor): point maybeRestartAfterTask at the timeout constant 2026-09-03 17:52:52 +02:00
Julien Neuhart
57b048c611 fix(supervisor): reset the request counter on a failed relaunch
restart() reset reqCounter only after a successful Launch, so a failed
one left it at the limit. maybeRestartAfterTask then re-fired on every
subsequent task, producing back-to-back restarts and, with planned
restarts now reporting healthy, a node that keeps restarting while
claiming health.

Reset on the attempt instead. A process that will not start is recovered
by ensureHealthy, which restarts synchronously and reports the failure.
2026-09-03 17:52:01 +02:00
Julien Neuhart
d79e174c6f fix(supervisor): bound the eager restart with a deadline
maybeRestartAfterTask ran its restart on a bare background context while
the drain loop in doRestartLocked selects only on ctx.Done(). A task that
never completed blocked the drain forever, pinning isRestarting. Since a
planned restart now reports healthy, that left the node claiming health
for good. LibreOffice escaped it because its concurrency of 1 drains no
slots, but Chromium defaults to 6.

Give that restart its own deadline. On expiry it aborts and the next task
retries it.
2026-09-03 16:10:16 +02:00
Julien Neuhart
88ddaed09b fix(supervisor): report healthy during planned process restarts
The eager restart fired after --chromium-restart-after or
--libreoffice-restart-after conversions made Healthy() report false,
so a client probing /health between two conversions got a 503 from an
otherwise serving node. Tasks arriving during that window are requeued
by acquireSlot, not rejected.

Track whether the in-flight restart is planned and keep reporting
healthy for those. Unplanned restarts still report unhealthy so load
balancers get honest information.

Closes #1648
2026-09-03 16:04:20 +02:00
Julien Neuhart
8944db131c fix(api): bound downloadFrom concurrency and entry count 2026-09-02 14:57:29 +02:00
Johannes Klein
676570074a fix(Dockerfile): re-add libreoffice-math so DOCX math formulas render
The v8.30.0 image slimming replaced the libreoffice metapackage with an
explicit component list that omits libreoffice-math. Writer imports OMML
equations as Math objects, so without the component every formula is
silently dropped from converted PDFs.

Fixes #1644
2026-09-02 11:55:30 +02:00
dependabot[bot]
11179cb271 chore(deps): bump google.golang.org/grpc from 1.83.0 to 1.83.1 (#1646)
Bumps [google.golang.org/grpc](https://github.com/grpc/grpc-go) from 1.83.0 to 1.83.1.
- [Release notes](https://github.com/grpc/grpc-go/releases)
- [Commits](https://github.com/grpc/grpc-go/compare/v1.83.0...v1.83.1)

---
updated-dependencies:
- dependency-name: google.golang.org/grpc
  dependency-version: 1.83.1
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-02 10:51:17 +02:00
Julien Neuhart
c2a85f92d0 chore(deps): update unoconverter to v0.5.0 2026-08-31 19:01:18 +02:00
Muhammad Haseeb
7dbff18e65 fix(chromium): fail fast with 503 when Chromium crashes (#1641) 2026-08-31 16:48:21 +02:00
Julien Neuhart
923e5f71eb style(api): remove redundant parentheses in type switch 2026-08-31 13:38:05 +02:00
Julien Neuhart
334f859d95 chore(deps): update golangci-lint to v2.13.2 2026-08-31 13:36:02 +02:00
Julien Neuhart
0819514b7b chore(deps): update Go to 1.27.0 2026-08-31 13:36:02 +02:00
Julien Neuhart
c636a52666 fix(otel): keep telemetry off unless an exporter is configured 2026-08-31 13:27:05 +02:00
dependabot[bot]
0b16b0ab34 chore(deps): bump github.com/stretchr/testify from 1.11.1 to 1.12.1 (#1636)
Bumps [github.com/stretchr/testify](https://github.com/stretchr/testify) from 1.11.1 to 1.12.1.
- [Release notes](https://github.com/stretchr/testify/releases)
- [Commits](https://github.com/stretchr/testify/compare/v1.11.1...v1.12.1)

---
updated-dependencies:
- dependency-name: github.com/stretchr/testify
  dependency-version: 1.12.1
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-21 11:57:55 +02:00
Julien Neuhart
c0f487e333 feat(api): add OIDC bearer token authentication 2026-08-14 17:04:20 +02:00
Julien Neuhart
3de6932279 feat(pdfengines): apply multiple stamps and watermarks on every route 2026-08-14 16:36:40 +02:00
Julien Neuhart
e9a67132ec feat(pdfengines): apply multiple stamps in one request (#1601) 2026-08-14 15:52:14 +02:00
Julien Neuhart
41b33fd6ad test(integration): allow repeated form fields in requests 2026-08-14 15:52:14 +02:00
Julien Neuhart
2521485bf8 test(libreoffice): add .ppsm and .ppsx to the Bad Request extension list 2026-08-14 15:20:46 +02:00
Julien Neuhart
65e5699b71 feat(pdfengines): generate document-title bookmarks when merging (#867) 2026-08-14 12:36:13 +02:00
Julien Neuhart
b3c06fb8ae test(chromium): cover thenable waitForExpression 2026-08-14 12:00:57 +02:00
Julien Neuhart
f90684e165 fix(chromium): add missing import for awaited waitForExpression 2026-08-14 11:58:19 +02:00
Niklas
90e614afa4 feat(chromium): await thenable waitForExpression (#1617) 2026-08-14 11:57:29 +02:00
Julien Neuhart
6a087cb1f7 test(libreoffice): cover OOXML PowerPoint show conversion (#1626) 2026-08-14 11:49:46 +02:00
Ilya Fedorov
38db892552 fix(libreoffice): accept the OOXML PowerPoint show extensions (#1626) 2026-08-14 11:43:12 +02:00
Julien Neuhart
ca9603bdf4 chore(test): format optimize-images testdata and integration README 2026-08-14 11:07:16 +02:00
Julien Neuhart
0c1e82c885 chore(deps): update Go and npm dependencies 2026-08-14 11:07:16 +02:00
Julien Neuhart
1303e0ebc9 refactor(pdfengines): remove qpdf embed fallback (#1628) 2026-08-14 11:07:16 +02:00
Julien Neuhart
cc1341c3cf chore(deps): update pdfcpu to v0.15.0 (#1628) 2026-08-13 20:39:11 +02:00
Julien Neuhart
1ac1d9887e feat(pdfengines): optimize PDF images to reduce file size (#359) 2026-08-13 19:39:09 +02:00
Julien Neuhart
7a730cfdc2 feat(chromium): add --chromium-clear-storage to clear local storage between conversions (#919) 2026-08-13 14:56:46 +02:00
Julien Neuhart
b213f2ffed fix(chromium): fit the page to content for landscape single-page (#1390) 2026-08-13 14:14:48 +02:00
Julien Neuhart
8d1eeaa73a feat(chromium): screenshot a single element via the selector form field (#947) 2026-08-13 13:56:17 +02:00
Julien Neuhart
2f6818d3df fix(chromium): hint at --chromium-start-timeout when Chromium fails to start 2026-08-13 12:47:45 +02:00
Julien Neuhart
38f6e466d4 chore(libreoffice): drop the now-unneeded gosec nolint 2026-08-13 12:32:55 +02:00
Julien Neuhart
e39840d5cf ci: bump golangci-lint to v2.12.2 2026-08-13 12:31:28 +02:00
Julien Neuhart
9b48d3f84e chore(libreoffice): silence a gosec G703 false positive on temp-file cleanup 2026-08-13 12:28:46 +02:00
Julien Neuhart
05bde96334 refactor(qpdf): make embed-metadata logging context-aware 2026-08-13 08:21:35 +02:00
Julien Neuhart
8b2c15d5de fix(api): classify client-cancelled requests as 499 instead of 500 2026-08-13 08:21:35 +02:00
Julien Neuhart
8d327a5196 fix(libreoffice): render scrolled workbooks in full for SinglePageSheets 2026-08-13 08:21:35 +02:00
Julien Neuhart
dc61c3631e test(libreoffice): rename the linked-content tag to libreoffice-ssrf 2026-08-13 08:21:35 +02:00
Julien Neuhart
dc7c68152c fix(chromium): filter WebSocket handshakes against the outbound policy 2026-08-13 08:21:35 +02:00
Julien Neuhart
357c3b4a59 fix(outbound): return 403 instead of 500 for an unresolvable outbound host 2026-08-13 08:21:35 +02:00
Julien Neuhart
91f2587fef Merge commit from fork
Co-authored-by: Jacob Brackett <jbrackett@makenotion.com>
2026-08-13 08:08:28 +02:00
115 changed files with 8488 additions and 1036 deletions

View File

@@ -47,6 +47,8 @@ body:multipart-form {
~splitUnify: false
~pdfa: PDF/A-1b
~pdfua: true
~optimizeImages: false
~imageQuality: 80
~metadata: {"Author":"Bruno","Title":"Test"}
~userPassword:
~ownerPassword:

View File

@@ -48,6 +48,8 @@ body:multipart-form {
~splitUnify: false
~pdfa: PDF/A-1b
~pdfua: true
~optimizeImages: false
~imageQuality: 80
~metadata: {"Author":"Bruno","Title":"Test"}
~userPassword:
~ownerPassword:

View File

@@ -47,6 +47,8 @@ body:multipart-form {
~splitUnify: false
~pdfa: PDF/A-1b
~pdfua: true
~optimizeImages: false
~imageQuality: 80
~metadata: {"Author":"Bruno","Title":"Test"}
~userPassword:
~ownerPassword:

View File

@@ -15,6 +15,7 @@ body:multipart-form {
~width: 800
~height: 600
~clip: false
~selector:
~format: png
~quality: 100
~optimizeForSpeed: false

View File

@@ -16,6 +16,7 @@ body:multipart-form {
~width: 800
~height: 600
~clip: false
~selector:
~format: png
~quality: 100
~optimizeForSpeed: false

View File

@@ -15,6 +15,7 @@ body:multipart-form {
~width: 800
~height: 600
~clip: false
~selector:
~format: png
~quality: 100
~optimizeForSpeed: false

View File

@@ -64,6 +64,8 @@ body:multipart-form {
~splitUnify: false
~pdfa: PDF/A-1b
~pdfua: true
~optimizeImages: false
~imageQuality: 80
~metadata: {"Author":"Bruno","Title":"Test"}
~userPassword:
~ownerPassword:

View File

@@ -15,8 +15,11 @@ body:multipart-form {
files: @file(../../test/integration/testdata/page_2.pdf)
~flatten: false
~autoIndexBookmarks: false
~titleBookmarks: false
~pdfa: PDF/A-1b
~pdfua: true
~optimizeImages: false
~imageQuality: 80
~metadata: {"Author":"Bruno","Title":"Test"}
~bookmarks: [{"title":"Page 1","page":1},{"title":"Page 2","page":2}]
~userPassword:

View File

@@ -0,0 +1,26 @@
meta {
name: Optimize PDF
type: http
seq: 1
}
post {
url: {{baseUrl}}/forms/pdfengines/optimize
body: multipartForm
auth: none
}
body:multipart-form {
files: @file(../../test/integration/testdata/page_1.pdf)
~imageQuality: 80
}
headers {
~Gotenberg-Output-Filename: optimized
~Gotenberg-Webhook-Url: http://localhost:8080/webhook
~Gotenberg-Webhook-Error-Url: http://localhost:8080/webhook/error
~Gotenberg-Webhook-Events-Url: http://localhost:8080/webhook/events
~Gotenberg-Webhook-Method: POST
~Gotenberg-Webhook-Error-Method: POST
~Gotenberg-Webhook-Extra-Http-Headers: {"X-Custom":"value"}
}

View File

@@ -18,6 +18,8 @@ body:multipart-form {
~flatten: false
~pdfa: PDF/A-1b
~pdfua: true
~optimizeImages: false
~imageQuality: 80
~metadata: {"Author":"Bruno","Title":"Test"}
~userPassword:
~ownerPassword:

View File

@@ -83,12 +83,16 @@ runs:
INPUT_PLATFORM: ${{ inputs.platform }}
INPUT_ALTERNATE_REPOSITORY: ${{ inputs.alternate_repository }}
INPUT_DRY_RUN: ${{ inputs.dry_run }}
# Exporting the build cache needs a registry login. Forks run without
# credentials, so they import the cache but never export it.
INPUT_CACHE_WRITABLE: ${{ inputs.docker_hub_username != '' }}
run: |
.github/actions/build-test-push/build.sh \
--version "$INPUT_VERSION" \
--platform "$INPUT_PLATFORM" \
--alternate-repository "$INPUT_ALTERNATE_REPOSITORY" \
--dry-run "$INPUT_DRY_RUN"
--dry-run "$INPUT_DRY_RUN" \
--cache-writable "$INPUT_CACHE_WRITABLE"
- name: Run integration tests
if: inputs.skip_integrations_tests != 'true'

View File

@@ -12,6 +12,7 @@ version=""
platform=""
alternate_repository=""
dry_run=""
cache_writable=""
while [[ $# -gt 0 ]]; do
case $1 in
@@ -31,6 +32,10 @@ while [[ $# -gt 0 ]]; do
dry_run="$2"
shift 2
;;
--cache-writable)
cache_writable="$2"
shift 2
;;
*)
echo "Unknown option $1"
exit 1
@@ -44,11 +49,41 @@ echo
echo "Gotenberg version: $version"
echo "Target platform: $platform"
# The build cache lives under the canonical repository, captured before the
# alternate-repository override below. Pull requests build into "snapshot", so
# deriving the cache ref after the override would give them a cache namespace
# of their own and they would never import what main published, which is the
# population that benefits most.
cache_image="$DOCKER_REGISTRY/$DOCKER_REPOSITORY"
# Layers are per-architecture, so each platform keeps its own cache manifest.
cache_platform="${platform//\//-}"
# Layers running "apt-get upgrade" install whatever versions are current at
# build time, and the packages are deliberately not pinned. A persistent cache
# would turn those into hits and freeze security patches into a published
# image until debian:13-slim itself changes digest. Keying them on the ISO week
# bounds that staleness to seven days while leaving every build within a week
# free to reuse the cache.
apt_snapshot="$(date -u +%G-W%V)"
# Only a build that is not redirected to an alternate repository writes the
# cache, so a pull request cannot make its own state the baseline for main.
# Reading stays enabled everywhere, including forks, since the cache ref is
# public and needs no credentials.
cache_to_enabled="false"
if [ "$cache_writable" = "true" ] && [ -z "$alternate_repository" ]; then
cache_to_enabled="true"
fi
if [ -n "$alternate_repository" ]; then
DOCKER_REPOSITORY=$alternate_repository
echo "⚠️ Using $alternate_repository for DOCKER_REPOSITORY"
fi
echo "Build cache: $cache_image:buildcache-<target>-$cache_platform (write: $cache_to_enabled)"
echo "APT snapshot: $apt_snapshot"
if [ "$dry_run" = "true" ]; then
echo "🚧 Dry run"
fi
@@ -189,12 +224,36 @@ join() {
echo "$*"
}
# cache_flags echoes the buildx cache arguments for a build target. Each target
# keeps its own manifest so that the Chromium and LibreOffice variants, which
# branch from common-stage rather than from each other, do not overwrite one
# another's entry.
#
# mode=max exports intermediate stages too, not just the final layers, which is
# what makes the expensive apt and jlink stages reusable. type=registry, not
# type=gha: the GitHub Actions cache is capped at 10 GB per repository and is
# already carrying the Go and golangci-lint caches that the lint and test jobs
# depend on. Multi-GB image layers across five platforms would evict them.
cache_flags() {
local target="$1"
local ref="$cache_image:buildcache-$target-$cache_platform"
local flags="--cache-from type=registry,ref=$ref"
if [ "$cache_to_enabled" = "true" ]; then
flags="$flags --cache-to type=registry,ref=$ref,mode=max"
fi
echo "$flags"
}
no_arch_tag="$DOCKER_REGISTRY/$DOCKER_REPOSITORY:$version"
# Full variant.
cmd="docker buildx build \
--target gotenberg \
--build-arg GOTENBERG_VERSION=$version \
--build-arg APT_SNAPSHOT=$apt_snapshot \
$(cache_flags gotenberg) \
--platform $platform \
--load \
${tags_flags[*]} \
@@ -207,6 +266,8 @@ run_cmd "$cmd"
cmd="docker buildx build \
--target gotenberg-chromium \
--build-arg GOTENBERG_VERSION=$version \
--build-arg APT_SNAPSHOT=$apt_snapshot \
$(cache_flags gotenberg-chromium) \
--platform $platform \
--load \
${tags_chromium_flags[*]} \
@@ -218,6 +279,8 @@ run_cmd "$cmd"
cmd="docker buildx build \
--target gotenberg-libreoffice \
--build-arg GOTENBERG_VERSION=$version \
--build-arg APT_SNAPSHOT=$apt_snapshot \
$(cache_flags gotenberg-libreoffice) \
--platform $platform \
--load \
${tags_libreoffice_flags[*]} \
@@ -230,6 +293,8 @@ if [ "$platform" = "linux/amd64" ]; then
cmd="docker buildx build \
--target gotenberg-cloudrun \
--build-arg GOTENBERG_VERSION=$version \
--build-arg APT_SNAPSHOT=$apt_snapshot \
$(cache_flags gotenberg-cloudrun) \
--platform $platform \
--load \
${tags_cloud_run_flags[*]} \
@@ -240,6 +305,8 @@ if [ "$platform" = "linux/amd64" ]; then
cmd="docker buildx build \
--target gotenberg-cloudrun-chromium \
--build-arg GOTENBERG_VERSION=$version \
--build-arg APT_SNAPSHOT=$apt_snapshot \
$(cache_flags gotenberg-cloudrun-chromium) \
--platform $platform \
--load \
${tags_cloud_run_chromium_flags[*]} \
@@ -250,6 +317,8 @@ if [ "$platform" = "linux/amd64" ]; then
cmd="docker buildx build \
--target gotenberg-cloudrun-libreoffice \
--build-arg GOTENBERG_VERSION=$version \
--build-arg APT_SNAPSHOT=$apt_snapshot \
$(cache_flags gotenberg-cloudrun-libreoffice) \
--platform $platform \
--load \
${tags_cloud_run_libreoffice_flags[*]} \
@@ -263,6 +332,8 @@ if [ "$platform" = "linux/amd64" ] || [ "$platform" = "linux/arm64" ]; then
cmd="docker buildx build \
--target gotenberg-aws-lambda \
--build-arg GOTENBERG_VERSION=$version \
--build-arg APT_SNAPSHOT=$apt_snapshot \
$(cache_flags gotenberg-aws-lambda) \
--platform $platform \
--load \
${tags_aws_lambda_flags[*]} \
@@ -273,6 +344,8 @@ if [ "$platform" = "linux/amd64" ] || [ "$platform" = "linux/arm64" ]; then
cmd="docker buildx build \
--target gotenberg-aws-lambda-chromium \
--build-arg GOTENBERG_VERSION=$version \
--build-arg APT_SNAPSHOT=$apt_snapshot \
$(cache_flags gotenberg-aws-lambda-chromium) \
--platform $platform \
--load \
${tags_aws_lambda_chromium_flags[*]} \
@@ -283,6 +356,8 @@ if [ "$platform" = "linux/amd64" ] || [ "$platform" = "linux/arm64" ]; then
cmd="docker buildx build \
--target gotenberg-aws-lambda-libreoffice \
--build-arg GOTENBERG_VERSION=$version \
--build-arg APT_SNAPSHOT=$apt_snapshot \
$(cache_flags gotenberg-aws-lambda-libreoffice) \
--platform $platform \
--load \
${tags_aws_lambda_libreoffice_flags[*]} \

View File

@@ -31,7 +31,7 @@ jobs:
- name: Run linters
uses: golangci/golangci-lint-action@v9
with:
version: v2.10.1
version: v2.13.2
lint-prettier:
name: Lint non-Golang codebase

View File

@@ -28,12 +28,21 @@ API_CORRELATION_ID_HEADER=Gotenberg-Trace
API_ENABLE_BASIC_AUTH=false
GOTENBERG_API_BASIC_AUTH_USERNAME=
GOTENBERG_API_BASIC_AUTH_PASSWORD=
API_ENABLE_OIDC_AUTH=false
API_OIDC_ISSUER=
API_OIDC_AUDIENCE=
API_OIDC_JWKS_URL=
API_DOWNLOAD_FROM_ALLOW_LIST=
API_DOWNLOAD_FROM_DENY_LIST=^https?://(10\.|172\.(1[6-9]|2[0-9]|3[01])\.|192\.168\.|169\.254\.|0\.0\.0\.0|127\.|localhost|\[::1\]|\[fd)
# Empty, like the flag default since 8.32.0. A textual deny-list cannot
# enumerate every way to write a private address, so *_DENY_PRIVATE_IPS is the
# control to reach for. Left false here so local testing can reach the host.
API_DOWNLOAD_FROM_DENY_LIST=
API_DOWNLOAD_FROM_DENY_PRIVATE_IPS=false
API_DOWNLOAD_FROM_DENY_PUBLIC_IPS=false
API_DOWNLOAD_FROM_ENABLE_ENVIRONMENT_PROXY=false
API_DOWNLOAD_FROM_MAX_RETRY=4
API_DOWNLOAD_FROM_MAX_CONCURRENCY=10
API_DOWNLOAD_FROM_MAX_ENTRIES=0
API_DISABLE_DOWNLOAD_FROM=false
API_DISABLE_HEALTH_CHECK_ROUTE_TELEMETRY=true
API_DISABLE_ROOT_ROUTE_TELEMETRY=true
@@ -78,10 +87,12 @@ LOG_STD_FORMAT=auto
LOG_STD_ENABLE_GCP_FIELDS=false
LOG_STD_LEVEL_CASE=lower
PDFENGINES_DISABLE_ROUTES=false
PDFENGINES_MAX_CONCURRENCY=1
PDFENGINES_MERGE_ENGINES=qpdf,pdfcpu,pdftk
PDFENGINES_SPLIT_ENGINES=pdfcpu,qpdf,pdftk
PDFENGINES_FLATTEN_ENGINES=qpdf
PDFENGINES_CONVERT_ENGINES=libreoffice-pdfengine
PDFENGINES_OPTIMIZE_IMAGES_ENGINES=pdfcpu
PDFENGINES_READ_METADATA_ENGINES=exiftool
PDFENGINES_WRITE_METADATA_ENGINES=exiftool
PDFENGINES_READ_BOOKMARKS_ENGINES=pdfcpu
@@ -90,7 +101,7 @@ PDFENGINES_WATERMARK_ENGINES=pdfcpu,pdftk
PDFENGINES_STAMP_ENGINES=pdfcpu,pdftk
PDFENGINES_ENCRYPT_ENGINES=qpdf,pdfcpu,pdftk
PDFENGINES_ROTATE_ENGINES=pdfcpu,pdftk
PDFENGINES_EMBED_ENGINES=qpdf,pdfcpu
PDFENGINES_EMBED_ENGINES=pdfcpu
PDFENGINES_EMBED_METADATA_ENGINES=qpdf
PDFENGINES_FACTUR_X_ENGINES=qpdf
PROMETHEUS_NAMESPACE=gotenberg
@@ -107,7 +118,8 @@ OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4317
OTEL_EXPORTER_OTLP_INSECURE=true
WEBHOOK_ENABLE_SYNC_MODE=false
WEBHOOK_ALLOW_LIST=
WEBHOOK_DENY_LIST=^https?://(10\.|172\.(1[6-9]|2[0-9]|3[01])\.|192\.168\.|169\.254\.|0\.0\.0\.0|127\.|localhost|\[::1\]|\[fd)
# See the note on API_DOWNLOAD_FROM_DENY_LIST.
WEBHOOK_DENY_LIST=
WEBHOOK_DENY_PRIVATE_IPS=false
WEBHOOK_DENY_PUBLIC_IPS=false
WEBHOOK_ENABLE_ENVIRONMENT_PROXY=false
@@ -147,10 +159,12 @@ NO_CONCURRENCY=false
# chromium-screenshot-html
# chromium-screenshot-markdown
# chromium-screenshot-url
# chromium-ssrf
# debug
# health
# libreoffice
# libreoffice-convert
# libreoffice-ssrf
# output-filename
# pdfengines
# pdfengines-convert
@@ -160,6 +174,8 @@ NO_CONCURRENCY=false
# encrypt
# pdfengines-flatten
# flatten
# pdfengines-optimize
# optimize
# pdfengines-merge
# merge
# pdfengines-metadata
@@ -201,11 +217,19 @@ lint-prettier: ## Lint non-Golang codebase
.PHONY: lint-todo
lint-todo: ## Find TODOs in Golang codebase
golangci-lint run --no-config --disable-all --enable godox
golangci-lint run --no-config --default=none --enable godox
# TODO: restore a plain "go fix ./..." once the errorsastype modernizer stops
# rewriting this codebase into code that does not compile. Re-check by dropping
# the flag and running "make fmt && make lint". Removing the analyzer upstream
# makes go fix fail with "flag provided but not defined", so this cannot rot
# silently.
# errorsastype rewrites errors.As to errors.AsType[T] without checking that T
# satisfies error, which breaks on api.HttpError since it does not embed
# error. Still broken as of Go 1.27.1.
.PHONY: fmt
fmt: ## Format Golang codebase and "optimize" the dependencies
go fix ./...
go fix -errorsastype=false ./...
golangci-lint fmt
go mod tidy

View File

@@ -1,7 +1,7 @@
# ARG instructions do not create additional layers. Instead, next layers will
# concatenate them. Also, we have to repeat ARG instructions in each build
# stage that uses them.
ARG GOLANG_VERSION=1.26.5
ARG GOLANG_VERSION=1.27.1
# ----------------------------------------------
# pdfcpu binary build stage
@@ -11,7 +11,7 @@ ARG GOLANG_VERSION=1.26.5
FROM golang:$GOLANG_VERSION AS pdfcpu-binary-stage
# See https://github.com/pdfcpu/pdfcpu/releases.
ARG PDFCPU_VERSION=v0.13.0
ARG PDFCPU_VERSION=v0.15.0
ENV CGO_ENABLED=0
# Define the working directory outside of $GOPATH (we're using go modules).
@@ -59,7 +59,14 @@ RUN go build -o gotenberg -ldflags "-s -w -X 'github.com/gotenberg/gotenberg/v8/
# ----------------------------------------------
FROM debian:13-slim AS custom-jre-stage
RUN apt-get update -qq \
# APT_SNAPSHOT busts every layer below it when CI rotates the value, weekly.
# Without it a persistent build cache turns the unpinned "apt-get upgrade" into
# a cache hit and the published image keeps shipping the package versions that
# were current when the cache was first populated.
ARG APT_SNAPSHOT=""
RUN echo "apt snapshot: $APT_SNAPSHOT" \
&& apt-get update -qq \
&& apt-get upgrade -yqq \
&& DEBIAN_FRONTEND=noninteractive apt-get install -y -qq --no-install-recommends default-jdk-headless binutils
@@ -88,7 +95,7 @@ RUN apt-get update -qq \
WORKDIR /downloads
RUN curl -Ls https://raw.githubusercontent.com/gotenberg/unoconverter/v0.4.0/unoconv -o unoconverter \
RUN curl -Ls https://raw.githubusercontent.com/gotenberg/unoconverter/v0.5.0/unoconv -o unoconverter \
&& chmod +x unoconverter
RUN curl -o pdftk-all.jar "https://gitlab.com/api/v4/projects/5024297/packages/generic/pdftk-java/$PDFTK_VERSION/pdftk-all.jar" \
@@ -114,9 +121,15 @@ FROM base-image-stage AS common-stage
ARG GOTENBERG_USER_GID=1001
ARG GOTENBERG_USER_UID=1001
# See the note on APT_SNAPSHOT in custom-jre-stage. Declaring it here covers
# every "apt-get upgrade" in the gotenberg, gotenberg-chromium and
# gotenberg-libreoffice targets too, since all three branch from this stage.
ARG APT_SNAPSHOT=""
# Create a non-root user.
# All processes in the Docker container will run with this dedicated user.
RUN groupadd --gid "$GOTENBERG_USER_GID" gotenberg \
RUN echo "apt snapshot: $APT_SNAPSHOT" \
&& groupadd --gid "$GOTENBERG_USER_GID" gotenberg \
&& useradd --uid "$GOTENBERG_USER_UID" --gid gotenberg --shell /bin/bash --home /home/gotenberg --no-create-home gotenberg \
&& mkdir /home/gotenberg \
&& chown gotenberg: /home/gotenberg
@@ -259,7 +272,7 @@ RUN echo "deb http://deb.debian.org/debian trixie-backports main" >> /etc/apt/so
hyphen-no hyphen-or hyphen-pa hyphen-pl hyphen-pt-br hyphen-pt-pt hyphen-ro hyphen-ru hyphen-sk hyphen-sl \
hyphen-sr hyphen-sv hyphen-ta hyphen-te hyphen-th hyphen-uk hyphen-zu \
&& DEBIAN_FRONTEND=noninteractive apt-get install -y -qq --no-install-recommends -t trixie-backports \
libreoffice-writer libreoffice-calc libreoffice-impress libreoffice-draw python3-uno \
libreoffice-writer libreoffice-calc libreoffice-impress libreoffice-draw libreoffice-math python3-uno \
# unoconverter will look for the Python binary, which has to be at version 3.
&& ln -s /usr/bin/python3 /usr/bin/python \
# Cleanup.
@@ -386,7 +399,7 @@ RUN echo "deb http://deb.debian.org/debian trixie-backports main" >> /etc/apt/so
hyphen-no hyphen-or hyphen-pa hyphen-pl hyphen-pt-br hyphen-pt-pt hyphen-ro hyphen-ru hyphen-sk hyphen-sl \
hyphen-sr hyphen-sv hyphen-ta hyphen-te hyphen-th hyphen-uk hyphen-zu \
&& DEBIAN_FRONTEND=noninteractive apt-get install -y -qq --no-install-recommends -t trixie-backports \
libreoffice-writer libreoffice-calc libreoffice-impress libreoffice-draw python3-uno \
libreoffice-writer libreoffice-calc libreoffice-impress libreoffice-draw libreoffice-math python3-uno \
# unoconverter will look for the Python binary, which has to be at version 3.
&& ln -s /usr/bin/python3 /usr/bin/python \
# Cleanup.

View File

@@ -80,7 +80,7 @@ func Run() {
// Override their values if the corresponding environment variables are
// set.
fs.VisitAll(func(f *flag.Flag) {
envName := strings.ToUpper(strings.ReplaceAll(f.Name, "-", "_"))
envName := gotenberg.EnvVarName(f.Name)
val, ok := os.LookupEnv(envName)
if !ok {
return

View File

@@ -35,6 +35,8 @@ services:
- "--api-download-from-deny-public-ips=${API_DOWNLOAD_FROM_DENY_PUBLIC_IPS}"
- "--api-download-from-enable-environment-proxy=${API_DOWNLOAD_FROM_ENABLE_ENVIRONMENT_PROXY}"
- "--api-download-from-max-retry=${API_DOWNLOAD_FROM_MAX_RETRY}"
- "--api-download-from-max-concurrency=${API_DOWNLOAD_FROM_MAX_CONCURRENCY}"
- "--api-download-from-max-entries=${API_DOWNLOAD_FROM_MAX_ENTRIES}"
- "--api-disable-download-from=${API_DISABLE_DOWNLOAD_FROM}"
- "--api-disable-health-check-route-telemetry=${API_DISABLE_HEALTH_CHECK_ROUTE_TELEMETRY}"
- "--api-disable-root-route-telemetry=${API_DISABLE_ROOT_ROUTE_TELEMETRY}"
@@ -82,6 +84,7 @@ services:
- "--pdfengines-split-engines=${PDFENGINES_SPLIT_ENGINES}"
- "--pdfengines-flatten-engines=${PDFENGINES_FLATTEN_ENGINES}"
- "--pdfengines-convert-engines=${PDFENGINES_CONVERT_ENGINES}"
- "--pdfengines-optimize-images-engines=${PDFENGINES_OPTIMIZE_IMAGES_ENGINES}"
- "--pdfengines-read-metadata-engines=${PDFENGINES_READ_METADATA_ENGINES}"
- "--pdfengines-write-metadata-engines=${PDFENGINES_WRITE_METADATA_ENGINES}"
- "--pdfengines-read-bookmarks-engines=${PDFENGINES_READ_BOOKMARKS_ENGINES}"
@@ -93,6 +96,7 @@ services:
- "--pdfengines-embed-engines=${PDFENGINES_EMBED_ENGINES}"
- "--pdfengines-embed-metadata-engines=${PDFENGINES_EMBED_METADATA_ENGINES}"
- "--pdfengines-factur-x-engines=${PDFENGINES_FACTUR_X_ENGINES}"
- "--pdfengines-max-concurrency=${PDFENGINES_MAX_CONCURRENCY}"
- "--pdfengines-disable-routes=${PDFENGINES_DISABLE_ROUTES}"
- "--prometheus-namespace=${PROMETHEUS_NAMESPACE}"
- "--prometheus-collect-interval=${PROMETHEUS_COLLECT_INTERVAL}"

131
go.mod
View File

@@ -1,40 +1,42 @@
module github.com/gotenberg/gotenberg/v8
go 1.26.5
go 1.27.1
require (
github.com/alexliesenfeld/health v0.8.1
github.com/chromedp/cdproto v0.0.0-20250803210736-d308e07a266d // pinned with chromedp v0.14.2, see below
github.com/chromedp/chromedp v0.14.2 // pinned: v0.15.x breaks the headless print-mode paint pipeline (rAF / ResizeObserver / IntersectionObserver stop firing, blank charts). See https://github.com/gotenberg/gotenberg/issues/1535.
github.com/coreos/go-oidc/v3 v3.21.0
github.com/cucumber/godog v0.16.0
github.com/dlclark/regexp2 v1.12.0
github.com/gomarkdown/markdown v0.0.0-20260614204949-e08cff860f76
github.com/gomarkdown/markdown v0.0.0-20260824154242-13c5cf49db8d
github.com/google/uuid v1.6.0
github.com/hashicorp/go-retryablehttp v0.7.8
github.com/labstack/echo/v4 v4.15.4
github.com/labstack/gommon v0.5.0
github.com/mholt/archives v0.1.5
github.com/microcosm-cc/bluemonday v1.0.27
github.com/moby/moby/api v1.55.0
github.com/moby/moby/client v0.5.1
github.com/moby/moby/api v1.56.0
github.com/moby/moby/client v0.6.0
github.com/prometheus/client_golang v1.24.1
github.com/shirou/gopsutil/v4 v4.26.7
github.com/shirou/gopsutil/v4 v4.26.8
github.com/spf13/pflag v1.0.10
github.com/stretchr/testify v1.11.1
github.com/testcontainers/testcontainers-go v0.43.0
go.opentelemetry.io/contrib/bridges/otelslog v0.19.0
go.opentelemetry.io/contrib/exporters/autoexport v0.69.0
go.opentelemetry.io/otel v1.45.0
go.opentelemetry.io/otel/log v0.20.0
go.opentelemetry.io/otel/metric v1.45.0
go.opentelemetry.io/otel/sdk v1.45.0
go.opentelemetry.io/otel/sdk/log v0.20.0
go.opentelemetry.io/otel/sdk/metric v1.45.0
go.opentelemetry.io/otel/trace v1.45.0
golang.org/x/net v0.57.0
golang.org/x/sync v0.22.0
github.com/stretchr/testify v1.12.1
github.com/testcontainers/testcontainers-go v0.44.0
go.opentelemetry.io/contrib/bridges/otelslog v0.20.1
go.opentelemetry.io/contrib/exporters/autoexport v0.71.0
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.71.0
go.opentelemetry.io/otel v1.46.0
go.opentelemetry.io/otel/log v0.22.0
go.opentelemetry.io/otel/metric v1.46.0
go.opentelemetry.io/otel/sdk v1.46.0
go.opentelemetry.io/otel/sdk/log v0.22.0
go.opentelemetry.io/otel/sdk/metric v1.46.0
go.opentelemetry.io/otel/trace v1.46.0
golang.org/x/net v0.58.0
golang.org/x/sync v0.23.0
golang.org/x/term v0.45.0
golang.org/x/text v0.40.0
golang.org/x/text v0.41.0
)
require (
@@ -42,11 +44,11 @@ require (
github.com/Azure/go-ansiterm v0.0.0-20250102033503-faa5f7b0171c // indirect
github.com/Microsoft/go-winio v0.6.2 // indirect
github.com/STARRY-S/zip v0.2.3 // indirect
github.com/andybalholm/brotli v1.2.1 // indirect
github.com/andybalholm/brotli v1.2.3 // indirect
github.com/aymerick/douceur v0.2.0 // indirect
github.com/beorn7/perks v1.0.1 // indirect
github.com/bodgit/plumbing v1.3.0 // indirect
github.com/bodgit/sevenzip v1.6.4 // indirect
github.com/bodgit/sevenzip v1.6.5 // indirect
github.com/bodgit/windows v1.0.1 // indirect
github.com/cenkalti/backoff/v4 v4.3.0 // indirect
github.com/cenkalti/backoff/v5 v5.0.3 // indirect
@@ -57,16 +59,16 @@ require (
github.com/containerd/log v0.1.0 // indirect
github.com/containerd/platforms v0.2.1 // indirect
github.com/cpuguy83/dockercfg v0.3.2 // indirect
github.com/cucumber/gherkin/go/v42 v42.0.0 // indirect
github.com/cucumber/messages/go/v34 v34.2.0 // indirect
github.com/davecgh/go-spew v1.1.1 // indirect
github.com/cucumber/gherkin/go/v42 v42.0.1 // indirect
github.com/cucumber/messages/go/v34 v34.2.1 // indirect
github.com/distribution/reference v0.6.0 // indirect
github.com/docker/go-connections v0.7.0 // indirect
github.com/docker/go-connections v0.8.1 // indirect
github.com/docker/go-units v0.5.0 // indirect
github.com/dsnet/compress v0.0.2-0.20230904184137-39efe44ab707 // indirect
github.com/ebitengine/purego v0.10.2 // indirect
github.com/ebitengine/purego v0.11.0 // indirect
github.com/felixge/httpsnoop v1.1.0 // indirect
github.com/go-json-experiment/json v0.0.0-20260601182631-00ed12fed2a6 // indirect
github.com/go-jose/go-jose/v4 v4.1.5 // indirect
github.com/go-json-experiment/json v0.0.0-20260820222146-c27c302e5fc3 // indirect
github.com/go-logr/logr v1.4.4 // indirect
github.com/go-logr/stdr v1.2.2 // indirect
github.com/go-ole/go-ole v1.3.0 // indirect
@@ -74,70 +76,69 @@ require (
github.com/gobwas/pool v0.2.1 // indirect
github.com/gobwas/ws v1.4.0 // indirect
github.com/gorilla/css v1.0.1 // indirect
github.com/grpc-ecosystem/grpc-gateway/v2 v2.29.0 // indirect
github.com/grpc-ecosystem/grpc-gateway/v2 v2.30.0 // indirect
github.com/hashicorp/go-cleanhttp v0.5.2 // indirect
github.com/hashicorp/go-immutable-radix v1.3.1 // indirect
github.com/hashicorp/go-memdb v1.3.5 // indirect
github.com/hashicorp/golang-lru v1.0.2 // indirect
github.com/hashicorp/golang-lru/v2 v2.0.7 // indirect
github.com/klauspost/compress v1.19.1 // indirect
github.com/klauspost/compress v1.20.0 // indirect
github.com/klauspost/pgzip v1.2.6 // indirect
github.com/lufia/plan9stats v0.0.0-20260330125221-c963978e514e // indirect
github.com/magiconair/properties v1.8.10 // indirect
github.com/lufia/plan9stats v0.0.0-20260802145828-341c2f0c90b5 // indirect
github.com/magiconair/properties v1.18.11 // indirect
github.com/mattn/go-colorable v0.1.15 // indirect
github.com/mattn/go-isatty v0.0.22 // indirect
github.com/mattn/go-isatty v0.0.24 // indirect
github.com/mikelolasagasti/xz v1.0.1 // indirect
github.com/minio/minlz v1.1.1 // indirect
github.com/minio/minlz v1.2.0 // indirect
github.com/moby/docker-image-spec v1.3.1 // indirect
github.com/moby/go-archive v0.2.0 // indirect
github.com/moby/go-archive v0.3.3 // indirect
github.com/moby/patternmatcher v0.6.1 // indirect
github.com/moby/sys/sequential v0.7.0 // indirect
github.com/moby/sys/user v0.4.0 // indirect
github.com/moby/sys/userns v0.1.0 // indirect
github.com/moby/sys/user v0.4.1 // indirect
github.com/moby/sys/userns v0.2.0 // indirect
github.com/moby/term v0.5.2 // indirect
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 // indirect
github.com/nwaples/rardecode/v2 v2.2.5 // indirect
github.com/nwaples/rardecode/v2 v2.4.1 // indirect
github.com/opencontainers/go-digest v1.0.0 // indirect
github.com/opencontainers/image-spec v1.1.1 // indirect
github.com/pierrec/lz4/v4 v4.1.27 // indirect
github.com/pmezard/go-difflib v1.0.0 // indirect
github.com/power-devops/perfstat v0.0.0-20240221224432-82ca36839d55 // indirect
github.com/prometheus/client_model v0.6.2 // indirect
github.com/prometheus/common v0.70.1 // indirect
github.com/pierrec/lz4/v4 v4.1.29 // indirect
github.com/power-devops/perfstat v0.0.0-20260805114148-88456608a4f6 // indirect
github.com/prometheus/client_model v0.6.3 // indirect
github.com/prometheus/common v0.71.0 // indirect
github.com/prometheus/otlptranslator v1.0.0 // indirect
github.com/prometheus/procfs v0.21.1 // indirect
github.com/sirupsen/logrus v1.9.4 // indirect
github.com/prometheus/procfs v0.22.0 // indirect
github.com/sirupsen/logrus v1.10.2 // indirect
github.com/sorairolake/lzip-go v0.3.8 // indirect
github.com/spf13/afero v1.15.0 // indirect
github.com/stangelandcl/ppmd v0.1.1 // indirect
github.com/tklauser/go-sysconf v0.4.0 // indirect
github.com/tklauser/numcpus v0.12.0 // indirect
github.com/ulikunitz/xz v0.5.15 // indirect
github.com/ulikunitz/xz v0.5.16 // indirect
github.com/valyala/bytebufferpool v1.0.0 // indirect
github.com/valyala/fasttemplate v1.2.2 // indirect
github.com/yusufpapurcu/wmi v1.2.4 // indirect
go.opentelemetry.io/auto/sdk v1.2.1 // indirect
go.opentelemetry.io/contrib/bridges/prometheus v0.69.0 // indirect
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlplog/otlploggrpc v0.20.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlplog/otlploghttp v0.20.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetricgrpc v1.44.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.44.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp v1.44.0 // indirect
go.opentelemetry.io/otel/exporters/prometheus v0.66.0 // indirect
go.opentelemetry.io/otel/exporters/stdout/stdoutlog v0.20.0 // indirect
go.opentelemetry.io/otel/exporters/stdout/stdoutmetric v1.44.0 // indirect
go.opentelemetry.io/otel/exporters/stdout/stdouttrace v1.44.0 // indirect
go.opentelemetry.io/proto/otlp v1.10.0 // indirect
go.opentelemetry.io/contrib/bridges/prometheus v0.71.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlplog/otlploggrpc v0.22.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlplog/otlploghttp v0.22.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetricgrpc v1.46.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.46.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.46.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.46.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp v1.46.0 // indirect
go.opentelemetry.io/otel/exporters/prometheus v0.68.0 // indirect
go.opentelemetry.io/otel/exporters/stdout/stdoutlog v0.22.0 // indirect
go.opentelemetry.io/otel/exporters/stdout/stdoutmetric v1.46.0 // indirect
go.opentelemetry.io/otel/exporters/stdout/stdouttrace v1.46.0 // indirect
go.opentelemetry.io/proto/otlp v1.11.0 // indirect
go.yaml.in/yaml/v3 v3.0.5 // indirect
go4.org v0.0.0-20260112195520-a5071408f32f // indirect
golang.org/x/crypto v0.54.0 // indirect
golang.org/x/crypto v0.56.0 // indirect
golang.org/x/oauth2 v0.36.0 // indirect
golang.org/x/sys v0.47.0 // indirect
golang.org/x/time v0.15.0 // indirect
google.golang.org/genproto/googleapis/api v0.0.0-20260615183401-62b3387ff324 // indirect
google.golang.org/genproto/googleapis/rpc v0.0.0-20260615183401-62b3387ff324 // indirect
google.golang.org/grpc v1.82.1 // indirect
google.golang.org/protobuf v1.36.11 // indirect
gopkg.in/yaml.v3 v3.0.1 // indirect
google.golang.org/genproto/googleapis/api v0.0.0-20260831171406-18b4a7587f8a // indirect
google.golang.org/genproto/googleapis/rpc v0.0.0-20260831171406-18b4a7587f8a // indirect
google.golang.org/grpc v1.83.2 // indirect
google.golang.org/protobuf v1.36.12 // indirect
)

275
go.sum
View File

@@ -10,16 +10,16 @@ github.com/STARRY-S/zip v0.2.3 h1:luE4dMvRPDOWQdeDdUxUoZkzUIpTccdKdhHHsQJ1fm4=
github.com/STARRY-S/zip v0.2.3/go.mod h1:lqJ9JdeRipyOQJrYSOtpNAiaesFO6zVDsE8GIGFaoSk=
github.com/alexliesenfeld/health v0.8.1 h1:wdE3vt+cbJotiR8DGDBZPKHDFoJbAoWEfQTcqrmedUg=
github.com/alexliesenfeld/health v0.8.1/go.mod h1:TfNP0f+9WQVWMQRzvMUjlws4ceXKEL3WR+6Hp95HUFc=
github.com/andybalholm/brotli v1.2.1 h1:R+f5xP285VArJDRgowrfb9DqL18yVK0gKAW/F+eTWro=
github.com/andybalholm/brotli v1.2.1/go.mod h1:rzTDkvFWvIrjDXZHkuS16NPggd91W3kUSvPlQ1pLaKY=
github.com/andybalholm/brotli v1.2.3 h1:8H1qwOkl2LPfjf3YezB90JnCliZb6SInJ/OJkEbA5NQ=
github.com/andybalholm/brotli v1.2.3/go.mod h1:rzTDkvFWvIrjDXZHkuS16NPggd91W3kUSvPlQ1pLaKY=
github.com/aymerick/douceur v0.2.0 h1:Mv+mAeH1Q+n9Fr+oyamOlAkUNPWPlA8PPGR0QAaYuPk=
github.com/aymerick/douceur v0.2.0/go.mod h1:wlT5vV2O3h55X9m7iVYN0TBM0NH/MmbLnd30/FjWUq4=
github.com/beorn7/perks v1.0.1 h1:VlbKKnNfV8bJzeqoa4cOKqO6bYr3WgKZxO8Z16+hsOM=
github.com/beorn7/perks v1.0.1/go.mod h1:G2ZrVWU2WbWT9wwq4/hrbKbnv/1ERSJQ0ibhJ6rlkpw=
github.com/bodgit/plumbing v1.3.0 h1:pf9Itz1JOQgn7vEOE7v7nlEfBykYqvUYioC61TwWCFU=
github.com/bodgit/plumbing v1.3.0/go.mod h1:JOTb4XiRu5xfnmdnDJo6GmSbSbtSyufrsyZFByMtKEs=
github.com/bodgit/sevenzip v1.6.4 h1:iHiVJfxbrB6RF4X+snI2MpVgNBKmVfGaTqZGNlMQIU0=
github.com/bodgit/sevenzip v1.6.4/go.mod h1:ZtNi5KNgHXeXg1G7WiF0IWSuFE2eG6lt/cTGlvuirO0=
github.com/bodgit/sevenzip v1.6.5 h1:7H7BxgmeX0j6UX42lH+KXQ92WgMQJ49DoocFdfHbCng=
github.com/bodgit/sevenzip v1.6.5/go.mod h1:GhuB6Lq1xCpP1sps+horjZ8lgiKPJcy2zUX3prla9wc=
github.com/bodgit/windows v1.0.1 h1:tF7K6KOluPYygXa3Z2594zxlkbKPAOvqr97etrGNIz4=
github.com/bodgit/windows v1.0.1/go.mod h1:a6JLwrB4KrTR5hBpp8FI9/9W9jJfeQ2h4XDXU74ZCdM=
github.com/cenkalti/backoff/v4 v4.3.0 h1:MyRJ/UdXutAwSAT+s3wNd7MfTIcy71VQueUuFK343L8=
@@ -42,38 +42,41 @@ github.com/containerd/log v0.1.0 h1:TCJt7ioM2cr/tfR8GPbGf9/VRAX8D2B4PjzCpfX540I=
github.com/containerd/log v0.1.0/go.mod h1:VRRf09a7mHDIRezVKTRCrOq78v577GXq3bSa3EhrzVo=
github.com/containerd/platforms v0.2.1 h1:zvwtM3rz2YHPQsF2CHYM8+KtB5dvhISiXh5ZpSBQv6A=
github.com/containerd/platforms v0.2.1/go.mod h1:XHCb+2/hzowdiut9rkudds9bE5yJ7npe7dG/wG+uFPw=
github.com/coreos/go-oidc/v3 v3.21.0 h1:wZo4Q9Pum8dYEj0eMUPrqR+kvuGkeUplbLpNCkBqoWM=
github.com/coreos/go-oidc/v3 v3.21.0/go.mod h1:DYCf24+ncYi+XkIH97GY1+dqoRlbaSI26KVTCI9SrY4=
github.com/cpuguy83/dockercfg v0.3.2 h1:DlJTyZGBDlXqUZ2Dk2Q3xHs/FtnooJJVaad2S9GKorA=
github.com/cpuguy83/dockercfg v0.3.2/go.mod h1:sugsbF4//dDlL/i+S+rtpIWp+5h0BHJHfjj5/jFyUJc=
github.com/creack/pty v1.1.24 h1:bJrF4RRfyJnbTJqzRLHzcGaZK1NeM5kTC9jGgovnR1s=
github.com/creack/pty v1.1.24/go.mod h1:08sCNb52WyoAwi2QDyzUCTgcvVFhUzewun7wtTfvcwE=
github.com/cucumber/gherkin/go/v42 v42.0.0 h1:Ulh3E2awUUSSja+wonP/IOQ+ycmiZwZbgmzqk5H8JNI=
github.com/cucumber/gherkin/go/v42 v42.0.0/go.mod h1:CsaumaO2dR9XvBc6ZyiGLMhWCKtTRDxgoxqJigSjSSg=
github.com/cucumber/gherkin/go/v42 v42.0.1 h1:ao9TVJmBb8uNLEcjMDFbhsoL3yC7gMpRLHgP1ZGfFWA=
github.com/cucumber/gherkin/go/v42 v42.0.1/go.mod h1:CsaumaO2dR9XvBc6ZyiGLMhWCKtTRDxgoxqJigSjSSg=
github.com/cucumber/godog v0.16.0 h1:ezQbgItuWqZrjPUQwLJ3muwIlvzXBOfZso5QZfG7efE=
github.com/cucumber/godog v0.16.0/go.mod h1:EDUX9yCqANK+GpbftMDeu61sUDtdLuo1JJgXD2n3bbM=
github.com/cucumber/messages/go/v34 v34.2.0 h1:VCbcNOMz+f8ccjjOOx1NLBNhwvE7/X49Atc8klJa+i8=
github.com/cucumber/messages/go/v34 v34.2.0/go.mod h1:LYUPjqlTS1kS0pdkdf6sS5uirnjwiIzEGyXPezXNhL8=
github.com/cucumber/messages/go/v34 v34.2.1 h1:qBPEl+HhNJuRX8Kjaw1Pm60KOODZf2/4WQAi0SR/neE=
github.com/cucumber/messages/go/v34 v34.2.1/go.mod h1:LYUPjqlTS1kS0pdkdf6sS5uirnjwiIzEGyXPezXNhL8=
github.com/davecgh/go-spew v1.1.0/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/distribution/reference v0.6.0 h1:0IXCQ5g4/QMHHkarYzh5l+u8T3t73zM5QvfrDyIgxBk=
github.com/distribution/reference v0.6.0/go.mod h1:BbU0aIcezP1/5jX/8MP0YiH4SdvB5Y4f/wlDRiLyi3E=
github.com/dlclark/regexp2 v1.12.0 h1:0j4c5qQmnC6XOWNjP3PIXURXN2gWx76rd3KvgdPkCz8=
github.com/dlclark/regexp2 v1.12.0/go.mod h1:DHkYz0B9wPfa6wondMfaivmHpzrQ3v9q8cnmRbL6yW8=
github.com/docker/go-connections v0.7.0 h1:6SsRfJddP22WMrCkj19x9WKjEDTB+ahsdiGYf0mN39c=
github.com/docker/go-connections v0.7.0/go.mod h1:no1qkHdjq7kLMGUXYAduOhYPSJxxvgWBh7ogVvptn3Q=
github.com/docker/go-connections v0.8.1 h1:JibmG5hULs5qXSr/cp/w3Pw5fZuStt4MOHMUExb29/M=
github.com/docker/go-connections v0.8.1/go.mod h1:no1qkHdjq7kLMGUXYAduOhYPSJxxvgWBh7ogVvptn3Q=
github.com/docker/go-units v0.5.0 h1:69rxXcBk27SvSaaxTtLh/8llcHD8vYHT7WSdRZ/jvr4=
github.com/docker/go-units v0.5.0/go.mod h1:fgPhTUdO+D/Jk86RDLlptpiXQzgHJF7gydDDbaIK4Dk=
github.com/dsnet/compress v0.0.2-0.20230904184137-39efe44ab707 h1:2tV76y6Q9BB+NEBasnqvs7e49aEBFI8ejC89PSnWH+4=
github.com/dsnet/compress v0.0.2-0.20230904184137-39efe44ab707/go.mod h1:qssHWj60/X5sZFNxpG4HBPDHVqxNm4DfnCKgrbZOT+s=
github.com/dsnet/golib v0.0.0-20171103203638-1ea166775780/go.mod h1:Lj+Z9rebOhdfkVLjJ8T6VcRQv3SXugXy999NBtR9aFY=
github.com/ebitengine/purego v0.10.2 h1:W809HbnvzAxgdm+aOvlSekrM16wGCdT/e76+9tS7gzE=
github.com/ebitengine/purego v0.10.2/go.mod h1:iIjxzd6CiRiOG0UyXP+V1+jWqUXVjPKLAI0mRfJZTmQ=
github.com/ebitengine/purego v0.11.0 h1:jhp/D+Nyv7UUW8HAcmcjt2N2rYrYi9m3SL21k0Ua/NI=
github.com/ebitengine/purego v0.11.0/go.mod h1:DCHPP08djqhNSoTfImcnHYQRZmd0qhakvrozqaEYhGQ=
github.com/fatih/color v1.16.0 h1:zmkK9Ngbjj+K0yRhTVONQh1p/HknKYSlNT+vZCzyokM=
github.com/fatih/color v1.16.0/go.mod h1:fL2Sau1YI5c0pdGEVCbKQbLXB6edEj1ZgiY4NijnWvE=
github.com/felixge/httpsnoop v1.1.0 h1:3YtUj32ZZkqZtt3sZZsClsymw/QDuVfpNhoA31zeORc=
github.com/felixge/httpsnoop v1.1.0/go.mod h1:Zqxgdd+1Rkcz8euOqdr7lqgCRJztwr5hp9vDSi5UZCE=
github.com/go-json-experiment/json v0.0.0-20260601182631-00ed12fed2a6 h1:nxP4pPoyqOAgX8lYDFCfl3DyKeXErCvSvhcyzwGV9CE=
github.com/go-json-experiment/json v0.0.0-20260601182631-00ed12fed2a6/go.mod h1:tphK2c80bpPhMOI4v6bIc2xWywPfbqi1Z06+RcrMkDg=
github.com/go-jose/go-jose/v4 v4.1.5 h1:RjgjO2LOtWOJKUC5wpwY9LR3B3vwVAz6JS2YHfYU6eA=
github.com/go-jose/go-jose/v4 v4.1.5/go.mod h1:x4oUasVrzR7071A4TnHLGSPpNOm2a21K9Kf04k1rs08=
github.com/go-json-experiment/json v0.0.0-20260820222146-c27c302e5fc3 h1:UADEEmDKgfXbtnGJZ97beY5XLo9ZechG1nlU4KnRrkE=
github.com/go-json-experiment/json v0.0.0-20260820222146-c27c302e5fc3/go.mod h1:tphK2c80bpPhMOI4v6bIc2xWywPfbqi1Z06+RcrMkDg=
github.com/go-logr/logr v1.2.2/go.mod h1:jdQByPbusPIv2/zmleS9BjJVeZ6kBagPoEUsqbVz/1A=
github.com/go-logr/logr v1.4.4 h1:tG4xh9yMsRCAiodLVTxyrkzSZ9+o0L1Kg/+cPVcbP/8=
github.com/go-logr/logr v1.4.4/go.mod h1:9T104GzyrTigFIr8wt5mBrctHMim0Nb2HLGrmQ40KvY=
@@ -90,8 +93,8 @@ github.com/gobwas/ws v1.4.0 h1:CTaoG1tojrh4ucGPcoJFiAQUAsEWekEWvLy7GsVNqGs=
github.com/gobwas/ws v1.4.0/go.mod h1:G3gNqMNtPppf5XUz7O4shetPpcZ1VJ7zt18dlUeakrc=
github.com/golang/protobuf v1.5.4 h1:i7eJL8qZTpSEXOPTxNKhASYpMn+8e5Q6AdndVa1dWek=
github.com/golang/protobuf v1.5.4/go.mod h1:lnTiLA8Wa4RWRcIUkrtSVa5nRhsEGBg48fD6rSs7xps=
github.com/gomarkdown/markdown v0.0.0-20260614204949-e08cff860f76 h1:Ltt9ldIaSYEsjA7sPY2c8r9dOmnKM1vlzhh3dxlhBHM=
github.com/gomarkdown/markdown v0.0.0-20260614204949-e08cff860f76/go.mod h1:JDGcbDT52eL4fju3sZ4TeHGsQwhG9nbDV21aMyhwPoA=
github.com/gomarkdown/markdown v0.0.0-20260824154242-13c5cf49db8d h1:8VtgBGEPLZ2Yn0Fuh6Pwmy3qF6indeaqy8mrBMbUKRQ=
github.com/gomarkdown/markdown v0.0.0-20260824154242-13c5cf49db8d/go.mod h1:JDGcbDT52eL4fju3sZ4TeHGsQwhG9nbDV21aMyhwPoA=
github.com/google/go-cmp v0.5.5/go.mod h1:v8dTdLbMG2kIc/vJvl+f65V22dbkXbowE6jgT/gNBxE=
github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8=
github.com/google/go-cmp v0.7.0/go.mod h1:pXiqmnSA92OHEEa9HXL2W4E7lf9JzCmGVUdgjX3N/iU=
@@ -99,8 +102,8 @@ github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0=
github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo=
github.com/gorilla/css v1.0.1 h1:ntNaBIghp6JmvWnxbZKANoLyuXTPZ4cAMlo6RyhlbO8=
github.com/gorilla/css v1.0.1/go.mod h1:BvnYkspnSzMmwRK+b8/xgNPLiIuNZr6vbZBTPQ2A3b0=
github.com/grpc-ecosystem/grpc-gateway/v2 v2.29.0 h1:5VipnvEpbqr2gA2VbM+nYVbkIF28c5ZQfqCBQ5g2xfk=
github.com/grpc-ecosystem/grpc-gateway/v2 v2.29.0/go.mod h1:Hyl3n6Twe1hvtd9XUXDec4pTvgMSEixRuQKPTMH2bNs=
github.com/grpc-ecosystem/grpc-gateway/v2 v2.30.0 h1:/Tnpcb2E0Pz/tN9s3bfEY2Q8ePCEX9iuS+cneUwncnw=
github.com/grpc-ecosystem/grpc-gateway/v2 v2.30.0/go.mod h1:zOBXOsUaBSjKgmH4OGzV1esUpR3oUSCPYVd2cUBjKYY=
github.com/hashicorp/go-cleanhttp v0.5.2 h1:035FKYIWjmULyFRBKPs8TBQoi0x6d9G4xc9neXJWAZQ=
github.com/hashicorp/go-cleanhttp v0.5.2/go.mod h1:kO/YDlP8L1346E6Sodw+PrpBSV4/SoxCXGY6BqNFT48=
github.com/hashicorp/go-hclog v1.6.3 h1:Qr2kF+eVWjTiYmU7Y31tYlP1h0q/X3Nl3tPGdaB11/k=
@@ -121,15 +124,11 @@ github.com/hashicorp/golang-lru v1.0.2/go.mod h1:iADmTwqILo4mZ8BN3D2Q6+9jd8WM5uG
github.com/hashicorp/golang-lru/v2 v2.0.7 h1:a+bsQ5rvGLjzHuww6tVxozPZFVghXaHOwFs4luLUK2k=
github.com/hashicorp/golang-lru/v2 v2.0.7/go.mod h1:QeFd9opnmA6QUJc5vARoKUSoFhyfM2/ZepoAG6RGpeM=
github.com/klauspost/compress v1.4.1/go.mod h1:RyIbtBH6LamlWaDj8nUwkbUhJ87Yi3uG0guNDohfE1A=
github.com/klauspost/compress v1.19.1 h1:VsB4HPswih7mmZ8WleSFQ75c/Ui1M4trX5oAsJnhSlk=
github.com/klauspost/compress v1.19.1/go.mod h1:cwPg85FWrGar70rWktvGQj8/hthj3wpl0PGDogxkrSQ=
github.com/klauspost/compress v1.20.0 h1:a3C1ke2ohxFymNlb2HWAHjDeKCI90scRskErZkR0ezA=
github.com/klauspost/compress v1.20.0/go.mod h1:LUdAzn7YLVvxLpc7y3V1m40wESHTgc1422pwwBSKYuI=
github.com/klauspost/cpuid v1.2.0/go.mod h1:Pj4uuM528wm8OyEC2QMXAi2YiTZ96dNQPGgoMS4s3ek=
github.com/klauspost/pgzip v1.2.6 h1:8RXeL5crjEUFnR2/Sn6GJNWtSQ3Dk8pq4CL3jvdDyjU=
github.com/klauspost/pgzip v1.2.6/go.mod h1:Ch1tH69qFZu15pkjo5kYi6mth2Zzwzt50oCQKQE9RUs=
github.com/kr/pretty v0.3.1 h1:flRD4NNwYAUpkphVc1HcthR4KEIFJ65n8Mw5qdRn3LE=
github.com/kr/pretty v0.3.1/go.mod h1:hoEshYVHaxMs3cyo3Yncou5ZscifuDolrwPKZanG3xk=
github.com/kr/text v0.2.0 h1:5Nx0Ya0ZqY2ygV366QzturHI13Jq95ApcVaJBhpS+AY=
github.com/kr/text v0.2.0/go.mod h1:eLer722TekiGuMkidMxC/pM04lWEeraHUUmBw8l2grE=
github.com/kylelemons/godebug v1.1.0 h1:RPNrshWIDI6G2gRW9EHilWtl7Z6Sb1BR0xunSBf0SNc=
github.com/kylelemons/godebug v1.1.0/go.mod h1:9/0rRGxNHcop5bhtWyNeEfOS8JIWk580+fNqagV/RAw=
github.com/labstack/echo/v4 v4.15.4 h1:DL45vVYa+BWE+XuW+zZNd9H0YEdZ80UAWJGcTVW4EVs=
@@ -138,72 +137,73 @@ github.com/labstack/gommon v0.5.0 h1:6VSQ2NOzsnEJ5W6+84E0RbcaDDmgB6NIAzWCczTEe6c
github.com/labstack/gommon v0.5.0/go.mod h1:Rzlg7HHy1maLfzBYGg9NZcVuz1sA68HHhLjhcEllYE0=
github.com/ledongthuc/pdf v0.0.0-20220302134840-0c2507a12d80 h1:6Yzfa6GP0rIo/kULo2bwGEkFvCePZ3qHDDTC3/J9Swo=
github.com/ledongthuc/pdf v0.0.0-20220302134840-0c2507a12d80/go.mod h1:imJHygn/1yfhB7XSJJKlFZKl/J+dCPAknuiaGOshXAs=
github.com/lufia/plan9stats v0.0.0-20260330125221-c963978e514e h1:Q6MvJtQK/iRcRtzAscm/zF23XxJlbECiGPyRicsX+Ak=
github.com/lufia/plan9stats v0.0.0-20260330125221-c963978e514e/go.mod h1:autxFIvghDt3jPTLoqZ9OZ7s9qTGNAWmYCjVFWPX/zg=
github.com/magiconair/properties v1.8.10 h1:s31yESBquKXCV9a/ScB3ESkOjUYYv+X0rg8SYxI99mE=
github.com/magiconair/properties v1.8.10/go.mod h1:Dhd985XPs7jluiymwWYZ0G4Z61jb3vdS329zhj2hYo0=
github.com/lufia/plan9stats v0.0.0-20260802145828-341c2f0c90b5 h1:eveIIGn4BGM3qknO74omf6HYr30/exH+eVUTuAgwjZ0=
github.com/lufia/plan9stats v0.0.0-20260802145828-341c2f0c90b5/go.mod h1:autxFIvghDt3jPTLoqZ9OZ7s9qTGNAWmYCjVFWPX/zg=
github.com/magiconair/properties v1.18.11 h1:j5ozYZl0zCjG7ahMDH0GWIobOvvUzT0BdAguG0ViKy0=
github.com/magiconair/properties v1.18.11/go.mod h1:Dhd985XPs7jluiymwWYZ0G4Z61jb3vdS329zhj2hYo0=
github.com/mattn/go-colorable v0.1.15 h1:+u9SLTRGnXv73cEsnsmoZBom+dMU88B2M0aDcWy0/jY=
github.com/mattn/go-colorable v0.1.15/go.mod h1:6LmQG8QLFO4G5z1gPvYEzlUgJ2wF+stgPZH1UqBm1s8=
github.com/mattn/go-isatty v0.0.22 h1:j8l17JJ9i6VGPUFUYoTUKPSgKe/83EYU2zBC7YNKMw4=
github.com/mattn/go-isatty v0.0.22/go.mod h1:ZXfXG4SQHsB/w3ZeOYbR0PrPwLy+n6xiMrJlRFqopa4=
github.com/mattn/go-isatty v0.0.24 h1:tGZZoVgT/KiqK1c8ocVLeDS8BSWMRd47J3Lbz7vsReI=
github.com/mattn/go-isatty v0.0.24/go.mod h1:nMCL3Zebbrt45jsMDgnfIwz6ydEQApk5oEI3HqDio6A=
github.com/mholt/archives v0.1.5 h1:Fh2hl1j7VEhc6DZs2DLMgiBNChUux154a1G+2esNvzQ=
github.com/mholt/archives v0.1.5/go.mod h1:3TPMmBLPsgszL+1As5zECTuKwKvIfj6YcwWPpeTAXF4=
github.com/microcosm-cc/bluemonday v1.0.27 h1:MpEUotklkwCSLeH+Qdx1VJgNqLlpY2KXwXFM08ygZfk=
github.com/microcosm-cc/bluemonday v1.0.27/go.mod h1:jFi9vgW+H7c3V0lb6nR74Ib/DIB5OBs92Dimizgw2cA=
github.com/mikelolasagasti/xz v1.0.1 h1:Q2F2jX0RYJUG3+WsM+FJknv+6eVjsjXNDV0KJXZzkD0=
github.com/mikelolasagasti/xz v1.0.1/go.mod h1:muAirjiOUxPRXwm9HdDtB3uoRPrGnL85XHtokL9Hcgc=
github.com/minio/minlz v1.1.1 h1:OGmft1V6AnI/Wme332U6bhG54nxEan+VFgkD7lat4KM=
github.com/minio/minlz v1.1.1/go.mod h1:qT0aEB35q79LLornSzeDH75LBf3aH1MV+jB5w9Wasec=
github.com/minio/minlz v1.2.0 h1:6IOBuiHg04QxvbFfgFLT/9sMaO/UhL7S+ApW1mK8q5A=
github.com/minio/minlz v1.2.0/go.mod h1:Ls9H7nlkASeCcdl5thjVD5Eraj6z+zGa7xtq57jIKD4=
github.com/moby/docker-image-spec v1.3.1 h1:jMKff3w6PgbfSa69GfNg+zN/XLhfXJGnEx3Nl2EsFP0=
github.com/moby/docker-image-spec v1.3.1/go.mod h1:eKmb5VW8vQEh/BAr2yvVNvuiJuY6UIocYsFu/DxxRpo=
github.com/moby/go-archive v0.2.0 h1:zg5QDUM2mi0JIM9fdQZWC7U8+2ZfixfTYoHL7rWUcP8=
github.com/moby/go-archive v0.2.0/go.mod h1:mNeivT14o8xU+5q1YnNrkQVpK+dnNe/K6fHqnTg4qPU=
github.com/moby/moby/api v1.55.0 h1:2/sexvQyqIWS8pRSCFddBfpW2qE7vR7FCL+vN8pxwMc=
github.com/moby/moby/api v1.55.0/go.mod h1:+RQ6wluLwtYaTd1WnPLykIDPekkuyD/ROWQClE83pzs=
github.com/moby/moby/client v0.5.1 h1:tYNaJno4c0HXz12y5BiqEDy0rVTYkWzI26lGvnTMiJw=
github.com/moby/moby/client v0.5.1/go.mod h1:odLstlZ6uSnfvAgVxMpvgmb8SUdd+siH2T0GBuxVAlM=
github.com/moby/go-archive v0.3.3 h1:OxxR9paxsluYi+zDUEXTTaIxtkK3viymW+Ka7vRhhME=
github.com/moby/go-archive v0.3.3/go.mod h1:Npdv43fFqlhZW7Xo8fbm3ZMYFvAGNviUPqX21VERbcE=
github.com/moby/moby/api v1.56.0 h1:GQzua3NA599ASSIICx0iFgiJeO9YkdDARvQsm23ZZuQ=
github.com/moby/moby/api v1.56.0/go.mod h1:sZ+THbVWkjOmBPPfbnzdD/G1LuIexWhqlSHHPTDQ1Uk=
github.com/moby/moby/client v0.6.0 h1:AJjEB21QPbXSXjDsZorFBoDZPhMrfbpaPLgSMAW9Bgs=
github.com/moby/moby/client v0.6.0/go.mod h1:OCo00wNRyA3m4lmJ228W3JbyCN4ZNNYjpOXiJydBdcQ=
github.com/moby/patternmatcher v0.6.1 h1:qlhtafmr6kgMIJjKJMDmMWq7WLkKIo23hsrpR3x084U=
github.com/moby/patternmatcher v0.6.1/go.mod h1:hDPoyOpDY7OrrMDLaYoY3hf52gNCR/YOUYxkhApJIxc=
github.com/moby/sys/mount v0.3.5 h1:eS3fsZTjHaBihwjp4/+5Z3jxqLXYsbwxqpVSfFv3M00=
github.com/moby/sys/mount v0.3.5/go.mod h1:WUQDO+/uCiCIkIztx8SrwIDVn2dtMFRBebRhpDFT71M=
github.com/moby/sys/mountinfo v0.7.2 h1:1shs6aH5s4o5H2zQLn796ADW1wMrIwHsyJ2v9KouLrg=
github.com/moby/sys/mountinfo v0.7.2/go.mod h1:1YOa8w8Ih7uW0wALDUgT1dTTSBrZ+HiBLGws92L2RU4=
github.com/moby/sys/sequential v0.7.0 h1:ASQNGNROJSuOO6LL6bPHbKvuZu6NU8P4ldPWk31zj/8=
github.com/moby/sys/sequential v0.7.0/go.mod h1:NfSTAp6V3fw4tmkD62PEcOKeZKquXT8VKCkf7aVR79o=
github.com/moby/sys/user v0.4.0 h1:jhcMKit7SA80hivmFJcbB1vqmw//wU61Zdui2eQXuMs=
github.com/moby/sys/user v0.4.0/go.mod h1:bG+tYYYJgaMtRKgEmuueC0hJEAZWwtIbZTB+85uoHjs=
github.com/moby/sys/userns v0.1.0 h1:tVLXkFOxVu9A64/yh59slHVv9ahO9UIev4JZusOLG/g=
github.com/moby/sys/userns v0.1.0/go.mod h1:IHUYgu/kao6N8YZlp9Cf444ySSvCmDlmzUcYfDHOl28=
github.com/moby/sys/user v0.4.1 h1:RgjRlaDKi/Xmyrz4t8lyzXT6v2ooFeO/7xtchmhVWE0=
github.com/moby/sys/user v0.4.1/go.mod h1:E9QsW5WRe1kUAf7kW8hXKwu1uhsZEAdPLYHYSDudF4Y=
github.com/moby/sys/userns v0.2.0 h1:nEtDtp7NCV/6dutSklNe8FrENPwFdc4mXnZqC/JWgXM=
github.com/moby/sys/userns v0.2.0/go.mod h1:IHUYgu/kao6N8YZlp9Cf444ySSvCmDlmzUcYfDHOl28=
github.com/moby/term v0.5.2 h1:6qk3FJAFDs6i/q3W/pQ97SX192qKfZgGjCQqfCJkgzQ=
github.com/moby/term v0.5.2/go.mod h1:d3djjFCrjnB+fl8NJux+EJzu0msscUP+f8it8hPkFLc=
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 h1:C3w9PqII01/Oq1c1nUAm88MOHcQC9l5mIlSMApZMrHA=
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822/go.mod h1:+n7T8mK8HuQTcFwEeznm/DIxMOiR9yIdICNftLE1DvQ=
github.com/nwaples/rardecode/v2 v2.2.5 h1:L5doqgGfQwI7qADJMqnkrSB86rpPsqQDrHeO0HWa5JY=
github.com/nwaples/rardecode/v2 v2.2.5/go.mod h1:7uz379lSxPe6j9nvzxUZ+n7mnJNgjsRNb6IbvGVHRmw=
github.com/nwaples/rardecode/v2 v2.4.1 h1:F7zNW2LdAuuBThHWXQaiFUGVD/sef299NfWSB1nHAl4=
github.com/nwaples/rardecode/v2 v2.4.1/go.mod h1:7uz379lSxPe6j9nvzxUZ+n7mnJNgjsRNb6IbvGVHRmw=
github.com/opencontainers/go-digest v1.0.0 h1:apOUWs51W5PlhuyGyz9FCeeBIOUDA/6nW8Oi/yOhh5U=
github.com/opencontainers/go-digest v1.0.0/go.mod h1:0JzlMkj0TRzQZfJkVvzbP0HBR3IKzErnv2BNG4W4MAM=
github.com/opencontainers/image-spec v1.1.1 h1:y0fUlFfIZhPF1W537XOLg0/fcx6zcHCJwooC2xJA040=
github.com/opencontainers/image-spec v1.1.1/go.mod h1:qpqAh3Dmcf36wStyyWU+kCeDgrGnAve2nCC8+7h8Q0M=
github.com/orisano/pixelmatch v0.0.0-20220722002657-fb0b55479cde h1:x0TT0RDC7UhAVbbWWBzr41ElhJx5tXPWkIHA2HWPRuw=
github.com/orisano/pixelmatch v0.0.0-20220722002657-fb0b55479cde/go.mod h1:nZgzbfBr3hhjoZnS66nKrHmduYNpc34ny7RK4z5/HM0=
github.com/pierrec/lz4/v4 v4.1.27 h1:+PhzhWDrjRj89TH2sw43nE3+4+W8lSxIuQadEHZyjUk=
github.com/pierrec/lz4/v4 v4.1.27/go.mod h1:EoQMVJgeeEOMsCqCzqFm2O0cJvljX2nGZjcRIPL34O4=
github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM=
github.com/pierrec/lz4/v4 v4.1.29 h1:CDQY6qZOLI4DW0Nx6R1vRrifrCeQHnNXkMb0hZWXFjg=
github.com/pierrec/lz4/v4 v4.1.29/go.mod h1:EoQMVJgeeEOMsCqCzqFm2O0cJvljX2nGZjcRIPL34O4=
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/power-devops/perfstat v0.0.0-20240221224432-82ca36839d55 h1:o4JXh1EVt9k/+g42oCprj/FisM4qX9L3sZB3upGN2ZU=
github.com/power-devops/perfstat v0.0.0-20240221224432-82ca36839d55/go.mod h1:OmDBASR4679mdNQnz2pUhc2G8CO2JrUAVFDRBDP/hJE=
github.com/power-devops/perfstat v0.0.0-20260805114148-88456608a4f6 h1:jL3a8soXdzuTCcRnKhOmtcsVOObdDTFf4O2B403HPRU=
github.com/power-devops/perfstat v0.0.0-20260805114148-88456608a4f6/go.mod h1:OmDBASR4679mdNQnz2pUhc2G8CO2JrUAVFDRBDP/hJE=
github.com/prometheus/client_golang v1.24.1 h1:JnJkREXzWxUdCuPFpIWZiPispT9xVV59uiuyR2bPlnU=
github.com/prometheus/client_golang v1.24.1/go.mod h1:F+oSRECHg4sse5ucfYpYDeIv/hu68Zo0uoHKetWnzcE=
github.com/prometheus/client_model v0.6.2 h1:oBsgwpGs7iVziMvrGhE53c/GrLUsZdHnqNwqPLxwZyk=
github.com/prometheus/client_model v0.6.2/go.mod h1:y3m2F6Gdpfy6Ut/GBsUqTWZqCUvMVzSfMLjcu6wAwpE=
github.com/prometheus/common v0.70.1 h1:1HvjP4D5oL3t8RsPlwxA9onvvStjtIHYE5XuuwOi/PY=
github.com/prometheus/common v0.70.1/go.mod h1:VdFUQDMZK3VLkurFUVhia6uys/0suUp86TJz5qbJRhc=
github.com/prometheus/client_model v0.6.3 h1:O0jaTVAYNxTHYInEPFJt5I3+sN8zqBtVMPTB1qyxiEo=
github.com/prometheus/client_model v0.6.3/go.mod h1:gpN5P9S7Rr6Yr92PiQ+Ixvhf6JZEkF1dnxsYL2aPBEM=
github.com/prometheus/common v0.71.0 h1:9KDAKb7Mj3HEVKyFCK6Dc/HIwlBzZIN2l7/lrHl3KK8=
github.com/prometheus/common v0.71.0/go.mod h1:CLJ5H8TEsGX8bl31BdMkfhIZ+QmZ9tBPPotUxUbfcmk=
github.com/prometheus/otlptranslator v1.0.0 h1:s0LJW/iN9dkIH+EnhiD3BlkkP5QVIUVEoIwkU+A6qos=
github.com/prometheus/otlptranslator v1.0.0/go.mod h1:vRYWnXvI6aWGpsdY/mOT/cbeVRBlPWtBNDb7kGR3uKM=
github.com/prometheus/procfs v0.21.1 h1:GljZCt+zSTS+NZq88cyQ1LjZ+RCHp3uVuabBWA5+OJI=
github.com/prometheus/procfs v0.21.1/go.mod h1:aB55Cww9pdSJVHk0hUf0inxWyyjPogFIjmHKYgMKmtY=
github.com/rogpeppe/go-internal v1.14.1 h1:UQB4HGPB6osV0SQTLymcB4TgvyWu6ZyliaW0tI/otEQ=
github.com/rogpeppe/go-internal v1.14.1/go.mod h1:MaRKkUm5W0goXpeCfT7UZI6fk/L7L7so1lCWt35ZSgc=
github.com/shirou/gopsutil/v4 v4.26.7 h1:IXzpHz/dkMRYAhKkOXr1HB6SuzWU3eoyyeWe7g3bNZc=
github.com/shirou/gopsutil/v4 v4.26.7/go.mod h1:5O9FjBiXoTDFatIWjZZosqj4pV0DRtLx598xGbBehzM=
github.com/sirupsen/logrus v1.9.4 h1:TsZE7l11zFCLZnZ+teH4Umoq5BhEIfIzfRDZ1Uzql2w=
github.com/sirupsen/logrus v1.9.4/go.mod h1:ftWc9WdOfJ0a92nsE2jF5u5ZwH8Bv2zdeOC42RjbV2g=
github.com/prometheus/procfs v0.22.0 h1:6q9+/JL9IKAPbCmBrv9n5O5Ty3NKnciV5X7YGw0oics=
github.com/prometheus/procfs v0.22.0/go.mod h1:CvmFr/GVhIjIvWJZW3tgkODBQMRIf0EyWMQLHCHab58=
github.com/shirou/gopsutil/v4 v4.26.8 h1:YQMTF/1J50B5+Y0vlo1eDRf5DoR7Gk69hY+8wjYkQeo=
github.com/shirou/gopsutil/v4 v4.26.8/go.mod h1:5O9FjBiXoTDFatIWjZZosqj4pV0DRtLx598xGbBehzM=
github.com/sirupsen/logrus v1.10.2 h1:G2SED73/qrAu6YwbdxOD6peLkCBI3z7L+ykJFTXJBBo=
github.com/sirupsen/logrus v1.10.2/go.mod h1:SLEg8TqYulVKKfIGHldVp2K2aYz2DKSVBq4g/H5bR7Q=
github.com/sorairolake/lzip-go v0.3.8 h1:j5Q2313INdTA80ureWYRhX+1K78mUXfMoPZCw/ivWik=
github.com/sorairolake/lzip-go v0.3.8/go.mod h1:JcBqGMV0frlxwrsE9sMWXDjqn3EeVf0/54YPsw66qkU=
github.com/spf13/afero v1.15.0 h1:b/YBCLWAJdFWJTN9cLhiXXcD7mzKn9Dm86dNnfyQw1I=
@@ -220,17 +220,17 @@ github.com/stretchr/objx v0.5.3/go.mod h1:rDQraq+vQZU7Fde9LOZLr8Tax6zZvy4kuNKF+Q
github.com/stretchr/testify v1.7.1/go.mod h1:6Fq8oRcR53rry900zMqJjRRixrwX3KX962/h/Wwjteg=
github.com/stretchr/testify v1.8.0/go.mod h1:yNjHg4UonilssWZ8iaSj1OCr/vHnekPRkoO+kdMU+MU=
github.com/stretchr/testify v1.8.1/go.mod h1:w2LPCIKwWwSfY2zedu0+kehJoqGctiVI29o6fzry7u4=
github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U=
github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U=
github.com/testcontainers/testcontainers-go v0.43.0 h1:oEQx5MW2DGd9z3AeEQfB2lPM0eLs7ztyaGRu75bFo5A=
github.com/testcontainers/testcontainers-go v0.43.0/go.mod h1:+VxkT2NQnKOZPKi6praMuMKYHYyOGXr0XSBSlSMCzFo=
github.com/stretchr/testify v1.12.1 h1:EuwCh5fleGS7H32xRwO3wRGT7DxrDhLAT6FF8MpWDWE=
github.com/stretchr/testify v1.12.1/go.mod h1:MDEgiDPPsNp5cuIrHPPCyornHKgEVbtFUmoNlxoYthg=
github.com/testcontainers/testcontainers-go v0.44.0 h1:/Fwh6HY1mIikhnm9e7HwoxGycx0lzRAE0f5VQpjFxzI=
github.com/testcontainers/testcontainers-go v0.44.0/go.mod h1:IcnwQrYTO86xHXu5bvMaBH7ATlbS3Qn1M1QWW3c66rE=
github.com/tklauser/go-sysconf v0.4.0 h1:7H0uAN+7RkwWRaxhYXDLqa5V3LPrJeV8wmD9dRUgPQU=
github.com/tklauser/go-sysconf v0.4.0/go.mod h1:8mTNWyog7H+MpKijp4VmKJAd2bbYQ2zuUwkYRbUArPI=
github.com/tklauser/numcpus v0.12.0 h1:NR85qdvHA9pFse3x3weVZ0r0ST8R6l5RHbZrlRaqob4=
github.com/tklauser/numcpus v0.12.0/go.mod h1:ABHeXzJnr/qqwguhClkZKT1/8VABcYrsyUiUGobwWJg=
github.com/ulikunitz/xz v0.5.8/go.mod h1:nbz6k7qbPmH4IRqmfOplQw/tblSgqTqBwxkY0oWt/14=
github.com/ulikunitz/xz v0.5.15 h1:9DNdB5s+SgV3bQ2ApL10xRc35ck0DuIX/isZvIk+ubY=
github.com/ulikunitz/xz v0.5.15/go.mod h1:nbz6k7qbPmH4IRqmfOplQw/tblSgqTqBwxkY0oWt/14=
github.com/ulikunitz/xz v0.5.16 h1:ld6NyySjx5lowVKwJvMRLnW5nxKX/xnpSiFYZ/Lxur0=
github.com/ulikunitz/xz v0.5.16/go.mod h1:H9Rt/W6/Qj27PGauhQc6nfCDy7vHpzsOThBSaYDoEhw=
github.com/valyala/bytebufferpool v1.0.0 h1:GqA5TC/0021Y/b9FG4Oi9Mr3q7XYx6KllzawFIhcdPw=
github.com/valyala/bytebufferpool v1.0.0/go.mod h1:6bBcMArwyJ5K/AmCkWv1jt77kVWyCJ6HpOuEn7z0Csc=
github.com/valyala/fasttemplate v1.2.2 h1:lxLXG0uE3Qnshl9QyaK6XJxMXlQZELvChBOCmQD0Loo=
@@ -241,68 +241,72 @@ github.com/yusufpapurcu/wmi v1.2.4 h1:zFUKzehAFReQwLys1b/iSMl+JQGSCSjtVqQn9bBrPo
github.com/yusufpapurcu/wmi v1.2.4/go.mod h1:SBZ9tNy3G9/m5Oi98Zks0QjeHVDvuK0qfxQmPyzfmi0=
go.opentelemetry.io/auto/sdk v1.2.1 h1:jXsnJ4Lmnqd11kwkBV2LgLoFMZKizbCi5fNZ/ipaZ64=
go.opentelemetry.io/auto/sdk v1.2.1/go.mod h1:KRTj+aOaElaLi+wW1kO/DZRXwkF4C5xPbEe3ZiIhN7Y=
go.opentelemetry.io/contrib/bridges/otelslog v0.19.0 h1:5RgvxieNq9tS3ewrV1vnODvbHPfKUIJcYtF9Cvz+6aQ=
go.opentelemetry.io/contrib/bridges/otelslog v0.19.0/go.mod h1:iTBIdNwx/xmUhfgJs6+84S4dIK059811cO1eUBjKcHY=
go.opentelemetry.io/contrib/bridges/prometheus v0.69.0 h1:saQoWg5845Q8TojpqeVStS7zGwVZ6bc5W2PJavTPiBM=
go.opentelemetry.io/contrib/bridges/prometheus v0.69.0/go.mod h1:AAaS6xs5AyqMdR3Ir0nSWK+QudL2XM8Vbw5INzUxNc8=
go.opentelemetry.io/contrib/exporters/autoexport v0.69.0 h1:R3jsCoTIzv0BiYNhW0axyswn/6SMJ8xL1OuGxvni1Kw=
go.opentelemetry.io/contrib/exporters/autoexport v0.69.0/go.mod h1:m07gqyr2QhQxKOKb5vqKCCBtLH3uqlNYR7PU/FISXVU=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 h1:8tvICD4vSTOOsNrsI4Ljf6C+6UKvpTEH5XY3JMoyPoo=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0/go.mod h1:z9+yiacE0IHRqM4qFfkbt/JYlmYXgss8GY/jXoNuPJI=
go.opentelemetry.io/otel v1.45.0 h1:pdrWmLHofpubmArBv1LgFSv1Z0Ie/ppdZzu+kUN5EeU=
go.opentelemetry.io/otel v1.45.0/go.mod h1:XZxIqPapzEYnhNSScF5DIqXhm/rYi0FzCe2XddAwZfQ=
go.opentelemetry.io/otel/exporters/otlp/otlplog/otlploggrpc v0.20.0 h1:rydZ9sxbcFdm/oWrVyfLTjHIygMgv0bEeMd+3B/BvoM=
go.opentelemetry.io/otel/exporters/otlp/otlplog/otlploggrpc v0.20.0/go.mod h1:earQ25dooT0Hhspq59DZ8YCC50jWfOlFEeWoxy/P444=
go.opentelemetry.io/otel/exporters/otlp/otlplog/otlploghttp v0.20.0 h1:owlhcJ3QO3X0YTDTCcDZ4V+6aVDkWbNmBoQ5NUp7Oww=
go.opentelemetry.io/otel/exporters/otlp/otlplog/otlploghttp v0.20.0/go.mod h1:MP4eemTiI9zC8fgg+DYynhYDYf3ba72S376TvP+Ye0Q=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetricgrpc v1.44.0 h1:SUplec5dp06reu1zaXmOXdvqH398taqrDXqUl99jxSc=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetricgrpc v1.44.0/go.mod h1:ho2g4N+ane+swq5I/VBkKWnRDY4kUINH3FuqyZqX/Ug=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.44.0 h1:RuynHbfU8JUEw7DyONgkVYg2SVtsoF28y0LGIr69jgA=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.44.0/go.mod h1:qZF+/lBs71APw8mlnEZcqZHMzqrYrsFiJOv83lX1OGo=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0 h1:4YsVu3B8+3qtWYYrsUYgn0OG78pN0rnNPRGX4SbokQI=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0/go.mod h1:+wnlSn0mD1ADVMe3v9Z/WIaiz6q6gL2J/ejaAmdmv80=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0 h1:qazEJlUOQzhCpzQpFETGby7EdqjI1wsd0W+6Gg1SCTU=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0/go.mod h1:fOD2Yefuxixkx3ahVNf0O/PERb6r4OlbxfATVnYvzCo=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp v1.44.0 h1:lgh3PiVrRUWMLOVSkQicxzZll5NjF1r+AtsX1XRIHw0=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp v1.44.0/go.mod h1:5Cnhth3m/AgOeTgE3ex12pPmiu/gGtZit03kSzx9X7s=
go.opentelemetry.io/otel/exporters/prometheus v0.66.0 h1:vkrK8PAznv2NKt2r+kdu252ccGzkEqLc2aSXbQIALYQ=
go.opentelemetry.io/otel/exporters/prometheus v0.66.0/go.mod h1:V/UB6D3vMF/UBOL5igAsAYnk1nG/bzYYTzvsB16cy7o=
go.opentelemetry.io/otel/exporters/stdout/stdoutlog v0.20.0 h1:aZfdmtI6QU/DAPD4b7YZ5zuJgewxO1EW9miOZklqleU=
go.opentelemetry.io/otel/exporters/stdout/stdoutlog v0.20.0/go.mod h1:isNl10/Om5CBWu9jj8WOb2+tJLbCVXDgqwzCaJMnJ6w=
go.opentelemetry.io/otel/exporters/stdout/stdoutmetric v1.44.0 h1:hqxVTu/GtBF+vJ8d1fzW7fRxZFvgoDjWcxwwCaFDYpU=
go.opentelemetry.io/otel/exporters/stdout/stdoutmetric v1.44.0/go.mod h1:z5fVEF4X5v0ESvlJqBrrFlBVoj5EQuefZpzsu7R+x5Q=
go.opentelemetry.io/otel/exporters/stdout/stdouttrace v1.44.0 h1:bl2S7Ubua0Nms+D/gAmznQTd4dxxMA93aKbcpKqiTCs=
go.opentelemetry.io/otel/exporters/stdout/stdouttrace v1.44.0/go.mod h1:L0hRV50XdVIODHUfWEqGRCXQvj2rV82STVo12FMFBU0=
go.opentelemetry.io/otel/log v0.20.0 h1:/5i0vuHxCLWUfChWG41K9wkM0jafruPw9NU1/RCJirs=
go.opentelemetry.io/otel/log v0.20.0/go.mod h1:wOcMcjsZpG8x7Bak7IhSi/lg8wscV2C1VdrKCLPlt0E=
go.opentelemetry.io/otel/metric v1.45.0 h1:7Eg1uH7CJ5cXv9is6tnBe1FI6rj1nwUdbFypRm3br/M=
go.opentelemetry.io/otel/metric v1.45.0/go.mod h1:HAPbm1nd3p1PmFH7v2dR+6BjXxw+Lq4a2+pndMAm08s=
go.opentelemetry.io/otel/metric/x v0.67.0 h1:PcicCNZFkZ4bXfSooXdo3WN7RBOVOtjVdo1wD358Uns=
go.opentelemetry.io/otel/metric/x v0.67.0/go.mod h1:FBjCWZe6wgcqxcMtjdGiClDKXb2YxxXii0CXftE4QtI=
go.opentelemetry.io/otel/sdk v1.45.0 h1:4VVSMgQ83dUgW2aoX5f6JgLvHwIvzcuLnF9lUdCSpCw=
go.opentelemetry.io/otel/sdk v1.45.0/go.mod h1:Sr40LgXV7DsKMMJMKOhUWOgMWTfAaqvm2kF0g7ilwuA=
go.opentelemetry.io/otel/sdk/log v0.20.0 h1:vM3xI7TQgKPiSghe6urZtAkyFY7SodrSpC83CffDFuY=
go.opentelemetry.io/otel/sdk/log v0.20.0/go.mod h1:Knej2nmsTUzN79T2eeXdRsjjPcoxoq2pUyUHz9TFyyU=
go.opentelemetry.io/otel/sdk/log/logtest v0.20.0 h1:OqdRZ1guyzamK3M6LlRsmGqRrjkHWw6WZOKKli5ELpg=
go.opentelemetry.io/otel/sdk/log/logtest v0.20.0/go.mod h1:PuMIlm7zAt7c3z8zfOI5ox4iT1Z87We+PF6YoINux/M=
go.opentelemetry.io/otel/sdk/metric v1.45.0 h1:oVFszMfyj1Am6s24Vtc7wBb8BKLcwepJjNEYILuiE3o=
go.opentelemetry.io/otel/sdk/metric v1.45.0/go.mod h1:vUWUxDZvu1WVRj8JA8S0AdhsPrZoDpA2DdZauIh4mDA=
go.opentelemetry.io/otel/trace v1.45.0 h1:l/mP6Uv7oNO7/TblbhpbgMidxhq1uO/rPsikOyVhxag=
go.opentelemetry.io/otel/trace v1.45.0/go.mod h1:qoJJA2xNMnxRrdISU/kLtfUH2wNeQbiv+jhs/CxI8bc=
go.opentelemetry.io/proto/otlp v1.10.0 h1:IQRWgT5srOCYfiWnpqUYz9CVmbO8bFmKcwYxpuCSL2g=
go.opentelemetry.io/proto/otlp v1.10.0/go.mod h1:/CV4QoCR/S9yaPj8utp3lvQPoqMtxXdzn7ozvvozVqk=
go.opentelemetry.io/contrib/bridges/otelslog v0.20.1 h1:5sHc4ToTFjfSZCtGAAM6jPunICAmJX73htv372T4ipc=
go.opentelemetry.io/contrib/bridges/otelslog v0.20.1/go.mod h1:oa6kgvyz/3GYW04dohd0++xJIH4xdQY8PAbpeCMaM8M=
go.opentelemetry.io/contrib/bridges/prometheus v0.71.0 h1:9qgxsFLskbDMXl8WMqThoF6w8yGJgCumn9qRc67OmnI=
go.opentelemetry.io/contrib/bridges/prometheus v0.71.0/go.mod h1:2rCjF4F2siiTeLCzJsaGZ3CK0XIoimCSKXEBPdv+Je0=
go.opentelemetry.io/contrib/exporters/autoexport v0.71.0 h1:VCsJbp0YLyPtx2tu5Vgv2a2/qLoaMCj8hT2uZ34+Mx0=
go.opentelemetry.io/contrib/exporters/autoexport v0.71.0/go.mod h1:qxZqn7e10f6ajmMCkg/47rMS7qQYfaOl2nj/4aytHUQ=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.71.0 h1:3g7B90UzBltIDKq1/5mrTGxTnOFDV0ICOhLoxiZ8jlg=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.71.0/go.mod h1:Ef8SuTh59BT7+ofpDxN9z+yOlc4t2GjLmKDgYNJL/NU=
go.opentelemetry.io/otel v1.46.0 h1:FHt5/CDyVxi/8IM1CH7VE/rRgq3kLHa2mSTVMO8AWyc=
go.opentelemetry.io/otel v1.46.0/go.mod h1:Gj3SEScelsNC45tp4nSxRYlS+f5iez7W8XPMCt905kE=
go.opentelemetry.io/otel/exporters/otlp/otlplog/otlploggrpc v0.22.0 h1:Bu39F5tzJct+f2IZbB8989fwyTps3c8e7EsUQsz+vs8=
go.opentelemetry.io/otel/exporters/otlp/otlplog/otlploggrpc v0.22.0/go.mod h1:dJUwod88EsFgYCqrDHaSPzhiY9pBUpt0d85/qSfua7k=
go.opentelemetry.io/otel/exporters/otlp/otlplog/otlploghttp v0.22.0 h1:lYk7RmxdLK865qLwibroNGldHa1U7SWKYYvNjlK7PIo=
go.opentelemetry.io/otel/exporters/otlp/otlplog/otlploghttp v0.22.0/go.mod h1:6GvlND0H0xdUJanOtIAn0xfwLkauh1tmsYEEVSMDdqY=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetricgrpc v1.46.0 h1:qkDYCAFiZXLcs1L4aY+tP2wguQ4kURANqHOQMA2et2s=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetricgrpc v1.46.0/go.mod h1:tkipS4DRzmpAmvg+Gw4++O1IdDq6TVDnvnYU6cmbQVs=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.46.0 h1:AP23h/mFgb/lc7tdck1Kfn9qxsM8TAeNPCU5C3pzaps=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.46.0/go.mod h1:K4EqCe1b4kGk5WR690ntg9LaBfsPoV32FwthbyoptuA=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.46.0 h1:OFnwLJr+pF3iHrlGSzbxyuo6/6HyBlnlN1CWEJmBVcw=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.46.0/go.mod h1:716wFneO0ov19A2beH5hjfh9AK5z/VWNAtDijp1Y0/g=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.46.0 h1:w53CDeOA/Kurp7yRsegSr6pbbr759dOvJ+yNmWM6Hxs=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.46.0/go.mod h1:BOmGMCbAtvcJiSJ+hLuhgPLdDbimnraSl8irz3iY8sY=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp v1.46.0 h1:KrC1YrQeSt46ITMWAbgQx1M1eV1/1TKzttrBzymPmss=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp v1.46.0/go.mod h1:zDSEzoEqsOrgBeGvH66KRgxh90VonFyJqBHA0Pk3+rM=
go.opentelemetry.io/otel/exporters/prometheus v0.68.0 h1:QOf2IftqQwITVRJpnn0M7M9ZCbgWfxz4P7i9C9yc2N4=
go.opentelemetry.io/otel/exporters/prometheus v0.68.0/go.mod h1:bgSvqu2TWGXiz7yr5UTMfObH8oqxJWHTnubQ3ef9BO4=
go.opentelemetry.io/otel/exporters/stdout/stdoutlog v0.22.0 h1:kvMAiLEudKmk+CSG+iYbU8vTUGNNDaf/V09OO5lrTwI=
go.opentelemetry.io/otel/exporters/stdout/stdoutlog v0.22.0/go.mod h1:L9Dlksri+MdT1cb2gIiA1cJJYW3Y92ipvDjNxYEyaDI=
go.opentelemetry.io/otel/exporters/stdout/stdoutmetric v1.46.0 h1:PR9eAf7o0dQs3hshZNZpE9aW2dXWX/KdDf6pJilVD3U=
go.opentelemetry.io/otel/exporters/stdout/stdoutmetric v1.46.0/go.mod h1:2Z4KyNdH1uuzivdinyfGsxzNNT/Rl45pwtVwfYVI0xk=
go.opentelemetry.io/otel/exporters/stdout/stdouttrace v1.46.0 h1:KdRxPiAoMptR3vfWzvjjvutTsSiwbC2uG0496rzZNfo=
go.opentelemetry.io/otel/exporters/stdout/stdouttrace v1.46.0/go.mod h1:K/qSA+3G7Eovxi4K09wzrAgkWRnosS0DAOZeEpve7sM=
go.opentelemetry.io/otel/log v0.22.0 h1:5DBNnfvaJ6CVdkJ+Jle8Tzs50aSSv49TXGj9XRsEYw0=
go.opentelemetry.io/otel/log v0.22.0/go.mod h1:gzOt/R67vF2GniAqWu8Qv0SXy89f71muHcrkz76PCdc=
go.opentelemetry.io/otel/metric v1.46.0 h1:yBnkXvgV7AXFILZc5K6IZe/CBFF3OS7BJ8ov6/lj0K8=
go.opentelemetry.io/otel/metric v1.46.0/go.mod h1:iPmdWqifKUdzziPkvvzIJXITl56fQx2mGM/DHLB3/2o=
go.opentelemetry.io/otel/metric/x v0.68.0 h1:TA/cBT23D3MnxYPwHL7YFOdYGdx0A0v+s7Mzotpd1dU=
go.opentelemetry.io/otel/metric/x v0.68.0/go.mod h1:agudOmvWhwUTjgibWDzxD2PoWYnpw5Ht5jISYOD2Hd4=
go.opentelemetry.io/otel/sdk v1.46.0 h1:h5CNQQjEbuQXY/JfZtgt3i7HVFV3aHPO2OAwO2eTYPI=
go.opentelemetry.io/otel/sdk v1.46.0/go.mod h1:GAERFXFt5SYCEB+YiKUbMBeza6UaDH7GmGOZEfh2gSM=
go.opentelemetry.io/otel/sdk/log v0.22.0 h1:PRL+s6P63XT4E/bheEflopPUpVxuvANqZwtt89yhoGk=
go.opentelemetry.io/otel/sdk/log v0.22.0/go.mod h1:JNp0sBELrjCTcu5W3GzABVypeU6vDJjBS+X0JISuz+g=
go.opentelemetry.io/otel/sdk/log/logtest v0.22.0 h1:infPnfNrhCNgOUZRs3gWUg8vhoBUHihq02gwK05gzlg=
go.opentelemetry.io/otel/sdk/log/logtest v0.22.0/go.mod h1:gkQZA3z15Bv3KU9vigBTi8dFechSozRP7v94X4VZv+s=
go.opentelemetry.io/otel/sdk/metric v1.46.0 h1:0piZ26EG4RBfebb2jhDH6ERCYHoVWduc3kLgPCwSnSE=
go.opentelemetry.io/otel/sdk/metric v1.46.0/go.mod h1:I1PbKrdVc8Qu8HYVDNtqVIwLwjNrhsV/uFuxfwg8mO4=
go.opentelemetry.io/otel/trace v1.46.0 h1:OULy7ccdJnZtJ0UDYFOIGaCmiWzJ8Vi2G/Rsu60qs1c=
go.opentelemetry.io/otel/trace v1.46.0/go.mod h1:J7GAXweO77XSFkB/rmAqk9D6ihszhFjLU+d9WuUxDLI=
go.opentelemetry.io/proto/otlp v1.11.0 h1:5rrYs0Ykyj50sdU/JU0x8etU+LubXWb+gED6TbEdMIk=
go.opentelemetry.io/proto/otlp v1.11.0/go.mod h1:SmVizdCOAm3XBtG1g1NnOdhW6jtddT72hLMhv8VwA8E=
go.uber.org/goleak v1.3.0 h1:2K3zAYmnTNqV73imy9J1T3WC+gmCePx2hEGkimedGto=
go.uber.org/goleak v1.3.0/go.mod h1:CoHD4mav9JJNrW/WLlf7HGZPjdw8EucARQHekz1X6bE=
go.yaml.in/yaml/v2 v2.4.4 h1:tuyd0P+2Ont/d6e2rl3be67goVK4R6deVxCUX5vyPaQ=
go.yaml.in/yaml/v2 v2.4.4/go.mod h1:gMZqIpDtDqOfM0uNfy0SkpRhvUryYH0Z6wdMYcacYXQ=
go.yaml.in/yaml/v3 v3.0.5 h1:N6y/pJk8buWs9NY5ERU2HSMfm+IuD/OtfdAnq6kESPw=
go.yaml.in/yaml/v3 v3.0.5/go.mod h1:HVTZu1O7/Vkt2N+BFy8Zza+lnLsABggaTM2ZpNIGuKg=
go4.org v0.0.0-20260112195520-a5071408f32f h1:ziUVAjmTPwQMBmYR1tbdRFJPtTcQUI12fH9QQjfb0Sw=
go4.org v0.0.0-20260112195520-a5071408f32f/go.mod h1:ZRJnO5ZI4zAwMFp+dS1+V6J6MSyAowhRqAE+DPa1Xp0=
golang.org/x/crypto v0.54.0 h1:YLIA59K4fiNzHzjnZt2tUJQjQtUWfWbeHBqKtk3eScw=
golang.org/x/crypto v0.54.0/go.mod h1:KWL8ny2AZdGR2cWmzeHrp2azQPGogOv+HeQaVEXC2dk=
golang.org/x/net v0.57.0 h1:K5+3DljvIuDG9/Jv9rvyMywYNFCQ9RSUY6OOTTkT+tE=
golang.org/x/net v0.57.0/go.mod h1:KpXc8iv+r3XplLAG/f7Jsf9RPszJzdR0f58q9vGOuEU=
golang.org/x/sync v0.22.0 h1:SZjpbeLmrCk4xhRSZFNZW5gFUeCeFgjekvI/+gfScek=
golang.org/x/sync v0.22.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
golang.org/x/crypto v0.56.0 h1:GUh5Ii4J5jtcseSMiRqr1jXCNHoxjeV9Fmekc2oLy6Y=
golang.org/x/crypto v0.56.0/go.mod h1:OMW5y6CY9l38uPLmxU6l6pwcXp1obtLo3e6gT7gQR2I=
golang.org/x/net v0.58.0 h1:ynWG7rqYi4ccpTEuPZ2QGWHktVEM9DMCj9yzDE0Q7To=
golang.org/x/net v0.58.0/go.mod h1:YwCddHnFlT7eLQqVprV19OnhLGtc5xOKgE0RyqgfWAU=
golang.org/x/oauth2 v0.36.0 h1:peZ/1z27fi9hUOFCAZaHyrpWG5lwe0RJEEEeH0ThlIs=
golang.org/x/oauth2 v0.36.0/go.mod h1:YDBUJMTkDnJS+A4BP4eZBjCqtokkg1hODuPjwiGPO7Q=
golang.org/x/sync v0.23.0 h1:KameEIfc1IkluZyXWLn39Wd4tURc6GbCiISGiZm2bQk=
golang.org/x/sync v0.23.0/go.mod h1:sUUOizhqBxiL6pEWpqNLUiaJn1ShEbZ6BBqskPbjZm0=
golang.org/x/sys v0.0.0-20190916202348-b4ddaad3f8a3/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20201204225414-ed752295db88/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20210616094352-59db8d763f22/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
@@ -312,26 +316,23 @@ golang.org/x/sys v0.47.0 h1:o7XGOvZQCADBQQ4Y7VNq2dRWQR7JmOUW8Kxx4ZsNgWs=
golang.org/x/sys v0.47.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/term v0.45.0 h1:NwWyBmoJCbfTHpxrWoZ9C6/VxOf7ic219I8xZZFdrf0=
golang.org/x/term v0.45.0/go.mod h1:9aqxs0blBcrm/n0L9QW0aRVD+ktan8ssZromtqJC43w=
golang.org/x/text v0.40.0 h1:Ub2Z6/xjgF1WrYQz2nuITOEegKFtiIy+rieRJ5lHZKs=
golang.org/x/text v0.40.0/go.mod h1:hpnzDAfGV753zIKo+wk3u1bVKCGPbrnF7+7LBF/UHVY=
golang.org/x/text v0.41.0 h1:vz/seA0lnX87Othu2f/0L24RcgrXD9/YFTSuGjj3rH8=
golang.org/x/text v0.41.0/go.mod h1:jvf1O8ajNzZqhSrQBPbutR/EB83Cc0CFrezNQIwbb5M=
golang.org/x/time v0.15.0 h1:bbrp8t3bGUeFOx08pvsMYRTCVSMk89u4tKbNOZbp88U=
golang.org/x/time v0.15.0/go.mod h1:Y4YMaQmXwGQZoFaVFk4YpCt4FLQMYKZe9oeV/f4MSno=
golang.org/x/xerrors v0.0.0-20191204190536-9bdfabe68543/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
gonum.org/v1/gonum v0.17.0 h1:VbpOemQlsSMrYmn7T2OUvQ4dqxQXU+ouZFQsZOx50z4=
gonum.org/v1/gonum v0.17.0/go.mod h1:El3tOrEuMpv2UdMrbNlKEh9vd86bmQ6vqIcDwxEOc1E=
google.golang.org/genproto/googleapis/api v0.0.0-20260615183401-62b3387ff324 h1:g0RAkxK/smSu/iRwC/KIX1mwUoVJtk2OjbgaeS4DmUM=
google.golang.org/genproto/googleapis/api v0.0.0-20260615183401-62b3387ff324/go.mod h1:Z4WJ5pJOYWFWcHEQUelD5QaZDknIQkpIL/+fyJOT9+A=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260615183401-62b3387ff324 h1:9HZDLIdYBJXAnaFOr9WHrKVycfpY+75s9HGadC0305A=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260615183401-62b3387ff324/go.mod h1:4Hqkh8ycfw05ld/3BWL7rJOSfebL2Q+DVDeRgYgxUU8=
google.golang.org/grpc v1.82.1 h1:NnAxzGRA0677vCa4BUkOAnO5+FfQqVl9iUXeD0IqcGE=
google.golang.org/grpc v1.82.1/go.mod h1:yzTZ1TB1Z3SG+LIYaI+WiE8D5+PZ3ArnrSp8zF3+/ZA=
google.golang.org/protobuf v1.36.11 h1:fV6ZwhNocDyBLK0dj+fg8ektcVegBBuEolpbTQyBNVE=
google.golang.org/protobuf v1.36.11/go.mod h1:HTf+CrKn2C3g5S8VImy6tdcUvCska2kB7j23XfzDpco=
google.golang.org/genproto/googleapis/api v0.0.0-20260831171406-18b4a7587f8a h1:i3TAXhpKc7TUP1VAPiBBrv45kamjoizCC3rOC0cAbOs=
google.golang.org/genproto/googleapis/api v0.0.0-20260831171406-18b4a7587f8a/go.mod h1:CvYJHpbzPlT0fb/PsgtAamdwru/GVxUsomFdXTpOTI8=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260831171406-18b4a7587f8a h1:3Dnd1cDaZlB68lziofO+bJXpjOy8UfRv8Unt+yH8tQ4=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260831171406-18b4a7587f8a/go.mod h1:DjtHYE8FKJLivXcBEjGwndXfIC23G0VpXiXKqG179uA=
google.golang.org/grpc v1.83.2 h1:EManeRomTObA0BU7I8vXgg/78uE5MJ9M8B39EX2WscU=
google.golang.org/grpc v1.83.2/go.mod h1:YPI1hK3kDked6iHvgX3tR0y+nX/qpMFKhPgFsokw1S8=
google.golang.org/protobuf v1.36.12 h1:pJOKDDOyeXErUroCihFAd5LQuwXBSpVnKGrj5o/fwxc=
google.golang.org/protobuf v1.36.12/go.mod h1:HTf+CrKn2C3g5S8VImy6tdcUvCska2kB7j23XfzDpco=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c h1:Hei/4ADfdWqJk1ZMxUNpqntNwaWcugrBjAiHlqqRiVk=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c/go.mod h1:JHkPIbrfpd72SG/EVd6muEfDQjcINNoR0C8j2r3qZ4Q=
gopkg.in/yaml.v3 v3.0.0-20200313102051-9f266ea9e77c/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
gotest.tools/v3 v3.5.2 h1:7koQfIKdy+I8UTetycgUqXWSDwpgv193Ka+qRsmBY8Q=
gotest.tools/v3 v3.5.2/go.mod h1:LtdLGcnqToBH83WByAAi/wiwSFCArdFIUV/xxN4pcjA=

51
package-lock.json generated
View File

@@ -6,31 +6,29 @@
"": {
"devDependencies": {
"prettier": "3.9.6",
"prettier-plugin-gherkin": "^3.1.3",
"prettier-plugin-gherkin": "^4.0.0",
"prettier-plugin-sh": "^0.19.0"
}
},
"node_modules/@cucumber/gherkin": {
"version": "32.2.0",
"resolved": "https://registry.npmjs.org/@cucumber/gherkin/-/gherkin-32.2.0.tgz",
"integrity": "sha512-X8xuVhSIqlUjxSRifRJ7t0TycVWyX58fygJH3wDNmHINLg9sYEkvQT0SO2G5YlRZnYc11TIFr4YPenscvdlBIw==",
"version": "39.1.0",
"resolved": "https://registry.npmjs.org/@cucumber/gherkin/-/gherkin-39.1.0.tgz",
"integrity": "sha512-pqmSO2bUWxJm3TbNrKXlDaHjL6c77+ez9kWmfCd9oRPeTRPEVH3spZvpAqdXYWOZYSNYwWFCAAeZ4RGpkauNoQ==",
"dev": true,
"license": "MIT",
"dependencies": {
"@cucumber/messages": ">=19.1.4 <28"
"@cucumber/messages": ">=31.0.0 <33"
}
},
"node_modules/@cucumber/messages": {
"version": "27.2.0",
"resolved": "https://registry.npmjs.org/@cucumber/messages/-/messages-27.2.0.tgz",
"integrity": "sha512-f2o/HqKHgsqzFLdq6fAhfG1FNOQPdBdyMGpKwhb7hZqg0yZtx9BVqkTyuoNk83Fcvk3wjMVfouFXXHNEk4nddA==",
"version": "32.3.1",
"resolved": "https://registry.npmjs.org/@cucumber/messages/-/messages-32.3.1.tgz",
"integrity": "sha512-yNQq1KoXRYaEKrWMFmpUQX7TdeQuU9jeGgJAZ3dArTsC/T4NpJ6DnqaJIIgwPnz/wtQIQTNX7/h0rOuF5xY4qQ==",
"dev": true,
"license": "MIT",
"dependencies": {
"@types/uuid": "10.0.0",
"class-transformer": "0.5.1",
"reflect-metadata": "0.2.2",
"uuid": "11.0.5"
"reflect-metadata": "0.2.2"
}
},
"node_modules/@reteps/dockerfmt": {
@@ -108,13 +106,6 @@
"linux"
]
},
"node_modules/@types/uuid": {
"version": "10.0.0",
"resolved": "https://registry.npmjs.org/@types/uuid/-/uuid-10.0.0.tgz",
"integrity": "sha512-7gqG38EyHgyP1S+7+xomFtL+ZNHcKv6DwNaCZmJmo1vgMugyF3TCnXVg4t1uk89mLNwnLtnY3TpOpCOyp1/xHQ==",
"dev": true,
"license": "MIT"
},
"node_modules/class-transformer": {
"version": "0.5.1",
"resolved": "https://registry.npmjs.org/class-transformer/-/class-transformer-0.5.1.tgz",
@@ -139,14 +130,14 @@
}
},
"node_modules/prettier-plugin-gherkin": {
"version": "3.1.3",
"resolved": "https://registry.npmjs.org/prettier-plugin-gherkin/-/prettier-plugin-gherkin-3.1.3.tgz",
"integrity": "sha512-w9uB413NlSi8ZQwpexyu+ttriJJ88eZLV0x88ZTkzkLZyHYEX5wrNtaCx/yFYviIu/tuwsBqDPM47VODnIV/hw==",
"version": "4.0.0",
"resolved": "https://registry.npmjs.org/prettier-plugin-gherkin/-/prettier-plugin-gherkin-4.0.0.tgz",
"integrity": "sha512-EBDwV1Ou9rG+seoh7jcwZ6sXXyHwsbUks5TeSyrlrLaTdz+/qYvnHuwY8rk7mb6/+M1ZGuY4/SiVc68ybwvJOA==",
"dev": true,
"license": "MIT",
"dependencies": {
"@cucumber/gherkin": "^32.0.0",
"@cucumber/messages": "^27.2.0",
"@cucumber/gherkin": "^39.1.0",
"@cucumber/messages": "^32.3.1",
"prettier": "^3.5.3"
}
},
@@ -189,20 +180,6 @@
"funding": {
"url": "https://opencollective.com/sh-syntax"
}
},
"node_modules/uuid": {
"version": "11.0.5",
"resolved": "https://registry.npmjs.org/uuid/-/uuid-11.0.5.tgz",
"integrity": "sha512-508e6IcKLrhxKdBbcA2b4KQZlLVp2+J5UwQ6F7Drckkc5N9ZJwFa4TgWtsww9UG8fGHbm6gbV19TdM5pQ4GaIA==",
"dev": true,
"funding": [
"https://github.com/sponsors/broofa",
"https://github.com/sponsors/ctavan"
],
"license": "MIT",
"bin": {
"uuid": "dist/esm/bin/uuid"
}
}
}
}

View File

@@ -1,7 +1,7 @@
{
"devDependencies": {
"prettier": "3.9.6",
"prettier-plugin-gherkin": "^3.1.3",
"prettier-plugin-gherkin": "^4.0.0",
"prettier-plugin-sh": "^0.19.0"
}
}

635
pkg/gotenberg/allowlist.go Normal file
View File

@@ -0,0 +1,635 @@
package gotenberg
import (
"strings"
)
// AllowListRisk classifies why an allow-list pattern is dangerous. A URL that
// matches an allow-list skips the private and public IP checks, so a pattern
// that matches more than its author intended silently widens outbound access.
// See [AuditAllowList].
type AllowListRisk string
const (
// AllowListRiskUnanchored marks a pattern with no leading "^". regexp2
// searches rather than matches, so the pattern hits anywhere in the URL,
// including the query string.
AllowListRiskUnanchored AllowListRisk = "unanchored"
// AllowListRiskUnanchoredBranch marks an alternation whose later branches
// have no leading "^". Anchoring only the first branch is a common slip.
AllowListRiskUnanchoredBranch AllowListRisk = "unanchored-branch"
// AllowListRiskCatchAll marks a pattern with no literal prefix, such as
// ".+", which matches every URL and disables filtering entirely.
AllowListRiskCatchAll AllowListRisk = "catch-all"
// AllowListRiskOpenHost marks a pattern whose host is not terminated, so
// it also matches attacker-chosen suffix hosts. For example
// "^https://trusted\.example\.com" matches
// "https://trusted.example.com.attacker.example/".
AllowListRiskOpenHost AllowListRisk = "open-host"
)
// AllowListFinding reports one risky entry of an allow-list.
type AllowListFinding struct {
// Index is the zero-based position of the pattern within the flag value.
Index int
// Pattern is the operator's pattern, verbatim.
Pattern string
// Risk is why the pattern is dangerous.
Risk AllowListRisk
}
// maxAuditedPatternLength bounds the patterns [AuditAllowList] inspects. A
// pathological pattern is not worth walking, and reporting nothing is better
// than reporting a partial verdict.
const maxAuditedPatternLength = 4096
// AuditAllowList reports the entries of an allow-list that match more URLs
// than their author is likely to intend. It is a lint over the pattern source,
// not a parser: it recognizes the shapes that are dangerous in practice and
// stays silent when it cannot be sure.
//
// Callers use the findings to warn operators. Never use them to reject a
// configuration: existing deployments rely on loose patterns, and a pattern
// this function does not flag is not thereby safe.
func AuditAllowList(patterns []string) []AllowListFinding {
var findings []AllowListFinding
for i, pattern := range patterns {
if pattern == "" || len(pattern) > maxAuditedPatternLength {
continue
}
risk, ok := auditPattern(pattern)
if ok {
findings = append(findings, AllowListFinding{Index: i, Pattern: pattern, Risk: risk})
}
}
return findings
}
// auditPattern classifies a single pattern, reporting the first risk found.
func auditPattern(pattern string) (AllowListRisk, bool) {
body := trimInlineFlags(pattern)
branches := splitTopLevelAlternation(body)
for i, branch := range branches {
branch = strings.TrimSpace(branch)
anchored := hasStartAnchor(branch)
rest := strings.TrimPrefix(strings.TrimPrefix(branch, `\A`), "^")
// Catch-all first: a pattern that constrains nothing matches every URL
// whether or not it is anchored, and saying so is more useful than
// telling the operator to anchor it.
if literalPrefix(rest) == "" {
return AllowListRiskCatchAll, true
}
if !anchored {
if i == 0 {
return AllowListRiskUnanchored, true
}
return AllowListRiskUnanchoredBranch, true
}
// A lookaround invalidates the token walk, so skip the host check for
// this branch rather than guess. The anchor and catch-all checks above
// still applied.
if containsLookaround(rest) {
continue
}
if hostIsOpen(rest) {
return AllowListRiskOpenHost, true
}
}
return "", false
}
// trimInlineFlags removes a leading inline flag group such as "(?i)" so that
// the anchor check sees the pattern proper.
func trimInlineFlags(pattern string) string {
if !strings.HasPrefix(pattern, "(?") {
return pattern
}
end := strings.Index(pattern, ")")
if end == -1 {
return pattern
}
// Only a flag group qualifies. "(?:", "(?=", "(?!" and "(?<" open a real
// group and must stay.
flags := pattern[2:end]
if flags == "" || strings.ContainsAny(flags, ":=!<") {
return pattern
}
for _, r := range flags {
if !strings.ContainsRune("imsUx-", r) {
return pattern
}
}
return pattern[end+1:]
}
// hasStartAnchor reports whether branch begins with a start-of-input anchor.
func hasStartAnchor(branch string) bool {
return strings.HasPrefix(branch, "^") || strings.HasPrefix(branch, `\A`)
}
// containsLookaround reports whether the pattern uses a lookaround, which the
// token walk in [hostIsOpen] cannot reason about.
func containsLookaround(s string) bool {
return strings.Contains(s, "(?=") || strings.Contains(s, "(?!") || strings.Contains(s, "(?<")
}
// splitTopLevelAlternation splits on "|" at paren depth zero, honoring escapes
// and character classes.
func splitTopLevelAlternation(s string) []string {
var (
parts []string
current strings.Builder
depth int
inClass bool
)
for i := 0; i < len(s); i++ {
c := s[i]
switch {
case c == '\\' && i+1 < len(s):
current.WriteByte(c)
current.WriteByte(s[i+1])
i++
continue
case c == '[' && !inClass:
inClass = true
case c == ']' && inClass:
inClass = false
case c == '(' && !inClass:
depth++
case c == ')' && !inClass:
depth--
case c == '|' && !inClass && depth == 0:
parts = append(parts, current.String())
current.Reset()
continue
}
current.WriteByte(c)
}
parts = append(parts, current.String())
return parts
}
// literalPrefix returns the characters a matching URL must start with. It
// stops at the first optional or non-literal token, and descends one level
// into a leading mandatory group so that "^(https|http)://" is not mistaken
// for a catch-all. An empty result means the pattern constrains nothing.
func literalPrefix(s string) string {
var prefix strings.Builder
for i := 0; i < len(s); {
// A group: descend once when it is mandatory, otherwise stop.
if s[i] == '(' {
end := matchingParen(s, i)
if end == -1 {
break
}
if isQuantified(s, end+1) {
break
}
inner := trimInlineFlags(s[i+1 : end])
branches := splitTopLevelAlternation(inner)
common := literalPrefix(branches[0])
for _, b := range branches[1:] {
common = commonPrefix(common, literalPrefix(b))
}
prefix.WriteString(common)
// Only the leading group is worth descending into.
break
}
var token string
switch {
case s[i] == '\\' && i+1 < len(s):
token = s[i : i+2]
case s[i] == '[':
end := matchingBracket(s, i)
if end == -1 {
return prefix.String()
}
token = s[i : end+1]
default:
token = s[i : i+1]
}
next := i + len(token)
if isQuantified(s, next) {
break
}
// Only a plain literal or an escaped literal contributes.
switch {
case len(token) == 2 && token[0] == '\\' && !isEscapeClass(token[1]):
prefix.WriteByte(token[1])
case len(token) == 1 && !strings.ContainsAny(token, `.[]()^$*+?{}|`):
prefix.WriteByte(token[0])
default:
return prefix.String()
}
i = next
}
return prefix.String()
}
// hostIsOpen reports whether the authority part of the pattern can be left
// without crossing a terminator, which means the pattern also matches
// attacker-chosen suffix hosts or userinfo.
//
// It walks the tokens after "://" and classifies each one. A terminator ends
// the authority, so the pattern is safe. A crosser can match "@", "?" or "#"
// and therefore lets a matching URL escape the authority, so the pattern is
// open. Reaching the end without a terminator is open too, which is the
// classic "^https://trusted\.example\.com" case.
func hostIsOpen(s string) bool {
_, after, ok := strings.Cut(s, "://")
if !ok {
// No authority to reason about, for example "^file:///tmp/".
return false
}
rest := after
for i := 0; i < len(rest); {
var token string
switch {
case rest[i] == '\\' && i+1 < len(rest):
token = rest[i : i+2]
case rest[i] == '[':
end := matchingBracket(rest, i)
if end == -1 {
return true
}
token = rest[i : end+1]
case rest[i] == '(':
end := matchingParen(rest, i)
if end == -1 {
return true
}
token = rest[i : end+1]
default:
token = rest[i : i+1]
}
next := i + len(token)
optional := isOptionalQuantifier(rest, next)
switch classifyHostToken(token) {
case hostTokenTerminator:
// An optional terminator does not end anything, since the URL may
// match without it.
if !optional {
return false
}
case hostTokenCrosser:
return true
case hostTokenNeutral:
// Part of the host itself, so keep walking.
}
i = next
for i < len(rest) && isQuantifierByte(rest[i]) {
if rest[i] == '{' {
end := strings.IndexByte(rest[i:], '}')
if end == -1 {
return true
}
i += end + 1
continue
}
i++
}
}
return true
}
// hostTokenKind is how a token affects the walk in [hostIsOpen].
type hostTokenKind int
const (
hostTokenNeutral hostTokenKind = iota
hostTokenTerminator
hostTokenCrosser
)
// hostTerminators are the characters that end the authority of a URL.
const hostTerminators = "/:#?"
// crosserClassChars are the characters that, if a class can match them, let a
// match escape the authority. "/" is deliberately absent: it ends the
// authority rather than escaping it, so a class such as "[:/]" is safe.
const crosserClassChars = "@?#"
// classifyHostToken classifies one token of the authority walk.
func classifyHostToken(token string) hostTokenKind {
switch {
case token == ".":
// The wildcard matches "@", "#" and "?", so a host built on it can be
// left without ever reaching a terminator.
return hostTokenCrosser
case token == "$":
return hostTokenTerminator
case len(token) == 1 && strings.Contains(hostTerminators, token):
return hostTokenTerminator
case len(token) == 2 && token[0] == '\\':
switch token[1] {
case 'S', 'D', 'W':
return hostTokenCrosser
case 'd', 'w', 's':
return hostTokenNeutral
case 'p', 'P':
return hostTokenCrosser
}
if strings.Contains(hostTerminators, token[1:]) {
return hostTokenTerminator
}
return hostTokenNeutral
case strings.HasPrefix(token, "["):
inner := strings.TrimSuffix(strings.TrimPrefix(token, "["), "]")
if strings.HasPrefix(inner, "^") {
// A negated class almost always admits "@".
return hostTokenCrosser
}
if classContainsAny(inner, crosserClassChars) {
return hostTokenCrosser
}
if classOnlyTerminators(inner) {
return hostTokenTerminator
}
return hostTokenNeutral
case strings.HasPrefix(token, "("):
return classifyGroup(token)
}
return hostTokenNeutral
}
// classifyGroup classifies a parenthesized group. A group whose every branch
// starts with a terminator ends the authority, which is what makes the
// idiomatic "(:|/|$)" safe. A group containing a crosser is a crosser.
func classifyGroup(token string) hostTokenKind {
inner := trimInlineFlags(strings.TrimSuffix(strings.TrimPrefix(token, "("), ")"))
inner = strings.TrimPrefix(inner, "?:")
branches := splitTopLevelAlternation(inner)
allTerminate := true
for _, branch := range branches {
if branch == "" {
allTerminate = false
continue
}
kind := classifyHostToken(firstToken(branch))
if kind == hostTokenCrosser {
return hostTokenCrosser
}
if kind != hostTokenTerminator {
allTerminate = false
}
// A crosser anywhere inside the branch still escapes the authority.
if branchHasCrosser(branch) {
return hostTokenCrosser
}
}
if allTerminate {
return hostTokenTerminator
}
return hostTokenNeutral
}
// branchHasCrosser reports whether any token of branch is a crosser.
func branchHasCrosser(branch string) bool {
for i := 0; i < len(branch); {
token := tokenAt(branch, i)
if token == "" {
return true
}
if classifyHostToken(token) == hostTokenCrosser {
return true
}
i += len(token)
}
return false
}
// firstToken returns the first regex token of s.
func firstToken(s string) string {
return tokenAt(s, 0)
}
// tokenAt returns the regex token starting at index i, or "" if it is
// malformed.
func tokenAt(s string, i int) string {
if i >= len(s) {
return ""
}
switch {
case s[i] == '\\' && i+1 < len(s):
return s[i : i+2]
case s[i] == '[':
end := matchingBracket(s, i)
if end == -1 {
return ""
}
return s[i : end+1]
case s[i] == '(':
end := matchingParen(s, i)
if end == -1 {
return ""
}
return s[i : end+1]
}
return s[i : i+1]
}
// classOnlyTerminators reports whether every character a class can match ends
// the authority, which makes the class itself a terminator. A range is never
// treated as one.
func classOnlyTerminators(class string) bool {
if class == "" {
return false
}
for i := 0; i < len(class); i++ {
if class[i] == '\\' && i+1 < len(class) {
if !strings.Contains(hostTerminators, class[i+1:i+2]) {
return false
}
i++
continue
}
if i+2 < len(class) && class[i+1] == '-' {
return false
}
if !strings.Contains(hostTerminators, class[i:i+1]) {
return false
}
}
return true
}
// classContainsAny reports whether a character class body can match any of the
// given characters, expanding simple ranges.
func classContainsAny(class, chars string) bool {
for i := 0; i < len(class); i++ {
if class[i] == '\\' && i+1 < len(class) {
// An escape class such as \S inside a class admits everything.
if strings.ContainsRune("SDW", rune(class[i+1])) {
return true
}
if strings.ContainsRune(chars, rune(class[i+1])) {
return true
}
i++
continue
}
if i+2 < len(class) && class[i+1] == '-' {
lo, hi := class[i], class[i+2]
for _, c := range []byte(chars) {
if c >= lo && c <= hi {
return true
}
}
i += 2
continue
}
if strings.ContainsRune(chars, rune(class[i])) {
return true
}
}
return false
}
// matchingParen returns the index of the ")" closing the "(" at start.
func matchingParen(s string, start int) int {
depth := 0
inClass := false
for i := start; i < len(s); i++ {
switch {
case s[i] == '\\' && i+1 < len(s):
i++
case s[i] == '[' && !inClass:
inClass = true
case s[i] == ']' && inClass:
inClass = false
case s[i] == '(' && !inClass:
depth++
case s[i] == ')' && !inClass:
depth--
if depth == 0 {
return i
}
}
}
return -1
}
// matchingBracket returns the index of the "]" closing the "[" at start.
func matchingBracket(s string, start int) int {
for i := start + 1; i < len(s); i++ {
switch {
case s[i] == '\\' && i+1 < len(s):
i++
case s[i] == ']':
// A "]" immediately after "[" or "[^" is a literal.
if i == start+1 || (i == start+2 && s[start+1] == '^') {
continue
}
return i
}
}
return -1
}
// isQuantifierByte reports whether c opens a quantifier.
func isQuantifierByte(c byte) bool {
return c == '?' || c == '*' || c == '+' || c == '{'
}
// isQuantified reports whether a quantifier starts at index i.
func isQuantified(s string, i int) bool {
return i < len(s) && isQuantifierByte(s[i])
}
// isOptionalQuantifier reports whether the quantifier at index i lets the
// preceding token match nothing.
func isOptionalQuantifier(s string, i int) bool {
if i >= len(s) {
return false
}
switch s[i] {
case '?', '*':
return true
case '{':
return strings.HasPrefix(s[i:], "{0")
}
return false
}
// isEscapeClass reports whether c after a backslash denotes a character class
// rather than a literal.
func isEscapeClass(c byte) bool {
return strings.ContainsRune("dDwWsSbBAzZpP", rune(c))
}
// commonPrefix returns the longest common prefix of a and b.
func commonPrefix(a, b string) string {
n := min(len(a), len(b))
for i := range n {
if a[i] != b[i] {
return a[:i]
}
}
return a[:n]
}

View File

@@ -0,0 +1,179 @@
package gotenberg
import (
"testing"
"github.com/dlclark/regexp2"
)
func TestAuditAllowList(t *testing.T) {
for _, tc := range []struct {
scenario string
pattern string
want AllowListRisk
}{
// Safe: the host is terminated before anything can leave it.
{"idiomatic terminator group", `^https?://internal\.svc(:|/|$)`, ""},
{"trailing slash", `^https://trusted\.example\.com/`, ""},
{"optional port then terminator", `^https://example\.com(:[0-9]+)?(/|$)`, ""},
{"positive class cannot leave authority", `^https://[a-z0-9.-]+\.s3\.amazonaws\.com/`, ""},
{"leading mandatory group", `^(https|http)://a\.example\.com/`, ""},
{"optional subdomain group", `^https://(www\.)?example\.com/`, ""},
{"port terminator", `^https://example\.com:8443/`, ""},
{"end anchor", `^https://example\.com$`, ""},
{"alternation both anchored and terminated", `^https://a\.example/|^https://b\.example/`, ""},
{"no authority to check", `^file:///tmp/`, ""},
{"digit class in host", `^https://node\d+\.example\.com/`, ""},
{"class of only terminators", `^https://example\.com[:/]`, ""},
{"feature file pattern, fixed", `^https?://host\.docker\.internal(:[0-9]+)?/`, ""},
// Unanchored: regexp2 searches, so these match anywhere in the URL.
{"no anchor", `trusted\.example\.com`, AllowListRiskUnanchored},
{"no anchor with scheme", `https://trusted\.example\.com/`, AllowListRiskUnanchored},
// Only the first branch anchored.
{"second branch unanchored", `^http://a\.example/|http://b\.example/`, AllowListRiskUnanchoredBranch},
// Catch-all: matches every URL.
{"dot plus", `.+`, AllowListRiskCatchAll},
{"dot star", `.*`, AllowListRiskCatchAll},
{"anchored dot star", `^.*`, AllowListRiskCatchAll},
{"anchored dot plus", `^.+`, AllowListRiskCatchAll},
// Open host: the reported vulnerability class.
{"advisory pattern", `^http://trusted\.example\.com`, AllowListRiskOpenHost},
{"gotenberg.dev internet-facing recipe", `^https?://[^/]+\.internal\.example\.com`, AllowListRiskOpenHost},
{"gotenberg.dev strict whitelist recipe", `^https://(api|cdn|images)\.internal\.example\.com`, AllowListRiskOpenHost},
{"gotenberg.dev hooks recipe", `^https?://hooks\.internal\.example\.com`, AllowListRiskOpenHost},
{"feature file pattern", `^https?://host.docker.internal.*`, AllowListRiskOpenHost},
{"scheme only", `^https?://`, AllowListRiskOpenHost},
{"wildcard subdomain", `^https://.+\.example\.com/`, AllowListRiskOpenHost},
{"escaped dot is not a terminator", `^https?://example\.com\.`, AllowListRiskOpenHost},
{"negated class in host", `^https://[^.]+\.example\.com/`, AllowListRiskOpenHost},
} {
t.Run(tc.scenario, func(t *testing.T) {
findings := AuditAllowList([]string{tc.pattern})
if tc.want == "" {
if len(findings) != 0 {
t.Fatalf("AuditAllowList(%q) = %+v, want no finding", tc.pattern, findings)
}
return
}
if len(findings) != 1 {
t.Fatalf("AuditAllowList(%q) returned %d findings, want 1", tc.pattern, len(findings))
}
if findings[0].Risk != tc.want {
t.Fatalf("AuditAllowList(%q) risk = %q, want %q", tc.pattern, findings[0].Risk, tc.want)
}
})
}
}
// TestAuditAllowList_FlaggedPatternsAreActuallyExploitable proves the audit is
// not merely syntactic: every pattern it flags as open-host really does admit
// a host the operator did not intend.
func TestAuditAllowList_FlaggedPatternsAreActuallyExploitable(t *testing.T) {
for _, tc := range []struct {
pattern string
attack string
}{
{`^http://trusted\.example\.com`, "http://trusted.example.com.attacker.example/"},
{`^https?://[^/]+\.internal\.example\.com`, "http://a.internal.example.com.attacker.example/"},
{`^https://(api|cdn|images)\.internal\.example\.com`, "https://api.internal.example.com.attacker.example/"},
{`^https?://hooks\.internal\.example\.com`, "http://hooks.internal.example.com.attacker.example/"},
{`^https?://host.docker.internal.*`, "http://host.docker.internal.attacker.example/"},
{`^https://.+\.example\.com/`, "https://attacker.example/#x.example.com/"},
{`^https?://example\.com\.`, "http://example.com.attacker.example/"},
} {
t.Run(tc.pattern, func(t *testing.T) {
findings := AuditAllowList([]string{tc.pattern})
if len(findings) == 0 {
t.Fatalf("pattern %q was not flagged", tc.pattern)
}
ok, err := regexp2.MustCompile(tc.pattern, 0).MatchString(tc.attack)
if err != nil {
t.Fatalf("match %q: %v", tc.attack, err)
}
if !ok {
t.Fatalf("pattern %q does not match %q, so the finding is a false positive", tc.pattern, tc.attack)
}
})
}
}
// TestAuditAllowList_SafePatternsRejectTheAttacks is the converse: the shapes
// the audit stays silent about really do reject the same attacks.
func TestAuditAllowList_SafePatternsRejectTheAttacks(t *testing.T) {
safe := []string{
`^https?://internal\.svc(:|/|$)`,
`^https://trusted\.example\.com/`,
`^https://example\.com(:[0-9]+)?(/|$)`,
`^https://[a-z0-9.-]+\.s3\.amazonaws\.com/`,
}
attacks := []string{
"https://internal.svc.attacker.example/",
"https://trusted.example.com.attacker.example/",
"https://trusted.example.com@169.254.169.254/",
"https://example.com.attacker.example/",
"https://example.com@10.0.0.5/",
"https://bucket.s3.amazonaws.com.attacker.example/",
"https://bucket.s3.amazonaws.com@127.0.0.1/",
}
for _, pattern := range safe {
t.Run(pattern, func(t *testing.T) {
if findings := AuditAllowList([]string{pattern}); len(findings) != 0 {
t.Fatalf("safe pattern %q was flagged as %q", pattern, findings[0].Risk)
}
re := regexp2.MustCompile(pattern, 0)
for _, attack := range attacks {
ok, err := re.MatchString(attack)
if err != nil {
t.Fatalf("match %q: %v", attack, err)
}
if ok {
t.Fatalf("pattern %q matches attack %q but was not flagged", pattern, attack)
}
}
})
}
}
func TestAuditAllowList_SkipsEmptyAndOversized(t *testing.T) {
oversized := make([]byte, maxAuditedPatternLength+1)
for i := range oversized {
oversized[i] = 'a'
}
findings := AuditAllowList([]string{"", string(oversized)})
if len(findings) != 0 {
t.Fatalf("AuditAllowList returned %+v, want no finding", findings)
}
}
func TestAuditAllowList_ReportsIndex(t *testing.T) {
findings := AuditAllowList([]string{
`^https://ok\.example\.com/`,
`^https://open\.example\.com`,
})
if len(findings) != 1 {
t.Fatalf("got %d findings, want 1", len(findings))
}
if findings[0].Index != 1 {
t.Fatalf("findings[0].Index = %d, want 1", findings[0].Index)
}
}
// TestAuditAllowList_ShippedChromiumDenyListIsNotAudited guards the rule that
// deny-lists are never audited. The shipped Chromium deny-list uses a
// lookaround and has no authority, so auditing it would produce noise.
func TestAuditAllowList_LookaroundIsNotFlaggedForHost(t *testing.T) {
findings := AuditAllowList([]string{`^file:(?!//\/tmp/).*`})
if len(findings) != 0 {
t.Fatalf("AuditAllowList returned %+v, want no finding", findings)
}
}

View File

@@ -1,70 +0,0 @@
package gotenberg
import (
"context"
"errors"
"fmt"
"time"
"github.com/dlclark/regexp2"
)
// ErrFiltered happens if a value is filtered by the [FilterDeadline] function.
var ErrFiltered = errors.New("value filtered")
// FilterDeadline checks if the given value is allowed and not denied according
// to regex patterns. The allowed list uses OR semantics (value must match at
// least one pattern). The denied list uses OR semantics (value is denied if it
// matches any pattern). It returns a [context.DeadlineExceeded] if it takes
// too long to process.
func FilterDeadline(allowed, denied []*regexp2.Regexp, s string, deadline time.Time) error {
if len(allowed) > 0 {
matched := false
for _, pattern := range allowed {
// FIXME: not ideal to compile everytime, but is there another way to create a clone?
clone := regexp2.MustCompile(pattern.String(), 0)
clone.MatchTimeout = time.Until(deadline)
ok, err := clone.MatchString(s)
if err != nil {
if time.Now().After(deadline) {
return context.DeadlineExceeded
}
return fmt.Errorf("'%s' cannot handle '%s': %w", clone.String(), s, err)
}
if ok {
matched = true
break
}
}
if !matched {
return fmt.Errorf("'%s' does not match any expression from the allowed list: %w", s, ErrFiltered)
}
}
if len(denied) > 0 {
for _, pattern := range denied {
clone := regexp2.MustCompile(pattern.String(), 0)
clone.MatchTimeout = time.Until(deadline)
ok, err := clone.MatchString(s)
if err != nil {
if time.Now().After(deadline) {
return context.DeadlineExceeded
}
return fmt.Errorf("'%s' cannot handle '%s': %w", clone.String(), s, err)
}
if ok {
return fmt.Errorf("'%s' matches the expression from the denied list: %w", s, ErrFiltered)
}
}
}
return nil
}

View File

@@ -1,117 +0,0 @@
package gotenberg
import (
"context"
"errors"
"testing"
"time"
"github.com/dlclark/regexp2"
)
func TestFilterDeadline(t *testing.T) {
for _, tc := range []struct {
scenario string
allowed []*regexp2.Regexp
denied []*regexp2.Regexp
s string
deadline time.Time
expectError bool
expectedError error
}{
{
scenario: "DeadlineExceeded (allowed)",
allowed: []*regexp2.Regexp{regexp2.MustCompile("foo", 0)},
denied: nil,
s: "foo",
deadline: time.Now().Add(time.Duration(-1) * time.Hour),
expectError: true,
expectedError: context.DeadlineExceeded,
},
{
scenario: "ErrFiltered (allowed, no match)",
allowed: []*regexp2.Regexp{regexp2.MustCompile("foo", 0)},
denied: nil,
s: "bar",
deadline: time.Now().Add(time.Duration(5) * time.Second),
expectError: true,
expectedError: ErrFiltered,
},
{
scenario: "DeadlineExceeded (denied)",
allowed: nil,
denied: []*regexp2.Regexp{regexp2.MustCompile("foo", 0)},
s: "foo",
deadline: time.Now().Add(time.Duration(-1) * time.Hour),
expectError: true,
expectedError: context.DeadlineExceeded,
},
{
scenario: "ErrFiltered (denied)",
allowed: nil,
denied: []*regexp2.Regexp{regexp2.MustCompile("foo", 0)},
s: "foo",
deadline: time.Now().Add(time.Duration(5) * time.Second),
expectError: true,
expectedError: ErrFiltered,
},
{
scenario: "success (empty lists)",
allowed: nil,
denied: nil,
s: "foo",
deadline: time.Now().Add(time.Duration(5) * time.Second),
expectError: false,
},
{
scenario: "multi-pattern allow list, second matches",
allowed: []*regexp2.Regexp{regexp2.MustCompile("^https://", 0), regexp2.MustCompile("^file:///tmp/", 0)},
denied: nil,
s: "file:///tmp/abc/index.html",
deadline: time.Now().Add(time.Duration(5) * time.Second),
expectError: false,
},
{
scenario: "multi-pattern allow list, none matches",
allowed: []*regexp2.Regexp{regexp2.MustCompile("^https://", 0), regexp2.MustCompile("^ftp://", 0)},
denied: nil,
s: "file:///tmp/abc/index.html",
deadline: time.Now().Add(time.Duration(5) * time.Second),
expectError: true,
expectedError: ErrFiltered,
},
{
scenario: "multi-pattern deny list, second matches",
allowed: nil,
denied: []*regexp2.Regexp{regexp2.MustCompile("^ftp://", 0), regexp2.MustCompile("^file:.*", 0)},
s: "file:///etc/passwd",
deadline: time.Now().Add(time.Duration(5) * time.Second),
expectError: true,
expectedError: ErrFiltered,
},
{
scenario: "https URL passes deny list targeting file://",
allowed: nil,
denied: []*regexp2.Regexp{regexp2.MustCompile("^file:.*", 0)},
s: "https://example.com",
deadline: time.Now().Add(time.Duration(5) * time.Second),
expectError: false,
},
} {
t.Run(tc.scenario, func(t *testing.T) {
err := FilterDeadline(tc.allowed, tc.denied, tc.s, tc.deadline)
if tc.expectError && err == nil {
t.Fatal("expected an error but got none")
}
if !tc.expectError && err != nil {
t.Fatalf("expected no error but got: %v", err)
}
if tc.expectedError != nil && !errors.Is(err, tc.expectedError) {
t.Fatalf("expected error %v but got: %v", tc.expectedError, err)
}
})
}
}

View File

@@ -1,11 +1,17 @@
package gotenberg
import (
"context"
"fmt"
"log/slog"
"strings"
"time"
"github.com/dlclark/regexp2"
"github.com/labstack/gommon/bytes"
flag "github.com/spf13/pflag"
"github.com/gotenberg/gotenberg/v8/pkg/gotenberg/internal/log"
)
// ParsedFlags wraps a [flag.FlagSet] so that retrieving the typed values is
@@ -201,15 +207,38 @@ func (f *ParsedFlags) MustDeprecatedHumanReadableBytes(deprecated string, newNam
return f.MustHumanReadableBytes(newName)
}
// PatternMatchTimeout bounds a single match against an operator-supplied
// allow-list or deny-list pattern.
//
// regexp2 backtracks, and the strings matched against these patterns are
// client-controlled: a request URL, a CONNECT host. A pattern that backtracks
// catastrophically would otherwise burn a core for as long as the caller's
// deadline allows, which is --api-timeout (env API_TIMEOUT), 30 seconds by
// default. The ceiling mirrors the one the Chromium module already applies to
// the per-request extraHttpHeaders scope pattern.
//
// [ParsedFlags.MustRegexp] and [ParsedFlags.MustRegexpSlice] stamp this onto
// every pattern they compile, which is how all four production lists are
// built. Patterns compiled any other way keep regexp2's default of
// math.MaxInt64, which it treats as no timeout at all, so a hand-built slice
// must set this itself before reaching [DecideOutbound].
const PatternMatchTimeout = 250 * time.Millisecond
// MustRegexp returns the regular expression of a flag given by name.
// It panics if an error occurs.
//
// The returned expression carries [PatternMatchTimeout] and is safe to match
// on concurrently: callers must not compile a private copy per match.
func (f *ParsedFlags) MustRegexp(name string) *regexp2.Regexp {
val, err := f.GetString(name)
if err != nil {
panic(err)
}
return regexp2.MustCompile(val, 0)
re := regexp2.MustCompile(val, 0)
re.MatchTimeout = PatternMatchTimeout
return re
}
// MustDeprecatedRegexp returns the regular expression of a deprecated flag if
@@ -226,21 +255,140 @@ func (f *ParsedFlags) MustDeprecatedRegexp(deprecated string, newName string) *r
// MustRegexpSlice returns a slice of compiled regular expressions from a
// string-slice flag given by name. Empty strings are skipped.
// It panics if an error occurs.
//
// Every allow-list and deny-list in Gotenberg is read through this method, so
// it is also where allow-list patterns are audited. See [AuditAllowList].
//
// The returned expressions carry [PatternMatchTimeout] and are safe to match
// on concurrently: callers must not compile a private copy per match.
func (f *ParsedFlags) MustRegexpSlice(name string) []*regexp2.Regexp {
vals := f.MustStringSlice(name)
f.warnRiskyAllowList(name, vals)
var regexps []*regexp2.Regexp
for _, val := range vals {
if val == "" {
continue
}
regexps = append(regexps, regexp2.MustCompile(val, 0))
re := regexp2.MustCompile(val, 0)
re.MatchTimeout = PatternMatchTimeout
regexps = append(regexps, re)
}
return regexps
}
// allowListFlagSuffix identifies the flags whose patterns grant an IP-check
// bypass. Deny-lists are never audited: they always apply, cannot be bypassed,
// and a loose deny-list is safe rather than dangerous.
const allowListFlagSuffix = "-allow-list"
// warnRiskyAllowList logs one warning per allow-list entry that matches more
// URLs than its author is likely to intend.
//
// It warns and never fails: operators depend on loose patterns today, and
// rejecting them at startup would break running deployments.
func (f *ParsedFlags) warnRiskyAllowList(name string, vals []string) {
if !strings.HasSuffix(name, allowListFlagSuffix) {
return
}
findings := AuditAllowList(vals)
if len(findings) == 0 {
return
}
// The logger is nil until the entry point initializes it, which happens
// before any module is provisioned. Tests and embedders that call this
// method directly get no logger, and must not panic for it.
logger := log.Logger()
if logger == nil {
return
}
for _, finding := range findings {
// Provision has no context.Context to propagate, so the trace-aware
// logging convention is satisfied with a background context.
logger.WarnContext(
context.Background(),
f.allowListWarning(name, finding),
slog.String("flag", "--"+name),
slog.String("env", EnvVarName(name)),
slog.Int("entry", finding.Index+1),
slog.String("reason", string(finding.Risk)),
)
}
}
// allowListWarning builds the operator-facing message for a finding. It names
// the flag and its environment variable, and, when they exist, the IP-check
// flags the entry silently disables.
func (f *ParsedFlags) allowListWarning(name string, finding AllowListFinding) string {
var b strings.Builder
// Print the pattern raw rather than quoted: %q escapes every backslash, so
// the operator would not recognize the value they set.
fmt.Fprintf(&b, "--%s (%s) entry %d '%s' ", name, EnvVarName(name), finding.Index+1, finding.Pattern)
switch finding.Risk {
case AllowListRiskUnanchored:
b.WriteString("is not anchored with ^, so it matches anywhere in the URL and a URL such as http://attacker.example/?u=trusted.example.com passes. ")
case AllowListRiskUnanchoredBranch:
b.WriteString("has an alternation branch that is not anchored with ^, and that branch matches anywhere in the URL. ")
case AllowListRiskCatchAll:
b.WriteString("matches every URL. ")
case AllowListRiskOpenHost:
b.WriteString("does not terminate the host, so it also matches suffix hosts such as http://trusted.example.com.attacker.example/. ")
}
b.WriteString(f.bypassSentence(name))
switch finding.Risk {
case AllowListRiskUnanchored, AllowListRiskUnanchoredBranch:
b.WriteString("Anchor every branch with ^ and end the host with /, :, or $.")
case AllowListRiskCatchAll:
b.WriteString("Restrict the entry to the hosts you trust, or unset the flag.")
case AllowListRiskOpenHost:
b.WriteString("End the host with /, :, $, or a group such as (:|/|$).")
}
return b.String()
}
// bypassSentence names the IP-check flags an allow-list match skips, when the
// module registers them.
func (f *ParsedFlags) bypassSentence(name string) string {
prefix := strings.TrimSuffix(name, allowListFlagSuffix)
private, public := prefix+"-deny-private-ips", prefix+"-deny-public-ips"
if f.Lookup(private) == nil || f.Lookup(public) == nil {
// A deprecated alias such as webhook-error-allow-list carries an extra
// segment that the IP-check flags do not have.
if i := strings.LastIndex(prefix, "-"); i != -1 {
private, public = prefix[:i]+"-deny-private-ips", prefix[:i]+"-deny-public-ips"
}
}
if f.Lookup(private) == nil || f.Lookup(public) == nil {
return "A URL that matches the allow-list skips the private and public IP checks. "
}
return fmt.Sprintf(
"A URL that matches the allow-list skips --%s (%s) and --%s (%s). ",
private, EnvVarName(private), public, EnvVarName(public),
)
}
// EnvVarName returns the environment variable that overrides the flag given by
// name. The entry point derives the same name when it applies environment
// overrides, so operator-facing messages can name both without drifting.
func EnvVarName(name string) string {
return strings.ToUpper(strings.ReplaceAll(name, "-", "_"))
}
// MustDeprecatedRegexpSlice returns the slice of compiled regular expressions
// of a deprecated flag if it was explicitly set or the slice of the new flag.
// It panics if an error occurs.

View File

@@ -3,6 +3,7 @@ package gotenberg
import (
"reflect"
"regexp"
"strings"
"testing"
"time"
@@ -951,3 +952,138 @@ func TestParsedFlags_MustDeprecatedRegexpSlice(t *testing.T) {
_ = regexp2.None // Keep import alive.
}
func TestParsedFlags_AllowListWarning(t *testing.T) {
fs := flag.NewFlagSet("tests", flag.ContinueOnError)
fs.StringSlice("chromium-allow-list", []string{}, "")
fs.Bool("chromium-deny-private-ips", false, "")
fs.Bool("chromium-deny-public-ips", false, "")
fs.StringSlice("standalone-allow-list", []string{}, "")
parsedFlags := ParsedFlags{FlagSet: fs}
for _, tc := range []struct {
scenario string
name string
finding AllowListFinding
contains []string
}{
{
scenario: "open host names both IP-check flags and their env vars",
name: "chromium-allow-list",
finding: AllowListFinding{Index: 0, Pattern: `^https://trusted\.example\.com`, Risk: AllowListRiskOpenHost},
contains: []string{
"--chromium-allow-list (CHROMIUM_ALLOW_LIST)",
"entry 1",
`^https://trusted\.example\.com`,
"does not terminate the host",
"--chromium-deny-private-ips (CHROMIUM_DENY_PRIVATE_IPS)",
"--chromium-deny-public-ips (CHROMIUM_DENY_PUBLIC_IPS)",
"End the host with",
},
},
{
scenario: "catch-all tells the operator to restrict or unset",
name: "chromium-allow-list",
finding: AllowListFinding{Index: 2, Pattern: ".+", Risk: AllowListRiskCatchAll},
contains: []string{"entry 3", "matches every URL", "Restrict the entry"},
},
{
scenario: "unanchored explains the search semantics",
name: "chromium-allow-list",
finding: AllowListFinding{Index: 0, Pattern: `trusted\.example\.com`, Risk: AllowListRiskUnanchored},
contains: []string{"is not anchored with ^", "Anchor every branch with ^"},
},
{
scenario: "module without IP-check flags falls back to a generic sentence",
name: "standalone-allow-list",
finding: AllowListFinding{Index: 0, Pattern: `^https://a\.example\.com`, Risk: AllowListRiskOpenHost},
contains: []string{"skips the private and public IP checks"},
},
} {
t.Run(tc.scenario, func(t *testing.T) {
msg := parsedFlags.allowListWarning(tc.name, tc.finding)
for _, want := range tc.contains {
if !strings.Contains(msg, want) {
t.Fatalf("message %q does not contain %q", msg, want)
}
}
if strings.Contains(msg, "—") {
t.Fatalf("message must not contain an em dash: %q", msg)
}
})
}
}
func TestParsedFlags_WarnRiskyAllowList_SkipsDenyLists(t *testing.T) {
fs := flag.NewFlagSet("tests", flag.ContinueOnError)
fs.StringSlice("chromium-deny-list", []string{}, "")
parsedFlags := ParsedFlags{FlagSet: fs}
// A deny-list is never audited: it always applies and cannot be bypassed,
// so a loose one is safe. This must also not panic on a nil logger.
parsedFlags.warnRiskyAllowList("chromium-deny-list", []string{".+", `^file:(?!//\/tmp/).*`})
}
func TestParsedFlags_WarnRiskyAllowList_NilLoggerDoesNotPanic(t *testing.T) {
fs := flag.NewFlagSet("tests", flag.ContinueOnError)
fs.StringSlice("chromium-allow-list", []string{}, "")
parsedFlags := ParsedFlags{FlagSet: fs}
// Provision runs after the entry point initializes the logger, but tests
// and embedders reach this path with no logger at all.
parsedFlags.warnRiskyAllowList("chromium-allow-list", []string{".+"})
}
func TestEnvVarName(t *testing.T) {
for _, tc := range []struct {
name string
want string
}{
{"chromium-allow-list", "CHROMIUM_ALLOW_LIST"},
{"api-download-from-deny-private-ips", "API_DOWNLOAD_FROM_DENY_PRIVATE_IPS"},
{"log-level", "LOG_LEVEL"},
} {
t.Run(tc.name, func(t *testing.T) {
if got := EnvVarName(tc.name); got != tc.want {
t.Fatalf("EnvVarName(%q) = %q, want %q", tc.name, got, tc.want)
}
})
}
}
func TestParsedFlags_RegexpMatchTimeout(t *testing.T) {
// [DecideOutbound] matches on these patterns directly instead of compiling
// a private copy per call, so the bound has to come from here. regexp2's
// own default is math.MaxInt64, which it treats as no
// timeout at all, so a pattern built without this stamp runs unbounded
// against a client-controlled string.
fs := flag.NewFlagSet("tests", flag.ContinueOnError)
fs.StringSlice("some-deny-list", []string{`^file:`, `^https?://`}, "")
fs.String("some-pattern", `^file:`, "")
err := fs.Parse(nil)
if err != nil {
t.Fatalf("expected no error but got: %v", err)
}
parsedFlags := ParsedFlags{FlagSet: fs}
regexps := parsedFlags.MustRegexpSlice("some-deny-list")
if len(regexps) != 2 {
t.Fatalf("expected 2 patterns but got %d", len(regexps))
}
for _, re := range regexps {
if re.MatchTimeout != PatternMatchTimeout {
t.Fatalf("pattern '%s' has MatchTimeout %s, expected %s", re.String(), re.MatchTimeout, PatternMatchTimeout)
}
}
if got := parsedFlags.MustRegexp("some-pattern").MatchTimeout; got != PatternMatchTimeout {
t.Fatalf("expected MustRegexp MatchTimeout %s but got %s", PatternMatchTimeout, got)
}
}

View File

@@ -5,6 +5,7 @@ import (
"fmt"
"log/slog"
"os"
"strings"
"sync/atomic"
"go.opentelemetry.io/contrib/bridges/otelslog"
@@ -74,6 +75,24 @@ func buildResource(ctx context.Context, logger *slog.Logger, serviceName, servic
return merged
}
// OTEL_*_EXPORTER select the exporter for each signal. autoexport treats an
// unset or empty value as a request for the OTLP exporter, which then fails
// against the default localhost:4318 endpoint when nothing listens there and,
// for metrics, keeps retrying on the periodic reader's timer. Gotenberg keeps
// telemetry opt-in: a signal with no exporter configured is built without one
// and stays inert. See https://github.com/gotenberg/gotenberg/issues/1643.
const (
tracesExporterEnvKey = "OTEL_TRACES_EXPORTER"
metricsExporterEnvKey = "OTEL_METRICS_EXPORTER"
logsExporterEnvKey = "OTEL_LOGS_EXPORTER"
)
// exporterConfigured reports whether the operator selected an exporter for the
// signal owning envKey. An unset or blank value keeps that signal off.
func exporterConfigured(envKey string) bool {
return strings.TrimSpace(os.Getenv(envKey)) != ""
}
// InitTracerProvider initializes the OpenTelemetry tracer provider.
func InitTracerProvider(logger *slog.Logger, serviceName, serviceVersion string) (shutdown func(context.Context) error, err error) {
initOtelLogger(logger)
@@ -86,13 +105,14 @@ func InitTracerProvider(logger *slog.Logger, serviceName, serviceVersion string)
trace.WithResource(res),
}
traceExporter, err := autoexport.NewSpanExporter(ctx)
if err != nil {
return nil, err
}
if !autoexport.IsNoneSpanExporter(traceExporter) {
traceOpts = append(traceOpts, trace.WithBatcher(traceExporter))
if exporterConfigured(tracesExporterEnvKey) {
traceExporter, err := autoexport.NewSpanExporter(ctx)
if err != nil {
return nil, err
}
if !autoexport.IsNoneSpanExporter(traceExporter) {
traceOpts = append(traceOpts, trace.WithBatcher(traceExporter))
}
}
traceProvider := trace.NewTracerProvider(traceOpts...)
@@ -119,13 +139,14 @@ func InitMeterProvider(logger *slog.Logger, serviceName, serviceVersion string)
}
metricOpts = append(metricOpts, exemplarFilterOptions()...)
metricReader, err := autoexport.NewMetricReader(ctx)
if err != nil {
return nil, err
}
if !autoexport.IsNoneMetricReader(metricReader) {
metricOpts = append(metricOpts, metric.WithReader(metricReader))
if exporterConfigured(metricsExporterEnvKey) {
metricReader, err := autoexport.NewMetricReader(ctx)
if err != nil {
return nil, err
}
if !autoexport.IsNoneMetricReader(metricReader) {
metricOpts = append(metricOpts, metric.WithReader(metricReader))
}
}
meterProvider := metric.NewMeterProvider(metricOpts...)
@@ -157,13 +178,14 @@ func InitLoggerProvider(logger *slog.Logger, serviceName, serviceVersion string)
log.WithResource(res),
}
logExporter, err := autoexport.NewLogExporter(ctx)
if err != nil {
return nil, nil, err
}
if !autoexport.IsNoneLogExporter(logExporter) {
logOpts = append(logOpts, log.WithProcessor(log.NewBatchProcessor(logExporter)))
if exporterConfigured(logsExporterEnvKey) {
logExporter, err := autoexport.NewLogExporter(ctx)
if err != nil {
return nil, nil, err
}
if !autoexport.IsNoneLogExporter(logExporter) {
logOpts = append(logOpts, log.WithProcessor(log.NewBatchProcessor(logExporter)))
}
}
loggerProvider := log.NewLoggerProvider(logOpts...)

View File

@@ -89,6 +89,49 @@ func TestInitTracerProvider_HonorsSamplerEnv(t *testing.T) {
}
}
// TestExporterConfigured pins the opt-in gate: an unset or blank
// OTEL_*_EXPORTER keeps the signal off, so Gotenberg never wires the OTLP
// exporter that autoexport would otherwise default to and fail to reach at
// localhost:4318. See https://github.com/gotenberg/gotenberg/issues/1643.
func TestExporterConfigured(t *testing.T) {
const key = "OTEL_METRICS_EXPORTER"
orig, had := os.LookupEnv(key)
t.Cleanup(func() {
if had {
os.Setenv(key, orig)
return
}
os.Unsetenv(key)
})
for _, tc := range []struct {
scenario string
unset bool
value string
want bool
}{
{"unset", true, "", false},
{"empty", false, "", false},
{"whitespace only", false, " ", false},
{"none", false, "none", true},
{"otlp", false, "otlp", true},
{"padded value", false, " otlp ", true},
} {
t.Run(tc.scenario, func(t *testing.T) {
if tc.unset {
os.Unsetenv(key)
} else {
os.Setenv(key, tc.value)
}
if got := exporterConfigured(key); got != tc.want {
t.Errorf("exporterConfigured(%q) = %v, want %v", key, got, tc.want)
}
})
}
}
func TestExemplarFilterOptions(t *testing.T) {
t.Run("default pins trace-based", func(t *testing.T) {
if v, ok := os.LookupEnv("OTEL_METRICS_EXEMPLAR_FILTER"); ok {

View File

@@ -49,6 +49,7 @@ type PdfEngineMock struct {
SplitMock func(ctx context.Context, logger *slog.Logger, mode SplitMode, inputPath, outputDirPath string) ([]string, error)
FlattenMock func(ctx context.Context, logger *slog.Logger, inputPath string) error
ConvertMock func(ctx context.Context, logger *slog.Logger, formats PdfFormats, inputPath, outputPath string) error
OptimizeImagesMock func(ctx context.Context, logger *slog.Logger, imageQuality int, inputPath string) error
ReadMetadataMock func(ctx context.Context, logger *slog.Logger, inputPath string) (map[string]any, error)
PageCountMock func(ctx context.Context, logger *slog.Logger, inputPath string) (int, error)
WriteMetadataMock func(ctx context.Context, logger *slog.Logger, metadata map[string]any, inputPath string) error
@@ -80,6 +81,10 @@ func (engine *PdfEngineMock) Convert(ctx context.Context, logger *slog.Logger, f
return engine.ConvertMock(ctx, logger, formats, inputPath, outputPath)
}
func (engine *PdfEngineMock) OptimizeImages(ctx context.Context, logger *slog.Logger, imageQuality int, inputPath string) error {
return engine.OptimizeImagesMock(ctx, logger, imageQuality, inputPath)
}
func (engine *PdfEngineMock) ReadMetadata(ctx context.Context, logger *slog.Logger, inputPath string) (map[string]any, error) {
return engine.ReadMetadataMock(ctx, logger, inputPath)
}

View File

@@ -16,6 +16,7 @@ import (
"time"
"github.com/dlclark/regexp2"
"github.com/hashicorp/go-retryablehttp"
"golang.org/x/net/http/httpproxy"
)
@@ -26,6 +27,12 @@ import (
// example [::ffff:127.0.0.1]).
var ErrNonPublicIP = errors.New("non-public IP")
// ErrFiltered happens when a value is rejected by an allow-list or a
// deny-list, or when it cannot be validated and [DecideOutbound] fails closed.
// Callers map it to a generic 403: the specific reason stays in the operator
// logs so a client cannot probe the lists.
var ErrFiltered = errors.New("value filtered")
// ErrPublicIP indicates that an outbound URL targets an IP address that is
// reachable on the public internet. It is returned when a caller opts
// into denying public destinations via [WithDenyPublicIPs]; typical use
@@ -78,6 +85,18 @@ var nonPublicIPv6Prefixes = []netip.Prefix{
netip.MustParsePrefix("100::/64"),
}
// nonPublicIPv4Prefixes lists IPv4 ranges that the [netip.Addr] helpers do
// not classify but that must not be considered public:
//
// - 100.64.0.0/10 Carrier-grade NAT (RFC 6598). Routable inside provider
// and cluster networks, and Alibaba Cloud serves instance metadata from
// 100.100.100.200.
// - 198.18.0.0/15 Benchmarking (RFC 2544). Never routed on the internet.
var nonPublicIPv4Prefixes = []netip.Prefix{
netip.MustParsePrefix("100.64.0.0/10"),
netip.MustParsePrefix("198.18.0.0/15"),
}
// IsPublicIP reports whether addr is reachable on the public internet. It
// returns false for loopback, private (RFC1918), link-local, unspecified,
// multicast, and unique-local addresses. IPv4-mapped IPv6 addresses are
@@ -88,7 +107,8 @@ var nonPublicIPv6Prefixes = []netip.Prefix{
// (6to4, Teredo, NAT64) are rejected wholesale rather than recursed into,
// because a host that routes them implicitly trusts the IPv4 mapping and
// the prefixes themselves are deprecated or translation-only. See
// [nonPublicIPv6Prefixes] for the full list and rationale.
// [nonPublicIPv6Prefixes] and [nonPublicIPv4Prefixes] for the full lists
// and rationale.
func IsPublicIP(addr netip.Addr) bool {
if !addr.IsValid() {
return false
@@ -104,6 +124,13 @@ func IsPublicIP(addr netip.Addr) bool {
addr.IsInterfaceLocalMulticast():
return false
}
if addr.Is4() {
for _, p := range nonPublicIPv4Prefixes {
if p.Contains(addr) {
return false
}
}
}
if addr.Is6() {
for _, p := range nonPublicIPv6Prefixes {
if p.Contains(addr) {
@@ -249,7 +276,9 @@ func httpLikeScheme(scheme string) bool {
//
// The semantics:
//
// 1. The URL is parsed and its scheme and host lowercased.
// 1. The URL is parsed, its scheme and host lowercased, and any userinfo
// dropped from the form the regexes see. The request still carries the
// credentials.
// 2. allowList and denyList apply against the normalized form with OR
// semantics. The deny-list always applies.
// 3. For http, https, ws, and wss, the host is resolved and every
@@ -268,26 +297,44 @@ func DecideOutbound(ctx context.Context, rawURL string, allowList, denyList []*r
opt(&cfg)
}
// Each match is bounded by [PatternMatchTimeout] rather than by the
// remaining budget, so an already-spent deadline no longer surfaces from
// the match itself. Schemes that resolve a host still learn about it from
// resolveHost, but a non-matching file:// or data: URL returns before that
// point, so check it here to keep failing closed on every path.
if !time.Now().Before(deadline) {
return OutboundDecision{}, context.DeadlineExceeded
}
parsed, err := url.Parse(rawURL)
if err != nil {
return OutboundDecision{}, fmt.Errorf("parse URL %q: %w", rawURL, ErrFiltered)
}
parsed.Scheme = strings.ToLower(parsed.Scheme)
parsed.Host = strings.ToLower(parsed.Host)
normalized := parsed.String()
// Match on a credential-free form. [url.URL.String] re-emits userinfo
// between "scheme://" and the host, so keeping it would let any
// "^https?://<host>" pattern be shifted past its own anchor:
// http://a@127.0.0.1/ escapes a deny-list anchored on 127\. and
// http://trusted.example.com@10.0.0.1/ satisfies an allow-list anchored on
// trusted\.example\.com. The host checks below already read
// [url.URL.Hostname], which ignores userinfo, so only the regex layer was
// affected. Dropping the credentials here also keeps them out of the error
// strings below, which reach operator logs and any OTEL log exporter.
matchable := *parsed
matchable.User = nil
normalized := matchable.String()
allowMatched := false
if len(allowList) > 0 {
for _, pattern := range allowList {
clone := regexp2.MustCompile(pattern.String(), 0)
clone.MatchTimeout = time.Until(deadline)
ok, err := clone.MatchString(normalized)
ok, err := pattern.MatchString(normalized)
if err != nil {
if time.Now().After(deadline) {
return OutboundDecision{}, context.DeadlineExceeded
}
return OutboundDecision{}, fmt.Errorf("'%s' cannot handle '%s': %w", clone.String(), normalized, err)
return OutboundDecision{}, fmt.Errorf("'%s' cannot handle '%s': %w", pattern.String(), normalized, err)
}
if ok {
@@ -302,15 +349,12 @@ func DecideOutbound(ctx context.Context, rawURL string, allowList, denyList []*r
}
for _, pattern := range denyList {
clone := regexp2.MustCompile(pattern.String(), 0)
clone.MatchTimeout = time.Until(deadline)
ok, err := clone.MatchString(normalized)
ok, err := pattern.MatchString(normalized)
if err != nil {
if time.Now().After(deadline) {
return OutboundDecision{}, context.DeadlineExceeded
}
return OutboundDecision{}, fmt.Errorf("'%s' cannot handle '%s': %w", clone.String(), normalized, err)
return OutboundDecision{}, fmt.Errorf("'%s' cannot handle '%s': %w", pattern.String(), normalized, err)
}
if ok {
@@ -338,8 +382,18 @@ func DecideOutbound(ctx context.Context, rawURL string, allowList, denyList []*r
return OutboundDecision{}, fmt.Errorf("'%s' targets a non-public address: %w", normalized, ErrFiltered)
case errors.Is(err, ErrPublicIP):
return OutboundDecision{}, fmt.Errorf("'%s' targets a public address: %w", normalized, ErrFiltered)
default:
case errors.Is(err, context.Canceled), errors.Is(err, context.DeadlineExceeded):
// A cancellation or timeout is not a policy decision; surface it
// as-is so callers do not report it as a filtered request.
return OutboundDecision{}, fmt.Errorf("validate '%s' host: %w", normalized, err)
default:
// The host could not be resolved, so its address class cannot be
// verified. Fail closed and treat it as filtered, the same as a
// host that resolves to a blocked address, so clients get a
// generic 403 rather than a 500. This also denies alternate IP
// encodings such as http://2130706433/ that the resolver rejects
// as a hostname but Chromium would read as a private IP.
return OutboundDecision{}, fmt.Errorf("validate '%s' host: %v: %w", normalized, err, ErrFiltered)
}
}
@@ -347,9 +401,8 @@ func DecideOutbound(ctx context.Context, rawURL string, allowList, denyList []*r
}
// FilterOutboundURL validates that rawURL is acceptable for an outbound
// request from Gotenberg. It is the URL-aware replacement for
// [FilterDeadline] and should be preferred for any new code that filters
// a URL before issuing or instructing an outbound request.
// request from Gotenberg. Prefer it for any new code that filters a URL
// before issuing or instructing an outbound request.
//
// The default behavior is permissive: the URL passes as long as it clears
// the regex allow-list and deny-list. Callers that need IP-class checks
@@ -426,6 +479,16 @@ func (rt *outboundRoundTripper) RoundTrip(req *http.Request) (*http.Response, er
// gate this behind their module's opt-in flag. See
// https://github.com/gotenberg/gotenberg/issues/1592.
func NewOutboundHttpClient(timeout time.Duration, allowList, denyList []*regexp2.Regexp, enableEnvironmentProxy bool, opts ...DecideOption) *http.Client {
// A negative timeout means the caller's budget is already spent, which
// happens when it derives one from a deadline that has passed. [http.Client]
// treats any non-positive Timeout as no deadline at all, so passing it
// through would silently produce an unbounded client. Fail closed instead.
// Zero keeps meaning unbounded: callers that own the connection lifetime
// themselves pass it deliberately.
if timeout < 0 {
timeout = time.Nanosecond
}
base := http.DefaultTransport.(*http.Transport).Clone()
var proxyFunc func(*url.URL) (*url.URL, error)
@@ -469,6 +532,25 @@ func NewOutboundHttpClient(timeout time.Duration, allowList, denyList []*regexp2
}
}
// ClampedBackoff is a [retryablehttp.Backoff] that honors max on every path.
//
// [retryablehttp.DefaultBackoff] returns a Retry-After header from the remote
// verbatim for 429 and 503, and returns it before applying its own max clamp.
// A hostile origin therefore decides how long Gotenberg waits, and the wait is
// not interruptible. Retry-After is still respected here, just never beyond
// the ceiling the caller set.
func ClampedBackoff(min, max time.Duration, attemptNum int, resp *http.Response) time.Duration {
wait := retryablehttp.DefaultBackoff(min, max, attemptNum, resp)
if wait > max {
return max
}
if wait < 0 {
return 0
}
return wait
}
// environmentProxyVariables are the variables golang.org/x/net/http/httpproxy
// reads, in the casing precedence it applies.
var environmentProxyVariables = []string{

View File

@@ -3,7 +3,9 @@ package gotenberg
import (
"context"
"errors"
"net/http"
"net/netip"
"strings"
"testing"
"time"
@@ -38,6 +40,24 @@ func TestIsPublicIP(t *testing.T) {
// Link-local.
{"169.254.169.254", false},
{"169.254.170.2", false},
// Carrier-grade NAT (RFC 6598). Alibaba Cloud serves instance
// metadata from 100.100.100.200.
{"100.64.0.0", false},
{"100.100.100.200", false},
{"100.127.255.255", false},
{"::ffff:100.100.100.200", false},
// Benchmarking (RFC 2544).
{"198.18.0.1", false},
{"198.19.255.255", false},
// Adjacent to the ranges above, and public.
{"100.63.255.255", true},
{"100.128.0.0", true},
{"198.17.255.255", true},
{"198.20.0.0", true},
{"fe80::1", false},
// Unique-local.
@@ -311,6 +331,39 @@ func TestFilterOutboundURL(t *testing.T) {
}
}
func TestDecideOutbound_UnresolvableHostFailsClosed(t *testing.T) {
withStubResolver(t, func(string) ([]netip.Addr, error) {
return nil, errors.New("no such host")
})
// An alternate IP encoding (decimal for 127.0.0.1) that the resolver
// rejects as a hostname must fail closed as filtered, not surface as a
// server error, so clients receive a generic 403.
_, err := DecideOutbound(context.Background(), "http://2130706433/", nil, nil, time.Now().Add(5*time.Second), WithDenyPrivateIPs(true))
if !errors.Is(err, ErrFiltered) {
t.Fatalf("expected ErrFiltered, got: %v", err)
}
}
func TestDecideOutbound_ResolverCancellationNotFiltered(t *testing.T) {
withStubResolver(t, func(string) ([]netip.Addr, error) {
return nil, context.Canceled
})
// A cancellation or timeout is not a policy decision and must not be
// reported as a filtered request.
_, err := DecideOutbound(context.Background(), "http://example.com/", nil, nil, time.Now().Add(5*time.Second), WithDenyPrivateIPs(true))
if err == nil {
t.Fatal("expected error, got nil")
}
if errors.Is(err, ErrFiltered) {
t.Fatalf("cancellation must not be filtered, got: %v", err)
}
if !errors.Is(err, context.Canceled) {
t.Fatalf("expected context.Canceled, got: %v", err)
}
}
func TestResolveAndCheckPublic_IPLiteralLoopback(t *testing.T) {
withStubResolver(t, func(host string) ([]netip.Addr, error) {
t.Fatalf("unexpected DNS lookup for %q", host)
@@ -487,3 +540,229 @@ func TestDecideOutbound_Permissive_AllowsPrivate(t *testing.T) {
t.Fatalf("decision.Pinned = %v, want [10.0.0.5]", decision.Pinned)
}
}
// privateIPsDenyList is the textual private-IP deny-list that shipped as the
// default for api-download-from-deny-list and webhook-deny-list in v8.31.0 and
// is still published as a migration recipe. Every alternative is anchored on
// "://", so userinfo used to slide the private address past the anchor.
const privateIPsDenyList = `^https?://(10\.|172\.(1[6-9]|2[0-9]|3[01])\.|192\.168\.|169\.254\.|0\.0\.0\.0|127\.|localhost|\[::1\]|\[fd)`
func TestDecideOutbound_UserinfoDoesNotEvadeDenyList(t *testing.T) {
for _, rawURL := range []string{
"http://127.0.0.1:9999/",
"http://a@127.0.0.1:9999/",
"http://@127.0.0.1:9999/",
"http://:@127.0.0.1:9999/",
"http://%61@127.0.0.1:9999/",
"http://user:pass@127.0.0.1:9999/",
"HTTP://A@127.0.0.1:9999/",
"http://a@169.254.169.254/latest/meta-data/",
// url.Parse takes the last "@" as the userinfo separator, so the host
// here is the second literal.
"http://a@127.0.0.1:9999@127.0.0.1:9999/",
} {
t.Run(rawURL, func(t *testing.T) {
withStubResolver(t, func(host string) ([]netip.Addr, error) {
t.Fatalf("unexpected DNS lookup for %q: the deny-list must reject before resolution", host)
return nil, nil
})
// Deny-list only, with the permissive IP defaults the modules ship.
_, err := DecideOutbound(
context.Background(),
rawURL,
nil,
[]*regexp2.Regexp{regexp2.MustCompile(privateIPsDenyList, 0)},
time.Now().Add(5*time.Second),
)
if !errors.Is(err, ErrFiltered) {
t.Fatalf("userinfo must not evade the deny-list, got: %v", err)
}
})
}
}
func TestDecideOutbound_UserinfoDoesNotSatisfyAllowList(t *testing.T) {
// A host-terminated allow-list, the shape the documentation recommends.
allowList := []*regexp2.Regexp{regexp2.MustCompile(`^https://trusted\.example\.com(:[0-9]+)?(/|$)`, 0)}
for _, rawURL := range []string{
"https://trusted.example.com@169.254.169.254/latest/meta-data/",
"https://trusted.example.com@10.0.0.5/",
"https://trusted.example.com:443@10.0.0.5/",
} {
t.Run(rawURL, func(t *testing.T) {
withStubResolver(t, func(host string) ([]netip.Addr, error) {
return mustAddrs(t, "10.0.0.5"), nil
})
decision, err := DecideOutbound(
context.Background(),
rawURL,
allowList, nil,
time.Now().Add(5*time.Second),
WithDenyPrivateIPs(true),
)
if err == nil {
t.Fatalf("userinfo must not satisfy the allow-list, got decision %+v", decision)
}
if decision.Bypass {
t.Fatal("userinfo must never produce a bypass")
}
})
}
}
func TestDecideOutbound_UserinfoKeptOutOfErrorMessages(t *testing.T) {
withStubResolver(t, func(host string) ([]netip.Addr, error) {
t.Fatalf("unexpected DNS lookup for %q", host)
return nil, nil
})
_, err := DecideOutbound(
context.Background(),
"http://alice:hunter2@127.0.0.1:9999/",
nil,
[]*regexp2.Regexp{regexp2.MustCompile(privateIPsDenyList, 0)},
time.Now().Add(5*time.Second),
)
if err == nil {
t.Fatal("expected the URL to be filtered")
}
if strings.Contains(err.Error(), "hunter2") || strings.Contains(err.Error(), "alice") {
t.Fatalf("error message must not leak URL credentials: %v", err)
}
}
func TestDecideOutbound_LegitimateCredentialsStillReachTheHost(t *testing.T) {
withStubResolver(t, func(host string) ([]netip.Addr, error) {
if host != "example.com" {
t.Fatalf("host = %q, want example.com: userinfo must not reach resolution", host)
}
return mustAddrs(t, "93.184.216.34"), nil
})
// Stripping userinfo is a matching concern only. A credentialed URL that
// breaks no rule must still be allowed through.
decision, err := DecideOutbound(
context.Background(),
"https://alice:hunter2@example.com/report.pdf",
[]*regexp2.Regexp{regexp2.MustCompile(`^https://example\.com(:[0-9]+)?(/|$)`, 0)},
nil,
time.Now().Add(5*time.Second),
WithDenyPrivateIPs(true),
)
if err != nil {
t.Fatalf("credentialed URL matching the allow-list must pass, got: %v", err)
}
if !decision.Bypass {
t.Fatalf("decision.Bypass = false, want true")
}
}
func TestClampedBackoff(t *testing.T) {
const (
min = 1 * time.Second
max = 30 * time.Second
)
retryAfter := func(status int, seconds string) *http.Response {
return &http.Response{StatusCode: status, Header: http.Header{"Retry-After": []string{seconds}}}
}
for _, tc := range []struct {
scenario string
resp *http.Response
want time.Duration
}{
{"429 with an hour is clamped", retryAfter(http.StatusTooManyRequests, "3600"), max},
{"429 with a day is clamped", retryAfter(http.StatusTooManyRequests, "86400"), max},
{"503 with an hour is clamped", retryAfter(http.StatusServiceUnavailable, "3600"), max},
{"429 under the ceiling is honored", retryAfter(http.StatusTooManyRequests, "5"), 5 * time.Second},
{"no response falls back to exponential", nil, min},
} {
t.Run(tc.scenario, func(t *testing.T) {
got := ClampedBackoff(min, max, 0, tc.resp)
if got != tc.want {
t.Fatalf("ClampedBackoff = %s, want %s", got, tc.want)
}
if got > max {
t.Fatalf("ClampedBackoff = %s, which exceeds max %s", got, max)
}
})
}
}
// A negative max means the caller's budget is spent. The backoff must not
// return a negative duration, which would make the retry loop spin.
func TestClampedBackoff_NegativeMaxIsNotNegative(t *testing.T) {
got := ClampedBackoff(1*time.Second, -5*time.Second, 0, nil)
if got < 0 {
t.Fatalf("ClampedBackoff = %s, want a non-negative duration", got)
}
}
func TestNewOutboundHttpClient_NonPositiveTimeout(t *testing.T) {
// Zero stays unbounded: the LibreOffice proxy owns its own lifetime and
// passes it deliberately.
if got := NewOutboundHttpClient(0, nil, nil, false).Timeout; got != 0 {
t.Fatalf("timeout for 0 = %s, want 0", got)
}
// Negative means an expired budget. http.Client reads any non-positive
// Timeout as no deadline at all, so it must not be passed through.
if got := NewOutboundHttpClient(-5*time.Second, nil, nil, false).Timeout; got <= 0 {
t.Fatalf("timeout for a negative budget = %s, want a positive value so the client fails closed", got)
}
}
func TestDecideOutboundExpiredDeadline(t *testing.T) {
// Patterns are matched under the fixed PatternMatchTimeout rather than
// under the caller's remaining budget, so an expired deadline no longer
// surfaces from the match itself. Every scheme must still fail closed,
// including the ones that return before a host is resolved.
expired := time.Now().Add(-time.Second)
for _, rawURL := range []string{
"https://example.com/",
"file:///tmp/foo.html",
"data:text/html,hello",
} {
_, err := DecideOutbound(context.Background(), rawURL, nil, nil, expired)
if !errors.Is(err, context.DeadlineExceeded) {
t.Fatalf("DecideOutbound(%q) with an expired deadline = %v, want context.DeadlineExceeded", rawURL, err)
}
}
}
func TestDecideOutboundBoundsCatastrophicPatterns(t *testing.T) {
// A deny-list pattern that backtracks catastrophically, matched against a
// client-controlled URL. Before PatternMatchTimeout the ceiling was the
// caller's whole budget, so a 30s API_TIMEOUT bought a 30s CPU burn.
// The trailing "!" makes the match fail only after the nested quantifier
// has explored every way to split the run of "a"s.
pattern := regexp2.MustCompile(`^https://example\.com/(a+)+$`, 0)
pattern.MatchTimeout = PatternMatchTimeout
rawURL := "https://example.com/" + strings.Repeat("a", 40) + "!"
start := time.Now()
_, err := DecideOutbound(
context.Background(),
rawURL,
nil,
[]*regexp2.Regexp{pattern},
time.Now().Add(30*time.Second),
)
elapsed := time.Since(start)
if err == nil {
t.Fatal("expected an error from a catastrophic deny-list pattern")
}
// Generous headroom over the 250ms ceiling, still far below the 30s
// deadline the match would otherwise have been allowed to consume.
if elapsed > 5*time.Second {
t.Fatalf("match took %s, want it aborted near PatternMatchTimeout (%s)", elapsed, PatternMatchTimeout)
}
}

View File

@@ -281,6 +281,12 @@ type PdfEngine interface {
// PdfFormats. If no format, it does nothing.
Convert(ctx context.Context, logger *slog.Logger, formats PdfFormats, inputPath, outputPath string) error
// OptimizeImages re-encodes the raster images of a PDF in place to shrink
// the file, leaving text, vectors, fonts and structure untouched.
// imageQuality is the JPEG quality (1 to 100) applied to each re-encoded
// image.
OptimizeImages(ctx context.Context, logger *slog.Logger, imageQuality int, inputPath string) error
// ReadMetadata extracts the metadata of a given PDF file.
ReadMetadata(ctx context.Context, logger *slog.Logger, inputPath string) (map[string]any, error)

View File

@@ -201,17 +201,13 @@ func TestNewServerRecordMetrics(t *testing.T) {
server.RecordMetrics(t.Context(), semconv.ServerMetricData{
ServerName: "stuff",
ResponseSize: 200,
MetricAttributes: semconv.MetricAttributes{
Req: req,
StatusCode: 301,
AdditionalAttributes: []attribute.KeyValue{
attribute.String("key", "value"),
},
},
MetricData: semconv.MetricData{
RequestSize: 100,
ElapsedTime: 300,
Req: req,
StatusCode: 301,
AdditionalAttributes: []attribute.KeyValue{
attribute.String("key", "value"),
},
RequestSize: 100,
ElapsedTime: 300,
})
rm := metricdata.ResourceMetrics{}

View File

@@ -59,8 +59,9 @@ type ProcessSupervisor interface {
// Healthy checks and returns the health status of the managed [Process].
//
// A non-started process is considered healthy (startup is deferred until
// the first request). Returns false if the process is currently restarting
// or is reported unhealthy by the underlying [Process].
// the first request), as is one going through a planned restart, since it
// keeps serving traffic. Returns false during an unplanned restart or when
// the underlying [Process] reports unhealthy.
Healthy() bool
// Run executes a provided task while managing the state of the [Process].
@@ -103,6 +104,28 @@ const healthCheckCacheTTL = 2 * time.Second
// this. See https://github.com/gotenberg/gotenberg/issues/1561.
const healthFailureThreshold = 2
// Restart reasons, also reported as the gotenberg.process.start.reason span
// attribute by [processSupervisor.tracedLaunch]. Only
// [restartReasonMaxRequests] is a planned restart: it fires on a healthy
// process that reached its conversion limit, so the node keeps serving
// traffic throughout. The others signal a process that cannot serve.
const (
restartReasonFirstStart = "first_start"
restartReasonUnhealthy = "unhealthy"
restartReasonMaxRequests = "max_requests"
)
// defaultEagerRestartTimeout bounds the restart triggered after the maximum
// request limit. That restart runs on a background context, unlike the one from
// ensureHealthy which inherits the request deadline, so without a deadline of
// its own the drain loop in [processSupervisor.doRestartLocked] would wait
// forever on a task that never completes. That would pin isRestarting and,
// with it, the health reported by [processSupervisor.Healthy]. Sized well above
// --api-timeout (30s by default) plus the engine start timeouts (20s by
// default) so it never fires while tasks are merely slow. The eager restart is
// opportunistic: on expiry it aborts, and the next task retries it.
const defaultEagerRestartTimeout = 2 * time.Minute
type processSupervisor struct {
logger *slog.Logger
engine string
@@ -118,11 +141,16 @@ type processSupervisor struct {
// transient failure (such as a cold-start timeout) must not poison the
// supervisor for the rest of the container's lifetime. See
// https://github.com/gotenberg/gotenberg/issues/1538.
firstStartMu sync.Mutex
reqCounter atomic.Int64
reqQueueSize atomic.Int64
restartsCounter atomic.Int64
isRestarting atomic.Bool
firstStartMu sync.Mutex
reqCounter atomic.Int64
reqQueueSize atomic.Int64
restartsCounter atomic.Int64
isRestarting atomic.Bool
// restartPlanned records whether the in-flight restart is a planned one
// (see [restartReasonMaxRequests]). Written before isRestarting and never
// cleared, so a reader that observed isRestarting always sees the matching
// kind. See [processSupervisor.Healthy].
restartPlanned atomic.Bool
activeTasks atomic.Int64
restartMutex sync.Mutex
idleShutdownTimeout time.Duration
@@ -135,6 +163,9 @@ type processSupervisor struct {
consecutiveHealthFailures atomic.Int64 // reset to 0 on every successful probe
idleMu sync.Mutex // protects idleStopChan
idleStopChan chan struct{} // signal to stop the idle ticker goroutine
// eagerRestartTimeout bounds the restart from maybeRestartAfterTask.
// Defaults to [defaultEagerRestartTimeout]; only tests shorten it.
eagerRestartTimeout time.Duration
}
// NewProcessSupervisor initializes a new [ProcessSupervisor]. engine names the
@@ -158,6 +189,7 @@ func NewProcessSupervisor(logger *slog.Logger, engine string, process Process, m
maxQueueSize: maxQueueSize,
maxConcurrency: maxConcurrency,
idleShutdownTimeout: idleShutdownTimeout,
eagerRestartTimeout: defaultEagerRestartTimeout,
}
b.reqCounter.Store(0)
b.reqQueueSize.Store(0)
@@ -212,12 +244,18 @@ func (s *processSupervisor) restart() error {
s.logger.WarnContext(context.Background(), fmt.Sprintf("stop process before restart: %s", err))
}
// Reset the counter on the attempt, not on its outcome. Leaving it at the
// limit after a failed launch re-triggers maybeRestartAfterTask on every
// subsequent task, producing back-to-back restarts. Recovering a process
// that will not start is ensureHealthy's job: it restarts synchronously
// before running a task, and reports the failure to the caller.
s.reqCounter.Store(0)
err = s.Launch()
if err != nil {
return fmt.Errorf("restart process: %w", err)
}
s.reqCounter.Store(0)
s.restartsCounter.Add(1)
s.logger.DebugContext(context.Background(), "process successfully restarted")
@@ -234,9 +272,17 @@ func (s *processSupervisor) Healthy() bool {
}
if s.isRestarting.Load() {
// A restarting process is not yet healthy. This gives load balancers
// honest information so they can avoid routing traffic to this node.
return false
// A planned restart is routine maintenance: the process reached the
// limit set by --chromium-restart-after (env CHROMIUM_RESTART_AFTER) or
// --libreoffice-restart-after (env LIBREOFFICE_RESTART_AFTER) while
// healthy. Tasks arriving during it are requeued by acquireSlot, not
// rejected, so the node still serves traffic and must report healthy. A
// probe sent between two conversions used to fail here.
// See https://github.com/gotenberg/gotenberg/issues/1648.
//
// An unplanned restart keeps reporting unhealthy, which gives load
// balancers honest information so they can avoid routing traffic here.
return s.restartPlanned.Load()
}
// Cache hit: a recent probe succeeded. Skip the CDP roundtrip so probe
@@ -484,7 +530,7 @@ func (s *processSupervisor) ensureStarted(ctx context.Context) error {
return nil
}
err := s.tracedLaunch(ctx, "first_start", func() error {
err := s.tracedLaunch(ctx, restartReasonFirstStart, func() error {
return s.runWithDeadline(ctx, s.Launch)
})
if err != nil {
@@ -525,7 +571,7 @@ func (s *processSupervisor) ensureHealthy(ctx context.Context) error {
s.logger.DebugContext(context.Background(), "process is unhealthy, cannot handle task, restarting...")
if err := s.doRestart(ctx, "unhealthy"); err != nil {
if err := s.doRestart(ctx, restartReasonUnhealthy); err != nil {
return fmt.Errorf("process restart before task: %w", err)
}
@@ -533,9 +579,10 @@ func (s *processSupervisor) ensureHealthy(ctx context.Context) error {
}
// maybeRestartAfterTask checks if the maximum request limit has been reached
// and, if so, triggers an asynchronous restart. If a restart is initiated, it
// takes ownership of the caller's semaphore slot (the caller must not release
// it). Returns true if ownership was taken.
// and, if so, triggers an asynchronous restart bounded by
// [defaultEagerRestartTimeout]. If a restart is initiated, it takes ownership
// of the caller's semaphore slot (the caller must not release it). Returns true
// if ownership was taken.
func (s *processSupervisor) maybeRestartAfterTask(logger *slog.Logger) bool {
if s.maxReqLimit <= 0 || s.reqCounter.Load() < s.maxReqLimit {
return false
@@ -548,7 +595,10 @@ func (s *processSupervisor) maybeRestartAfterTask(logger *slog.Logger) bool {
s.logger.DebugContext(context.Background(), "max request limit reached, restarting eagerly...")
go func() {
restartErr := s.doRestartLocked(context.Background(), "max_requests")
ctx, cancel := context.WithTimeout(context.Background(), s.eagerRestartTimeout)
defer cancel()
restartErr := s.doRestartLocked(ctx, restartReasonMaxRequests)
s.restartMutex.Unlock()
if restartErr != nil {
s.logger.ErrorContext(context.Background(), fmt.Sprintf("process restart after task: %v", restartErr))
@@ -571,6 +621,10 @@ func (s *processSupervisor) doRestart(ctx context.Context, reason string) error
// doRestartLocked performs the restart drain logic. The caller must hold restartMutex.
func (s *processSupervisor) doRestartLocked(ctx context.Context, reason string) error {
// Publish the kind before raising the flag. [processSupervisor.Healthy]
// reads restartPlanned only after it observes isRestarting, so this
// ordering keeps it from pairing a new restart with a stale kind.
s.restartPlanned.Store(reason == restartReasonMaxRequests)
s.isRestarting.Store(true)
defer s.isRestarting.Store(false)

View File

@@ -160,13 +160,49 @@ func TestProcessSupervisor_restart(t *testing.T) {
}
}
// TestProcessSupervisor_restart_ResetsCounterOnFailedLaunch verifies that a
// restart whose launch fails still clears the request counter. Leaving it at
// the limit makes maybeRestartAfterTask re-fire on every subsequent task.
func TestProcessSupervisor_restart_ResetsCounterOnFailedLaunch(t *testing.T) {
logger := slog.New(slog.DiscardHandler)
const maxReqLimit = 5
process := &ProcessMock{
StartMock: func(_ *slog.Logger) error { return errors.New("start error") },
StopMock: func(_ *slog.Logger) error { return nil },
HealthyMock: func(_ *slog.Logger) bool { return true },
}
ps := NewProcessSupervisor(logger, "test", process, maxReqLimit, 0, 1, 0).(*processSupervisor)
ps.reqCounter.Store(maxReqLimit)
err := ps.restart()
if err == nil {
t.Fatal("expected error but got none")
}
if got := ps.reqCounter.Load(); got != 0 {
t.Fatalf("expected the request counter to be reset but got %d", got)
}
if got := ps.restartsCounter.Load(); got != 0 {
t.Fatalf("expected the restarts counter to stay at 0 but got %d", got)
}
if ps.maybeRestartAfterTask(logger) {
t.Fatal("expected no further eager restart to be triggered")
}
}
func TestProcessSupervisor_Healthy(t *testing.T) {
for _, tc := range []struct {
scenario string
initiallyStarted bool
initiallyRestarting bool
processHealthy bool
expectHealthy bool
scenario string
initiallyStarted bool
initiallyRestarting bool
initiallyRestartPlanned bool
processHealthy bool
expectHealthy bool
}{
{
scenario: "non-started process is healthy",
@@ -179,6 +215,13 @@ func TestProcessSupervisor_Healthy(t *testing.T) {
initiallyRestarting: true,
expectHealthy: false,
},
{
scenario: "process going through a planned restart is healthy",
initiallyStarted: true,
initiallyRestarting: true,
initiallyRestartPlanned: true,
expectHealthy: true,
},
{
scenario: "process reports as healthy",
initiallyStarted: true,
@@ -208,6 +251,9 @@ func TestProcessSupervisor_Healthy(t *testing.T) {
if tc.initiallyRestarting {
ps.isRestarting.Store(true)
}
if tc.initiallyRestartPlanned {
ps.restartPlanned.Store(true)
}
healthy := ps.Healthy()
@@ -257,6 +303,204 @@ func TestProcessSupervisor_Healthy_ConsecutiveFailures(t *testing.T) {
}
}
// TestProcessSupervisor_Healthy_PlannedRestart reproduces
// https://github.com/gotenberg/gotenberg/issues/1648. It drives the real
// Run() path until the maximum request limit triggers the eager restart, then
// asserts the supervisor reports healthy while that restart is in flight.
// Tasks arriving during it are requeued by acquireSlot, not rejected, so the
// node still serves traffic.
func TestProcessSupervisor_Healthy_PlannedRestart(t *testing.T) {
logger := slog.New(slog.DiscardHandler)
const maxReqLimit = 10
restarting := make(chan struct{})
release := make(chan struct{})
var (
starts atomic.Int64
signalOne sync.Once
)
process := &ProcessMock{
StartMock: func(_ *slog.Logger) error {
// Hold the restart open so the assertions below run inside the
// window that used to report unhealthy.
if starts.Add(1) > 1 {
signalOne.Do(func() { close(restarting) })
<-release
}
return nil
},
StopMock: func(_ *slog.Logger) error { return nil },
HealthyMock: func(_ *slog.Logger) bool { return true },
}
ps := NewProcessSupervisor(logger, "test", process, maxReqLimit, 0, 1, 0).(*processSupervisor)
for i := range maxReqLimit {
err := ps.Run(context.Background(), logger, func() error { return nil })
if err != nil {
t.Fatalf("task %d: expected no error but got: %v", i+1, err)
}
}
select {
case <-restarting:
case <-time.After(10 * time.Second):
t.Fatalf("expected an eager restart after %d tasks", maxReqLimit)
}
if !ps.isRestarting.Load() {
t.Fatal("expected the supervisor to be restarting")
}
if !ps.restartPlanned.Load() {
t.Fatal("expected the restart to be flagged as planned")
}
if !ps.Healthy() {
t.Fatal("expected a planned restart to report healthy")
}
close(release)
}
// TestProcessSupervisor_Healthy_UnplannedRestart verifies the counterpart of
// [TestProcessSupervisor_Healthy_PlannedRestart]: a restart triggered by an
// unhealthy process keeps reporting unhealthy, so load balancers get honest
// information.
func TestProcessSupervisor_Healthy_UnplannedRestart(t *testing.T) {
logger := slog.New(slog.DiscardHandler)
restarting := make(chan struct{})
release := make(chan struct{})
var signalOne sync.Once
process := &ProcessMock{
StartMock: func(_ *slog.Logger) error {
signalOne.Do(func() { close(restarting) })
<-release
return nil
},
StopMock: func(_ *slog.Logger) error { return nil },
HealthyMock: func(_ *slog.Logger) bool { return false },
}
ps := NewProcessSupervisor(logger, "test", process, 0, 0, 1, 0).(*processSupervisor)
ps.firstStart.Store(true)
go func() {
_ = ps.ensureHealthy(context.Background())
}()
select {
case <-restarting:
case <-time.After(10 * time.Second):
t.Fatal("expected an unhealthy restart to be triggered")
}
if !ps.isRestarting.Load() {
t.Fatal("expected the supervisor to be restarting")
}
if ps.restartPlanned.Load() {
t.Fatal("expected the restart not to be flagged as planned")
}
if ps.Healthy() {
t.Fatal("expected an unplanned restart to report unhealthy")
}
close(release)
}
// TestProcessSupervisor_doRestartLocked_DrainDeadline verifies that a drain
// unable to acquire every slot gives up on the context deadline and clears
// isRestarting. Without a deadline on the eager restart, a task that never
// completes would pin the flag and, since a planned restart reports healthy,
// leave the supervisor claiming health forever.
func TestProcessSupervisor_doRestartLocked_DrainDeadline(t *testing.T) {
logger := slog.New(slog.DiscardHandler)
var starts atomic.Int64
process := &ProcessMock{
StartMock: func(_ *slog.Logger) error {
starts.Add(1)
return nil
},
StopMock: func(_ *slog.Logger) error { return nil },
HealthyMock: func(_ *slog.Logger) bool { return true },
}
// A concurrency of 2 makes the drain acquire one slot on top of the one the
// triggering task hands over. Fill the semaphore so it never can, mimicking
// a concurrent task that never completes.
ps := NewProcessSupervisor(logger, "test", process, 1, 0, 2, 0).(*processSupervisor)
ps.semaphore <- struct{}{}
ps.semaphore <- struct{}{}
ctx, cancel := context.WithTimeout(context.Background(), 50*time.Millisecond)
defer cancel()
err := ps.doRestartLocked(ctx, restartReasonMaxRequests)
if err == nil {
t.Fatal("expected the drain to fail on the context deadline")
}
if ps.isRestarting.Load() {
t.Fatal("expected isRestarting to be cleared after a failed drain")
}
if starts.Load() != 0 {
t.Fatalf("expected no restart attempt after a failed drain but got %d", starts.Load())
}
}
// TestProcessSupervisor_maybeRestartAfterTask_Bounded verifies that the eager
// restart runs under a deadline. A concurrent task that never completes blocks
// the drain, and without a bound the restart goroutine would wait forever with
// isRestarting pinned, leaving Healthy() reporting a planned restart for good.
func TestProcessSupervisor_maybeRestartAfterTask_Bounded(t *testing.T) {
logger := slog.New(slog.DiscardHandler)
process := &ProcessMock{
StartMock: func(_ *slog.Logger) error { return nil },
StopMock: func(_ *slog.Logger) error { return nil },
HealthyMock: func(_ *slog.Logger) bool { return true },
}
ps := NewProcessSupervisor(logger, "test", process, 1, 0, 2, 0).(*processSupervisor)
ps.eagerRestartTimeout = 100 * time.Millisecond
ps.firstStart.Store(true)
// Wedge one slot so the drain, which needs one on top of the slot the
// triggering task hands over, can never complete.
ps.semaphore <- struct{}{}
err := ps.Run(context.Background(), logger, func() error { return nil })
if err != nil {
t.Fatalf("expected no error but got: %v", err)
}
waitFor := func(what string, want bool) {
t.Helper()
deadline := time.Now().Add(5 * time.Second)
for ps.isRestarting.Load() != want {
if time.Now().After(deadline) {
t.Fatalf("timed out waiting for the eager restart to %s", what)
}
time.Sleep(5 * time.Millisecond)
}
}
waitFor("start", true)
waitFor("give up on its deadline", false)
}
// TestProcessSupervisor_Healthy_CachesPositiveResult verifies that a
// successful probe is cached for [healthCheckCacheTTL] so subsequent
// supervisor.Healthy() calls do not re-issue the underlying process

View File

@@ -39,6 +39,10 @@ type Api struct {
correlationIdHeader string
basicAuthUsername string
basicAuthPassword string
oidcEnabled bool
oidcIssuer string
oidcAudience string
oidcJwksUrl string
downloadFromCfg downloadFromConfig
disableHealthCheckRouteTelemetry bool
disableRootRouteTelemetry bool
@@ -63,6 +67,8 @@ type downloadFromConfig struct {
denyPublicIPs bool
enableEnvironmentProxy bool
maxRetry int
maxConcurrency int
maxEntries int
disable bool
}
@@ -198,12 +204,18 @@ func (a *Api) Descriptor() gotenberg.ModuleDescriptor {
fs.String("api-root-path", "/", "Set the root path of the API - for service discovery via URL paths")
fs.String("api-correlation-id-header", "Gotenberg-Trace", "Set the header name to use for identifying requests")
fs.Bool("api-enable-basic-auth", false, "Enable basic authentication - will look for the GOTENBERG_API_BASIC_AUTH_USERNAME and GOTENBERG_API_BASIC_AUTH_PASSWORD environment variables")
fs.StringSlice("api-download-from-allow-list", []string{}, "Set the allowed URLs for the download from feature using regular expressions - supports multiple values")
fs.Bool("api-enable-oidc-auth", false, "Enable OIDC bearer token authentication - mutually exclusive with basic authentication")
fs.String("api-oidc-issuer", "", "Set the OIDC issuer URL, e.g. https://tenant.example.com/ - the token 'iss' claim must match")
fs.String("api-oidc-audience", "", "Set the expected OIDC audience - the token 'aud' claim must contain it")
fs.String("api-oidc-jwks-url", "", "Set the OIDC JWKS URL - discovered from the issuer's well-known configuration when empty")
fs.StringSlice("api-download-from-allow-list", []string{}, `Set the allowed URLs for the download from feature using regular expressions - supports multiple values. A match bypasses --api-download-from-deny-private-ips (API_DOWNLOAD_FROM_DENY_PRIVATE_IPS) and --api-download-from-deny-public-ips (API_DOWNLOAD_FROM_DENY_PUBLIC_IPS), so terminate the host or the pattern also matches suffix hosts, for example ^https?://internal\.svc(:|/|$)`)
fs.StringSlice("api-download-from-deny-list", []string{}, "Set the denied URLs for the download from feature using regular expressions - supports multiple values")
fs.Bool("api-download-from-deny-private-ips", false, "Reject downloadFrom URLs whose host resolves to a non-public IP address (loopback, RFC1918, link-local, unique-local). Enable on deployments that accept untrusted downloadFrom sources to mitigate SSRF against internal services")
fs.Bool("api-download-from-deny-public-ips", false, "Reject downloadFrom URLs whose host resolves to a public IP address. Enable on air-gapped or data-governed deployments to prevent downloads from reaching the public internet")
fs.Bool("api-download-from-enable-environment-proxy", false, "Route downloadFrom fetches through the proxy defined by the standard HTTP_PROXY, HTTPS_PROXY, and NO_PROXY variables, including credentials")
fs.Int("api-download-from-max-retry", 4, "Set the maximum number of retries for the download from feature")
fs.Int("api-download-from-max-concurrency", 10, "Set the maximum number of downloadFrom entries fetched concurrently per request - bounds the outbound fan-out. Set to 0 to disable this feature")
fs.Int("api-download-from-max-entries", 0, "Set the maximum number of downloadFrom entries allowed per request. Set to 0 to disable this feature")
fs.Bool("api-disable-download-from", false, "Disable the download from feature")
fs.Bool("api-disable-health-check-route-telemetry", true, "Disable telemetry for health check route")
fs.Bool("api-disable-root-route-telemetry", true, "Disable telemetry for the root route")
@@ -247,6 +259,8 @@ func (a *Api) Provision(ctx *gotenberg.Context) error {
denyPublicIPs: flags.MustBool("api-download-from-deny-public-ips"),
enableEnvironmentProxy: flags.MustBool("api-download-from-enable-environment-proxy"),
maxRetry: flags.MustInt("api-download-from-max-retry"),
maxConcurrency: flags.MustInt("api-download-from-max-concurrency"),
maxEntries: flags.MustInt("api-download-from-max-entries"),
disable: flags.MustBool("api-disable-download-from"),
}
a.disableHealthCheckRouteTelemetry = flags.MustDeprecatedBool("api-disable-health-check-logging", "api-disable-health-check-route-telemetry")
@@ -280,6 +294,15 @@ func (a *Api) Provision(ctx *gotenberg.Context) error {
a.basicAuthPassword = basicAuthPassword
}
// Enable OIDC auth? The flags are populated from their API_OIDC_* env vars
// by the CLI, so no manual environment lookup is needed here.
a.oidcEnabled = flags.MustBool("api-enable-oidc-auth")
if a.oidcEnabled {
a.oidcIssuer = flags.MustString("api-oidc-issuer")
a.oidcAudience = flags.MustString("api-oidc-audience")
a.oidcJwksUrl = flags.MustString("api-oidc-jwks-url")
}
// Get routes from modules.
mods, err := ctx.Modules(new(Router))
if err != nil {
@@ -360,12 +383,37 @@ func (a *Api) Provision(ctx *gotenberg.Context) error {
// Logger.
a.logger = gotenberg.Logger(a)
a.warnInsecureDebugRoute()
// File system.
a.fs = gotenberg.NewFileSystem(new(gotenberg.OsMkdirAll))
return nil
}
// warnInsecureDebugRoute logs a warning when the debug route is reachable
// without authentication.
//
// The route reports the resolved configuration of every module, which is
// useful to an operator and equally useful to anyone else who can reach it.
// This warns rather than refuses: an operator may sit behind a gateway that
// authenticates on Gotenberg's behalf, and failing startup would break them.
func (a *Api) warnInsecureDebugRoute() {
if !a.enableDebugRoute || a.basicAuthUsername != "" || a.oidcEnabled {
return
}
if a.logger == nil {
return
}
a.logger.WarnContext(
context.Background(),
"--api-enable-debug-route (API_ENABLE_DEBUG_ROUTE) is enabled but no authentication is configured, so anyone who can reach Gotenberg can read its configuration. Set --api-enable-basic-auth (API_ENABLE_BASIC_AUTH) with GOTENBERG_API_BASIC_AUTH_USERNAME and GOTENBERG_API_BASIC_AUTH_PASSWORD, set --api-enable-oidc-auth (API_ENABLE_OIDC_AUTH), or disable the route.",
slog.String("flag", "--api-enable-debug-route"),
slog.String("env", "API_ENABLE_DEBUG_ROUTE"),
)
}
// Validate validates the module properties.
func (a *Api) Validate() error {
var err error
@@ -387,6 +435,18 @@ func (a *Api) Validate() error {
}
}
if a.downloadFromCfg.maxConcurrency < 0 {
err = errors.Join(err,
fmt.Errorf("download from max concurrency must not be negative, got %d; set --api-download-from-max-concurrency (env API_DOWNLOAD_FROM_MAX_CONCURRENCY) to 0 to disable the limit", a.downloadFromCfg.maxConcurrency),
)
}
if a.downloadFromCfg.maxEntries < 0 {
err = errors.Join(err,
fmt.Errorf("download from max entries must not be negative, got %d; set --api-download-from-max-entries (env API_DOWNLOAD_FROM_MAX_ENTRIES) to 0 to disable the limit", a.downloadFromCfg.maxEntries),
)
}
if (a.tlsCertFile != "" && a.tlsKeyFile == "") || (a.tlsCertFile == "" && a.tlsKeyFile != "") {
err = errors.Join(err,
errors.New("both TLS certificate and key files must be set"),
@@ -411,6 +471,25 @@ func (a *Api) Validate() error {
)
}
if a.basicAuthUsername != "" && a.oidcEnabled {
err = errors.Join(err,
errors.New("basic authentication and OIDC authentication cannot both be enabled"),
)
}
if a.oidcEnabled {
if a.oidcIssuer == "" {
err = errors.Join(err,
errors.New("OIDC issuer must not be empty when OIDC auth is enabled; set --api-oidc-issuer"),
)
}
if a.oidcAudience == "" {
err = errors.Join(err,
errors.New("OIDC audience must not be empty when OIDC auth is enabled; set --api-oidc-audience"),
)
}
}
if err != nil {
return err
}
@@ -517,11 +596,18 @@ func (a *Api) Start() error {
hardTimeout := a.timeout + (time.Duration(5) * time.Second)
// Basic auth?
// Authentication?
var securityMiddleware echo.MiddlewareFunc
if a.basicAuthUsername != "" {
switch {
case a.basicAuthUsername != "":
securityMiddleware = basicAuthMiddleware(a.basicAuthUsername, a.basicAuthPassword)
} else {
case a.oidcEnabled:
verifier, err := a.buildOidcVerifier()
if err != nil {
return fmt.Errorf("build OIDC verifier: %w", err)
}
securityMiddleware = oidcAuthMiddleware(verifier)
default:
securityMiddleware = func(next echo.HandlerFunc) echo.HandlerFunc {
return func(c echo.Context) error {
return next(c)

134
pkg/modules/api/api_test.go Normal file
View File

@@ -0,0 +1,134 @@
package api
import (
"bytes"
"log/slog"
"strings"
"testing"
)
func TestApi_Validate_Auth(t *testing.T) {
base := func() *Api {
return &Api{port: 3000, rootPath: "/", correlationIdHeader: "Gotenberg-Trace"}
}
for _, tc := range []struct {
scenario string
mutate func(*Api)
wantErr string // substring expected in the error, "" means no error
}{
{"no auth", func(*Api) {}, ""},
{"basic auth only", func(a *Api) { a.basicAuthUsername = "foo" }, ""},
{
"oidc auth valid",
func(a *Api) {
a.oidcEnabled = true
a.oidcIssuer = "https://tenant.example.com/"
a.oidcAudience = "gotenberg"
},
"",
},
{
"basic and oidc are mutually exclusive",
func(a *Api) {
a.basicAuthUsername = "foo"
a.oidcEnabled = true
a.oidcIssuer = "https://tenant.example.com/"
a.oidcAudience = "gotenberg"
},
"cannot both be enabled",
},
{
"oidc missing issuer",
func(a *Api) { a.oidcEnabled = true; a.oidcAudience = "gotenberg" },
"issuer must not be empty",
},
{
"oidc missing audience",
func(a *Api) { a.oidcEnabled = true; a.oidcIssuer = "https://tenant.example.com/" },
"audience must not be empty",
},
} {
t.Run(tc.scenario, func(t *testing.T) {
a := base()
tc.mutate(a)
err := a.Validate()
if tc.wantErr == "" {
if err != nil {
t.Fatalf("expected no error, got %v", err)
}
return
}
if err == nil || !strings.Contains(err.Error(), tc.wantErr) {
t.Fatalf("error = %v, want a substring %q", err, tc.wantErr)
}
})
}
}
func TestApi_warnInsecureDebugRoute(t *testing.T) {
for _, tc := range []struct {
scenario string
api Api
expectWarn bool
expectFields []string
}{
{
scenario: "debug route on with no auth warns",
api: Api{enableDebugRoute: true},
expectWarn: true,
expectFields: []string{"--api-enable-debug-route", "API_ENABLE_DEBUG_ROUTE", "--api-enable-basic-auth", "API_ENABLE_BASIC_AUTH", "--api-enable-oidc-auth", "API_ENABLE_OIDC_AUTH"},
},
{
scenario: "debug route off is silent",
api: Api{enableDebugRoute: false},
expectWarn: false,
},
{
scenario: "basic auth silences it",
api: Api{enableDebugRoute: true, basicAuthUsername: "foo"},
expectWarn: false,
},
{
scenario: "oidc silences it",
api: Api{enableDebugRoute: true, oidcEnabled: true},
expectWarn: false,
},
} {
t.Run(tc.scenario, func(t *testing.T) {
buf := new(bytes.Buffer)
tc.api.logger = slog.New(slog.NewJSONHandler(buf, &slog.HandlerOptions{Level: slog.LevelWarn}))
tc.api.warnInsecureDebugRoute()
logged := buf.String()
if !tc.expectWarn {
if logged != "" {
t.Fatalf("expected no warning, got: %s", logged)
}
return
}
if logged == "" {
t.Fatal("expected a warning, got none")
}
// Every flag named must carry its environment variable.
for _, want := range tc.expectFields {
if !strings.Contains(logged, want) {
t.Fatalf("warning does not mention %q: %s", want, logged)
}
}
if strings.Contains(logged, "—") {
t.Fatalf("warning must not contain an em dash: %s", logged)
}
})
}
}
// The warning reads a.logger, which is nil until Provision assigns it.
func TestApi_warnInsecureDebugRoute_NilLoggerDoesNotPanic(t *testing.T) {
api := Api{enableDebugRoute: true}
api.warnInsecureDebugRoute()
}

View File

@@ -51,6 +51,24 @@ type Context struct {
outputPaths []string
cancelled bool
// fileOrder records the order files were received in, keyed by disk path.
// It breaks ties when two uploads share an original filename, so that
// de-duplicated files keep their upload order instead of being ordered by
// the suffix uniqueFilename added.
fileOrder map[string]int
// fileBase maps a disk path to the original filename as received, before
// de-duplication. Sorting on it keeps a de-duplicated file next to its
// twin rather than wherever its numbered name would land.
fileBase map[string]string
// outputFilename is the sanitized Gotenberg-Output-Filename header,
// snapshotted while the [echo.Context] is still live. Echo returns that
// context to a pool as soon as the handler returns, and an asynchronous
// conversion outlives it, so reading the header from the pooled store later
// yields whichever request happens to own it by then.
outputFilename string
logger *slog.Logger
echoCtx echo.Context
mkdirAll gotenberg.MkdirAll
@@ -80,6 +98,48 @@ func (t *trackingReader) Read(p []byte) (int, error) {
return n, nil
}
// errTooManyDownloadFromEntries is returned by [decodeDownloadFrom] when the
// array holds more entries than the configured maximum.
var errTooManyDownloadFromEntries = errors.New("too many downloadFrom entries")
// decodeDownloadFrom decodes the downloadFrom form field, refusing to
// accumulate more than maxEntries. A maxEntries of 0 means no limit.
//
// It decodes element by element rather than calling [json.Unmarshal] on the
// whole value. A compact array such as "[{},{},{}]" costs three bytes per
// entry on the wire and expands to roughly seventy times that once
// unmarshalled, so counting the entries afterwards is too late to bound the
// allocation. Streaming keeps the cost proportional to maxEntries no matter
// how long the array is.
func decodeDownloadFrom(raw string, maxEntries int) ([]downloadFrom, error) {
dec := json.NewDecoder(strings.NewReader(raw))
token, err := dec.Token()
if err != nil {
return nil, err
}
if delim, ok := token.(json.Delim); !ok || delim != '[' {
return nil, fmt.Errorf("expected a JSON array, got '%v'", token)
}
var dls []downloadFrom
for dec.More() {
if maxEntries > 0 && len(dls) >= maxEntries {
return nil, errTooManyDownloadFromEntries
}
var dl downloadFrom
err = dec.Decode(&dl)
if err != nil {
return nil, err
}
dls = append(dls, dl)
}
return dls, nil
}
type downloadFrom struct {
// Url is the URL to download a file from.
Url string `json:"url"`
@@ -117,14 +177,18 @@ func newContext(echoCtx echo.Context, logger *slog.Logger, fs *gotenberg.FileSys
return nil
}
// Snapshot now, while echoCtx still belongs to this request.
outputFilename, _ := echoCtx.Get("outputFilename").(string)
ctx := &Context{
outputPaths: make([]string, 0),
cancelled: false,
logger: logger,
echoCtx: echoCtx,
mkdirAll: new(gotenberg.OsMkdirAll),
pathRename: new(gotenberg.OsPathRename),
Context: processCtx,
outputPaths: make([]string, 0),
cancelled: false,
outputFilename: outputFilename,
logger: logger,
echoCtx: echoCtx,
mkdirAll: new(gotenberg.OsMkdirAll),
pathRename: new(gotenberg.OsPathRename),
Context: processCtx,
}
// A custom cancel function which removes the context's working directory
@@ -178,6 +242,12 @@ func newContext(echoCtx echo.Context, logger *slog.Logger, fs *gotenberg.FileSys
return nil, cancel, fmt.Errorf("get multipart form: %w", err)
}
defer func() {
err := form.RemoveAll()
if err != nil {
logger.ErrorContext(context.Background(), fmt.Sprintf("remove multipart temporary files: %s", err))
}
}()
// This will ensure we do not exceed the body limit.
var formValuesSize int64
@@ -207,8 +277,13 @@ func newContext(echoCtx echo.Context, logger *slog.Logger, fs *gotenberg.FileSys
// any.
raw, ok := ctx.values["downloadFrom"]
if !downloadFromCfg.disable && ok {
var dls []downloadFrom
err = json.Unmarshal([]byte(raw[0]), &dls)
dls, err := decodeDownloadFrom(raw[0], downloadFromCfg.maxEntries)
if errors.Is(err, errTooManyDownloadFromEntries) {
return nil, cancel, WrapError(
fmt.Errorf("decode downloadFrom: %w", err),
NewSentinelHttpError(http.StatusBadRequest, fmt.Sprintf("Invalid 'downloadFrom' form field value: too many entries, the maximum is %d", downloadFromCfg.maxEntries)),
)
}
if err != nil {
return nil, cancel, WrapError(
fmt.Errorf("unmarshal json: %w", err),
@@ -226,6 +301,13 @@ func newContext(echoCtx echo.Context, logger *slog.Logger, fs *gotenberg.FileSys
results := make([]downloadFromResult, len(dls))
eg, _ := errgroup.WithContext(ctx)
// Bound the number of in-flight downloads. Each entry allocates a
// retryable client, an outbound transport, a span, and logger state,
// so an unbounded array would otherwise exhaust process memory. A
// value of 0 keeps the fan-out unbounded.
if downloadFromCfg.maxConcurrency > 0 {
eg.SetLimit(downloadFromCfg.maxConcurrency)
}
for i, dl := range dls {
eg.Go(func() error {
deadline, ok := ctx.Deadline()
@@ -257,7 +339,12 @@ func newContext(echoCtx echo.Context, logger *slog.Logger, fs *gotenberg.FileSys
logger.DebugContext(dlCtx, fmt.Sprintf("download file from '%s'", dl.Url))
req, err := retryablehttp.NewRequest(http.MethodGet, dl.Url, nil)
// The request must carry dlCtx: retryablehttp.NewRequest builds
// on context.Background(), and its wait between attempts is a
// select on the request context, so a contextless request cannot
// be interrupted by --api-timeout (env API_TIMEOUT) or by the
// caller going away.
req, err := retryablehttp.NewRequestWithContext(dlCtx, http.MethodGet, dl.Url, nil)
if err != nil {
dlSpan.RecordError(err)
dlSpan.SetStatus(codes.Error, err.Error())
@@ -280,14 +367,28 @@ func newContext(echoCtx echo.Context, logger *slog.Logger, fs *gotenberg.FileSys
}
}
// Entries are serialized by the concurrency limit above, so a
// late one can start after the deadline has already passed.
// Fail closed rather than derive a non-positive timeout, which
// [http.Client] reads as no deadline at all.
remaining := time.Until(deadline)
if remaining <= 0 {
dlSpan.RecordError(context.DeadlineExceeded)
dlSpan.SetStatus(codes.Error, context.DeadlineExceeded.Error())
dlSpan.End()
return fmt.Errorf("download file from '%s': %w", dl.Url, context.DeadlineExceeded)
}
client := &retryablehttp.Client{
HTTPClient: gotenberg.NewOutboundHttpClient(time.Until(deadline), downloadFromCfg.allowList, downloadFromCfg.denyList, downloadFromCfg.enableEnvironmentProxy, ipOpts...),
HTTPClient: gotenberg.NewOutboundHttpClient(remaining, downloadFromCfg.allowList, downloadFromCfg.denyList, downloadFromCfg.enableEnvironmentProxy, ipOpts...),
RetryMax: downloadFromCfg.maxRetry,
RetryWaitMin: time.Duration(1) * time.Second,
RetryWaitMax: time.Until(deadline),
RetryWaitMax: remaining,
Logger: gotenberg.NewLeveledLogger(logger),
CheckRetry: retryablehttp.DefaultRetryPolicy,
Backoff: retryablehttp.DefaultBackoff,
// Not DefaultBackoff: it hands a hostile origin control of
// the wait via Retry-After.
Backoff: gotenberg.ClampedBackoff,
}
resp, err := client.Do(req)
@@ -295,6 +396,17 @@ func newContext(echoCtx echo.Context, logger *slog.Logger, fs *gotenberg.FileSys
dlSpan.RecordError(err)
dlSpan.SetStatus(codes.Error, err.Error())
dlSpan.End()
// A redirect target is filtered inside the client, so the
// policy verdict surfaces here rather than from the
// pre-flight above. Keep it out of the response: the first
// hop answers a filtered URL with a generic 403, and a
// later hop must not describe the allow-list, the deny-list
// or the IP policy instead.
if errors.Is(err, gotenberg.ErrFiltered) {
return fmt.Errorf("download file from '%s': %w", dl.Url, err)
}
return WrapError(
fmt.Errorf("download file from to '%s': %w", dl.Url, err),
NewSentinelHttpError(http.StatusBadRequest, fmt.Sprintf("Unable to download file from '%s': %s", dl.Url, err)),
@@ -368,7 +480,7 @@ func newContext(echoCtx echo.Context, logger *slog.Logger, fs *gotenberg.FileSys
// Use a UUID-based name on disk to avoid filesystem
// NAME_MAX limits with long filenames.
// See: https://github.com/gotenberg/gotenberg/issues/1500.
safeName := uuid.New().String() + filepath.Ext(filename)
safeName := uuid.New().String() + safeExt(filename)
path := fmt.Sprintf("%s/%s", ctx.dirPath, safeName)
out, err := os.Create(path)
@@ -422,18 +534,20 @@ func newContext(echoCtx echo.Context, logger *slog.Logger, fs *gotenberg.FileSys
}
for _, r := range results {
ctx.files[r.filename] = r.path
ctx.diskToOriginal[r.path] = r.filename
filename := ctx.uniqueFilename(r.filename)
ctx.files[filename] = r.path
ctx.diskToOriginal[r.path] = filename
ctx.trackFileOrder(r.path, r.filename)
if r.formField != "" {
ctx.filesByField[r.formField] = append(ctx.filesByField[r.formField], r.path)
}
}
}
copyToDisk := func(fh *multipart.FileHeader) error {
copyToDisk := func(fh *multipart.FileHeader) (string, error) {
in, err := fh.Open()
if err != nil {
return fmt.Errorf("open multipart file: %w", err)
return "", fmt.Errorf("open multipart file: %w", err)
}
defer func() {
@@ -456,12 +570,12 @@ func newContext(echoCtx echo.Context, logger *slog.Logger, fs *gotenberg.FileSys
// Use a UUID-based name on disk to avoid filesystem
// NAME_MAX limits with long filenames.
// See: https://github.com/gotenberg/gotenberg/issues/1500.
safeName := uuid.New().String() + filepath.Ext(filename)
safeName := uuid.New().String() + safeExt(filename)
path := fmt.Sprintf("%s/%s", ctx.dirPath, safeName)
out, err := os.Create(path)
if err != nil {
return fmt.Errorf("create local file: %w", err)
return "", fmt.Errorf("create local file: %w", err)
}
defer func() {
err := out.Close()
@@ -472,26 +586,28 @@ func newContext(echoCtx echo.Context, logger *slog.Logger, fs *gotenberg.FileSys
_, err = io.Copy(out, reader)
if err != nil {
return fmt.Errorf("copy multipart file to local file: %w", err)
return "", fmt.Errorf("copy multipart file to local file: %w", err)
}
base := filename
filename = ctx.uniqueFilename(filename)
ctx.files[filename] = path
ctx.diskToOriginal[path] = filename
ctx.trackFileOrder(path, base)
return nil
return filename, nil
}
// Then, copy the form files, if any.
for fieldName, files := range form.File {
for _, fh := range files {
err = copyToDisk(fh)
if err != nil {
return ctx, cancel, fmt.Errorf("copy to disk: %w", err)
filename, errCopy := copyToDisk(fh)
if errCopy != nil {
return ctx, cancel, fmt.Errorf("copy to disk: %w", errCopy)
}
// Track files by field name
filename := sanitizeFilename(fh.Filename)
filePath := ctx.files[filename]
ctx.filesByField[fieldName] = append(ctx.filesByField[fieldName], filePath)
// Track files by field name, under the name copyToDisk actually
// stored, which may be a de-duplicated variant.
ctx.filesByField[fieldName] = append(ctx.filesByField[fieldName], ctx.files[filename])
}
}
@@ -506,9 +622,9 @@ func newContext(echoCtx echo.Context, logger *slog.Logger, fs *gotenberg.FileSys
if symlinkPath == diskPath {
continue
}
err = os.Symlink(filepath.Base(diskPath), symlinkPath)
if err != nil {
logger.DebugContext(context.Background(), fmt.Sprintf("skip symlink for '%s': %s", originalName, err))
errSymlink := os.Symlink(filepath.Base(diskPath), symlinkPath)
if errSymlink != nil {
logger.DebugContext(context.Background(), fmt.Sprintf("skip symlink for '%s': %s", originalName, errSymlink))
}
}
@@ -517,7 +633,10 @@ func newContext(echoCtx echo.Context, logger *slog.Logger, fs *gotenberg.FileSys
ctx.Log().DebugContext(ctx, fmt.Sprintf("form files by field: %+v", ctx.filesByField))
ctx.Log().DebugContext(ctx, fmt.Sprintf("total bytes: %d", totalBytesRead.Load()))
return ctx, cancel, err
// Explicitly nil: the best-effort symlink loop above must not decide the
// outcome of the request. Its failure used to escape here as a bare 500,
// non-deterministically, because ctx.files iterates in random order.
return ctx, cancel, nil
}
// Request returns the [http.Request].
@@ -532,6 +651,8 @@ func (ctx *Context) FormData() *FormData {
files: ctx.files,
filesByField: ctx.filesByField,
diskToOriginal: ctx.diskToOriginal,
fileOrder: ctx.fileOrder,
fileBase: ctx.fileBase,
errors: nil,
}
}
@@ -572,7 +693,7 @@ func (ctx *Context) GeneratePath(extension string) string {
// limits but registers the given filename so that [Context.OriginalFilename]
// can resolve it. It does not create a file.
func (ctx *Context) GeneratePathFromFilename(filename string) string {
safeName := uuid.New().String() + filepath.Ext(filename)
safeName := uuid.New().String() + safeExt(filename)
path := fmt.Sprintf("%s/%s", ctx.dirPath, safeName)
ctx.diskToOriginal[path] = filename
return path
@@ -675,15 +796,72 @@ func (ctx *Context) BuildOutputFile() (string, error) {
// OutputFilename returns the filename based on the given output path or the
// "Gotenberg-Output-Filename" header's value.
func (ctx *Context) OutputFilename(outputPath string) string {
filename := ctx.echoCtx.Get("outputFilename").(string)
if filename == "" {
if ctx.outputFilename == "" {
return ctx.OriginalFilename(outputPath)
}
filename := ctx.outputFilename
return fmt.Sprintf("%s%s", filename, filepath.Ext(outputPath))
}
// maxDiskExtLength bounds the extension copied onto a UUID-based disk name.
// The UUID stem is 36 characters, so a longer extension risks NAME_MAX, which
// is 255 on ext4 and overlayfs. The untruncated name is kept in
// [Context.diskToOriginal], which never reaches the filesystem.
const maxDiskExtLength = 32
// safeExt returns the extension to append to a UUID-based disk name. It drops
// an extension too long to be safe rather than let [os.Create] fail with
// ENAMETOOLONG, which surfaced to the caller as a bare 500.
func safeExt(filename string) string {
ext := filepath.Ext(filename)
if len(ext) > maxDiskExtLength {
return ""
}
return ext
}
// uniqueFilename returns filename, or a numbered variant of it when the
// request already carries a file by that name.
//
// Uploads are keyed by their sanitized original filename, so two files sharing
// one name used to collide: the second overwrote the first and only one
// reached the conversion, while both stayed on disk and counted against the
// body limit. Sanitizing strips directories, so "a/doc.pdf" and "b/doc.pdf"
// collide too.
func (ctx *Context) uniqueFilename(filename string) string {
_, exists := ctx.files[filename]
if !exists {
return filename
}
ext := filepath.Ext(filename)
stem := strings.TrimSuffix(filename, ext)
for i := 2; ; i++ {
candidate := fmt.Sprintf("%s (%d)%s", stem, i, ext)
_, exists = ctx.files[candidate]
if !exists {
return candidate
}
}
}
// trackFileOrder records where a file arrived in the request and the filename
// it arrived under, so [FormData.paths] can order it the way the caller sent
// it.
func (ctx *Context) trackFileOrder(path, base string) {
if ctx.fileOrder == nil {
ctx.fileOrder = make(map[string]int)
ctx.fileBase = make(map[string]string)
}
ctx.fileOrder[path] = len(ctx.fileOrder)
ctx.fileBase[path] = base
}
// sanitizeFilename strips path separators (including backslashes, which
// [filepath.Base] ignores on Linux) and control characters from a
// caller-supplied filename, then NFC-normalizes the result. This prevents a

View File

@@ -4,15 +4,22 @@ import (
"bytes"
"context"
"encoding/json"
"errors"
"fmt"
"log/slog"
"mime/multipart"
"net/http"
"net/http/httptest"
"os"
"regexp"
"runtime"
"strings"
"sync"
"sync/atomic"
"testing"
"time"
"github.com/dlclark/regexp2"
"github.com/labstack/echo/v4"
"github.com/gotenberg/gotenberg/v8/pkg/gotenberg"
@@ -73,6 +80,64 @@ func TestNewContext_Cancellation(t *testing.T) {
}
}
func TestNewContext_RemovesMultipartTemporaryFiles(t *testing.T) {
body := new(bytes.Buffer)
writer := multipart.NewWriter(body)
part, err := writer.CreateFormFile("files", "input.odt")
if err != nil {
t.Fatalf("create multipart file: %v", err)
}
_, err = part.Write(bytes.Repeat([]byte("x"), 1024))
if err != nil {
t.Fatalf("write multipart file: %v", err)
}
err = writer.Close()
if err != nil {
t.Fatalf("close multipart writer: %v", err)
}
req := httptest.NewRequest(http.MethodPost, "/forms/libreoffice/convert", body)
req.Header.Set("Content-Type", writer.FormDataContentType())
err = req.ParseMultipartForm(1)
if err != nil {
t.Fatalf("parse multipart form: %v", err)
}
defer func() {
_ = req.MultipartForm.RemoveAll()
}()
upload, err := req.MultipartForm.File["files"][0].Open()
if err != nil {
t.Fatalf("open disk-backed multipart file: %v", err)
}
temporaryFile, ok := upload.(*os.File)
if !ok {
_ = upload.Close()
t.Fatal("multipart upload is not disk-backed")
}
temporaryPath := temporaryFile.Name()
err = temporaryFile.Close()
if err != nil {
t.Fatalf("close disk-backed multipart file: %v", err)
}
echoCtx := echo.New().NewContext(req, httptest.NewRecorder())
logger := slog.New(slog.DiscardHandler)
fs := gotenberg.NewFileSystem(new(gotenberg.OsMkdirAll))
downloadFromCfg := downloadFromConfig{disable: true}
_, cancel, err := newContext(echoCtx, logger, fs, 10*time.Second, 0, downloadFromCfg)
if err != nil {
t.Fatalf("newContext returned error: %v", err)
}
defer cancel()
_, err = os.Stat(temporaryPath)
if !os.IsNotExist(err) {
t.Fatalf("multipart temporary file still exists: %s", temporaryPath)
}
}
// Concurrent downloadFrom entries must not race on the shared maps
// (ctx.files, ctx.diskToOriginal, ctx.filesByField). Run under -race
// to catch the data race; without -race a sufficient number of entries
@@ -150,6 +215,137 @@ func TestNewContext_DownloadFromConcurrentMapWrites(t *testing.T) {
}
}
// An oversized downloadFrom array must be rejected at the trust boundary with
// a 400, before any download goroutine is spawned.
// https://github.com/gotenberg/gotenberg/security/advisories/GHSA-6vqw-2jgm-4x88
func TestNewContext_DownloadFromMaxEntries(t *testing.T) {
var hits atomic.Int64
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
hits.Add(1)
w.Header().Set("Content-Disposition", `attachment; filename="download.txt"`)
_, _ = w.Write([]byte("downloaded"))
}))
defer server.Close()
dls := make([]downloadFrom, 3)
for i := range dls {
dls[i] = downloadFrom{Url: fmt.Sprintf("%s/file?i=%d", server.URL, i)}
}
payload, err := json.Marshal(dls)
if err != nil {
t.Fatalf("marshal downloadFrom payload: %v", err)
}
body := new(bytes.Buffer)
writer := multipart.NewWriter(body)
err = writer.WriteField("downloadFrom", string(payload))
if err != nil {
t.Fatalf("write downloadFrom field: %v", err)
}
err = writer.Close()
if err != nil {
t.Fatalf("close multipart writer: %v", err)
}
req := httptest.NewRequest(http.MethodPost, "/forms/libreoffice/convert", body)
req.Header.Set("Content-Type", writer.FormDataContentType())
echoCtx := echo.New().NewContext(req, httptest.NewRecorder())
logger := slog.New(slog.DiscardHandler)
fs := gotenberg.NewFileSystem(new(gotenberg.OsMkdirAll))
downloadFromCfg := downloadFromConfig{maxEntries: 2}
_, cancel, err := newContext(echoCtx, logger, fs, 10*time.Second, 0, downloadFromCfg)
if cancel != nil {
defer cancel()
}
if err == nil {
t.Fatal("newContext returned no error, want a 400 for too many entries")
}
var httpErr HttpError
if !errors.As(err, &httpErr) {
t.Fatalf("error %v is not an HttpError", err)
}
if status, _ := httpErr.HttpError(); status != http.StatusBadRequest {
t.Fatalf("HTTP status = %d, want %d", status, http.StatusBadRequest)
}
if got := hits.Load(); got != 0 {
t.Fatalf("server hits = %d, want 0 (rejected before any download)", got)
}
}
// The number of in-flight downloadFrom fetches must never exceed the
// configured concurrency limit, regardless of array length.
// https://github.com/gotenberg/gotenberg/security/advisories/GHSA-6vqw-2jgm-4x88
func TestNewContext_DownloadFromMaxConcurrency(t *testing.T) {
const (
downloads = 8
maxConcurrency = 2
)
var current, peak atomic.Int64
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
inFlight := current.Add(1)
for {
observed := peak.Load()
if inFlight <= observed || peak.CompareAndSwap(observed, inFlight) {
break
}
}
time.Sleep(20 * time.Millisecond)
current.Add(-1)
filename := fmt.Sprintf("download-%s.txt", r.URL.Query().Get("i"))
w.Header().Set("Content-Disposition", fmt.Sprintf(`attachment; filename="%s"`, filename))
_, _ = w.Write([]byte("downloaded"))
}))
defer server.Close()
dls := make([]downloadFrom, downloads)
for i := range dls {
dls[i] = downloadFrom{Url: fmt.Sprintf("%s/file?i=%d", server.URL, i)}
}
payload, err := json.Marshal(dls)
if err != nil {
t.Fatalf("marshal downloadFrom payload: %v", err)
}
body := new(bytes.Buffer)
writer := multipart.NewWriter(body)
err = writer.WriteField("downloadFrom", string(payload))
if err != nil {
t.Fatalf("write downloadFrom field: %v", err)
}
err = writer.Close()
if err != nil {
t.Fatalf("close multipart writer: %v", err)
}
req := httptest.NewRequest(http.MethodPost, "/forms/libreoffice/convert", body)
req.Header.Set("Content-Type", writer.FormDataContentType())
echoCtx := echo.New().NewContext(req, httptest.NewRecorder())
logger := slog.New(slog.DiscardHandler)
fs := gotenberg.NewFileSystem(new(gotenberg.OsMkdirAll))
downloadFromCfg := downloadFromConfig{maxConcurrency: maxConcurrency}
ctx, cancel, err := newContext(echoCtx, logger, fs, 10*time.Second, 0, downloadFromCfg)
if err != nil {
t.Fatalf("newContext returned error: %v", err)
}
defer cancel()
if got := len(ctx.files); got != downloads {
t.Fatalf("downloaded files = %d, want %d", got, downloads)
}
if got := peak.Load(); got > maxConcurrency {
t.Fatalf("peak concurrency = %d, want <= %d", got, maxConcurrency)
}
}
func TestSanitizeFilename(t *testing.T) {
for _, tc := range []struct {
scenario string
@@ -222,3 +418,487 @@ func TestContext_FileCount(t *testing.T) {
t.Errorf("expected 3 files, got %d", got)
}
}
// A hostile origin must not choose how long Gotenberg waits.
// [retryablehttp.DefaultBackoff] returns a Retry-After header verbatim for 429
// and 503, and the wait between attempts is a select on the request context.
// Building the request without a context therefore pinned the goroutine, its
// connection, and its working directory for the attacker's chosen duration,
// well past --api-timeout (env API_TIMEOUT).
func TestNewContext_DownloadFromHostileRetryAfterIsBounded(t *testing.T) {
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Retry-After", "3600")
w.WriteHeader(http.StatusTooManyRequests)
}))
defer server.Close()
payload, err := json.Marshal([]downloadFrom{{Url: server.URL + "/file"}})
if err != nil {
t.Fatalf("marshal downloadFrom payload: %v", err)
}
body := new(bytes.Buffer)
writer := multipart.NewWriter(body)
err = writer.WriteField("downloadFrom", string(payload))
if err != nil {
t.Fatalf("write downloadFrom field: %v", err)
}
err = writer.Close()
if err != nil {
t.Fatalf("close multipart writer: %v", err)
}
req := httptest.NewRequest(http.MethodPost, "/forms/libreoffice/convert", body)
req.Header.Set("Content-Type", writer.FormDataContentType())
echoCtx := echo.New().NewContext(req, httptest.NewRecorder())
logger := slog.New(slog.DiscardHandler)
fs := gotenberg.NewFileSystem(new(gotenberg.OsMkdirAll))
const timeout = 500 * time.Millisecond
start := time.Now()
_, cancel, err := newContext(echoCtx, logger, fs, timeout, 0, downloadFromConfig{maxRetry: 2})
elapsed := time.Since(start)
if cancel != nil {
defer cancel()
}
if err == nil {
t.Fatal("expected newContext to fail against an origin that only answers 429")
}
// Generous: the deadline is 500ms and Retry-After asks for an hour. Any
// value in seconds means the remote is still in control.
if elapsed > 10*time.Second {
t.Fatalf("newContext took %s with Retry-After 3600; --api-timeout must bound it", elapsed)
}
}
// An entry that starts after the deadline has passed must fail closed. It used
// to derive a negative client timeout, which [http.Client] reads as no
// deadline at all, leaving the download unbounded.
func TestNewContext_DownloadFromExpiredBudgetFailsClosed(t *testing.T) {
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
<-r.Context().Done()
}))
defer server.Close()
// Two entries, serialized by the concurrency limit, so the second one
// starts once the first has burned the whole budget.
payload, err := json.Marshal([]downloadFrom{
{Url: server.URL + "/first"},
{Url: server.URL + "/second"},
})
if err != nil {
t.Fatalf("marshal downloadFrom payload: %v", err)
}
body := new(bytes.Buffer)
writer := multipart.NewWriter(body)
err = writer.WriteField("downloadFrom", string(payload))
if err != nil {
t.Fatalf("write downloadFrom field: %v", err)
}
err = writer.Close()
if err != nil {
t.Fatalf("close multipart writer: %v", err)
}
req := httptest.NewRequest(http.MethodPost, "/forms/libreoffice/convert", body)
req.Header.Set("Content-Type", writer.FormDataContentType())
echoCtx := echo.New().NewContext(req, httptest.NewRecorder())
logger := slog.New(slog.DiscardHandler)
fs := gotenberg.NewFileSystem(new(gotenberg.OsMkdirAll))
done := make(chan error, 1)
go func() {
_, cancel, err := newContext(echoCtx, logger, fs, 400*time.Millisecond, 0, downloadFromConfig{
maxRetry: 0,
maxConcurrency: 1,
})
if cancel != nil {
cancel()
}
done <- err
}()
select {
case err := <-done:
if err == nil {
t.Fatal("expected newContext to fail against a stalling origin")
}
case <-time.After(15 * time.Second):
t.Fatal("newContext never returned: an entry starting past the deadline built an unbounded client")
}
}
func TestDecodeDownloadFrom(t *testing.T) {
for _, tc := range []struct {
scenario string
raw string
maxEntries int
expectErr error
expectLen int
}{
{"empty array", `[]`, 10, nil, 0},
{"under the limit", `[{"url":"http://a"},{"url":"http://b"}]`, 10, nil, 2},
{"exactly the limit", `[{"url":"http://a"},{"url":"http://b"}]`, 2, nil, 2},
{"over the limit", `[{"url":"http://a"},{"url":"http://b"}]`, 1, errTooManyDownloadFromEntries, 0},
{"no limit", `[{"url":"http://a"},{"url":"http://b"}]`, 0, nil, 2},
{"not an array", `{"url":"http://a"}`, 10, nil, 0},
{"malformed", `[{"url":`, 10, nil, 0},
{"not json", `nope`, 10, nil, 0},
} {
t.Run(tc.scenario, func(t *testing.T) {
dls, err := decodeDownloadFrom(tc.raw, tc.maxEntries)
if tc.expectErr != nil {
if !errors.Is(err, tc.expectErr) {
t.Fatalf("error = %v, want %v", err, tc.expectErr)
}
return
}
if tc.scenario == "not an array" || tc.scenario == "malformed" || tc.scenario == "not json" {
if err == nil {
t.Fatalf("expected an error for %q", tc.raw)
}
return
}
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
if len(dls) != tc.expectLen {
t.Fatalf("decoded %d entries, want %d", len(dls), tc.expectLen)
}
})
}
}
// A compact array costs three bytes per entry on the wire and expands by
// roughly seventy times once unmarshalled. Decoding must stop at the limit
// rather than materialize the whole array and count afterwards.
func TestDecodeDownloadFrom_StopsBeforeMaterializingTheArray(t *testing.T) {
const entries = 2_000_000
raw := "[" + strings.Repeat("{},", entries) + "{}]"
var before, after runtime.MemStats
runtime.GC()
runtime.ReadMemStats(&before)
_, err := decodeDownloadFrom(raw, 1000)
runtime.ReadMemStats(&after)
if !errors.Is(err, errTooManyDownloadFromEntries) {
t.Fatalf("error = %v, want errTooManyDownloadFromEntries", err)
}
// json.Unmarshal on the same input allocates hundreds of MiB. Bounded
// decoding should stay in the low single-digit MiB, so this threshold is
// deliberately loose and still fails loudly on a regression.
allocated := after.TotalAlloc - before.TotalAlloc
if allocated > 32<<20 {
t.Fatalf("decoding allocated %d MiB for a %d-entry array, want the limit to bound it", allocated>>20, entries)
}
t.Logf("allocated %d KiB decoding a %d-entry array with a limit of 1000", allocated>>10, entries)
}
// An asynchronous conversion outlives the [echo.Context]. Echo returns that
// context to a sync.Pool as soon as the handler returns, and
// outputFilenameMiddleware runs in srv.Pre on every request, including
// /health, so a later request overwrites the store. Reading the output
// filename from it after the fact returned another caller's value.
func TestContext_OutputFilename_SurvivesEchoContextRecycling(t *testing.T) {
body := new(bytes.Buffer)
writer := multipart.NewWriter(body)
err := writer.Close()
if err != nil {
t.Fatalf("close multipart writer: %v", err)
}
req := httptest.NewRequest(http.MethodPost, "/forms/libreoffice/convert", body)
req.Header.Set("Content-Type", writer.FormDataContentType())
echoCtx := echo.New().NewContext(req, httptest.NewRecorder())
// What outputFilenameMiddleware does for this request.
echoCtx.Set("outputFilename", "victim")
logger := slog.New(slog.DiscardHandler)
fs := gotenberg.NewFileSystem(new(gotenberg.OsMkdirAll))
ctx, cancel, err := newContext(echoCtx, logger, fs, 10*time.Second, 0, downloadFromConfig{disable: true})
if err != nil {
t.Fatalf("newContext returned error: %v", err)
}
defer cancel()
// Echo recycles the context and another request claims the store.
echoCtx.Set("outputFilename", "attacker-controlled")
if got := ctx.OutputFilename("/tmp/out.pdf"); got != "victim.pdf" {
t.Fatalf("OutputFilename = %q, want %q", got, "victim.pdf")
}
}
// A recycled context has a nil store, so the previous unguarded type assertion
// could panic. The snapshot must tolerate an absent value.
func TestContext_OutputFilename_NoHeader(t *testing.T) {
body := new(bytes.Buffer)
writer := multipart.NewWriter(body)
err := writer.Close()
if err != nil {
t.Fatalf("close multipart writer: %v", err)
}
req := httptest.NewRequest(http.MethodPost, "/forms/libreoffice/convert", body)
req.Header.Set("Content-Type", writer.FormDataContentType())
// No Set call at all: the store holds nothing for "outputFilename".
echoCtx := echo.New().NewContext(req, httptest.NewRecorder())
logger := slog.New(slog.DiscardHandler)
fs := gotenberg.NewFileSystem(new(gotenberg.OsMkdirAll))
ctx, cancel, err := newContext(echoCtx, logger, fs, 10*time.Second, 0, downloadFromConfig{disable: true})
if err != nil {
t.Fatalf("newContext returned error: %v", err)
}
defer cancel()
if got := ctx.OutputFilename("/tmp/out.pdf"); got != "out.pdf" {
t.Fatalf("OutputFilename = %q, want the original filename %q", got, "out.pdf")
}
}
func TestSafeExt(t *testing.T) {
for _, tc := range []struct {
scenario string
filename string
want string
}{
{"ordinary extension", "report.pdf", ".pdf"},
{"no extension", "report", ""},
{"at the limit", "a." + strings.Repeat("x", maxDiskExtLength-1), "." + strings.Repeat("x", maxDiskExtLength-1)},
{"over the limit is dropped", "a." + strings.Repeat("x", 300), ""},
} {
t.Run(tc.scenario, func(t *testing.T) {
got := safeExt(tc.filename)
if got != tc.want {
t.Fatalf("safeExt(%q) = %q, want %q", tc.filename, got, tc.want)
}
// A UUID stem is 36 characters. The whole disk name must stay
// under NAME_MAX.
if len(got)+36 > 255 {
t.Fatalf("disk name would be %d characters, over NAME_MAX", len(got)+36)
}
})
}
}
// An upload whose extension exceeds NAME_MAX used to fail os.Create and return
// a bare 500. The extension is bounded, and the original name survives in
// diskToOriginal.
func TestNewContext_LongExtensionIsAccepted(t *testing.T) {
filename := "invoice." + strings.Repeat("x", 300)
body := new(bytes.Buffer)
writer := multipart.NewWriter(body)
part, err := writer.CreateFormFile("files", filename)
if err != nil {
t.Fatalf("create multipart file: %v", err)
}
_, err = part.Write([]byte("%PDF-1.4"))
if err != nil {
t.Fatalf("write multipart file: %v", err)
}
err = writer.Close()
if err != nil {
t.Fatalf("close multipart writer: %v", err)
}
req := httptest.NewRequest(http.MethodPost, "/forms/libreoffice/convert", body)
req.Header.Set("Content-Type", writer.FormDataContentType())
echoCtx := echo.New().NewContext(req, httptest.NewRecorder())
logger := slog.New(slog.DiscardHandler)
fs := gotenberg.NewFileSystem(new(gotenberg.OsMkdirAll))
ctx, cancel, err := newContext(echoCtx, logger, fs, 10*time.Second, 0, downloadFromConfig{disable: true})
if cancel != nil {
defer cancel()
}
if err != nil {
t.Fatalf("newContext returned error for a long extension: %v", err)
}
if got := len(ctx.files); got != 1 {
t.Fatalf("files = %d, want 1", got)
}
}
// A filename that cannot become a symlink (too long, "..", "/") must not fail
// the request. The symlink loop is best-effort, but its error escaped through
// the shared err variable, and ctx.files iterates randomly, so byte-identical
// requests gave different HTTP outcomes.
func TestNewContext_UnsymlinkableFilenameStillSucceeds(t *testing.T) {
for _, filename := range []string{
strings.Repeat("a", 300) + ".txt",
"..",
"/",
} {
t.Run(filename[:min(len(filename), 12)], func(t *testing.T) {
body := new(bytes.Buffer)
writer := multipart.NewWriter(body)
part, err := writer.CreateFormFile("files", filename)
if err != nil {
t.Fatalf("create multipart file: %v", err)
}
_, err = part.Write([]byte("%PDF-1.4"))
if err != nil {
t.Fatalf("write multipart file: %v", err)
}
err = writer.Close()
if err != nil {
t.Fatalf("close multipart writer: %v", err)
}
req := httptest.NewRequest(http.MethodPost, "/forms/libreoffice/convert", body)
req.Header.Set("Content-Type", writer.FormDataContentType())
echoCtx := echo.New().NewContext(req, httptest.NewRecorder())
logger := slog.New(slog.DiscardHandler)
fs := gotenberg.NewFileSystem(new(gotenberg.OsMkdirAll))
_, cancel, err := newContext(echoCtx, logger, fs, 10*time.Second, 0, downloadFromConfig{disable: true})
if cancel != nil {
defer cancel()
}
if err != nil {
t.Fatalf("newContext failed on a best-effort symlink for %q: %v", filename, err)
}
})
}
}
// Two uploads sharing a filename must both reach the conversion. The second
// used to overwrite the first in ctx.files, so one file was silently dropped
// while both stayed on disk and counted against the body limit.
func TestNewContext_DuplicateFilenamesAreBothKept(t *testing.T) {
body := new(bytes.Buffer)
writer := multipart.NewWriter(body)
for _, content := range []string{"FIRST", "SECOND"} {
part, err := writer.CreateFormFile("files", "doc.pdf")
if err != nil {
t.Fatalf("create multipart file: %v", err)
}
_, err = part.Write([]byte(content))
if err != nil {
t.Fatalf("write multipart file: %v", err)
}
}
err := writer.Close()
if err != nil {
t.Fatalf("close multipart writer: %v", err)
}
req := httptest.NewRequest(http.MethodPost, "/forms/pdfengines/merge", body)
req.Header.Set("Content-Type", writer.FormDataContentType())
echoCtx := echo.New().NewContext(req, httptest.NewRecorder())
logger := slog.New(slog.DiscardHandler)
fs := gotenberg.NewFileSystem(new(gotenberg.OsMkdirAll))
ctx, cancel, err := newContext(echoCtx, logger, fs, 10*time.Second, 0, downloadFromConfig{disable: true})
if err != nil {
t.Fatalf("newContext returned error: %v", err)
}
defer cancel()
if got := len(ctx.files); got != 2 {
t.Fatalf("ctx.files = %d entries, want 2: a duplicate filename dropped a file", got)
}
if got := len(ctx.filesByField["files"]); got != 2 {
t.Fatalf("filesByField[files] = %d entries, want 2", got)
}
// The two maps must agree, and both files must be distinct on disk.
seen := make(map[string]struct{})
for _, path := range ctx.files {
if _, ok := ctx.diskToOriginal[path]; !ok {
t.Fatalf("path %q has no diskToOriginal entry", path)
}
seen[path] = struct{}{}
}
if len(seen) != 2 {
t.Fatalf("distinct disk paths = %d, want 2", len(seen))
}
}
// A redirect target is filtered inside the HTTP client, so the policy verdict
// surfaces from client.Do rather than from the pre-flight check. It used to be
// interpolated into the response body, so a redirect described the allow-list,
// the deny-list or the IP policy where the first hop returns a generic 403.
func TestNewContext_DownloadFromRedirectVerdictStaysGeneric(t *testing.T) {
private := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Disposition", `attachment; filename="secret.txt"`)
_, _ = w.Write([]byte("internal"))
}))
defer private.Close()
redirector := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.Redirect(w, r, private.URL+"/secret", http.StatusFound)
}))
defer redirector.Close()
payload, err := json.Marshal([]downloadFrom{{Url: redirector.URL + "/start"}})
if err != nil {
t.Fatalf("marshal downloadFrom payload: %v", err)
}
body := new(bytes.Buffer)
writer := multipart.NewWriter(body)
err = writer.WriteField("downloadFrom", string(payload))
if err != nil {
t.Fatalf("write downloadFrom field: %v", err)
}
err = writer.Close()
if err != nil {
t.Fatalf("close multipart writer: %v", err)
}
req := httptest.NewRequest(http.MethodPost, "/forms/libreoffice/convert", body)
req.Header.Set("Content-Type", writer.FormDataContentType())
echoCtx := echo.New().NewContext(req, httptest.NewRecorder())
logger := slog.New(slog.DiscardHandler)
fs := gotenberg.NewFileSystem(new(gotenberg.OsMkdirAll))
// The first hop is allowed, the redirect target is denied by the deny-list.
denyList := []*regexp2.Regexp{regexp2.MustCompile("^"+regexp.QuoteMeta(private.URL), 0)}
_, cancel, err := newContext(echoCtx, logger, fs, 10*time.Second, 0, downloadFromConfig{
denyList: denyList,
maxRetry: 0,
})
if cancel != nil {
defer cancel()
}
if err == nil {
t.Fatal("expected the redirect to a denied host to fail")
}
status, message := ParseError(err)
if status != http.StatusForbidden {
t.Fatalf("status = %d, want %d: a filtered redirect must answer like a filtered first hop", status, http.StatusForbidden)
}
if message != http.StatusText(http.StatusForbidden) {
t.Fatalf("message = %q, want the generic %q", message, http.StatusText(http.StatusForbidden))
}
// The response must not name the policy, the pattern, or the blocked host.
for _, leak := range []string{"denied list", "allowed list", "non-public", "expression", private.URL} {
if strings.Contains(message, leak) {
t.Fatalf("response message %q leaks %q", message, leak)
}
}
}

View File

@@ -41,6 +41,8 @@ type FormData struct {
files map[string]string
filesByField map[string][]string
diskToOriginal map[string]string
fileOrder map[string]int
fileBase map[string]string
errors error
}
@@ -479,6 +481,51 @@ func (form *FormData) Stamp(target *string) *FormData {
return form
}
// Stamps binds the absolute paths of every file uploaded with the "stamp"
// field name, in submission order. Unlike [FormData.Stamp], it keeps all of
// them so a route can apply several stamps in a single request.
func (form *FormData) Stamps(target *[]string) *FormData {
if form.errors != nil {
return form
}
if paths, ok := form.filesByField[StampFormField]; ok {
*target = paths
}
return form
}
// Watermarks binds the absolute paths of every file uploaded with the
// "watermark" field name, in submission order. Unlike [FormData.Watermark], it
// keeps all of them so a route can apply several watermarks in a single request.
func (form *FormData) Watermarks(target *[]string) *FormData {
if form.errors != nil {
return form
}
if paths, ok := form.filesByField[WatermarkFormField]; ok {
*target = paths
}
return form
}
// Strings binds every value submitted for key, in submission order. A field
// repeated in the multipart body (e.g. multiple "stampSource") contributes one
// entry per occurrence, which lets a route read parallel field arrays.
func (form *FormData) Strings(key string, target *[]string) *FormData {
if form.errors != nil {
return form
}
if values, ok := form.values[key]; ok {
*target = values
}
return form
}
// FacturXXml binds the absolute path of the uploaded Factur-X CII invoice
// XML. Only a file uploaded with the "facturxXml" field name is included.
func (form *FormData) FacturXXml(target *string) *FormData {
@@ -537,24 +584,42 @@ func (form *FormData) paths(extensions []string, target *[]string) *FormData {
}
// See https://github.com/gotenberg/gotenberg/issues/139.
originals := make(gotenberg.AlphanumericSort, len(entries))
for i, e := range entries {
originals[i] = e.original
}
sort.Sort(originals)
//
// Sort on the filename as received rather than on the map key. The key
// carries the suffix uniqueFilename adds when two uploads share a name,
// and that suffix would otherwise decide the order: "doc (2).pdf" sorts
// before "doc.pdf". Ordering on the received name keeps the pair adjacent,
// and the arrival index breaks the tie, so duplicates merge in the order
// the caller sent them. A file with a unique name is unaffected, since its
// received name and its key are the same string.
sort.SliceStable(entries, func(i, j int) bool {
nameI := form.receivedName(entries[i].disk, entries[i].original)
nameJ := form.receivedName(entries[j].disk, entries[j].original)
if nameI != nameJ {
return gotenberg.AlphanumericSort{nameI, nameJ}.Less(0, 1)
}
return form.fileOrder[entries[i].disk] < form.fileOrder[entries[j].disk]
})
// Build a lookup from original name to disk path.
lookup := make(map[string]string, len(entries))
for _, e := range entries {
lookup[e.original] = e.disk
}
for _, o := range originals {
*target = append(*target, lookup[o])
*target = append(*target, e.disk)
}
return form
}
// receivedName returns the filename the file at disk arrived under, before
// de-duplication, falling back to fallback.
func (form *FormData) receivedName(disk, fallback string) string {
base, ok := form.fileBase[disk]
if ok {
return base
}
return fallback
}
// append adds an error to the list of errors.
func (form *FormData) append(err error) {
form.errors = errors.Join(form.errors, err)
@@ -567,7 +632,7 @@ func (form *FormData) mustValue(key string, target any, defaultValue any) *FormD
val, ok := form.values[key]
if !ok || val[0] == "" {
switch t := (target).(type) {
switch t := target.(type) {
case *string:
*t = defaultValue.(string)
case *bool:
@@ -614,7 +679,7 @@ func (form *FormData) mustMandatoryField(key string, target any) *FormData {
func (form *FormData) mustAssign(key, value string, target any) *FormData {
var err error
switch t := (target).(type) {
switch t := target.(type) {
case *string:
*t = value
case *bool:

View File

@@ -1837,3 +1837,139 @@ func TestFormData_paths_excludesFacturXXml(t *testing.T) {
t.Errorf("expected only the non-Factur-X .xml document, got %+v", paths)
}
}
func TestFormData_Strings(t *testing.T) {
form := &FormData{
values: map[string][]string{
"foo": {"a", "b", "c"},
},
}
var got []string
form.Strings("foo", &got)
if want := []string{"a", "b", "c"}; !reflect.DeepEqual(got, want) {
t.Errorf("expected %+v, got %+v", want, got)
}
var missing []string
form.Strings("bar", &missing)
if missing != nil {
t.Errorf("expected nil for a missing key, got %+v", missing)
}
}
func TestFormData_Stamps(t *testing.T) {
form := &FormData{
filesByField: map[string][]string{
StampFormField: {"/tmp/abc/a.png", "/tmp/abc/b.pdf"},
},
}
var got []string
form.Stamps(&got)
if want := []string{"/tmp/abc/a.png", "/tmp/abc/b.pdf"}; !reflect.DeepEqual(got, want) {
t.Errorf("expected %+v, got %+v", want, got)
}
empty := &FormData{}
var none []string
empty.Stamps(&none)
if none != nil {
t.Errorf("expected nil when no stamp file was uploaded, got %+v", none)
}
}
func TestFormData_Watermarks(t *testing.T) {
form := &FormData{
filesByField: map[string][]string{
WatermarkFormField: {"/tmp/abc/a.png"},
},
}
var got []string
form.Watermarks(&got)
if want := []string{"/tmp/abc/a.png"}; !reflect.DeepEqual(got, want) {
t.Errorf("expected %+v, got %+v", want, got)
}
}
// De-duplicating a repeated filename must not change merge order. Files with
// unique names keep exactly the order they had before de-duplication existed,
// and two files sharing a name merge in the order the caller sent them.
func TestFormData_paths_DuplicateFilenamesKeepUploadOrder(t *testing.T) {
for _, tc := range []struct {
scenario string
files map[string]string
fileBase map[string]string
order map[string]int
want []string
}{
{
scenario: "unique names sort exactly as before",
files: map[string]string{"b.pdf": "/w/2", "a.pdf": "/w/1", "c.pdf": "/w/3"},
fileBase: map[string]string{"/w/1": "a.pdf", "/w/2": "b.pdf", "/w/3": "c.pdf"},
order: map[string]int{"/w/1": 0, "/w/2": 1, "/w/3": 2},
want: []string{"/w/1", "/w/2", "/w/3"},
},
{
scenario: "numeric prefixes still win",
files: map[string]string{"10_x.pdf": "/w/3", "2_x.pdf": "/w/2", "1_x.pdf": "/w/1"},
fileBase: map[string]string{"/w/1": "1_x.pdf", "/w/2": "2_x.pdf", "/w/3": "10_x.pdf"},
order: map[string]int{"/w/1": 0, "/w/2": 1, "/w/3": 2},
want: []string{"/w/1", "/w/2", "/w/3"},
},
{
scenario: "duplicates merge in upload order, not suffix order",
files: map[string]string{"doc.pdf": "/w/1", "doc (2).pdf": "/w/2"},
fileBase: map[string]string{"/w/1": "doc.pdf", "/w/2": "doc.pdf"},
order: map[string]int{"/w/1": 0, "/w/2": 1},
want: []string{"/w/1", "/w/2"},
},
{
scenario: "duplicates stay adjacent and in position",
files: map[string]string{
"a.pdf": "/w/1", "doc.pdf": "/w/2", "doc (2).pdf": "/w/3", "z.pdf": "/w/4",
},
fileBase: map[string]string{
"/w/1": "a.pdf", "/w/2": "doc.pdf", "/w/3": "doc.pdf", "/w/4": "z.pdf",
},
order: map[string]int{"/w/1": 0, "/w/2": 1, "/w/3": 2, "/w/4": 3},
want: []string{"/w/1", "/w/2", "/w/3", "/w/4"},
},
{
scenario: "three copies keep their order",
files: map[string]string{"r.pdf": "/w/1", "r (2).pdf": "/w/2", "r (3).pdf": "/w/3"},
fileBase: map[string]string{"/w/1": "r.pdf", "/w/2": "r.pdf", "/w/3": "r.pdf"},
order: map[string]int{"/w/1": 0, "/w/2": 1, "/w/3": 2},
want: []string{"/w/1", "/w/2", "/w/3"},
},
} {
t.Run(tc.scenario, func(t *testing.T) {
form := &FormData{
files: tc.files,
filesByField: map[string][]string{},
fileBase: tc.fileBase,
fileOrder: tc.order,
}
// Map iteration is randomised, so run it repeatedly: an unstable
// comparator shows up as a differing result across runs.
for range 50 {
var got []string
form.paths([]string{".pdf"}, &got)
if len(got) != len(tc.want) {
t.Fatalf("paths() returned %d entries, want %d", len(got), len(tc.want))
}
for i := range got {
if got[i] != tc.want[i] {
t.Fatalf("paths() = %v, want %v", got, tc.want)
}
}
}
})
}
}

View File

@@ -10,9 +10,11 @@ import (
"strings"
"time"
"github.com/coreos/go-oidc/v3/oidc"
"github.com/google/uuid"
"github.com/labstack/echo/v4"
"github.com/labstack/echo/v4/middleware"
"go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp"
"go.opentelemetry.io/otel"
"go.opentelemetry.io/otel/attribute"
"go.opentelemetry.io/otel/propagation"
@@ -86,13 +88,37 @@ func ParseError(err error) (int, string) {
return http.StatusInternalServerError, http.StatusText(http.StatusInternalServerError)
}
// statusClientClosedRequest is the non-standard 499 status (nginx convention)
// recorded when the client aborts the request before it completes. It keeps
// such outcomes out of the 5xx range in the access log and Prometheus metrics.
const statusClientClosedRequest = 499
// requestCanceled reports whether err is the result of the client aborting the
// request rather than a server-side failure. It requires both that err wraps
// [context.Canceled] and that the request context itself was canceled, so a
// context.Canceled originating elsewhere still surfaces as an internal error.
// A server-side timeout is [context.DeadlineExceeded], mapped to 503 by
// [ParseError], and is deliberately not treated as a client abort.
// See https://github.com/gotenberg/gotenberg/issues/1627.
func requestCanceled(c echo.Context, err error) bool {
return errors.Is(err, context.Canceled) && errors.Is(c.Request().Context().Err(), context.Canceled)
}
// httpErrorHandler is the centralized HTTP error handler. It parses the error,
// returns a response as "text/plain; charset=UTF-8".
func httpErrorHandler() echo.HTTPErrorHandler {
return func(err error, c echo.Context) {
logger := c.Get("logger").(*slog.Logger)
status, message := ParseError(err)
if requestCanceled(c, err) {
// The client is gone, so writing a body would only fail and add
// noise. Record the status so the access log and metrics classify
// it as a client abort rather than an internal error.
c.Response().WriteHeader(statusClientClosedRequest)
return
}
status, message := ParseError(err)
c.Response().Header().Add(echo.HeaderContentType, echo.MIMETextPlainCharsetUTF8)
err = c.String(status, message)
@@ -263,9 +289,15 @@ func telemetryMiddleware(logger *slog.Logger, serverName, correlationIdHeader st
finishTime := time.Now()
status := c.Response().Status
canceled := false
if err != nil {
parsedStatus, _ := ParseError(err)
status = parsedStatus
canceled = requestCanceled(c, err)
if canceled {
status = statusClientClosedRequest
} else {
parsedStatus, _ := ParseError(err)
status = parsedStatus
}
span.SetAttributes(attribute.String("error", err.Error()))
c.Error(err)
@@ -277,44 +309,59 @@ func telemetryMiddleware(logger *slog.Logger, serverName, correlationIdHeader st
WriteBytes: c.Response().Size,
})...)
accessLogger := logger.
With(slog.String("log_type", "access")).
With(slog.String("correlation_id", correlationId)).
With(slog.String("remote_ip", c.RealIP())).
With(slog.String("host", c.Request().Host)).
With(slog.String("uri", c.Request().RequestURI)).
With(slog.String("method", c.Request().Method)).
With(slog.String("path", routePath)).
With(slog.String("referer", c.Request().Referer())).
With(slog.String("user_agent", c.Request().UserAgent())).
With(slog.Int("status", c.Response().Status)).
With(slog.Int64("latency", int64(finishTime.Sub(startTime)))).
With(slog.String("latency_human", finishTime.Sub(startTime).String())).
With(slog.Int64("bytes_in", c.Request().ContentLength)).
With(slog.Int64("bytes_out", c.Response().Size))
// Pick the level and message before building the record: err.Error
// walks a joined error chain, and the nil-error branch has no use
// for it.
level := slog.LevelInfo
msg := "request handled"
if err != nil {
accessLogger.ErrorContext(ctx, err.Error())
} else {
accessLogger.InfoContext(ctx, "request handled")
switch {
case err == nil:
case canceled:
// A client abort is expected, not a server failure; keep it
// visible but out of the error stream.
msg = err.Error()
default:
level = slog.LevelError
msg = err.Error()
}
// One record rather than a chain of With calls. Each With clones
// the whole handler chain, and this logger fans out to a JSON
// handler and an OpenTelemetry bridge that is wired in even when no
// exporter is configured, so a 14-deep chain clones both sub-chains
// 14 times to emit a single line.
latency := finishTime.Sub(startTime)
logger.LogAttrs(ctx, level, msg,
slog.String("log_type", "access"),
slog.String("correlation_id", correlationId),
slog.String("remote_ip", c.RealIP()),
slog.String("host", c.Request().Host),
slog.String("uri", c.Request().RequestURI),
slog.String("method", c.Request().Method),
slog.String("path", routePath),
slog.String("referer", c.Request().Referer()),
slog.String("user_agent", c.Request().UserAgent()),
slog.Int("status", c.Response().Status),
slog.Int64("latency", int64(latency)),
slog.String("latency_human", latency.String()),
slog.Int64("bytes_in", c.Request().ContentLength),
slog.Int64("bytes_out", c.Response().Size),
)
additionalAttributes := []attribute.KeyValue{
semconvSrv.Route(routePath),
}
semconvSrv.RecordMetrics(ctx, semconvutil.ServerMetricData{
ServerName: serverName,
ResponseSize: c.Response().Size,
MetricAttributes: semconvutil.MetricAttributes{
Req: request,
StatusCode: status,
AdditionalAttributes: additionalAttributes,
},
MetricData: semconvutil.MetricData{
RequestSize: request.ContentLength,
ElapsedTime: float64(time.Since(startTime)) / float64(time.Millisecond),
},
ServerName: serverName,
ResponseSize: c.Response().Size,
Req: request,
StatusCode: status,
AdditionalAttributes: additionalAttributes,
RequestSize: request.ContentLength,
ElapsedTime: float64(time.Since(startTime)) / float64(time.Millisecond),
})
return nil
@@ -333,6 +380,63 @@ func basicAuthMiddleware(username, password string) echo.MiddlewareFunc {
})
}
// buildOidcVerifier constructs an OIDC ID token verifier. When oidcJwksUrl is
// set, the keys are fetched from that URL lazily, so there is no network call at
// startup; otherwise the provider is discovered from its issuer, which does one.
// Both paths use an OTEL-instrumented HTTP client, so the JWKS and discovery
// fetches produce client spans.
func (a *Api) buildOidcVerifier() (*oidc.IDTokenVerifier, error) {
httpClient := &http.Client{
Timeout: 10 * time.Second,
Transport: otelhttp.NewTransport(http.DefaultTransport),
}
ctx := oidc.ClientContext(context.Background(), httpClient)
cfg := &oidc.Config{
ClientID: a.oidcAudience,
SupportedSigningAlgs: []string{oidc.RS256, oidc.ES256},
}
if a.oidcJwksUrl != "" {
keySet := oidc.NewRemoteKeySet(ctx, a.oidcJwksUrl)
return oidc.NewVerifier(a.oidcIssuer, keySet, cfg), nil
}
provider, err := oidc.NewProvider(ctx, a.oidcIssuer)
if err != nil {
return nil, fmt.Errorf("discover OIDC provider '%s': %w", a.oidcIssuer, err)
}
return provider.Verifier(cfg), nil
}
// oidcAuthMiddleware validates the Bearer token in the Authorization header with
// the OIDC verifier, which checks the signature against the provider's rotating
// JWKS and the issuer, audience and expiry claims. It answers 401 for a missing
// or invalid token, logging the underlying reason at debug level without leaking
// it to the client.
func oidcAuthMiddleware(verifier *oidc.IDTokenVerifier) echo.MiddlewareFunc {
return func(next echo.HandlerFunc) echo.HandlerFunc {
return func(c echo.Context) error {
rawToken, ok := strings.CutPrefix(c.Request().Header.Get("Authorization"), "Bearer ")
if !ok || rawToken == "" {
return echo.NewHTTPError(http.StatusUnauthorized, "a Bearer token is required in the Authorization header")
}
_, err := verifier.Verify(c.Request().Context(), rawToken)
if err != nil {
if logger, ok := c.Get("logger").(*slog.Logger); ok && logger != nil {
logger.DebugContext(c.Request().Context(), "OIDC token verification failed", slog.Any("error", err))
}
return echo.NewHTTPError(http.StatusUnauthorized, "the Bearer token is invalid")
}
return next(c)
}
}
}
// contextMiddleware, middleware for "multipart/form-data" requests, sets the
// [Context] and related context.CancelFunc in the [echo.Context] under
// "context" and "cancel". If the process is synchronous, it also handles the

View File

@@ -1,15 +1,87 @@
package api
import (
"context"
"crypto/rand"
"crypto/rsa"
"errors"
"fmt"
"log/slog"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
"github.com/coreos/go-oidc/v3/oidc"
"github.com/coreos/go-oidc/v3/oidc/oidctest"
"github.com/labstack/echo/v4"
)
// TestRequestCanceled pins the client-abort discriminator: only a
// context.Canceled that stems from the request context counts, so a server
// timeout or an unrelated cancellation still surfaces as an internal failure.
// See https://github.com/gotenberg/gotenberg/issues/1627.
func TestRequestCanceled(t *testing.T) {
canceled, cancel := context.WithCancel(context.Background())
cancel()
timedOut, cancelTimeout := context.WithDeadline(context.Background(), time.Now().Add(-time.Second))
defer cancelTimeout()
for _, tc := range []struct {
name string
reqCtx context.Context
err error
want bool
}{
{"client abort", canceled, context.Canceled, true},
{"wrapped client abort", canceled, fmt.Errorf("convert to PDF: %w", context.Canceled), true},
{"canceled error but live request", context.Background(), context.Canceled, false},
{"canceled request but unrelated error", canceled, errors.New("boom"), false},
{"server timeout is not a client abort", timedOut, context.DeadlineExceeded, false},
{"no error", canceled, nil, false},
} {
t.Run(tc.name, func(t *testing.T) {
req := httptest.NewRequest(http.MethodPost, "/", nil).WithContext(tc.reqCtx)
c := echo.New().NewContext(req, httptest.NewRecorder())
if got := requestCanceled(c, tc.err); got != tc.want {
t.Fatalf("requestCanceled = %v, want %v", got, tc.want)
}
})
}
}
// TestHttpErrorHandler_ClientClosedRequest ensures a client abort is recorded
// as 499 rather than 500, and that a genuine failure keeps its status.
func TestHttpErrorHandler_ClientClosedRequest(t *testing.T) {
canceled, cancel := context.WithCancel(context.Background())
cancel()
for _, tc := range []struct {
name string
reqCtx context.Context
err error
wantStatus int
}{
{"client abort", canceled, fmt.Errorf("convert to PDF: %w", context.Canceled), statusClientClosedRequest},
{"internal failure", context.Background(), errors.New("boom"), http.StatusInternalServerError},
} {
t.Run(tc.name, func(t *testing.T) {
req := httptest.NewRequest(http.MethodPost, "/", nil).WithContext(tc.reqCtx)
rec := httptest.NewRecorder()
c := echo.New().NewContext(req, rec)
c.Set("logger", slog.New(slog.DiscardHandler))
httpErrorHandler()(tc.err, c)
if rec.Code != tc.wantStatus {
t.Fatalf("status = %d, want %d", rec.Code, tc.wantStatus)
}
})
}
}
// TestOutputFilenameMiddleware pins the sanitizing of the
// "Gotenberg-Output-Filename" header. The value reaches archive entry names and
// a Content-Disposition header, so a path separator must never survive it.
@@ -84,3 +156,90 @@ func TestHardTimeoutMiddleware_MissingLoggerReturnsErrorInsteadOfPanicking(t *te
t.Fatalf("error = %q, want a message mentioning logger", err)
}
}
func TestOidcAuthMiddleware(t *testing.T) {
privateKey, err := rsa.GenerateKey(rand.Reader, 2048)
if err != nil {
t.Fatalf("generate key: %v", err)
}
const (
keyID = "test-key"
audience = "gotenberg"
)
oidcServer := &oidctest.Server{
PublicKeys: []oidctest.PublicKey{
{PublicKey: privateKey.Public(), KeyID: keyID, Algorithm: oidc.RS256},
},
}
srv := httptest.NewServer(oidcServer)
defer srv.Close()
oidcServer.SetIssuer(srv.URL)
// Building through the module's own helper exercises the discovery path too.
a := &Api{oidcIssuer: srv.URL, oidcAudience: audience}
verifier, err := a.buildOidcVerifier()
if err != nil {
t.Fatalf("build verifier: %v", err)
}
claims := func(issuer, aud string, expiresIn time.Duration) string {
now := time.Now()
return fmt.Sprintf(`{"iss":%q,"aud":%q,"sub":"user","exp":%d,"iat":%d}`,
issuer, aud, now.Add(expiresIn).Unix(), now.Unix())
}
sign := func(claims string) string {
return oidctest.SignIDToken(privateKey, keyID, oidc.RS256, claims)
}
otherKey, err := rsa.GenerateKey(rand.Reader, 2048)
if err != nil {
t.Fatalf("generate other key: %v", err)
}
for _, tc := range []struct {
scenario string
authHeader string
wantStatus int
}{
{"valid token", "Bearer " + sign(claims(srv.URL, audience, time.Hour)), http.StatusOK},
{"missing header", "", http.StatusUnauthorized},
{"wrong scheme", "Basic Zm9vOmJhcg==", http.StatusUnauthorized},
{"empty bearer", "Bearer ", http.StatusUnauthorized},
{"malformed token", "Bearer not-a-jwt", http.StatusUnauthorized},
{"wrong issuer", "Bearer " + sign(claims("https://evil.example/", audience, time.Hour)), http.StatusUnauthorized},
{"wrong audience", "Bearer " + sign(claims(srv.URL, "someone-else", time.Hour)), http.StatusUnauthorized},
{"expired token", "Bearer " + sign(claims(srv.URL, audience, -time.Hour)), http.StatusUnauthorized},
{"unknown signing key", "Bearer " + oidctest.SignIDToken(otherKey, "unknown", oidc.RS256, claims(srv.URL, audience, time.Hour)), http.StatusUnauthorized},
} {
t.Run(tc.scenario, func(t *testing.T) {
req := httptest.NewRequest(http.MethodGet, "/", nil)
if tc.authHeader != "" {
req.Header.Set("Authorization", tc.authHeader)
}
c := echo.New().NewContext(req, httptest.NewRecorder())
handler := oidcAuthMiddleware(verifier)(func(c echo.Context) error {
return c.NoContent(http.StatusOK)
})
err := handler(c)
if tc.wantStatus == http.StatusOK {
if err != nil {
t.Fatalf("expected the request to pass, got error: %v", err)
}
return
}
var httpErr *echo.HTTPError
if !errors.As(err, &httpErr) {
t.Fatalf("expected an *echo.HTTPError, got %T (%v)", err, err)
}
if httpErr.Code != tc.wantStatus {
t.Fatalf("status = %d, want %d", httpErr.Code, tc.wantStatus)
}
})
}
}

View File

@@ -23,6 +23,25 @@ import (
"github.com/gotenberg/gotenberg/v8/pkg/gotenberg"
)
// chromiumDisableFeatures is the value of Chromium's --disable-features
// switch.
//
// It restates the "site-per-process,Translate,BlinkGenPropertyTrees" default
// from chromedp.DefaultExecAllocatorOptions (chromedp v0.14.2) on purpose:
// chromedp.Flag keys its flags by switch name, so a second --disable-features
// replaces chromedp's value instead of merging with it. Revisit this list when
// bumping chromedp.
//
// WebUIOmniboxPopup and WebUIOmniboxAimPopup became enabled by default in
// Chromium 151.0.7922.132. Their presenters build the address-bar popup WebUI
// at browser start, headless included, which leaves a renderer process holding
// ~85 MB of anonymous memory for a UI a PDF service can never show. Chromium
// silently ignores feature names it does not know, so both stay harmless on
// older builds (they exist but default to disabled on the Chromium pinned for
// ppc64el) and once upstream eventually removes them.
// See https://github.com/gotenberg/gotenberg/issues/1656.
const chromiumDisableFeatures = "site-per-process,Translate,BlinkGenPropertyTrees,WebUIOmniboxPopup,WebUIOmniboxAimPopup"
type browser interface {
gotenberg.Process
pdf(ctx context.Context, logger *slog.Logger, url, outputPath string, options PdfOptions, aggregate *networkAggregate) error
@@ -49,6 +68,7 @@ type browserArguments struct {
denyPublicIPs bool
clearCache bool
clearCookies bool
clearStorage bool
disableJavaScript bool
}
@@ -132,6 +152,8 @@ func (b *chromiumBrowser) Start(logger *slog.Logger) error {
chromedp.Flag("disable-dev-shm-usage", true),
// See https://github.com/gotenberg/gotenberg/issues/1293.
chromedp.Flag("disable-component-update", false),
// See https://github.com/gotenberg/gotenberg/issues/1656.
chromedp.Flag("disable-features", chromiumDisableFeatures),
)
if b.arguments.allowInsecureLocalhost {
@@ -179,6 +201,28 @@ func (b *chromiumBrowser) Start(logger *slog.Logger) error {
return fmt.Errorf("start pinning proxy: %w", err)
}
opts = append(opts, chromedp.ProxyServer(b.pinningProxy.URL()))
if b.arguments.denyPrivateIPs || b.arguments.denyPublicIPs {
// Chromium implicitly bypasses the proxy for loopback and
// link-local destinations. A WebSocket handshake is never surfaced
// as a fetch.EventRequestPaused, so listenForEventRequestPaused
// cannot filter it; the pinning proxy is the only layer that sees
// it. Left alone, a page could open a WebSocket to 127.0.0.1, ::1,
// localhost, or the link-local cloud metadata endpoint
// (169.254.169.254) and reach it unfiltered. "<-loopback>" removes
// the implicit bypass so those handshakes also traverse the pinning
// proxy and go through [gotenberg.DecideOutbound] like every other
// request.
//
// Gated on the IP-class policy: it is the control this closes, and
// under it loopback and link-local HTTP sub-resources are already
// blocked by listenForEventRequestPaused before they would reach
// the proxy, so this adds only the missing WebSocket coverage. When
// the policy is off, loopback is not restricted, and routing it
// through the proxy would merely change how an unreachable loopback
// sub-resource reports its failure.
opts = append(opts, chromedp.Flag("proxy-bypass-list", "<-loopback>"))
}
}
// See https://github.com/gotenberg/gotenberg/issues/524.
@@ -200,7 +244,7 @@ func (b *chromiumBrowser) Start(logger *slog.Logger) error {
if stopErr != nil {
logger.ErrorContext(context.Background(), fmt.Sprintf("stop pinning proxy after failed start: %s", stopErr))
}
return fmt.Errorf("run exec allocator: %w", err)
return fmt.Errorf("run exec allocator: %w; if Chromium is slow to start, raise --chromium-start-timeout (currently %s)", err, b.arguments.wsUrlReadTimeout)
}
b.ctxMu.Lock()
@@ -345,6 +389,7 @@ func (b *chromiumBrowser) pdf(ctx context.Context, logger *slog.Logger, url, out
runtime.Enable(),
clearCacheActionFunc(logger, b.arguments.clearCache),
clearCookiesActionFunc(logger, b.arguments.clearCookies),
clearStorageActionFunc(logger, b.arguments.clearStorage, url),
disableJavaScriptActionFunc(logger, b.arguments.disableJavaScript),
setCookiesActionFunc(logger, options.Cookies),
userAgentOverride(logger, options.UserAgent),
@@ -371,6 +416,7 @@ func (b *chromiumBrowser) screenshot(ctx context.Context, logger *slog.Logger, u
runtime.Enable(),
clearCacheActionFunc(logger, b.arguments.clearCache),
clearCookiesActionFunc(logger, b.arguments.clearCookies),
clearStorageActionFunc(logger, b.arguments.clearStorage, url),
disableJavaScriptActionFunc(logger, b.arguments.disableJavaScript),
setCookiesActionFunc(logger, options.Cookies),
userAgentOverride(logger, options.UserAgent),
@@ -434,6 +480,17 @@ func (b *chromiumBrowser) do(ctx context.Context, logger *slog.Logger, url strin
extraHttpHeaders: options.ExtraHttpHeaders,
})
// WebSocket handshakes never surface as fetch.EventRequestPaused, so
// listenForEventRequestPaused above cannot filter them. Validate them
// against the same allow / deny lists and IP-class policy.
// See https://github.com/gotenberg/gotenberg/issues/1011.
listenForEventWebSocketCreated(taskCtx, logger, eventWebSocketCreatedOptions{
allowList: b.arguments.allowList,
denyList: b.arguments.denyList,
denyPrivateIPs: b.arguments.denyPrivateIPs,
denyPublicIPs: b.arguments.denyPublicIPs,
})
var (
invalidHttpStatusCode error
invalidHttpStatusCodeMu sync.RWMutex
@@ -487,8 +544,45 @@ func (b *chromiumBrowser) do(ctx context.Context, logger *slog.Logger, url strin
cancelOnMainPageError: taskCancel,
})
var (
crashed error
crashedMu sync.RWMutex
)
// See https://github.com/gotenberg/gotenberg/issues/1640.
listenForEventTargetCrashed(taskCtx, logger, eventTargetCrashedOptions{
crashed: &crashed,
crashedMu: &crashedMu,
cancel: taskCancel,
})
runErr := chromedp.Run(taskCtx, tasks...)
// A crashed renderer is the root cause of every other failure this
// conversion may have recorded, so check it first.
// See https://github.com/gotenberg/gotenberg/issues/1640.
crashedMu.RLock()
defer crashedMu.RUnlock()
if crashed != nil {
return fmt.Errorf("handle tasks: %w", crashed)
}
// The browser context is only ever canceled when the browser process
// dies or is stopped, never on a request timeout. If the run failed
// and the browser context is done, the conversion failed because the
// browser went away mid-flight; fail fast with the same crash error
// instead of letting the error fall through as a generic context
// cancellation. The check is gated on runErr so a successful
// conversion is never discarded by a browser death that lands right
// after it.
// See https://github.com/gotenberg/gotenberg/issues/1640.
if runErr != nil {
if err := b.ctx.Err(); err != nil {
return fmt.Errorf("handle tasks: %w", ErrChromiumCrashed)
}
}
// Check event-driven errors first — they take priority over chromedp.Run
// errors because they carry the actual root cause (e.g., HTTP 500 from
// the main page). When we cancel taskCtx on a main page error,

View File

@@ -3,6 +3,7 @@ package chromium
import (
"context"
"log/slog"
"slices"
"strings"
"testing"
)
@@ -37,3 +38,24 @@ func TestChromiumBrowser_Start_rejectsOverlappingStart(t *testing.T) {
t.Fatal("expected the browser to stay not started")
}
}
// TestChromiumDisableFeatures guards the override described in
// https://github.com/gotenberg/gotenberg/issues/1656. Gotenberg replaces
// chromedp's --disable-features value rather than extending it, as
// chromedp.Flag keys its flags by switch name. Dropping one of chromedp's own
// entries while editing this list would silently re-enable it.
func TestChromiumDisableFeatures(t *testing.T) {
for _, feature := range []string{
// chromedp.DefaultExecAllocatorOptions.
"site-per-process",
"Translate",
"BlinkGenPropertyTrees",
// The address-bar popup WebUI, built even in headless.
"WebUIOmniboxPopup",
"WebUIOmniboxAimPopup",
} {
if !slices.Contains(strings.Split(chromiumDisableFeatures, ","), feature) {
t.Errorf("expected %q to be disabled, got %q", feature, chromiumDisableFeatures)
}
}
}

View File

@@ -43,6 +43,10 @@ var (
// or undefined.
ErrInvalidSelectorQuery = errors.New("invalid selector query")
// ErrScreenshotSelectorNotFound happens when the CSS selector of a
// screenshot matches no element with a rendered box.
ErrScreenshotSelectorNotFound = errors.New("screenshot selector not found")
// ErrRpccMessageTooLarge happens when the messages received by
// ChromeDevTools are larger than 100 MB.
ErrRpccMessageTooLarge = errors.New("rpcc message too large")
@@ -67,6 +71,10 @@ var (
// ErrResourceLoadingFailed happens when one or more resources failed to load.
ErrResourceLoadingFailed = errors.New("resource loading failed")
// ErrChromiumCrashed happens when the Chromium renderer crashes during a
// conversion.
ErrChromiumCrashed = errors.New("chromium crashed")
// PDF specific.
// ErrOmitBackgroundWithoutPrintBackground happens if
@@ -347,6 +355,11 @@ type ScreenshotOptions struct {
// dimensions.
Clip bool
// Selector clips the screenshot to the bounding box of the first element
// matching this CSS selector. Empty captures the whole page. Takes
// precedence over Clip.
Selector string
// Format is the image compression format, either "png" or "jpeg" or
// "webp".
Format string
@@ -370,6 +383,7 @@ func DefaultScreenshotOptions() ScreenshotOptions {
Width: 800,
Height: 600,
Clip: false,
Selector: "",
Format: "png",
Quality: 100,
OptimizeForSpeed: false,
@@ -461,12 +475,13 @@ func (mod *Chromium) Descriptor() gotenberg.ModuleDescriptor {
fs.String("chromium-host-resolver-rules", "", "Set custom mappings to the host resolver")
fs.String("chromium-proxy-server", "", "Set the outbound proxy server; this switch only affects HTTP and HTTPS requests")
fs.Bool("chromium-enable-environment-proxy", false, "Route Chromium's outbound requests through the proxy defined by the standard HTTP_PROXY, HTTPS_PROXY, and NO_PROXY variables, including credentials. Use this instead of --chromium-proxy-server for authenticated proxies, and leave --chromium-proxy-server and --chromium-host-resolver-rules unset")
fs.StringSlice("chromium-allow-list", []string{}, "Set the allowed URLs for Chromium using regular expressions - supports multiple values")
fs.StringSlice("chromium-allow-list", []string{}, `Set the allowed URLs for Chromium using regular expressions - supports multiple values. A match bypasses --chromium-deny-private-ips (CHROMIUM_DENY_PRIVATE_IPS) and --chromium-deny-public-ips (CHROMIUM_DENY_PUBLIC_IPS), so terminate the host or the pattern also matches suffix hosts, for example ^https?://internal\.svc(:|/|$)`)
fs.StringSlice("chromium-deny-list", []string{`^file:(?!//\/tmp/).*`}, "Set the denied URLs for Chromium using regular expressions - supports multiple values")
fs.Bool("chromium-deny-private-ips", false, "Reject URLs whose host resolves to a non-public IP address (loopback, RFC1918, link-local, unique-local). Enable on deployments that accept untrusted form input to mitigate SSRF against internal services")
fs.Bool("chromium-deny-public-ips", false, "Reject URLs whose host resolves to a public IP address. Enable on air-gapped or data-governed deployments to prevent outbound traffic from leaving a private network")
fs.Bool("chromium-clear-cache", false, "Clear Chromium cache between each conversion")
fs.Bool("chromium-clear-cookies", false, "Clear Chromium cookies between each conversion")
fs.Bool("chromium-clear-storage", false, "Clear Chromium local storage between each conversion (session storage is already isolated per conversion)")
fs.Bool("chromium-disable-javascript", false, "Disable JavaScript")
fs.Bool("chromium-disable-routes", false, "Disable the routes")
@@ -518,6 +533,7 @@ func (mod *Chromium) Provision(ctx *gotenberg.Context) error {
denyPublicIPs: flags.MustBool("chromium-deny-public-ips"),
clearCache: flags.MustBool("chromium-clear-cache"),
clearCookies: flags.MustBool("chromium-clear-cookies"),
clearStorage: flags.MustBool("chromium-clear-storage"),
disableJavaScript: flags.MustBool("chromium-disable-javascript"),
}
@@ -1094,6 +1110,8 @@ func chromiumErrorType(err error, queueReason string) string {
errors.Is(err, ErrInvalidEvaluationExpression),
errors.Is(err, ErrInvalidSelectorQuery):
return gotenberg.ErrorTypeInvalidInput
case errors.Is(err, ErrChromiumCrashed):
return "chromium_unavailable"
case errors.Is(err, gotenberg.ErrMaximumQueueSizeExceeded):
return queueReason
case errors.Is(err, gotenberg.ErrProcessAlreadyRestarting):

View File

@@ -21,6 +21,7 @@ func TestChromiumErrorType(t *testing.T) {
{"invalid resource http status", ErrInvalidResourceHttpStatusCode, "chromium_unavailable", "invalid_input"},
{"loading failed", ErrLoadingFailed, "chromium_unavailable", "invalid_input"},
{"resource loading failed", ErrResourceLoadingFailed, "chromium_unavailable", "invalid_input"},
{"crashed", ErrChromiumCrashed, "chromium_unavailable", "chromium_unavailable"},
{"invalid evaluation expression", ErrInvalidEvaluationExpression, "chromium_unavailable", "invalid_input"},
{"invalid selector query", ErrInvalidSelectorQuery, "chromium_unavailable", "invalid_input"},
{"pdf queue", gotenberg.ErrMaximumQueueSizeExceeded, "chromium_unavailable", "chromium_unavailable"},

View File

@@ -14,6 +14,7 @@ import (
"github.com/chromedp/cdproto/cdp"
"github.com/chromedp/cdproto/fetch"
"github.com/chromedp/cdproto/inspector"
"github.com/chromedp/cdproto/network"
"github.com/chromedp/cdproto/page"
"github.com/chromedp/cdproto/runtime"
@@ -44,6 +45,54 @@ func listenForNetworkActivity(ctx context.Context, aggregate *networkAggregate)
})
}
type eventWebSocketCreatedOptions struct {
allowList, denyList []*regexp2.Regexp
denyPrivateIPs bool
denyPublicIPs bool
}
// listenForEventWebSocketCreated validates the target of every WebSocket
// handshake against the same allow / deny lists and IP-class policy as
// [listenForEventRequestPaused]. Chromium never surfaces a WebSocket
// handshake as a fetch.EventRequestPaused, so without this listener a page
// could open a WebSocket to an address the outbound filter would otherwise
// block. See https://github.com/gotenberg/gotenberg/issues/1011.
//
// This listener records an operator-visible warning with the full ws:// URL.
// The connection itself is severed by the pinning proxy, which every
// WebSocket handshake traverses once the implicit loopback / link-local proxy
// bypass is removed (see the "<-loopback>" flag in browser.go). When the
// operator configures a custom proxy or host-resolver mappings, the pinning
// proxy is not started; the WebSocket then follows the operator's egress path
// and this warning is the remaining safeguard, since a WebSocket handshake
// cannot be aborted through the CDP Network domain.
func listenForEventWebSocketCreated(ctx context.Context, logger *slog.Logger, options eventWebSocketCreatedOptions) {
chromedp.ListenTarget(ctx, func(ev any) {
e, ok := ev.(*network.EventWebSocketCreated)
if !ok {
return
}
go func() {
logger.DebugContext(ctx, fmt.Sprintf("event EventWebSocketCreated fired for '%s'", e.URL))
deadline, ok := ctx.Deadline()
if !ok {
logger.ErrorContext(ctx, "context has no deadline, cannot filter WebSocket URL")
return
}
err := gotenberg.FilterOutboundURL(ctx, e.URL, options.allowList, options.denyList, deadline,
gotenberg.WithDenyPrivateIPs(options.denyPrivateIPs),
gotenberg.WithDenyPublicIPs(options.denyPublicIPs),
)
if err != nil {
logger.WarnContext(ctx, err.Error())
}
}()
})
}
type eventRequestPausedOptions struct {
allowList, denyList []*regexp2.Regexp
denyPrivateIPs bool
@@ -274,7 +323,15 @@ func listenForEventResponseReceived(
return
}
logger.DebugContext(ctx, fmt.Sprintf("event EventResponseReceived fired for a resource: %+v", ev.Response))
// Formatting the whole response is the most expensive thing this
// listener does, and it runs per sub-resource on chromedp's single
// per-target event goroutine while that goroutine holds the mutex
// it also takes to dispatch command responses. At the default log
// level the result is discarded, so gate it on the level rather
// than let slog drop it after the fact.
if logger.Enabled(ctx, slog.LevelDebug) {
logger.DebugContext(ctx, fmt.Sprintf("event EventResponseReceived fired for a resource: %+v", ev.Response))
}
if slices.Contains(options.failOnResourceOnHttpStatusCode, ev.Response.Status) {
if !shouldCheckResourceHttpStatusCode(ev.Response.URL, normalizedIgnoreDomains) {
@@ -473,6 +530,38 @@ func listenForEventExceptionThrown(ctx context.Context, logger *slog.Logger, con
})
}
type eventTargetCrashedOptions struct {
crashed *error
crashedMu *sync.RWMutex
cancel context.CancelFunc
}
// listenForEventTargetCrashed listens for the Inspector.targetCrashed event,
// which Chromium sends when the renderer serving the conversion's tab
// crashes. chromedp enables the Inspector domain on every target but does
// not handle this event: left alone, the in-flight CDP command never
// receives a response and the conversion blocks until the request deadline.
// Record the crash and cancel the task context so the conversion fails fast
// instead.
// See https://github.com/gotenberg/gotenberg/issues/1640.
func listenForEventTargetCrashed(ctx context.Context, logger *slog.Logger, options eventTargetCrashedOptions) {
chromedp.ListenTarget(ctx, func(ev any) {
if _, ok := ev.(*inspector.EventTargetCrashed); ok {
logger.DebugContext(ctx, "event EventTargetCrashed fired")
options.crashedMu.Lock()
defer options.crashedMu.Unlock()
*options.crashed = ErrChromiumCrashed
// Cancel the task context so the in-flight CDP command aborts
// immediately instead of waiting for a response the crashed
// renderer can never send.
options.cancel()
}
})
}
// waitForEventDomContentEventFired registers a listener for the
// DomContentEventFired event and returns a waiter that blocks until the
// event fires or ctx is done. The listener registers at call time, not

View File

@@ -11,6 +11,15 @@ import (
// pathological page cannot grow the set without limit.
const maxTrackedOrigins = 64
// maxTrackedRequests bounds the request id to URL map. Entries are dropped as
// soon as the request settles, so the map normally holds only what is in
// flight, but a request that never reports a loading-finished or
// loading-failed event never settles. Without a cap, a page that opens
// requests it never resolves would grow the map for the whole conversion, at
// the cost of one full response URL per entry. Losing an entry only costs the
// heaviest-resource URL attribution for that request.
const maxTrackedRequests = 1024
// networkAggregate accumulates per-conversion network activity from Chromium
// DevTools events. It is safe for concurrent use by the chromedp event listener
// goroutine and the conversion goroutine that reads the snapshot afterwards.
@@ -60,7 +69,9 @@ func (a *networkAggregate) onResponseReceived(ev *network.EventResponseReceived)
a.origins[origin] = struct{}{}
}
}
a.requestURLByID[ev.RequestID] = ev.Response.URL
if len(a.requestURLByID) < maxTrackedRequests {
a.requestURLByID[ev.RequestID] = ev.Response.URL
}
}
// onLoadingFinished records a successfully completed request and its size,
@@ -81,6 +92,11 @@ func (a *networkAggregate) onLoadingFinished(ev *network.EventLoadingFinished) {
a.heaviestBytes = size
a.heaviestURL = a.requestURLByID[ev.RequestID]
}
// The request has settled and nothing reads its URL again. Dropping it
// keeps the map proportional to the requests in flight rather than to
// every request the page ever made.
delete(a.requestURLByID, ev.RequestID)
}
// onLoadingFailed records a request that failed to complete.
@@ -94,6 +110,9 @@ func (a *networkAggregate) onLoadingFailed(ev *network.EventLoadingFailed) {
a.requestCount++
a.failedCount++
// Settled, like a finished request: its URL is never read again.
delete(a.requestURLByID, ev.RequestID)
}
func (a *networkAggregate) snapshot() networkStats {

View File

@@ -98,3 +98,86 @@ func TestNetworkAggregate_ConcurrentSafe(t *testing.T) {
t.Errorf("requestCount = %d, want 100", got)
}
}
// TestNetworkAggregate_SettledRequestsAreDropped covers the growth where every
// response URL stayed in the map for the whole conversion even though nothing
// reads it again once the request settles.
func TestNetworkAggregate_SettledRequestsAreDropped(t *testing.T) {
a := newNetworkAggregate()
for i := range 500 {
id := network.RequestID(fmt.Sprintf("r%d", i))
a.onResponseReceived(&network.EventResponseReceived{
RequestID: id,
Response: &network.Response{URL: fmt.Sprintf("https://host.example.com/%d", i)},
})
if i%2 == 0 {
a.onLoadingFinished(&network.EventLoadingFinished{RequestID: id, EncodedDataLength: 10})
continue
}
a.onLoadingFailed(&network.EventLoadingFailed{RequestID: id})
}
a.mu.Lock()
tracked := len(a.requestURLByID)
a.mu.Unlock()
if tracked != 0 {
t.Errorf("tracked requests = %d, want 0: settled requests must not be retained", tracked)
}
// The bookkeeping the map feeds must survive the pruning.
got := a.snapshot()
if got.requestCount != 500 {
t.Errorf("requestCount = %d, want 500", got.requestCount)
}
if got.failedCount != 250 {
t.Errorf("failedCount = %d, want 250", got.failedCount)
}
}
// TestNetworkAggregate_UnsettledRequestCap verifies the ceiling that applies
// when requests never settle, which is the only way the map can still grow.
func TestNetworkAggregate_UnsettledRequestCap(t *testing.T) {
a := newNetworkAggregate()
for i := range maxTrackedRequests + 500 {
a.onResponseReceived(&network.EventResponseReceived{
RequestID: network.RequestID(fmt.Sprintf("r%d", i)),
Response: &network.Response{URL: fmt.Sprintf("https://host.example.com/%d", i)},
})
}
a.mu.Lock()
tracked := len(a.requestURLByID)
a.mu.Unlock()
if tracked != maxTrackedRequests {
t.Errorf("tracked requests = %d, want %d (capped)", tracked, maxTrackedRequests)
}
}
// TestNetworkAggregate_HeaviestURLSurvivesPruning guards the attribution the
// map exists for: the URL must still be resolved before the entry is dropped.
func TestNetworkAggregate_HeaviestURLSurvivesPruning(t *testing.T) {
a := newNetworkAggregate()
a.onResponseReceived(&network.EventResponseReceived{
RequestID: "small",
Response: &network.Response{URL: "https://example.com/small.css"},
})
a.onLoadingFinished(&network.EventLoadingFinished{RequestID: "small", EncodedDataLength: 10})
a.onResponseReceived(&network.EventResponseReceived{
RequestID: "big",
Response: &network.Response{URL: "https://example.com/big.png"},
})
a.onLoadingFinished(&network.EventLoadingFinished{RequestID: "big", EncodedDataLength: 4096})
got := a.snapshot()
if got.heaviestURL != "https://example.com/big.png" || got.heaviestBytes != 4096 {
t.Errorf("heaviest = (%q, %d), want (%q, 4096)", got.heaviestURL, got.heaviestBytes, "https://example.com/big.png")
}
}

View File

@@ -11,6 +11,7 @@ import (
"net/netip"
"net/url"
"sync"
"sync/atomic"
"time"
"github.com/dlclark/regexp2"
@@ -57,6 +58,18 @@ type pinningProxy struct {
server *http.Server
wg sync.WaitGroup
// closing is closed by Stop to force in-flight CONNECT tunnels shut.
// [http.Server.Shutdown] cannot do it: net/http untracks a connection once
// a handler hijacks it, so a tunnel would otherwise outlive the proxy that
// created it. Recreated on every Start.
closing chan struct{}
// maxTunnels ceilings the CONNECT handlers in flight. Tests may lower it.
maxTunnels int64
// tunnels counts the CONNECT handlers in flight.
tunnels atomic.Int64
logger *slog.Logger
started bool
mu sync.Mutex
@@ -82,6 +95,7 @@ func newPinningProxy(allowList, denyList []*regexp2.Regexp, denyPrivateIPs, deny
dialer := &net.Dialer{Timeout: 10 * time.Second}
return dialer.DialContext(ctx, network, addr)
},
maxTunnels: maxConcurrentTunnels,
}
if enableEnvironmentProxy {
@@ -110,6 +124,7 @@ func (p *pinningProxy) Start(logger *slog.Logger) error {
}
p.listener = l
p.closing = make(chan struct{})
p.logger = logger.With(slog.String("logger", "pinning-proxy"))
p.server = &http.Server{
Handler: http.HandlerFunc(p.serveHTTP),
@@ -140,9 +155,18 @@ func (p *pinningProxy) Stop(logger *slog.Logger) error {
return nil
}
srv := p.server
closing := p.closing
p.closing = nil
p.started = false
p.mu.Unlock()
// Force in-flight tunnels shut before draining the server. Shutdown does
// not reach them, so a tunnel whose upstream never answers would otherwise
// survive the proxy, and with it every Chromium restart.
if closing != nil {
close(closing)
}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
@@ -181,6 +205,25 @@ func (p *pinningProxy) serveHTTP(w http.ResponseWriter, req *http.Request) {
// Chromium then negotiates TLS end-to-end with the original hostname in
// SNI.
func (p *pinningProxy) handleConnect(w http.ResponseWriter, req *http.Request) {
// A ceiling, not a tuning knob: it bounds what a tunnel that refuses to end
// can accumulate, whatever keeps it alive. [spliceIdleTimeout] ends a silent
// tunnel, but a peer trickling a byte just under it stays "active" forever,
// and a compromised renderer can hold the client side open to match.
//
// Chromium caps itself well below this. Its socket pool manager allows 128
// sockets per proxy chain for normal traffic plus 128 for WebSocket
// traffic, and every request Gotenberg's Chromium makes traverses this one
// proxy chain, so an honest browser cannot exceed 256 tunnels here. At
// double that, a real page never meets the ceiling and a hostile one stops
// at it.
if !p.acquireTunnel() {
p.logger.WarnContext(req.Context(), fmt.Sprintf("CONNECT to '%s' refused: %d tunnels already in flight", req.Host, p.maxTunnels))
http.Error(w, "too many tunnels", http.StatusServiceUnavailable)
return
}
defer p.releaseTunnel()
_, port, err := net.SplitHostPort(req.Host)
if err != nil {
http.Error(w, "bad CONNECT target", http.StatusBadRequest)
@@ -264,24 +307,142 @@ func (p *pinningProxy) handleConnect(w http.ResponseWriter, req *http.Request) {
return
}
// Splice bytes in both directions until either side closes.
var splice sync.WaitGroup
splice.Add(2)
p.mu.Lock()
closing := p.closing
p.mu.Unlock()
spliceTunnel(client, upstream, closing, spliceIdleTimeout)
}
// maxConcurrentTunnels is the default for [pinningProxy.maxTunnels]. See
// [pinningProxy.handleConnect] for how the value is derived.
const maxConcurrentTunnels = 512
// acquireTunnel reserves a slot for one CONNECT handler, reporting false when
// the proxy is already at [pinningProxy.maxTunnels]. The compare-and-swap loop
// keeps the check and the increment atomic, so concurrent handlers cannot
// overshoot the ceiling between them.
func (p *pinningProxy) acquireTunnel() bool {
for {
current := p.tunnels.Load()
if current >= p.maxTunnels {
return false
}
if p.tunnels.CompareAndSwap(current, current+1) {
return true
}
}
}
// releaseTunnel returns a slot taken by [pinningProxy.acquireTunnel].
func (p *pinningProxy) releaseTunnel() {
p.tunnels.Add(-1)
}
// spliceIdleTimeout bounds a CONNECT tunnel in which no byte has moved in
// either direction.
//
// Nothing else bounds one. The hijacked connections carry no deadline: the
// server clears the header read deadline once the request line is in, and
// net.Dialer.Timeout only covers the connect. net/http also untracks a
// connection once it is hijacked, so neither Server.Shutdown nor a Chromium
// restart reaps it. Left alone, an upstream that accepts the tunnel and then
// answers nothing holds two goroutines and two sockets until the process dies.
//
// Sized well above any legitimate pause between a request and its response, so
// a slow origin is never cut off. A transfer that keeps making progress
// refreshes the deadline and runs for as long as it needs.
const spliceIdleTimeout = 2 * time.Minute
// spliceTunnel copies bytes between the two ends of a CONNECT tunnel until
// both directions finish, the tunnel sits idle for idleTimeout, or closing is
// closed because the proxy is shutting down. Callers pass
// [spliceIdleTimeout]; only tests shorten it.
//
// Each direction half-closes its destination once its source reaches EOF, so a
// peer that waits for the request to end before answering still sees the EOF.
// Idleness is tracked across both directions rather than per direction: the
// client sends nothing for the length of a download, and half-closing its write
// side then would tell the origin the client had gone away.
func spliceTunnel(client, upstream net.Conn, closing <-chan struct{}, idleTimeout time.Duration) {
var lastActivity atomic.Int64
lastActivity.Store(time.Now().UnixNano())
var wg sync.WaitGroup
wg.Add(2)
go func() {
defer splice.Done()
_, _ = io.Copy(upstream, client)
defer wg.Done()
copyTracking(upstream, client, &lastActivity, idleTimeout)
if cw, ok := upstream.(interface{ CloseWrite() error }); ok {
_ = cw.CloseWrite()
}
}()
go func() {
defer splice.Done()
_, _ = io.Copy(client, upstream)
defer wg.Done()
copyTracking(client, upstream, &lastActivity, idleTimeout)
if cw, ok := client.(interface{ CloseWrite() error }); ok {
_ = cw.CloseWrite()
}
}()
splice.Wait()
done := make(chan struct{})
go func() {
wg.Wait()
close(done)
}()
ticker := time.NewTicker(idleTimeout / 4)
defer ticker.Stop()
for {
select {
case <-done:
return
case <-closing:
case <-ticker.C:
if time.Since(time.Unix(0, lastActivity.Load())) < idleTimeout {
continue
}
}
// Closing both ends unblocks whichever copy is still reading. The
// caller's own deferred Close calls then become no-ops.
_ = client.Close()
_ = upstream.Close()
<-done
return
}
}
// copyTracking copies src into dst, recording the time of every chunk that
// moves so [spliceTunnel] can tell a busy tunnel from an idle one.
func copyTracking(dst, src net.Conn, lastActivity *atomic.Int64, writeTimeout time.Duration) {
buf := make([]byte, 32*1024)
for {
n, readErr := src.Read(buf)
if n > 0 {
lastActivity.Store(time.Now().UnixNano())
// Bound the write. A destination that has gone away accepts the
// first chunk into its send buffer and only fails on the next one,
// so without a deadline this direction keeps a dead tunnel alive
// for one more chunk. A destination that stops reading altogether
// would block here forever.
_ = dst.SetWriteDeadline(time.Now().Add(writeTimeout))
_, writeErr := dst.Write(buf[:n])
if writeErr != nil {
return
}
lastActivity.Store(time.Now().UnixNano())
}
if readErr != nil {
return
}
}
}
// handleForward handles plain HTTP requests sent to the proxy as absolute

View File

@@ -898,3 +898,280 @@ func TestPinningProxy_StopIdempotent(t *testing.T) {
t.Fatalf("second Stop on stopped proxy: %v", err)
}
}
// tcpPair returns the two ends of a connected loopback TCP connection. Both
// ends are closed when the test finishes.
func tcpPair(t *testing.T) (net.Conn, net.Conn) {
t.Helper()
listener, err := net.Listen("tcp", "127.0.0.1:0")
if err != nil {
t.Fatalf("listen: %v", err)
}
defer func() { _ = listener.Close() }()
type accepted struct {
conn net.Conn
err error
}
acceptChan := make(chan accepted, 1)
go func() {
conn, acceptErr := listener.Accept()
acceptChan <- accepted{conn: conn, err: acceptErr}
}()
dialed, err := net.Dial("tcp", listener.Addr().String())
if err != nil {
t.Fatalf("dial: %v", err)
}
res := <-acceptChan
if res.err != nil {
t.Fatalf("accept: %v", res.err)
}
t.Cleanup(func() {
_ = dialed.Close()
_ = res.conn.Close()
})
return dialed, res.conn
}
// TestSpliceTunnel_IdleTunnelIsClosed covers the leak where an upstream that
// accepted a CONNECT tunnel and then never spoke pinned both splice goroutines
// and both sockets for the lifetime of the process. A hijacked connection
// carries no deadline and net/http stops tracking it, so the idle bound in
// spliceTunnel is the only thing that ends such a tunnel.
func TestSpliceTunnel_IdleTunnelIsClosed(t *testing.T) {
// The peers are kept open by the pair's cleanup: the tunnel is silent, not
// finished.
client, _ := tcpPair(t)
upstream, _ := tcpPair(t)
done := make(chan struct{})
go func() {
spliceTunnel(client, upstream, nil, 100*time.Millisecond)
close(done)
}()
select {
case <-done:
case <-time.After(10 * time.Second):
t.Fatal("spliceTunnel did not return on an idle tunnel")
}
}
// TestSpliceTunnel_ClosingShutsTunnelDown verifies that stopping the proxy
// reaps in-flight tunnels. http.Server.Shutdown cannot: it stops tracking a
// connection once a handler hijacks it, so without this signal a tunnel would
// outlive the proxy and every Chromium restart after it.
func TestSpliceTunnel_ClosingShutsTunnelDown(t *testing.T) {
client, _ := tcpPair(t)
upstream, _ := tcpPair(t)
closing := make(chan struct{})
done := make(chan struct{})
go func() {
// An idle timeout far beyond the test: only closing can end this.
spliceTunnel(client, upstream, closing, time.Hour)
close(done)
}()
close(closing)
select {
case <-done:
case <-time.After(10 * time.Second):
t.Fatal("spliceTunnel did not return when the proxy shut down")
}
}
// TestSpliceTunnel_ActiveTransferOutlivesIdleTimeout guards the idle bound
// against cutting a healthy transfer. Idleness is tracked across both
// directions, so a download that keeps making progress must survive well past
// the timeout even though the client sends nothing throughout.
func TestSpliceTunnel_ActiveTransferOutlivesIdleTimeout(t *testing.T) {
const (
idleTimeout = 100 * time.Millisecond
chunks = 10
interval = 30 * time.Millisecond
)
client, clientPeer := tcpPair(t)
upstream, upstreamPeer := tcpPair(t)
done := make(chan struct{})
go func() {
spliceTunnel(client, upstream, nil, idleTimeout)
close(done)
}()
// Trickle a response for well over the idle timeout, then finish.
go func() {
for range chunks {
_, _ = upstreamPeer.Write([]byte("x"))
time.Sleep(interval)
}
_ = upstreamPeer.Close()
}()
received := 0
buf := make([]byte, chunks)
for received < chunks {
err := clientPeer.SetReadDeadline(time.Now().Add(10 * time.Second))
if err != nil {
t.Fatalf("set read deadline: %v", err)
}
n, readErr := clientPeer.Read(buf)
received += n
if readErr != nil {
break
}
}
if received != chunks {
t.Fatalf("received %d bytes, want %d: the tunnel was cut while still transferring", received, chunks)
}
select {
case <-done:
case <-time.After(10 * time.Second):
t.Fatal("spliceTunnel did not return after the upstream closed")
}
}
// TestPinningProxy_CONNECT_TunnelCeiling verifies the ceiling that bounds what
// tunnels refusing to end can accumulate. The idle bound cannot cover a peer
// that trickles just under it, so the count is what stops the growth.
func TestPinningProxy_CONNECT_TunnelCeiling(t *testing.T) {
// An upstream that accepts and then says nothing: the tunnel stays open.
upstreamAddr, stop := newRawTCPServer(t, func(c net.Conn) {
<-make(chan struct{})
})
t.Cleanup(stop)
p := newPinningProxy(nil, nil, false, false, false)
p.maxTunnels = 1
p.decide = func(_ context.Context, _ string, _, _ []*regexp2.Regexp, _ time.Time) (gotenberg.OutboundDecision, error) {
return gotenberg.OutboundDecision{Pinned: []netip.Addr{netip.MustParseAddr("127.0.0.1")}}, nil
}
p.dialPinned = func(_ context.Context, network string, _ []netip.Addr, _ string) (net.Conn, error) {
return net.Dial(network, upstreamAddr)
}
proxyURL := newProxyForTest(t, p)
proxyAddr := strings.TrimPrefix(proxyURL, "http://")
connect := func(t *testing.T) *bufio.Reader {
t.Helper()
conn, err := net.Dial("tcp", proxyAddr)
if err != nil {
t.Fatalf("dial proxy: %v", err)
}
t.Cleanup(func() { _ = conn.Close() })
err = conn.SetDeadline(time.Now().Add(10 * time.Second))
if err != nil {
t.Fatalf("set deadline: %v", err)
}
_, err = fmt.Fprintf(conn, "CONNECT example.com:443 HTTP/1.1\r\nHost: example.com:443\r\n\r\n")
if err != nil {
t.Fatalf("write CONNECT: %v", err)
}
return bufio.NewReader(conn)
}
first := connect(t)
statusLine, err := first.ReadString('\n')
if err != nil {
t.Fatalf("read first status: %v", err)
}
if !strings.Contains(statusLine, " 200 ") {
t.Fatalf("first CONNECT status = %q, want 200", statusLine)
}
// The first tunnel now holds the only slot.
second := connect(t)
resp, err := http.ReadResponse(second, nil)
if err != nil {
t.Fatalf("read second response: %v", err)
}
defer func() { _ = resp.Body.Close() }()
if resp.StatusCode != http.StatusServiceUnavailable {
t.Fatalf("second CONNECT status = %d, want %d", resp.StatusCode, http.StatusServiceUnavailable)
}
if got := p.tunnels.Load(); got != 1 {
t.Errorf("tunnels in flight = %d, want 1: a refused CONNECT must not consume a slot", got)
}
}
// TestPinningProxy_TunnelSlotIsReleased verifies a completed tunnel gives its
// slot back, so the ceiling bounds concurrency rather than lifetime totals.
func TestPinningProxy_TunnelSlotIsReleased(t *testing.T) {
upstreamAddr, stop := newRawTCPServer(t, func(c net.Conn) {
defer c.Close()
_, _ = c.Write([]byte("HI"))
})
t.Cleanup(stop)
p := newPinningProxy(nil, nil, false, false, false)
p.maxTunnels = 1
p.decide = func(_ context.Context, _ string, _, _ []*regexp2.Regexp, _ time.Time) (gotenberg.OutboundDecision, error) {
return gotenberg.OutboundDecision{Pinned: []netip.Addr{netip.MustParseAddr("127.0.0.1")}}, nil
}
p.dialPinned = func(_ context.Context, network string, _ []netip.Addr, _ string) (net.Conn, error) {
return net.Dial(network, upstreamAddr)
}
proxyURL := newProxyForTest(t, p)
proxyAddr := strings.TrimPrefix(proxyURL, "http://")
for attempt := range 3 {
conn, err := net.Dial("tcp", proxyAddr)
if err != nil {
t.Fatalf("attempt %d dial proxy: %v", attempt, err)
}
err = conn.SetDeadline(time.Now().Add(10 * time.Second))
if err != nil {
t.Fatalf("attempt %d set deadline: %v", attempt, err)
}
_, err = fmt.Fprintf(conn, "CONNECT example.com:443 HTTP/1.1\r\nHost: example.com:443\r\n\r\n")
if err != nil {
t.Fatalf("attempt %d write CONNECT: %v", attempt, err)
}
br := bufio.NewReader(conn)
statusLine, err := br.ReadString('\n')
if err != nil {
t.Fatalf("attempt %d read status: %v", attempt, err)
}
if !strings.Contains(statusLine, " 200 ") {
t.Fatalf("attempt %d CONNECT status = %q, want 200: the slot was not released", attempt, statusLine)
}
// Drain until the upstream's close ends the tunnel, then release it.
_, _ = io.ReadAll(br)
_ = conn.Close()
// The handler returns just after the splice ends.
for range 100 {
if p.tunnels.Load() == 0 {
break
}
time.Sleep(20 * time.Millisecond)
}
if got := p.tunnels.Load(); got != 0 {
t.Fatalf("attempt %d: tunnels in flight = %d, want 0", attempt, got)
}
}
}

View File

@@ -366,6 +366,7 @@ func FormDataChromiumScreenshotOptions(ctx *api.Context) (*api.FormData, Screens
var (
width, height int
clip bool
selector string
format string
quality int
optimizeForSpeed bool
@@ -376,6 +377,7 @@ func FormDataChromiumScreenshotOptions(ctx *api.Context) (*api.FormData, Screens
Int("width", &width, defaultScreenshotOptions.Width).
Int("height", &height, defaultScreenshotOptions.Height).
Bool("clip", &clip, defaultScreenshotOptions.Clip).
String("selector", &selector, defaultScreenshotOptions.Selector).
Custom("format", func(value string) error {
if value == "" {
format = defaultScreenshotOptions.Format
@@ -420,6 +422,7 @@ func FormDataChromiumScreenshotOptions(ctx *api.Context) (*api.FormData, Screens
Width: width,
Height: height,
Clip: clip,
Selector: selector,
Format: format,
Quality: quality,
OptimizeForSpeed: optimizeForSpeed,
@@ -471,11 +474,18 @@ func convertUrlRoute(chromium Api, engine gotenberg.PdfEngine) api.Route {
metadata := pdfengines.FormDataPdfMetadata(form, false)
encrypt := pdfengines.FormDataPdfEncrypt(form)
embedPaths := pdfengines.FormDataPdfEmbeds(form)
watermark := pdfengines.FormDataPdfWatermark(form, false)
watermarkFile := pdfengines.FormDataPdfWatermarkFile(form)
stamp := pdfengines.FormDataPdfStamp(form, false)
stampFile := pdfengines.FormDataPdfStampFile(form)
watermarks, wErr := pdfengines.FormDataPdfWatermarks(form)
if wErr != nil {
return fmt.Errorf("form data watermarks: %w", wErr)
}
stamps, sErr := pdfengines.FormDataPdfStamps(form)
if sErr != nil {
return fmt.Errorf("form data stamps: %w", sErr)
}
var watermarkFiles, stampFiles []string
form.Watermarks(&watermarkFiles).Stamps(&stampFiles)
rotateAngle, rotatePages := pdfengines.FormDataPdfRotate(form, false)
optimizeImages, imageQuality := pdfengines.FormDataPdfOptimize(form)
embedsMetadata := pdfengines.FormDataPdfEmbedsMetadata(form)
facturX, facturxXmlPath := pdfengines.FormDataPdfFacturX(form)
@@ -492,16 +502,16 @@ func convertUrlRoute(chromium Api, engine gotenberg.PdfEngine) api.Route {
return fmt.Errorf("reject URL scheme: %w", err)
}
err = pdfengines.EnsureWatermarkFile(&watermark, watermarkFile)
err = pdfengines.BindWatermarkFiles(watermarks, watermarkFiles)
if err != nil {
return fmt.Errorf("validate watermark: %w", err)
return fmt.Errorf("bind watermark files: %w", err)
}
err = pdfengines.EnsureStampFile(&stamp, stampFile)
err = pdfengines.BindStampFiles(stamps, stampFiles)
if err != nil {
return fmt.Errorf("validate stamp: %w", err)
return fmt.Errorf("bind stamp files: %w", err)
}
err = convertUrl(ctx, chromium, engine, url, options, mode, pdfFormats, metadata, encrypt, embedPaths, embedsMetadata, facturX, facturxXmlPath, watermark, stamp, rotateAngle, rotatePages)
err = convertUrl(ctx, chromium, engine, url, options, mode, pdfFormats, metadata, encrypt, embedPaths, embedsMetadata, facturX, facturxXmlPath, watermarks, stamps, rotateAngle, rotatePages, optimizeImages, imageQuality)
if err != nil {
return fmt.Errorf("convert URL to PDF: %w", err)
}
@@ -560,11 +570,18 @@ func convertHtmlRoute(chromium Api, engine gotenberg.PdfEngine) api.Route {
metadata := pdfengines.FormDataPdfMetadata(form, false)
encrypt := pdfengines.FormDataPdfEncrypt(form)
embedPaths := pdfengines.FormDataPdfEmbeds(form)
watermark := pdfengines.FormDataPdfWatermark(form, false)
watermarkFile := pdfengines.FormDataPdfWatermarkFile(form)
stamp := pdfengines.FormDataPdfStamp(form, false)
stampFile := pdfengines.FormDataPdfStampFile(form)
watermarks, wErr := pdfengines.FormDataPdfWatermarks(form)
if wErr != nil {
return fmt.Errorf("form data watermarks: %w", wErr)
}
stamps, sErr := pdfengines.FormDataPdfStamps(form)
if sErr != nil {
return fmt.Errorf("form data stamps: %w", sErr)
}
var watermarkFiles, stampFiles []string
form.Watermarks(&watermarkFiles).Stamps(&stampFiles)
rotateAngle, rotatePages := pdfengines.FormDataPdfRotate(form, false)
optimizeImages, imageQuality := pdfengines.FormDataPdfOptimize(form)
embedsMetadata := pdfengines.FormDataPdfEmbedsMetadata(form)
facturX, facturxXmlPath := pdfengines.FormDataPdfFacturX(form)
@@ -576,18 +593,18 @@ func convertHtmlRoute(chromium Api, engine gotenberg.PdfEngine) api.Route {
return fmt.Errorf("validate form data: %w", err)
}
err = pdfengines.EnsureWatermarkFile(&watermark, watermarkFile)
err = pdfengines.BindWatermarkFiles(watermarks, watermarkFiles)
if err != nil {
return fmt.Errorf("validate watermark: %w", err)
return fmt.Errorf("bind watermark files: %w", err)
}
err = pdfengines.EnsureStampFile(&stamp, stampFile)
err = pdfengines.BindStampFiles(stamps, stampFiles)
if err != nil {
return fmt.Errorf("validate stamp: %w", err)
return fmt.Errorf("bind stamp files: %w", err)
}
url := fmt.Sprintf("file://%s", inputPath)
options.AllowedFilePrefixes = []string{ctx.DirPath()}
err = convertUrl(ctx, chromium, engine, url, options, mode, pdfFormats, metadata, encrypt, embedPaths, embedsMetadata, facturX, facturxXmlPath, watermark, stamp, rotateAngle, rotatePages)
err = convertUrl(ctx, chromium, engine, url, options, mode, pdfFormats, metadata, encrypt, embedPaths, embedsMetadata, facturX, facturxXmlPath, watermarks, stamps, rotateAngle, rotatePages, optimizeImages, imageQuality)
if err != nil {
return fmt.Errorf("convert HTML to PDF: %w", err)
}
@@ -643,11 +660,18 @@ func convertMarkdownRoute(chromium Api, engine gotenberg.PdfEngine) api.Route {
metadata := pdfengines.FormDataPdfMetadata(form, false)
encrypt := pdfengines.FormDataPdfEncrypt(form)
embedPaths := pdfengines.FormDataPdfEmbeds(form)
watermark := pdfengines.FormDataPdfWatermark(form, false)
watermarkFile := pdfengines.FormDataPdfWatermarkFile(form)
stamp := pdfengines.FormDataPdfStamp(form, false)
stampFile := pdfengines.FormDataPdfStampFile(form)
watermarks, wErr := pdfengines.FormDataPdfWatermarks(form)
if wErr != nil {
return fmt.Errorf("form data watermarks: %w", wErr)
}
stamps, sErr := pdfengines.FormDataPdfStamps(form)
if sErr != nil {
return fmt.Errorf("form data stamps: %w", sErr)
}
var watermarkFiles, stampFiles []string
form.Watermarks(&watermarkFiles).Stamps(&stampFiles)
rotateAngle, rotatePages := pdfengines.FormDataPdfRotate(form, false)
optimizeImages, imageQuality := pdfengines.FormDataPdfOptimize(form)
embedsMetadata := pdfengines.FormDataPdfEmbedsMetadata(form)
facturX, facturxXmlPath := pdfengines.FormDataPdfFacturX(form)
@@ -664,13 +688,13 @@ func convertMarkdownRoute(chromium Api, engine gotenberg.PdfEngine) api.Route {
return fmt.Errorf("validate form data: %w", err)
}
err = pdfengines.EnsureWatermarkFile(&watermark, watermarkFile)
err = pdfengines.BindWatermarkFiles(watermarks, watermarkFiles)
if err != nil {
return fmt.Errorf("validate watermark: %w", err)
return fmt.Errorf("bind watermark files: %w", err)
}
err = pdfengines.EnsureStampFile(&stamp, stampFile)
err = pdfengines.BindStampFiles(stamps, stampFiles)
if err != nil {
return fmt.Errorf("validate stamp: %w", err)
return fmt.Errorf("bind stamp files: %w", err)
}
url, err := markdownToHtml(ctx, inputPath, markdownPaths)
@@ -679,7 +703,7 @@ func convertMarkdownRoute(chromium Api, engine gotenberg.PdfEngine) api.Route {
}
options.AllowedFilePrefixes = []string{ctx.DirPath()}
err = convertUrl(ctx, chromium, engine, url, options, mode, pdfFormats, metadata, encrypt, embedPaths, embedsMetadata, facturX, facturxXmlPath, watermark, stamp, rotateAngle, rotatePages)
err = convertUrl(ctx, chromium, engine, url, options, mode, pdfFormats, metadata, encrypt, embedPaths, embedsMetadata, facturX, facturxXmlPath, watermarks, stamps, rotateAngle, rotatePages, optimizeImages, imageQuality)
if err != nil {
return fmt.Errorf("convert markdown to PDF: %w", err)
}
@@ -804,7 +828,7 @@ func markdownToHtml(ctx *api.Context, inputPath string, markdownPaths []string)
return fmt.Sprintf("file://%s", inputPath), nil
}
func convertUrl(ctx *api.Context, chromium Api, engine gotenberg.PdfEngine, url string, options PdfOptions, mode gotenberg.SplitMode, pdfFormats gotenberg.PdfFormats, metadata map[string]any, encrypt gotenberg.EncryptOptions, embedPaths []string, embedsMetadata map[string]map[string]string, facturX gotenberg.FacturX, facturxXmlPath string, watermark, stamp gotenberg.Stamp, rotateAngle int, rotatePages string) error {
func convertUrl(ctx *api.Context, chromium Api, engine gotenberg.PdfEngine, url string, options PdfOptions, mode gotenberg.SplitMode, pdfFormats gotenberg.PdfFormats, metadata map[string]any, encrypt gotenberg.EncryptOptions, embedPaths []string, embedsMetadata map[string]map[string]string, facturX gotenberg.FacturX, facturxXmlPath string, watermarks, stamps []gotenberg.Stamp, rotateAngle int, rotatePages string, optimizeImages bool, imageQuality int) error {
outputPath := ctx.GeneratePath(".pdf")
// See https://github.com/gotenberg/gotenberg/issues/1130.
filename := ctx.OutputFilename(outputPath)
@@ -886,12 +910,12 @@ func convertUrl(ctx *api.Context, chromium Api, engine gotenberg.PdfEngine, url
return fmt.Errorf("split PDF: %w", err)
}
err = pdfengines.WatermarkStub(ctx, engine, watermark, outputPaths)
err = pdfengines.WatermarkStub(ctx, engine, watermarks, outputPaths)
if err != nil {
return fmt.Errorf("watermark PDFs: %w", err)
}
err = pdfengines.StampStub(ctx, engine, stamp, outputPaths)
err = pdfengines.StampStub(ctx, engine, stamps, outputPaths)
if err != nil {
return fmt.Errorf("stamp PDFs: %w", err)
}
@@ -901,6 +925,11 @@ func convertUrl(ctx *api.Context, chromium Api, engine gotenberg.PdfEngine, url
return fmt.Errorf("rotate PDFs: %w", err)
}
err = pdfengines.OptimizeStub(ctx, engine, optimizeImages, imageQuality, outputPaths)
if err != nil {
return fmt.Errorf("optimize PDF images: %w", err)
}
pdfFormats = pdfengines.FacturXPdfFormats(ctx, engine, facturX, pdfFormats, true, nil)
convertOutputPaths, err := pdfengines.ConvertStub(ctx, engine, pdfFormats, outputPaths)
@@ -963,7 +992,17 @@ func screenshotUrl(ctx *api.Context, chromium Api, url string, options Screensho
outputPath := ctx.GeneratePath(ext)
err := chromium.Screenshot(ctx, ctx.Log(), url, outputPath, options)
err = handleChromiumError(err, options.Options)
if errors.Is(err, ErrScreenshotSelectorNotFound) {
err = api.WrapError(
err,
api.NewSentinelHttpError(
http.StatusBadRequest,
fmt.Sprintf("The selector '%s' (selector) matched no element with a visible box", options.Selector),
),
)
} else {
err = handleChromiumError(err, options.Options)
}
if err != nil {
return fmt.Errorf("screenshot: %w", err)
}
@@ -981,6 +1020,16 @@ func handleChromiumError(err error, options Options) error {
return nil
}
if errors.Is(err, ErrChromiumCrashed) {
return api.WrapError(
err,
api.NewSentinelHttpError(
http.StatusServiceUnavailable,
"Chromium crashed while processing the request. Retry, or reduce the workload if the problem persists.",
),
)
}
if errors.Is(err, ErrInvalidEvaluationExpression) {
if options.WaitForExpression == "" {
// We do not expect the 'waitWindowStatus' form field to return

View File

@@ -0,0 +1,51 @@
package chromium
import (
"fmt"
"net/http"
"testing"
"github.com/gotenberg/gotenberg/v8/pkg/modules/api"
)
// TestHandleChromiumError_Crashed pins the mapping of a Chromium renderer
// crash to a 503 Service Unavailable. When the renderer crashes mid-conversion,
// the request must fail fast with 503 rather than hang until the deadline and
// surface as a generic timeout.
// See https://github.com/gotenberg/gotenberg/issues/1640.
func TestHandleChromiumError_Crashed(t *testing.T) {
// Mirror the wrapping done by [chromiumBrowser.do].
err := handleChromiumError(fmt.Errorf("handle tasks: %w", ErrChromiumCrashed), Options{})
if err == nil {
t.Fatal("expected an error, got none")
}
status, message := api.ParseError(err)
if status != http.StatusServiceUnavailable {
t.Errorf("status = %d, want %d (message: %s)", status, http.StatusServiceUnavailable, message)
}
want := "Chromium crashed while processing the request. Retry, or reduce the workload if the problem persists."
if message != want {
t.Errorf("message = %q, want %q", message, want)
}
}
// TestHandleChromiumError_CrashedTakesPrecedence guards the ordering in
// [handleChromiumError]: a crash is a server-side failure and must map to 503
// even when the error chain also carries a marker that another branch would
// map to a client-error status.
func TestHandleChromiumError_CrashedTakesPrecedence(t *testing.T) {
err := handleChromiumError(
fmt.Errorf("handle tasks: %w; %w", ErrChromiumCrashed, ErrInvalidHttpStatusCode),
Options{},
)
if err == nil {
t.Fatal("expected an error, got none")
}
status, _ := api.ParseError(err)
if status != http.StatusServiceUnavailable {
t.Errorf("status = %d, want %d", status, http.StatusServiceUnavailable)
}
}

View File

@@ -6,13 +6,17 @@ import (
"errors"
"fmt"
"log/slog"
"net/url"
"os"
"strconv"
"time"
"github.com/chromedp/cdproto/cdp"
"github.com/chromedp/cdproto/emulation"
"github.com/chromedp/cdproto/network"
"github.com/chromedp/cdproto/page"
"github.com/chromedp/cdproto/runtime"
"github.com/chromedp/cdproto/storage"
"github.com/chromedp/chromedp"
"go.opentelemetry.io/otel/attribute"
"go.opentelemetry.io/otel/codes"
@@ -54,6 +58,7 @@ func printToPdfActionFunc(reqCtx context.Context, logger *slog.Logger, outputPat
defer span.End()
err := func() error {
paperWidth := options.PaperWidth
paperHeight := options.PaperHeight
pageRanges := options.PageRanges
@@ -67,9 +72,20 @@ func printToPdfActionFunc(reqCtx context.Context, logger *slog.Logger, outputPat
// There are 96 CSS pixels per inch.
// See https://issues.chromium.org/issues/40267771#comment14.
// We add top and bottom margins so that the content area
// is large enough to fit the entire content.
paperHeight = (cssContentSize.Height / 96) + options.MarginTop + options.MarginBottom
if options.Landscape {
// Landscape swaps the paper dimensions, so the page is
// WithPaperHeight wide by WithPaperWidth tall. Size both to
// the content so the width expands to fit a wide document
// (e.g. a table) instead of the height-only expansion
// landing on the width axis and truncating it.
// See https://github.com/gotenberg/gotenberg/issues/1390.
paperWidth = (cssContentSize.Height / 96) + options.MarginTop + options.MarginBottom
paperHeight = (cssContentSize.Width / 96) + options.MarginLeft + options.MarginRight
} else {
// We add top and bottom margins so that the content area
// is large enough to fit the entire content.
paperHeight = (cssContentSize.Height / 96) + options.MarginTop + options.MarginBottom
}
pageRanges = "1" // little dirty hack to avoid leftovers.
}
@@ -78,7 +94,7 @@ func printToPdfActionFunc(reqCtx context.Context, logger *slog.Logger, outputPat
WithLandscape(options.Landscape).
WithPrintBackground(options.PrintBackground).
WithScale(options.Scale).
WithPaperWidth(options.PaperWidth).
WithPaperWidth(paperWidth).
WithPaperHeight(paperHeight).
WithMarginTop(options.MarginTop).
WithMarginBottom(options.MarginBottom).
@@ -188,7 +204,16 @@ func captureScreenshotActionFunc(logger *slog.Logger, outputPath string, options
WithOptimizeForSpeed(options.OptimizeForSpeed).
WithFormat(page.CaptureScreenshotFormat(options.Format))
if options.Clip {
switch {
case options.Selector != "":
clip, err := elementClip(ctx, options.Selector)
if err != nil {
return err
}
logger.DebugContext(ctx, fmt.Sprintf("clip screenshot to selector '%s'", options.Selector))
captureScreenshot = captureScreenshot.WithClip(clip)
case options.Clip:
captureScreenshot = captureScreenshot.WithClip(&page.Viewport{
Width: float64(options.Width),
Height: float64(options.Height),
@@ -229,6 +254,49 @@ func captureScreenshotActionFunc(logger *slog.Logger, outputPath string, options
}
}
// elementClip resolves the first element matching selector to a page-space clip
// rectangle for Page.captureScreenshot.
//
// getBoundingClientRect reports viewport-relative CSS pixels; adding the scroll
// offset puts the rectangle in the document coordinate space that
// WithCaptureBeyondViewport expects. It fails with
// [ErrScreenshotSelectorNotFound] when nothing matches or the match has no
// rendered box (display:none or a zero area), so the caller can answer 400.
func elementClip(ctx context.Context, selector string) (*page.Viewport, error) {
var rect struct {
Found bool `json:"found"`
X float64 `json:"x"`
Y float64 `json:"y"`
Width float64 `json:"width"`
Height float64 `json:"height"`
}
expr := fmt.Sprintf(`(() => {
const el = document.querySelector(%s);
if (!el) {
return { found: false };
}
const r = el.getBoundingClientRect();
return { found: true, x: r.left + window.scrollX, y: r.top + window.scrollY, width: r.width, height: r.height };
})()`, strconv.Quote(selector))
err := chromedp.Evaluate(expr, &rect).Do(ctx)
if err != nil {
return nil, fmt.Errorf("evaluate selector box: %v: %w", err, ErrScreenshotSelectorNotFound)
}
if !rect.Found || rect.Width <= 0 || rect.Height <= 0 {
return nil, fmt.Errorf("selector %q matched no element with a visible box: %w", selector, ErrScreenshotSelectorNotFound)
}
return &page.Viewport{
X: rect.X,
Y: rect.Y,
Width: rect.Width,
Height: rect.Height,
Scale: 1,
}, nil
}
func setDeviceMetricsOverride(logger *slog.Logger, width, height int, deviceScaleFactor float64) chromedp.ActionFunc {
return func(ctx context.Context) error {
logger.DebugContext(ctx, "set device metrics override")
@@ -280,6 +348,54 @@ func clearCookiesActionFunc(logger *slog.Logger, clear bool) chromedp.ActionFunc
}
}
// clearStorageActionFunc clears the converted origin's local storage before the
// page loads, so state written by a previous conversion of the same origin does
// not leak into this one. See https://github.com/gotenberg/gotenberg/issues/919.
//
// Session storage is not touched: each conversion runs in its own browsing
// context (a fresh tab), so it is already isolated and cannot leak. Local
// storage is per-origin and shared across tabs of the long-lived browser, so it
// is the only web storage that carries over.
func clearStorageActionFunc(logger *slog.Logger, clear bool, rawURL string) chromedp.ActionFunc {
return func(ctx context.Context) error {
if !clear {
logger.DebugContext(ctx, "local storage not cleared")
return nil
}
origin, ok := httpOrigin(rawURL)
if !ok {
// A file:// upload gets an opaque, per-request origin that is not
// shared between conversions, so there is nothing to clear.
logger.DebugContext(ctx, "local storage not cleared: non-http(s) origin is already isolated")
return nil
}
logger.DebugContext(ctx, fmt.Sprintf("clear local storage for %s", origin))
err := storage.ClearDataForOrigin(origin, string(storage.TypeLocalStorage)).Do(ctx)
if err == nil {
return nil
}
return fmt.Errorf("clear local storage: %w", err)
}
}
// httpOrigin returns the http(s) security origin (scheme://host[:port]) of
// rawURL, and false when rawURL is not http(s). A non-http(s) URL such as a
// file:// upload has an opaque origin that no other conversion shares.
func httpOrigin(rawURL string) (string, bool) {
parsed, err := url.Parse(rawURL)
if err != nil {
return "", false
}
if parsed.Scheme != "http" && parsed.Scheme != "https" {
return "", false
}
return fmt.Sprintf("%s://%s", parsed.Scheme, parsed.Host), true
}
func disableJavaScriptActionFunc(logger *slog.Logger, disable bool) chromedp.ActionFunc {
return func(ctx context.Context) error {
// See https://github.com/gotenberg/gotenberg/issues/175.
@@ -593,7 +709,13 @@ func waitForExpressionBeforePrintActionFunc(logger *slog.Logger, disableJavaScri
return fmt.Errorf("context done while evaluating '%s': %w", expression, ctx.Err())
case <-ticker.C:
var ok bool
evaluate := chromedp.Evaluate(expression, &ok)
// Await the result so a thenable expression (an async function
// returning a Promise) resolves before its value is read. A
// non-promise result is unaffected.
// See https://github.com/gotenberg/gotenberg/pull/1617.
evaluate := chromedp.Evaluate(expression, &ok, func(p *runtime.EvaluateParams) *runtime.EvaluateParams {
return p.WithAwaitPromise(true)
})
err := evaluate.Do(ctx)
if err != nil {

View File

@@ -80,6 +80,37 @@ var dangerousTags = []string{
"FilePermissions", // Writing this changes the file's permissions
}
// controlOptions lists ExifTool command-line option names that collide with a
// tag assignment. A metadata key of "csv" becomes the argv entry "-csv=value",
// which exiftool reads as its own option rather than as a tag, so the value
// becomes a filename exiftool opens. Only an unprefixed key can collide:
// "-XMP:csv=value" is unambiguously a tag.
//
// See https://exiftool.org/exiftool_pod.html.
var controlOptions = []string{
"api", "argfile", "charset", "common_args", "config", "csv", "diff",
"echo", "efile", "execute", "ext", "fileorder", "geotag", "geosync",
"htmldump", "if", "json", "lang", "listitem", "o", "out", "p", "php",
"require", "srcfile", "stay_open", "tagsfromfile", "textout", "use", "w",
"wm", "xmlformat",
}
// isControlOption reports whether an unprefixed key would reach exiftool as
// one of its own options instead of as a tag assignment.
func isControlOption(key string) bool {
if strings.Contains(key, ":") {
return false
}
for _, option := range controlOptions {
if strings.EqualFold(key, option) {
return true
}
}
return false
}
// isDangerousTag reports whether key matches one of the [dangerousTags]
// after case-insensitive comparison with any group prefix stripped.
func isDangerousTag(key string) bool {
@@ -114,19 +145,33 @@ func buildExifToolWriteArgs(metadata map[string]any) ([]string, error) {
if !safeKeyPattern.MatchString(key) {
return nil, fmt.Errorf("write PDF metadata with ExifTool: invalid metadata key %q: %w", key, gotenberg.ErrPdfEngineMetadataValueNotSupported)
}
if isControlOption(key) {
return nil, fmt.Errorf("write PDF metadata with ExifTool: metadata key %q is an ExifTool option, prefix it with a group such as %q: %w", key, "XMP:"+key, gotenberg.ErrPdfEngineMetadataValueNotSupported)
}
tag := key
if key == "Trapped" {
// ExifTool writes the document info /Trapped entry as a malformed
// name-in-a-string, e.g. "(/Unknown)". pdfcpu's stricter validation
// (as of v0.15) rejects it, which breaks later pdfcpu operations on
// the file such as embedding. Writing Trapped to XMP keeps the value
// readable without the invalid document info entry.
// See https://github.com/gotenberg/gotenberg/issues/1628.
tag = "XMP-pdf:Trapped"
}
switch val := value.(type) {
case string:
if err := validateMetadataValue(key, val); err != nil {
return nil, err
}
args = append(args, fmt.Sprintf("-%s=%s", key, val))
args = append(args, fmt.Sprintf("-%s=%s", tag, val))
case []string:
for _, s := range val {
if err := validateMetadataValue(key, s); err != nil {
return nil, err
}
args = append(args, fmt.Sprintf("-%s=%s", key, s))
args = append(args, fmt.Sprintf("-%s=%s", tag, s))
}
case []any:
// See https://github.com/gotenberg/gotenberg/issues/1048.
@@ -138,18 +183,18 @@ func buildExifToolWriteArgs(metadata map[string]any) ([]string, error) {
if err := validateMetadataValue(key, s); err != nil {
return nil, err
}
args = append(args, fmt.Sprintf("-%s=%s", key, s))
args = append(args, fmt.Sprintf("-%s=%s", tag, s))
}
case bool:
args = append(args, fmt.Sprintf("-%s=%t", key, val))
args = append(args, fmt.Sprintf("-%s=%t", tag, val))
case int:
args = append(args, fmt.Sprintf("-%s=%d", key, val))
args = append(args, fmt.Sprintf("-%s=%d", tag, val))
case int64:
args = append(args, fmt.Sprintf("-%s=%d", key, val))
args = append(args, fmt.Sprintf("-%s=%d", tag, val))
case float32:
args = append(args, fmt.Sprintf("-%s=%g", key, val))
args = append(args, fmt.Sprintf("-%s=%g", tag, val))
case float64:
args = append(args, fmt.Sprintf("-%s=%g", key, val))
args = append(args, fmt.Sprintf("-%s=%g", tag, val))
default:
return nil, fmt.Errorf("write PDF metadata with ExifTool: unsupported type %T for key %q: %w", value, key, gotenberg.ErrPdfEngineMetadataValueNotSupported)
}
@@ -296,6 +341,20 @@ func (engine *ExifTool) Convert(ctx context.Context, logger *slog.Logger, format
return err
}
// OptimizeImages is not available in this implementation.
func (engine *ExifTool) OptimizeImages(ctx context.Context, logger *slog.Logger, imageQuality int, inputPath string) error {
_, span := gotenberg.Tracer().Start(ctx, "exiftool.OptimizeImages",
trace.WithSpanKind(trace.SpanKindClient),
trace.WithAttributes(engine.spanAttrs()...),
)
defer span.End()
err := fmt.Errorf("optimize PDF images with ExifTool: %w", gotenberg.ErrPdfEngineMethodNotSupported)
span.RecordError(err)
span.SetStatus(codes.Error, err.Error())
return err
}
// ReadMetadata extracts the metadata of a given PDF file by invoking
// the exiftool binary with "-j" (JSON output) and parsing the result.
func (engine *ExifTool) ReadMetadata(ctx context.Context, logger *slog.Logger, inputPath string) (map[string]any, error) {

View File

@@ -2,6 +2,7 @@ package exiftool
import (
"errors"
"fmt"
"slices"
"testing"
@@ -211,3 +212,48 @@ func TestSafeKeyPattern(t *testing.T) {
}
}
}
// A metadata key that collides with an ExifTool option becomes a bare argv
// entry such as "-csv=/etc/passwd", which exiftool reads as its own option and
// treats the value as a filename to open.
func TestBuildExifToolWriteArgs_RejectsControlOptions(t *testing.T) {
for _, key := range []string{"csv", "CSV", "json", "geotag", "config", "tagsFromFile", "execute", "stay_open", "o", "w", "if", "p"} {
t.Run(key, func(t *testing.T) {
_, err := buildExifToolWriteArgs(map[string]any{key: "/etc/passwd"})
if err == nil {
t.Fatalf("buildExifToolWriteArgs accepted the control option %q", key)
}
if !errors.Is(err, gotenberg.ErrPdfEngineMetadataValueNotSupported) {
t.Fatalf("error %v does not wrap ErrPdfEngineMetadataValueNotSupported", err)
}
})
}
}
// A group prefix makes the key unambiguous, so it must still be accepted.
func TestBuildExifToolWriteArgs_AcceptsPrefixedOptionNames(t *testing.T) {
for _, key := range []string{"XMP:csv", "XMP-dc:json", "IPTC:p"} {
t.Run(key, func(t *testing.T) {
args, err := buildExifToolWriteArgs(map[string]any{key: "value"})
if err != nil {
t.Fatalf("buildExifToolWriteArgs rejected the prefixed key %q: %v", key, err)
}
want := fmt.Sprintf("-%s=value", key)
if len(args) != 1 || args[0] != want {
t.Fatalf("args = %v, want [%s]", args, want)
}
})
}
}
// Ordinary tags must be unaffected.
func TestBuildExifToolWriteArgs_AcceptsOrdinaryTags(t *testing.T) {
for _, key := range []string{"Author", "Title", "Subject", "Keywords", "Producer", "Creator"} {
t.Run(key, func(t *testing.T) {
_, err := buildExifToolWriteArgs(map[string]any{key: "value"})
if err != nil {
t.Fatalf("buildExifToolWriteArgs rejected the ordinary tag %q: %v", key, err)
}
})
}
}

View File

@@ -353,7 +353,7 @@ func (a *Api) Descriptor() gotenberg.ModuleDescriptor {
fs.Duration("libreoffice-idle-shutdown-timeout", 0, "Shutdown LibreOffice after being idle for the given duration. Set to 0 to disable this feature")
fs.Bool("libreoffice-auto-start", false, "Automatically launch LibreOffice upon initialization if set to true; otherwise, LibreOffice will start at the time of the first conversion")
fs.Duration("libreoffice-start-timeout", time.Duration(20)*time.Second, "Maximum duration to wait for LibreOffice to start or restart")
fs.StringSlice("libreoffice-allow-list", []string{}, "Set the allowed URLs for LibreOffice outbound fetches (embedded images, linked content) using regular expressions - supports multiple values")
fs.StringSlice("libreoffice-allow-list", []string{}, `Set the allowed URLs for LibreOffice outbound fetches (embedded images, linked content) using regular expressions - supports multiple values. A match bypasses --libreoffice-deny-private-ips (LIBREOFFICE_DENY_PRIVATE_IPS) and --libreoffice-deny-public-ips (LIBREOFFICE_DENY_PUBLIC_IPS), so terminate the host or the pattern also matches suffix hosts, for example ^https?://internal\.svc(:|/|$)`)
fs.StringSlice("libreoffice-deny-list", []string{}, "Set the denied URLs for LibreOffice outbound fetches using regular expressions - supports multiple values")
fs.Bool("libreoffice-deny-private-ips", false, "Reject LibreOffice outbound URLs whose host resolves to a non-public IP address (loopback, RFC1918, link-local, unique-local). Enable on deployments that accept untrusted documents to mitigate SSRF against internal services")
fs.Bool("libreoffice-deny-public-ips", false, "Reject LibreOffice outbound URLs whose host resolves to a public IP address. Enable on air-gapped or data-governed deployments to prevent outbound traffic from leaving a private network")
@@ -879,6 +879,8 @@ func (a *Api) Extensions() []string {
".potx",
".ppm",
".pps",
".ppsm",
".ppsx",
".ppt",
".pptm",
".pptx",

View File

@@ -68,6 +68,59 @@ func (p *libreOfficeProcess) Start(logger *slog.Logger) error {
userProfileDirPath := p.fs.NewDirPath()
var (
cmd *gotenberg.Cmd
success bool
)
// Registered here, right after the proxy starts listening, so that every
// failure below tears it down. A return between the proxy start and this
// point strands it: its listener stays bound, its Serve goroutine and HTTP
// client stay alive, and p.proxy is only assigned on success, so nothing
// could ever reach it to stop it. exec.Cmd.Start fails precisely under fd
// or memory pressure, and the supervisor retries the launch on the next
// request, so each stranded proxy compounds the condition that caused it.
defer func() {
if success {
p.cfgMu.Lock()
defer p.cfgMu.Unlock()
p.socketPort = port
p.userProfileDirPath = userProfileDirPath
p.cmd = cmd
p.proxy = proxy
p.isStarted.Store(true)
return
}
// LibreOffice failed to start; tear the proxy down too.
stopErr := proxy.Stop(context.Background())
if stopErr != nil {
logger.WarnContext(context.Background(), fmt.Sprintf("stop LibreOffice outbound proxy after failed start: %s", stopErr))
}
// Let's make sure the process is killed. It is nil when the failure
// happened before the command was built.
if cmd != nil {
killErr := cmd.Kill()
if killErr != nil {
logger.DebugContext(context.Background(), fmt.Sprintf("kill LibreOffice process: %v", killErr))
}
}
// And the user profile directory is deleted. It may never have been
// created, which RemoveAll reports as success.
removeErr := os.RemoveAll(userProfileDirPath)
if removeErr != nil {
logger.ErrorContext(context.Background(), fmt.Sprintf("remove LibreOffice's user profile directory: %v", removeErr))
return
}
logger.DebugContext(context.Background(), fmt.Sprintf("'%s' LibreOffice's user profile directory removed", userProfileDirPath))
}()
// LibreOffice fetches external content (OOXML images via
// TargetMode=External, RTF INCLUDEPICTURE, ODT linked images) inside
// its own libcurl. The profile config routes those fetches through the
@@ -75,7 +128,6 @@ func (p *libreOfficeProcess) Start(logger *slog.Logger) error {
// blocks content linked from untrusted locations so absolute-path
// (file://) and direct fetches are dropped at the source.
if err := writeSofficeProfileConfig(userProfileDirPath, proxy.Addr()); err != nil {
_ = proxy.Stop(context.Background())
return fmt.Errorf("write soffice profile config: %w", err)
}
sofficeEnv := sofficeProxyEnv(os.Environ(), proxy.Addr())
@@ -95,9 +147,8 @@ func (p *libreOfficeProcess) Start(logger *slog.Logger) error {
ctx, cancel := context.WithTimeout(context.Background(), p.arguments.startTimeout)
defer cancel()
cmd, err := gotenberg.CommandContext(ctx, logger, p.arguments.binPath, args...)
cmd, err = gotenberg.CommandContext(ctx, logger, p.arguments.binPath, args...)
if err != nil {
_ = proxy.Stop(context.Background())
return fmt.Errorf("create LibreOffice command: %w", err)
}
cmd.SetEnv(sofficeEnv)
@@ -106,7 +157,6 @@ func (p *libreOfficeProcess) Start(logger *slog.Logger) error {
// able to run as a daemon.
exitCode, err := cmd.Exec()
if err != nil && exitCode != 81 {
_ = proxy.Stop(context.Background())
return fmt.Errorf("execute LibreOffice: %w", err)
}
@@ -155,43 +205,6 @@ func (p *libreOfficeProcess) Start(logger *slog.Logger) error {
}
}()
var success bool
defer func() {
if success {
p.cfgMu.Lock()
defer p.cfgMu.Unlock()
p.socketPort = port
p.userProfileDirPath = userProfileDirPath
p.cmd = cmd
p.proxy = proxy
p.isStarted.Store(true)
return
}
// LibreOffice failed to start; tear the proxy down too.
stopErr := proxy.Stop(context.Background())
if stopErr != nil {
logger.WarnContext(context.Background(), fmt.Sprintf("stop LibreOffice outbound proxy after failed start: %s", stopErr))
}
// Let's make sure the process is killed.
err = cmd.Kill()
if err != nil {
logger.DebugContext(context.Background(), fmt.Sprintf("kill LibreOffice process: %v", err))
}
// And the user profile directory is deleted.
err = os.RemoveAll(userProfileDirPath)
if err != nil {
logger.ErrorContext(context.Background(), fmt.Sprintf("remove LibreOffice's user profile directory: %v", err))
}
logger.DebugContext(context.Background(), fmt.Sprintf("'%s' LibreOffice's user profile directory removed", userProfileDirPath))
}()
logger.DebugContext(context.Background(), "waiting for the LibreOffice socket to be available...")
for {
@@ -293,6 +306,14 @@ func (p *libreOfficeProcess) pdf(ctx context.Context, logger *slog.Logger, input
return errors.New("LibreOffice not started, cannot handle PDF conversion")
}
// SinglePageSheets starts each sheet's single page at the workbook's saved
// scroll position, truncating everything above and to the left of it.
// Render a copy with that position reset to the top-left cell instead.
// See https://github.com/gotenberg/gotenberg/issues/1222.
if options.SinglePageSheets {
inputPath = resetCalcScrollPosition(ctx, logger, inputPath)
}
args := []string{
"--no-launch",
"--format",

View File

@@ -35,12 +35,13 @@ var (
zipMagic = []byte{0x50, 0x4b, 0x03, 0x04}
// An unencrypted OOXML document is always a ZIP package, so any of these
// extensions over a compound file means the payload is encrypted. Legacy
// binary formats (.doc, .xls, .ppt) are compound files either way and are
// deliberately absent.
// extensions over a compound file means the payload is encrypted. A .xlsb
// workbook stores binary parts inside that same ZIP package, so it belongs
// here too. Legacy binary formats (.doc, .xls, .ppt) are compound files
// either way and are deliberately absent.
ooxmlExtensions = map[string]struct{}{
".docx": {}, ".docm": {}, ".dotx": {}, ".dotm": {},
".xlsx": {}, ".xlsm": {}, ".xltx": {}, ".xltm": {},
".xlsx": {}, ".xlsm": {}, ".xltx": {}, ".xltm": {}, ".xlsb": {},
".pptx": {}, ".pptm": {}, ".potx": {}, ".potm": {},
".ppsx": {}, ".ppsm": {},
}

View File

@@ -76,6 +76,11 @@ func TestDetectPasswordProtection(t *testing.T) {
path: ole2("encrypted.xlsx"),
want: PasswordProtectionRequired,
},
{
name: "encrypted binary workbook",
path: ole2("encrypted.xlsb"),
want: PasswordProtectionRequired,
},
{
name: "legacy binary document is inconclusive",
path: ole2("legacy.doc"),

View File

@@ -361,7 +361,7 @@ func TestSofficeProxyEnv_OverridesExisting(t *testing.T) {
// Old proxy values must be gone, not duplicated. Count exact-case keys.
counts := map[string]int{}
for _, kv := range out {
key := strings.SplitN(kv, "=", 2)[0]
key, _, _ := strings.Cut(kv, "=")
counts[key]++
}
for _, key := range []string{"http_proxy", "HTTP_PROXY", "https_proxy", "HTTPS_PROXY", "no_proxy", "NO_PROXY"} {

View File

@@ -0,0 +1,230 @@
package api
import (
"archive/zip"
"bytes"
"context"
"fmt"
"io"
"log/slog"
"os"
"path/filepath"
"regexp"
"strings"
)
// topLeftCellAttr matches the topLeftCell attribute that an OOXML worksheet
// uses (on <sheetView> and, for frozen panes, <pane>) to store the cell that
// was at the top-left of the window when the workbook was saved.
var topLeftCellAttr = regexp.MustCompile(` topLeftCell="[^"]*"`)
// maxDecompressedWorksheet bounds how much a single worksheet may decompress to
// while rewriting it. It guards against a decompression bomb and keeps memory
// predictable. A worksheet larger than this is left untouched, so a pathological
// workbook falls back to the original file rather than being rewritten.
const maxDecompressedWorksheet = 128 << 20 // 128 MiB
// resetCalcScrollPosition returns a path to a copy of inputPath whose worksheet
// scroll positions have been reset to the top-left cell, or inputPath unchanged
// when the reset does not apply or cannot be performed safely.
//
// LibreOffice's SinglePageSheets export starts each single page at the sheet's
// saved topLeftCell, dropping every row and column above and to the left of it.
// A workbook saved scrolled away from A1 therefore renders truncated. Removing
// the attribute before the conversion makes the whole used range render.
// See https://github.com/gotenberg/gotenberg/issues/1222.
//
// The function never fails the conversion. On a non-xlsx input, a workbook that
// carries no scroll position, or any read, rewrite or validation error, it
// returns the original path so a malformed rewrite can never reach LibreOffice.
func resetCalcScrollPosition(ctx context.Context, logger *slog.Logger, inputPath string) string {
// Resolve the extension to a literal so the sanitized filename is never
// derived from the (user-controlled) upload name.
var ext string
switch strings.ToLower(filepath.Ext(inputPath)) {
case ".xlsx":
ext = ".xlsx"
case ".xlsm":
ext = ".xlsm"
default:
return inputPath
}
src, err := os.ReadFile(inputPath)
if err != nil {
logger.WarnContext(ctx, fmt.Sprintf("reset calc scroll position: read input: %s; using the original file", err))
return inputPath
}
out, changed, err := stripWorksheetScrollPosition(src)
if err != nil {
logger.WarnContext(ctx, fmt.Sprintf("reset calc scroll position: %s; using the original file", err))
return inputPath
}
if !changed {
// The common case: nothing was saved scrolled, so nothing to do.
return inputPath
}
// A rewrite that dropped, renamed or corrupted an entry must never reach
// LibreOffice; fall back to the original workbook if it does not round-trip.
if err = validateWorkbook(src, out); err != nil {
logger.WarnContext(ctx, fmt.Sprintf("reset calc scroll position: %s; using the original file", err))
return inputPath
}
// Write the sanitized copy alongside the input, inside the request working
// directory that LibreOffice already reads from. The pattern is constant,
// so the resulting name carries no user-controlled path component.
dst, err := os.CreateTemp(filepath.Dir(inputPath), "singlepagesheets-*"+ext)
if err != nil {
logger.WarnContext(ctx, fmt.Sprintf("reset calc scroll position: create sanitized file: %s; using the original file", err))
return inputPath
}
defer dst.Close()
_, err = dst.Write(out)
if err != nil {
_ = os.Remove(dst.Name())
logger.WarnContext(ctx, fmt.Sprintf("reset calc scroll position: write sanitized file: %s; using the original file", err))
return inputPath
}
logger.DebugContext(ctx, "reset calc scroll position: cleared worksheet topLeftCell for SinglePageSheets export")
return dst.Name()
}
// stripWorksheetScrollPosition rewrites the worksheet XML entries of an xlsx
// workbook, removing the topLeftCell attribute, and reports whether anything
// changed. Every non-worksheet entry, and every worksheet that does not carry
// the attribute, is copied byte-for-byte without recompression.
func stripWorksheetScrollPosition(src []byte) ([]byte, bool, error) {
reader, err := zip.NewReader(bytes.NewReader(src), int64(len(src)))
if err != nil {
return nil, false, fmt.Errorf("open workbook: %w", err)
}
var buf bytes.Buffer
writer := zip.NewWriter(&buf)
changed := false
for _, file := range reader.File {
rewritten, ok, err := rewriteWorksheet(file)
if err != nil {
return nil, false, err
}
if ok {
// Recompress only the worksheets that actually changed.
header := file.FileHeader
header.Method = zip.Deflate
w, err := writer.CreateHeader(&header)
if err != nil {
return nil, false, fmt.Errorf("write worksheet %q: %w", file.Name, err)
}
_, err = w.Write(rewritten)
if err != nil {
return nil, false, fmt.Errorf("write worksheet %q: %w", file.Name, err)
}
changed = true
continue
}
err = copyZipEntry(writer, file)
if err != nil {
return nil, false, err
}
}
err = writer.Close()
if err != nil {
return nil, false, fmt.Errorf("finalize workbook: %w", err)
}
if !changed {
return nil, false, nil
}
return buf.Bytes(), true, nil
}
// rewriteWorksheet returns file's contents with topLeftCell removed, and
// whether file is a worksheet that carried the attribute. A worksheet without
// the attribute, or any other entry, returns ok false so the caller copies it
// verbatim.
func rewriteWorksheet(file *zip.File) ([]byte, bool, error) {
if !strings.HasPrefix(file.Name, "xl/worksheets/") || !strings.HasSuffix(strings.ToLower(file.Name), ".xml") {
return nil, false, nil
}
rc, err := file.Open()
if err != nil {
return nil, false, fmt.Errorf("open worksheet %q: %w", file.Name, err)
}
defer rc.Close()
// Read at most maxDecompressedWorksheet+1 bytes so a decompression bomb
// cannot exhaust memory; a genuine overflow aborts the rewrite.
data, err := io.ReadAll(io.LimitReader(rc, maxDecompressedWorksheet+1))
if err != nil {
return nil, false, fmt.Errorf("read worksheet %q: %w", file.Name, err)
}
if len(data) > maxDecompressedWorksheet {
return nil, false, fmt.Errorf("worksheet %q exceeds %d bytes", file.Name, maxDecompressedWorksheet)
}
if !bytes.Contains(data, []byte("topLeftCell")) {
return nil, false, nil
}
return topLeftCellAttr.ReplaceAll(data, nil), true, nil
}
// copyZipEntry writes file into writer without decompressing and recompressing
// it, preserving its exact bytes.
func copyZipEntry(writer *zip.Writer, file *zip.File) error {
w, err := writer.CreateRaw(&file.FileHeader)
if err != nil {
return fmt.Errorf("copy entry %q: %w", file.Name, err)
}
rc, err := file.OpenRaw()
if err != nil {
return fmt.Errorf("open entry %q: %w", file.Name, err)
}
_, err = io.Copy(w, rc)
if err != nil {
return fmt.Errorf("copy entry %q: %w", file.Name, err)
}
return nil
}
// validateWorkbook checks that out reopens as a zip holding exactly the same
// entry names as src, rejecting a rewrite that lost, renamed or added an entry
// or produced a broken central directory. The entry payloads themselves are
// not re-read: unchanged entries are copied byte-for-byte from a workbook that
// already parsed, and rewritten worksheets are produced by the standard library
// writer, so re-decompressing everything would only add a decompression-bomb
// surface without catching a failure this transform can introduce.
func validateWorkbook(src, out []byte) error {
original, err := zip.NewReader(bytes.NewReader(src), int64(len(src)))
if err != nil {
return fmt.Errorf("reopen original workbook: %w", err)
}
rewritten, err := zip.NewReader(bytes.NewReader(out), int64(len(out)))
if err != nil {
return fmt.Errorf("reopen rewritten workbook: %w", err)
}
if len(rewritten.File) != len(original.File) {
return fmt.Errorf("entry count changed from %d to %d", len(original.File), len(rewritten.File))
}
names := make(map[string]struct{}, len(original.File))
for _, file := range original.File {
names[file.Name] = struct{}{}
}
for _, file := range rewritten.File {
_, ok := names[file.Name]
if !ok {
return fmt.Errorf("unexpected entry %q", file.Name)
}
}
return nil
}

View File

@@ -0,0 +1,216 @@
package api
import (
"archive/zip"
"bytes"
"context"
"io"
"log/slog"
"os"
"path/filepath"
"strings"
"testing"
)
// buildWorkbook packs entries into an in-memory xlsx-like zip.
func buildWorkbook(t *testing.T, entries map[string]string) []byte {
t.Helper()
var buf bytes.Buffer
w := zip.NewWriter(&buf)
for name, content := range entries {
f, err := w.Create(name)
if err != nil {
t.Fatalf("create entry %q: %v", name, err)
}
_, err = f.Write([]byte(content))
if err != nil {
t.Fatalf("write entry %q: %v", name, err)
}
}
err := w.Close()
if err != nil {
t.Fatalf("close workbook: %v", err)
}
return buf.Bytes()
}
func readEntry(t *testing.T, workbook []byte, name string) string {
t.Helper()
r, err := zip.NewReader(bytes.NewReader(workbook), int64(len(workbook)))
if err != nil {
t.Fatalf("open workbook: %v", err)
}
for _, f := range r.File {
if f.Name != name {
continue
}
rc, err := f.Open()
if err != nil {
t.Fatalf("open entry %q: %v", name, err)
}
defer rc.Close()
data, err := io.ReadAll(rc)
if err != nil {
t.Fatalf("read entry %q: %v", name, err)
}
return string(data)
}
t.Fatalf("entry %q not found", name)
return ""
}
const scrolledSheet = `<?xml version="1.0"?><worksheet><dimension ref="A1:B83"/>` +
`<sheetViews><sheetView tabSelected="1" topLeftCell="A37" workbookViewId="0">` +
`<pane topLeftCell="A37"/><selection activeCell="A1" sqref="A1"/></sheetView></sheetViews>` +
`<sheetData><row r="1"><c r="A1"><v>1</v></c></row></sheetData></worksheet>`
const topSheet = `<?xml version="1.0"?><worksheet><dimension ref="A1:B83"/>` +
`<sheetViews><sheetView tabSelected="1" workbookViewId="0"/></sheetViews>` +
`<sheetData><row r="1"><c r="A1"><v>1</v></c></row></sheetData></worksheet>`
func TestStripWorksheetScrollPosition(t *testing.T) {
for _, tc := range []struct {
scenario string
workbook []byte
expectErr bool
expectChange bool
}{
{
scenario: "removes topLeftCell from sheetView and pane",
workbook: buildWorkbook(t, map[string]string{
"[Content_Types].xml": "<Types/>",
"xl/worksheets/sheet1.xml": scrolledSheet,
"xl/sharedStrings.xml": "<sst/>",
}),
expectChange: true,
},
{
scenario: "leaves a workbook without a saved scroll position untouched",
workbook: buildWorkbook(t, map[string]string{
"[Content_Types].xml": "<Types/>",
"xl/worksheets/sheet1.xml": topSheet,
}),
expectChange: false,
},
{
scenario: "only rewrites worksheet entries",
workbook: buildWorkbook(t, map[string]string{
"xl/worksheets/sheet1.xml": scrolledSheet,
// A stray topLeftCell elsewhere must not be touched.
"xl/workbook.xml": `<workbook topLeftCell="A9"/>`,
}),
expectChange: true,
},
{
scenario: "rejects a non-zip input",
workbook: []byte("not a zip file"),
expectErr: true,
},
} {
t.Run(tc.scenario, func(t *testing.T) {
out, changed, err := stripWorksheetScrollPosition(tc.workbook)
if tc.expectErr {
if err == nil {
t.Fatalf("expected error, got nil")
}
return
}
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
if changed != tc.expectChange {
t.Fatalf("expected changed=%v, got %v", tc.expectChange, changed)
}
if !changed {
return
}
// The rewrite must round-trip and hold the same entries.
if err = validateWorkbook(tc.workbook, out); err != nil {
t.Fatalf("rewritten workbook did not validate: %v", err)
}
if strings.Contains(readEntry(t, out, "xl/worksheets/sheet1.xml"), "topLeftCell") {
t.Fatalf("worksheet still contains topLeftCell")
}
// Non-worksheet entries are copied verbatim.
if _, ok := entryNames(t, out)["xl/workbook.xml"]; ok {
if got := readEntry(t, out, "xl/workbook.xml"); got != `<workbook topLeftCell="A9"/>` {
t.Fatalf("non-worksheet entry was modified: %q", got)
}
}
})
}
}
func entryNames(t *testing.T, workbook []byte) map[string]struct{} {
t.Helper()
r, err := zip.NewReader(bytes.NewReader(workbook), int64(len(workbook)))
if err != nil {
t.Fatalf("open workbook: %v", err)
}
names := make(map[string]struct{}, len(r.File))
for _, f := range r.File {
names[f.Name] = struct{}{}
}
return names
}
func TestResetCalcScrollPosition(t *testing.T) {
logger := slog.New(slog.NewTextHandler(io.Discard, nil))
ctx := context.Background()
writeFile := func(t *testing.T, name string, content []byte) string {
t.Helper()
path := filepath.Join(t.TempDir(), name)
err := os.WriteFile(path, content, 0o600)
if err != nil {
t.Fatalf("write %q: %v", name, err)
}
return path
}
t.Run("non-xlsx input is returned unchanged", func(t *testing.T) {
path := writeFile(t, "input.docx", []byte("whatever"))
if got := resetCalcScrollPosition(ctx, logger, path); got != path {
t.Fatalf("expected %q, got %q", path, got)
}
})
t.Run("workbook without a scroll position is returned unchanged", func(t *testing.T) {
path := writeFile(t, "input.xlsx", buildWorkbook(t, map[string]string{
"xl/worksheets/sheet1.xml": topSheet,
}))
if got := resetCalcScrollPosition(ctx, logger, path); got != path {
t.Fatalf("expected original path %q, got %q", path, got)
}
})
t.Run("corrupt xlsx falls back to the original path", func(t *testing.T) {
path := writeFile(t, "input.xlsx", []byte("PK\x03\x04 not really a zip"))
if got := resetCalcScrollPosition(ctx, logger, path); got != path {
t.Fatalf("expected fallback to %q, got %q", path, got)
}
})
t.Run("scrolled workbook yields a sanitized copy", func(t *testing.T) {
path := writeFile(t, "input.xlsx", buildWorkbook(t, map[string]string{
"[Content_Types].xml": "<Types/>",
"xl/worksheets/sheet1.xml": scrolledSheet,
}))
got := resetCalcScrollPosition(ctx, logger, path)
if got == path {
t.Fatalf("expected a sanitized copy, got the original path")
}
if filepath.Dir(got) != filepath.Dir(path) {
t.Fatalf("sanitized copy escaped the working directory: %q", got)
}
sanitized, err := os.ReadFile(got)
if err != nil {
t.Fatalf("read sanitized file: %v", err)
}
if strings.Contains(readEntry(t, sanitized, "xl/worksheets/sheet1.xml"), "topLeftCell") {
t.Fatalf("sanitized worksheet still contains topLeftCell")
}
})
}

View File

@@ -60,6 +60,11 @@ func (engine *LibreOfficePdfEngine) Flatten(ctx context.Context, logger *slog.Lo
return fmt.Errorf("flatten PDF with LibreOffice: %w", gotenberg.ErrPdfEngineMethodNotSupported)
}
// OptimizeImages is not available in this implementation.
func (engine *LibreOfficePdfEngine) OptimizeImages(ctx context.Context, logger *slog.Logger, imageQuality int, inputPath string) error {
return fmt.Errorf("optimize PDF images with LibreOffice: %w", gotenberg.ErrPdfEngineMethodNotSupported)
}
// Convert converts the given PDF to a specific PDF format. Currently, only the
// PDF/A-1b, PDF/A-2b, PDF/A-3b and PDF/UA formats are available. If another
// PDF format is requested, it returns a [gotenberg.ErrPdfFormatNotSupported]

View File

@@ -37,13 +37,20 @@ func convertRoute(libreOffice libreofficeapi.Uno, engine gotenberg.PdfEngine) ap
metadata := pdfengines.FormDataPdfMetadata(form, false)
encrypt := pdfengines.FormDataPdfEncrypt(form)
embedPaths := pdfengines.FormDataPdfEmbeds(form)
watermark := pdfengines.FormDataPdfWatermark(form, false)
watermarkFile := pdfengines.FormDataPdfWatermarkFile(form)
stamp := pdfengines.FormDataPdfStamp(form, false)
stampFile := pdfengines.FormDataPdfStampFile(form)
watermarks, wErr := pdfengines.FormDataPdfWatermarks(form)
if wErr != nil {
return fmt.Errorf("form data watermarks: %w", wErr)
}
stamps, sErr := pdfengines.FormDataPdfStamps(form)
if sErr != nil {
return fmt.Errorf("form data stamps: %w", sErr)
}
var watermarkFiles, stampFiles []string
form.Watermarks(&watermarkFiles).Stamps(&stampFiles)
angle, rotatePages := pdfengines.FormDataPdfRotate(form, false)
embedsMetadata := pdfengines.FormDataPdfEmbedsMetadata(form)
facturX, facturxXmlPath := pdfengines.FormDataPdfFacturX(form)
optimizeImages, imageQuality := pdfengines.FormDataPdfOptimize(form)
zeroValuedSplitMode := gotenberg.SplitMode{}
@@ -310,13 +317,13 @@ func convertRoute(libreOffice libreofficeapi.Uno, engine gotenberg.PdfEngine) ap
return fmt.Errorf("validate form data: %w", err)
}
err = pdfengines.EnsureWatermarkFile(&watermark, watermarkFile)
err = pdfengines.BindWatermarkFiles(watermarks, watermarkFiles)
if err != nil {
return fmt.Errorf("validate watermark: %w", err)
return fmt.Errorf("bind watermark files: %w", err)
}
err = pdfengines.EnsureStampFile(&stamp, stampFile)
err = pdfengines.BindStampFiles(stamps, stampFiles)
if err != nil {
return fmt.Errorf("validate stamp: %w", err)
return fmt.Errorf("bind stamp files: %w", err)
}
err = pdfengines.ValidatePdfFormatsCompat(pdfFormats, encrypt.UserPassword, embedPaths)
@@ -338,7 +345,7 @@ func convertRoute(libreOffice libreofficeapi.Uno, engine gotenberg.PdfEngine) ap
// requested. The conversion runs as a post-processing step below.
pdfFormats = pdfengines.FacturXPdfFormats(ctx, engine, facturX, pdfFormats, true, nil)
hasPostProcessing := watermark.Source != "" || stamp.Source != "" || angle != 0 ||
hasPostProcessing := len(watermarks) > 0 || len(stamps) > 0 || angle != 0 ||
len(embedPaths) > 0 || len(metadata) > 0 || flatten || facturX.ConformanceLevel != ""
outputPaths := make([]string, len(inputPaths))
@@ -493,12 +500,12 @@ func convertRoute(libreOffice libreofficeapi.Uno, engine gotenberg.PdfEngine) ap
}
}
err = pdfengines.WatermarkStub(ctx, engine, watermark, outputPaths)
err = pdfengines.WatermarkStub(ctx, engine, watermarks, outputPaths)
if err != nil {
return fmt.Errorf("watermark PDFs: %w", err)
}
err = pdfengines.StampStub(ctx, engine, stamp, outputPaths)
err = pdfengines.StampStub(ctx, engine, stamps, outputPaths)
if err != nil {
return fmt.Errorf("stamp PDFs: %w", err)
}
@@ -515,6 +522,11 @@ func convertRoute(libreOffice libreofficeapi.Uno, engine gotenberg.PdfEngine) ap
}
}
err = pdfengines.OptimizeStub(ctx, engine, optimizeImages, imageQuality, outputPaths)
if err != nil {
return fmt.Errorf("optimize PDF images: %w", err)
}
needsConvertStub := !nativePdfFormats ||
(nativePdfFormats && splitMode != zeroValuedSplitMode) ||
(nativePdfFormats && hasPostProcessing)

View File

@@ -0,0 +1,321 @@
package pdfcpu
import (
"bytes"
"context"
"fmt"
"image"
"image/jpeg"
_ "image/png" // Register the PNG decoder: pdfcpu extracts FlateDecode images as PNG.
"log/slog"
"os"
"os/exec"
"path/filepath"
"regexp"
"strconv"
"strings"
"syscall"
"go.opentelemetry.io/otel/codes"
"go.opentelemetry.io/otel/trace"
"github.com/gotenberg/gotenberg/v8/pkg/gotenberg"
)
// minOptimizeImageSize is the smallest encoded image worth re-encoding. Smaller
// images (thumbnails, icons, and line art that FlateDecode already keeps tiny)
// are left untouched: a JPEG pass would add artifacts for little or no gain.
const minOptimizeImageSize = 30 << 10 // 30 KiB
// pdfcpuListRowID matches the image Id (e.g. "X6") in a `pdfcpu images extract`
// filename such as "input_1_X6.png".
var pdfcpuListRowID = regexp.MustCompile(`_(X\d+)\.`)
// pdfcpuImage is one raster image XObject as reported by `pdfcpu images list`.
type pdfcpuImage struct {
obj int
id string
masked bool
comp int
bytes int64
filter string
}
// OptimizeImages re-encodes the raster images of inputPath to JPEG in place,
// shrinking image-heavy PDFs (a common case for Chromium output, which embeds
// non-JPEG images losslessly) while leaving text, vectors, fonts and structure
// untouched. See https://github.com/gotenberg/gotenberg/issues/359.
//
// Only lossless (FlateDecode), non-CMYK, non-masked images at or above
// [minOptimizeImageSize] are touched. Already-compressed, transparent, CMYK and
// small images are skipped so the pass never enlarges a file or corrupts
// transparency. It never fails the conversion for a single unreadable image; it
// logs and moves on, and returns the input unchanged when nothing qualifies.
func (engine *PdfCpu) OptimizeImages(ctx context.Context, logger *slog.Logger, imageQuality int, inputPath string) error {
ctx, span := gotenberg.Tracer().Start(ctx, "pdfcpu.OptimizeImages",
trace.WithSpanKind(trace.SpanKindClient),
trace.WithAttributes(engine.spanAttrs()...),
)
defer span.End()
fail := func(err error) error {
span.RecordError(err)
span.SetStatus(codes.Error, err.Error())
return err
}
images, err := engine.listImages(ctx, inputPath)
if err != nil {
return fail(fmt.Errorf("optimize PDF images with pdfcpu: %w", err))
}
var targets []pdfcpuImage
for _, img := range images {
if optimizableImage(img) {
targets = append(targets, img)
}
}
if len(targets) == 0 {
logger.DebugContext(ctx, "no images to optimize")
span.SetStatus(codes.Ok, "")
return nil
}
workDir, err := os.MkdirTemp(filepath.Dir(inputPath), "optimize-images-*")
if err != nil {
return fail(fmt.Errorf("optimize PDF images with pdfcpu: create work directory: %w", err))
}
defer func() {
if err := os.RemoveAll(workDir); err != nil {
logger.ErrorContext(ctx, fmt.Sprintf("remove image optimization work directory: %v", err))
}
}()
extracted, err := engine.extractImages(ctx, logger, inputPath, workDir)
if err != nil {
return fail(fmt.Errorf("optimize PDF images with pdfcpu: %w", err))
}
// Chain one update per image. Each update writes a fresh file; current holds
// the latest successful output, so a single failed image is skipped without
// discarding the ones already done. The input is only replaced on success.
current := inputPath
optimized := 0
for _, img := range targets {
src, ok := extracted[img.id]
if !ok {
logger.WarnContext(ctx, fmt.Sprintf("optimize images: image %s was not extracted, leaving it as is", img.id))
continue
}
reencoded := filepath.Join(workDir, img.id+".jpg")
err = reencodeToJpeg(src, reencoded, imageQuality)
if err != nil {
logger.WarnContext(ctx, fmt.Sprintf("optimize images: re-encode %s: %v; leaving it as is", img.id, err))
continue
}
next := filepath.Join(workDir, fmt.Sprintf("optimized-%d.pdf", optimized))
err = engine.updateImage(ctx, logger, current, reencoded, next, img.obj)
if err != nil {
logger.WarnContext(ctx, fmt.Sprintf("optimize images: update %s: %v; leaving it as is", img.id, err))
continue
}
current = next
optimized++
}
if optimized == 0 {
span.SetStatus(codes.Ok, "")
return nil
}
err = os.Rename(current, inputPath)
if err != nil {
return fail(fmt.Errorf("optimize PDF images with pdfcpu: replace input: %w", err))
}
logger.DebugContext(ctx, fmt.Sprintf("optimized %d image(s) at quality %d", optimized, imageQuality))
span.SetStatus(codes.Ok, "")
return nil
}
// optimizableImage reports whether an image is a safe, worthwhile target: a
// lossless (FlateDecode), non-CMYK, non-masked image at or above the size
// threshold. Everything else is left untouched.
func optimizableImage(img pdfcpuImage) bool {
switch {
case img.filter != "FlateDecode":
return false // Already compressed (JPEG/JPX); re-encoding would only add loss.
case img.comp == 4:
return false // CMYK; a JPEG round-trip is unsafe.
case img.masked:
return false // Soft mask, image mask or alpha; JPEG has no transparency.
case img.bytes < minOptimizeImageSize:
return false
default:
return true
}
}
// listImages runs `pdfcpu images list` and parses its table. The command writes
// to stdout, so it is run directly to capture the output.
func (engine *PdfCpu) listImages(ctx context.Context, inputPath string) ([]pdfcpuImage, error) {
cmd := exec.CommandContext(ctx, engine.binPath, "images", "list", inputPath) //nolint:gosec // binPath is validated at Provision; inputPath is a Gotenberg working file.
cmd.SysProcAttr = &syscall.SysProcAttr{Setpgid: true}
var stdout bytes.Buffer
cmd.Stdout = &stdout
err := cmd.Run()
if err != nil {
return nil, fmt.Errorf("run pdfcpu images list: %w", err)
}
return parseImagesList(stdout.String()), nil
}
// parseImagesList parses the fixed-column table of `pdfcpu images list`. Columns
// are separated by U+2502; the header and separator rows are skipped because
// their second column is not a numeric object number.
func parseImagesList(output string) []pdfcpuImage {
var images []pdfcpuImage
for line := range strings.SplitSeq(output, "\n") {
cols := strings.Split(line, "│")
if len(cols) < 9 {
continue
}
obj, err := strconv.Atoi(strings.TrimSpace(cols[1]))
if err != nil {
continue
}
comp := 0
if fields := strings.Fields(cols[6]); len(fields) >= 2 {
comp, _ = strconv.Atoi(fields[1])
}
images = append(images, pdfcpuImage{
obj: obj,
id: strings.TrimSpace(cols[2]),
masked: strings.TrimSpace(cols[3]) != "image",
comp: comp,
bytes: parseHumanSize(cols[7]),
filter: strings.TrimSpace(cols[8]),
})
}
return images
}
// parseHumanSize converts a pdfcpu size cell such as "4.4 MB" or "194 KB" into
// a byte count.
func parseHumanSize(cell string) int64 {
fields := strings.Fields(cell)
if len(fields) == 0 {
return 0
}
value, err := strconv.ParseFloat(fields[0], 64)
if err != nil {
return 0
}
multiplier := float64(1)
if len(fields) > 1 {
switch strings.ToUpper(fields[1]) {
case "KB":
multiplier = 1 << 10
case "MB":
multiplier = 1 << 20
case "GB":
multiplier = 1 << 30
}
}
return int64(value * multiplier)
}
// extractImages extracts every image of inputPath into dir and returns a map of
// image Id (e.g. "X6") to the extracted file path.
func (engine *PdfCpu) extractImages(ctx context.Context, logger *slog.Logger, inputPath, dir string) (map[string]string, error) {
args := []string{"images", "extract", inputPath, dir}
cmd, err := gotenberg.CommandContext(ctx, logger, engine.binPath, args...)
if err != nil {
return nil, fmt.Errorf("create command: %w", err)
}
_, err = cmd.Exec()
if err != nil {
return nil, fmt.Errorf("extract images: %w", err)
}
entries, err := os.ReadDir(dir)
if err != nil {
return nil, fmt.Errorf("read extracted images: %w", err)
}
extracted := make(map[string]string, len(entries))
for _, entry := range entries {
if match := pdfcpuListRowID.FindStringSubmatch(entry.Name()); match != nil {
extracted[match[1]] = filepath.Join(dir, entry.Name())
}
}
return extracted, nil
}
// updateImage replaces the image object objNr of inFile with the image at
// imagePath, writing the result to outFile. The replacement must share the
// original image dimensions, which reencodeToJpeg preserves.
func (engine *PdfCpu) updateImage(ctx context.Context, logger *slog.Logger, inFile, imagePath, outFile string, objNr int) error {
args := []string{"images", "update", inFile, imagePath, outFile, strconv.Itoa(objNr)}
cmd, err := gotenberg.CommandContext(ctx, logger, engine.binPath, args...)
if err != nil {
return fmt.Errorf("create command: %w", err)
}
_, err = cmd.Exec()
if err != nil {
return fmt.Errorf("update image: %w", err)
}
return nil
}
// reencodeToJpeg decodes the image at src and writes it to dst as JPEG at the
// given quality, keeping the original pixel dimensions (pdfcpu requires the
// replacement to match). quality is clamped to the valid 1 to 100 range.
func reencodeToJpeg(src, dst string, quality int) error {
if quality < 1 {
quality = 1
}
if quality > 100 {
quality = 100
}
in, err := os.Open(src) //nolint:gosec // src is a file this package extracted into its own temp dir.
if err != nil {
return fmt.Errorf("open image: %w", err)
}
defer in.Close()
img, _, err := image.Decode(in)
if err != nil {
return fmt.Errorf("decode image: %w", err)
}
out, err := os.Create(dst) //nolint:gosec // dst is a file in this package's own temp dir.
if err != nil {
return fmt.Errorf("create re-encoded image: %w", err)
}
defer out.Close()
err = jpeg.Encode(out, img, &jpeg.Options{Quality: quality})
if err != nil {
return fmt.Errorf("encode JPEG: %w", err)
}
return nil
}

View File

@@ -0,0 +1,83 @@
package pdfcpu
import "testing"
const sampleImagesList = `pages: all
/tmp/multi.pdf:
4 images available (8.9 MB)
Page │ Obj# │ Id │ Type SoftMask ImgMask │ Width │ Height │ ColorSpace Comp bpc Interp │ Size │ Filters
━━━━━┿━━━━━━┿━━━━━┿━━━━━━━━━━━━━━━━━━━━━━━━┿━━━━━━━┿━━━━━━━━┿━━━━━━━━━━━━━━━━━━━━━━━━━━━━┿━━━━━━━━┿━━━━━━━━━━━━
1 │ 6 │ X6 │ image │ 2400 │ 1800 │ ICCBased 3 8 │ 5.5 MB │ FlateDecode
│ 8 │ X8 │ image * │ 1400 │ 1000 │ ICCBased 3 8 │ 63 KB │ FlateDecode
│ 9 │ X9 │ image │ 2400 │ 1800 │ ICCBased 3 8 │ 194 KB │ DCTDecode
│ 10 │ X10 │ image │ 120 │ 90 │ DeviceCMYK 4 8 │ 14 KB │ FlateDecode
`
func TestParseImagesList(t *testing.T) {
images := parseImagesList(sampleImagesList)
if len(images) != 4 {
t.Fatalf("expected 4 images, got %d", len(images))
}
for _, tc := range []struct {
index int
obj int
id string
masked bool
comp int
filter string
}{
{0, 6, "X6", false, 3, "FlateDecode"},
{1, 8, "X8", true, 3, "FlateDecode"},
{2, 9, "X9", false, 3, "DCTDecode"},
{3, 10, "X10", false, 4, "FlateDecode"},
} {
img := images[tc.index]
if img.obj != tc.obj || img.id != tc.id || img.masked != tc.masked || img.comp != tc.comp || img.filter != tc.filter {
t.Errorf("image %d = %+v, want obj=%d id=%s masked=%v comp=%d filter=%s",
tc.index, img, tc.obj, tc.id, tc.masked, tc.comp, tc.filter)
}
}
}
func TestParseHumanSize(t *testing.T) {
for _, tc := range []struct {
cell string
want int64
}{
{"5.5 MB", int64(5.5 * (1 << 20))},
{"194 KB", 194 << 10},
{" 14 KB ", 14 << 10},
{"512 B", 512},
{"2 GB", 2 << 30},
{"", 0},
{"garbage", 0},
} {
if got := parseHumanSize(tc.cell); got != tc.want {
t.Errorf("parseHumanSize(%q) = %d, want %d", tc.cell, got, tc.want)
}
}
}
func TestOptimizableImage(t *testing.T) {
base := pdfcpuImage{obj: 1, id: "X1", masked: false, comp: 3, bytes: 1 << 20, filter: "FlateDecode"}
for _, tc := range []struct {
scenario string
mutate func(pdfcpuImage) pdfcpuImage
want bool
}{
{"lossless RGB above threshold", func(i pdfcpuImage) pdfcpuImage { return i }, true},
{"already compressed", func(i pdfcpuImage) pdfcpuImage { i.filter = "DCTDecode"; return i }, false},
{"CMYK", func(i pdfcpuImage) pdfcpuImage { i.comp = 4; return i }, false},
{"masked", func(i pdfcpuImage) pdfcpuImage { i.masked = true; return i }, false},
{"below threshold", func(i pdfcpuImage) pdfcpuImage { i.bytes = minOptimizeImageSize - 1; return i }, false},
{"grayscale above threshold", func(i pdfcpuImage) pdfcpuImage { i.comp = 1; return i }, true},
} {
if got := optimizableImage(tc.mutate(base)); got != tc.want {
t.Errorf("%s: optimizableImage = %v, want %v", tc.scenario, got, tc.want)
}
}
}

View File

@@ -0,0 +1,172 @@
package pdfengines
import (
"sync"
"github.com/gotenberg/gotenberg/v8/pkg/modules/api"
)
// defaultMaxConcurrency is the number of PDF files a single stub processes at
// once when --pdfengines-max-concurrency (env PDFENGINES_MAX_CONCURRENCY) is
// not set. Each unit of work forks an external binary (qpdf, pdfcpu, pdftk or
// exiftool), so the ceiling trades wall clock against process count and RSS.
//
// It defaults to one, which processes files exactly as the sequential loops
// this package used to run did. Raising it only ever affects a request that
// carries several files, or one that splits into several outputs: a
// single-file request never reaches the concurrent path at all. Operators who
// send multi-file batches and have the memory headroom opt in.
//
// This never covers LibreOffice. libreoffice-pdfengine implements Convert and
// nothing else, every other [gotenberg.PdfEngine] method on it returns
// [gotenberg.ErrPdfEngineMethodNotSupported], and [ConvertStub] deliberately
// does not use this package's helpers. A soffice instance costs too much
// memory to run several of per container, so LibreOffice throughput is scaled
// by adding Gotenberg containers, not by raising this number.
const defaultMaxConcurrency = 1
// maxFileConcurrency is how many files one request may have in flight at once.
// It is replaced during [PdfEngines.Provision].
var maxFileConcurrency = defaultMaxConcurrency
// engineExtraSlots bounds the concurrency this package ADDS, across the whole
// process rather than per request, and holds one fewer slot than
// [maxFileConcurrency] because every request already owns one unit of its own.
//
// Bounding the added concurrency rather than the total is what keeps the
// ceiling from becoming a throughput regression. The sequential loops this
// helper replaced had no ceiling at all: X concurrent requests ran X engine
// binaries, one apiece. A pool covering the total would cut those X requests
// down to the ceiling, so an operator raising the flag to speed up a single
// multi-file request would slow the server down under real load. Reserving
// each request the unit it always had makes the worst case "what happened
// before, plus at most maxFileConcurrency-1".
//
// A per-request limit would have the opposite failure: X simultaneous requests
// forking X times the limit, trading the timeouts this exists to prevent for
// memory exhaustion.
var engineExtraSlots = make(chan struct{}, defaultMaxConcurrency-1)
// acquireEngineSlot waits for the first unit of capacity to become available,
// either this request's reserved unit or a slot from the shared pool, and
// returns the function that gives it back.
//
// Waiting on both at once is the whole point. Committing to one source and
// blocking on it strands the other: a goroutine parked on an exhausted pool
// cannot pick up its own request's reserved unit when the file before it
// finishes, so the reserved units sit idle while every file queues on the
// pool, which is slower than having no pool at all.
func acquireEngineSlot(ctx *api.Context, reserved chan struct{}) (func(), error) {
// An [api.Context] carries a request context in production, but one built
// as a literal, which the unit tests do, embeds a nil [context.Context]
// and would panic on Done. A nil channel never fires, which correctly
// leaves the two capacity sources as the only things to wait on.
var done <-chan struct{}
if ctx != nil && ctx.Context != nil {
done = ctx.Done()
}
// A select whose cancellation and capacity cases are both ready picks
// between them at random, so an already-dead request would start more
// files on a coin flip. Check first and stop taking on work.
if done != nil {
select {
case <-done:
return nil, ctx.Err()
default:
}
}
select {
case <-reserved:
return func() { reserved <- struct{}{} }, nil
case engineExtraSlots <- struct{}{}:
return func() { <-engineExtraSlots }, nil
case <-done:
return nil, ctx.Err()
}
}
// forEachInputPath runs fn against every input path, up to
// --pdfengines-max-concurrency (env PDFENGINES_MAX_CONCURRENCY) at a time.
//
// The stubs mutate each PDF in place, so distinct input paths never touch the
// same file and may run together. Callers that layer operations on one file,
// like [WatermarkStub] applying several watermarks in order, must keep that
// outer sequence and parallelize only the file dimension.
//
// Every path is attempted even after one fails, and the error returned is the
// first in input order rather than the first to arrive. That keeps the failing
// filename in the error message identical to what the sequential form
// reported, which the integration scenarios assert on.
func forEachInputPath(ctx *api.Context, inputPaths []string, fn func(inputPath string) error) error {
return forEachInputPathIndexed(ctx, inputPaths, func(_ int, inputPath string) error {
return fn(inputPath)
})
}
// forEachInputPathIndexed is [forEachInputPath] with the input path's index,
// for callers collecting a result per file. Writing into a preallocated slice
// at the given index needs no further synchronization; writing into a shared
// map does and must not be done from fn.
func forEachInputPathIndexed(ctx *api.Context, inputPaths []string, fn func(i int, inputPath string) error) error {
if len(inputPaths) == 0 {
return nil
}
// The common case is a single file. Skip the goroutine and the slot: the
// caller is already inside whatever bound its own route applies.
if len(inputPaths) == 1 {
return fn(0, inputPaths[0])
}
// At the default ceiling of one, run the plain sequential loop this helper
// replaced. Racing goroutines for a single slot would serialize the work
// just the same, but the order files are picked up in would be down to the
// scheduler, and every file would be attempted even once one has failed.
// Taking the old path keeps the default a genuine no-op: same order, same
// early return, no goroutines.
if maxFileConcurrency < 2 {
for i, inputPath := range inputPaths {
err := fn(i, inputPath)
if err != nil {
return err
}
}
return nil
}
// The unit this request would have had all to itself before any of this
// existed. Whichever file claims it runs without touching the shared pool,
// so concurrent requests can never throttle each other below the
// one-binary-apiece they already got. See [engineExtraSlots].
reserved := make(chan struct{}, 1)
reserved <- struct{}{}
errs := make([]error, len(inputPaths))
var wg sync.WaitGroup
for i, inputPath := range inputPaths {
wg.Go(func() {
release, err := acquireEngineSlot(ctx, reserved)
if err != nil {
errs[i] = err
return
}
defer release()
errs[i] = fn(i, inputPath)
})
}
wg.Wait()
for _, err := range errs {
if err != nil {
return err
}
}
return nil
}

View File

@@ -0,0 +1,341 @@
package pdfengines
import (
"context"
"errors"
"fmt"
"sync"
"sync/atomic"
"testing"
"time"
"github.com/gotenberg/gotenberg/v8/pkg/modules/api"
)
func TestForEachInputPath(t *testing.T) {
for _, tc := range []struct {
scenario string
inputPaths []string
fn func(inputPath string) error
expectErr string
}{
{
scenario: "no input path",
inputPaths: nil,
fn: func(string) error { return errors.New("must not run") },
},
{
scenario: "single input path",
inputPaths: []string{"a.pdf"},
fn: func(string) error { return nil },
},
{
scenario: "single input path with error",
inputPaths: []string{"a.pdf"},
fn: func(p string) error { return fmt.Errorf("boom %s", p) },
expectErr: "boom a.pdf",
},
{
scenario: "many input paths",
inputPaths: []string{"a.pdf", "b.pdf", "c.pdf", "d.pdf", "e.pdf"},
fn: func(string) error { return nil },
},
{
scenario: "error is the first in input order, not the first to arrive",
inputPaths: []string{"a.pdf", "b.pdf", "c.pdf"},
fn: func(p string) error {
// "c.pdf" fails without delay so that it lands well before
// "b.pdf"; the reported error must still be "b.pdf".
if p == "b.pdf" {
var counter int
for i := range 5_000_000 {
counter += i
}
return fmt.Errorf("slow failure %s (%d)", p, counter%1)
}
if p == "c.pdf" {
return fmt.Errorf("fast failure %s", p)
}
return nil
},
expectErr: "slow failure b.pdf (0)",
},
} {
t.Run(tc.scenario, func(t *testing.T) {
err := forEachInputPath(new(api.Context), tc.inputPaths, tc.fn)
if tc.expectErr == "" {
if err != nil {
t.Fatalf("expected no error but got: %v", err)
}
return
}
if err == nil {
t.Fatalf("expected error %q but got none", tc.expectErr)
}
if err.Error() != tc.expectErr {
t.Fatalf("expected error %q but got %q", tc.expectErr, err.Error())
}
})
}
}
func TestForEachInputPathRunsEveryPath(t *testing.T) {
inputPaths := make([]string, 50)
for i := range inputPaths {
inputPaths[i] = fmt.Sprintf("%d.pdf", i)
}
var (
mu sync.Mutex
seen = make(map[string]int)
)
err := forEachInputPath(new(api.Context), inputPaths, func(inputPath string) error {
mu.Lock()
defer mu.Unlock()
seen[inputPath]++
return nil
})
if err != nil {
t.Fatalf("expected no error but got: %v", err)
}
if len(seen) != len(inputPaths) {
t.Fatalf("expected %d distinct paths but got %d", len(inputPaths), len(seen))
}
for path, count := range seen {
if count != 1 {
t.Fatalf("expected '%s' to run once but it ran %d times", path, count)
}
}
}
func TestForEachInputPathRespectsTheSlotCeiling(t *testing.T) {
previousSlots, previousMax := engineExtraSlots, maxFileConcurrency
defer func() { engineExtraSlots, maxFileConcurrency = previousSlots, previousMax }()
// One request may run its reserved unit plus ceiling-1 borrowed ones.
const ceiling = 3
maxFileConcurrency = ceiling
engineExtraSlots = make(chan struct{}, ceiling-1)
inputPaths := make([]string, 40)
for i := range inputPaths {
inputPaths[i] = fmt.Sprintf("%d.pdf", i)
}
var inFlight, peak atomic.Int64
err := forEachInputPath(new(api.Context), inputPaths, func(string) error {
current := inFlight.Add(1)
defer inFlight.Add(-1)
for {
observed := peak.Load()
if current <= observed || peak.CompareAndSwap(observed, current) {
break
}
}
// Hold the slot long enough that the ceiling would be exceeded if it
// were not enforced.
var counter int
for i := range 200_000 {
counter += i
}
_ = counter
return nil
})
if err != nil {
t.Fatalf("expected no error but got: %v", err)
}
if peak.Load() > ceiling {
t.Fatalf("expected at most %d concurrent runs but observed %d", ceiling, peak.Load())
}
}
func TestForEachInputPathHonorsCancellation(t *testing.T) {
previousSlots, previousMax := engineExtraSlots, maxFileConcurrency
defer func() { engineExtraSlots, maxFileConcurrency = previousSlots, previousMax }()
// Concurrent path, with the shared pool exhausted by another request, so
// only this request's reserved unit is available.
maxFileConcurrency = 3
engineExtraSlots = make(chan struct{}, 2)
engineExtraSlots <- struct{}{}
engineExtraSlots <- struct{}{}
cancelledCtx, cancel := context.WithCancel(context.Background())
cancel()
// Neither source of capacity is available: the pool is exhausted by other
// requests and this request's reserved unit is already in use by one of its
// own files. A waiter must observe the cancelled context rather than block
// forever. Driving acquireEngineSlot directly keeps that deterministic:
// through forEachInputPath the reserved unit is reusable, so whether a
// given file waits at all depends on how fast the file before it finishes.
inUse := make(chan struct{}, 1)
release, err := acquireEngineSlot(&api.Context{Context: cancelledCtx}, inUse)
if !errors.Is(err, context.Canceled) {
t.Fatalf("expected context.Canceled but got: %v", err)
}
if release != nil {
t.Fatal("expected no release function when acquisition fails")
}
// A cancelled request stops taking on work even when capacity is free,
// rather than deciding on the coin flip a ready select would give.
inUse <- struct{}{}
_, err = acquireEngineSlot(&api.Context{Context: cancelledCtx}, inUse)
if !errors.Is(err, context.Canceled) {
t.Fatalf("expected context.Canceled with the reserved unit free but got: %v", err)
}
// Live request, free reserved unit: acquired and handed back.
release, err = acquireEngineSlot(&api.Context{Context: context.Background()}, inUse)
if err != nil {
t.Fatalf("expected the reserved unit to be acquired but got: %v", err)
}
release()
if len(inUse) != 1 {
t.Fatalf("expected the reserved unit to be returned but the channel holds %d", len(inUse))
}
}
func TestForEachInputPathCompletesWithACancelledContext(t *testing.T) {
previousSlots, previousMax := engineExtraSlots, maxFileConcurrency
defer func() { engineExtraSlots, maxFileConcurrency = previousSlots, previousMax }()
maxFileConcurrency = 3
engineExtraSlots = make(chan struct{}, 2)
engineExtraSlots <- struct{}{}
engineExtraSlots <- struct{}{}
cancelledCtx, cancel := context.WithCancel(context.Background())
cancel()
done := make(chan struct{})
go func() {
defer close(done)
_ = forEachInputPath(&api.Context{Context: cancelledCtx}, []string{"a.pdf", "b.pdf", "c.pdf"}, func(string) error {
return nil
})
}()
select {
case <-done:
case <-time.After(10 * time.Second):
t.Fatal("forEachInputPath hung on a cancelled context with the shared pool exhausted")
}
}
func TestForEachInputPathNeverThrottlesBelowOnePerRequest(t *testing.T) {
previousSlots, previousMax := engineExtraSlots, maxFileConcurrency
defer func() { engineExtraSlots, maxFileConcurrency = previousSlots, previousMax }()
// A small ceiling against far more concurrent requests than it covers.
// Before the shared pool existed each of these ran a binary of its own, so
// the pool must not drop aggregate concurrency below one per request.
const (
ceiling = 2
requests = 8
)
maxFileConcurrency = ceiling
engineExtraSlots = make(chan struct{}, ceiling-1)
// Every runner announces itself and then blocks, so the count of arrivals
// is the true simultaneous concurrency rather than whatever the scheduler
// happened to overlap.
arrived := make(chan struct{}, requests*3)
release := make(chan struct{})
var wg sync.WaitGroup
for range requests {
wg.Go(func() {
_ = forEachInputPath(new(api.Context), []string{"a.pdf", "b.pdf", "c.pdf"}, func(string) error {
arrived <- struct{}{}
<-release
return nil
})
})
}
// One runner per request must be able to start without waiting on the
// shared pool. If the pool governed the total instead of the surplus, only
// `ceiling` runners would ever arrive and this would time out.
for i := range requests {
select {
case <-arrived:
case <-time.After(10 * time.Second):
t.Fatalf("only %d runners started concurrently, expected at least one per request (%d); the shared pool is throttling requests against each other", i, requests)
}
}
close(release)
wg.Wait()
}
func TestForEachInputPathIsSequentialAtTheDefaultCeiling(t *testing.T) {
if defaultMaxConcurrency != 1 {
t.Fatalf("this test pins the default as a no-op, but defaultMaxConcurrency is %d", defaultMaxConcurrency)
}
previousSlots, previousMax := engineExtraSlots, maxFileConcurrency
defer func() { engineExtraSlots, maxFileConcurrency = previousSlots, previousMax }()
maxFileConcurrency = defaultMaxConcurrency
engineExtraSlots = make(chan struct{}, defaultMaxConcurrency-1)
// At the default the helper must behave exactly like the sequential loops
// it replaced: files in input order, and no file attempted once one has
// failed.
var order []string
err := forEachInputPath(new(api.Context), []string{"a.pdf", "b.pdf", "c.pdf", "d.pdf"}, func(inputPath string) error {
order = append(order, inputPath)
if inputPath == "b.pdf" {
return errors.New("boom")
}
return nil
})
if err == nil || err.Error() != "boom" {
t.Fatalf("expected error \"boom\" but got: %v", err)
}
if len(order) != 2 || order[0] != "a.pdf" || order[1] != "b.pdf" {
t.Fatalf("expected the run to stop after b.pdf in input order but got %v", order)
}
}
func TestForEachInputPathIndexed(t *testing.T) {
inputPaths := []string{"a.pdf", "b.pdf", "c.pdf", "d.pdf"}
collected := make([]string, len(inputPaths))
err := forEachInputPathIndexed(new(api.Context), inputPaths, func(i int, inputPath string) error {
collected[i] = inputPath
return nil
})
if err != nil {
t.Fatalf("expected no error but got: %v", err)
}
for i, inputPath := range inputPaths {
if collected[i] != inputPath {
t.Fatalf("expected index %d to hold '%s' but got '%s'", i, inputPath, collected[i])
}
}
}

View File

@@ -18,6 +18,7 @@ type multiPdfEngines struct {
splitEngines []gotenberg.PdfEngine
flattenEngines []gotenberg.PdfEngine
convertEngines []gotenberg.PdfEngine
optimizeImagesEngines []gotenberg.PdfEngine
readMetadataEngines []gotenberg.PdfEngine
writeMetadataEngines []gotenberg.PdfEngine
passwordEngines []gotenberg.PdfEngine
@@ -36,6 +37,7 @@ func newMultiPdfEngines(
splitEngines,
flattenEngines,
convertEngines,
optimizeImagesEngines,
readMetadataEngines,
writeMetadataEngines,
passwordEngines,
@@ -53,6 +55,7 @@ func newMultiPdfEngines(
splitEngines: splitEngines,
flattenEngines: flattenEngines,
convertEngines: convertEngines,
optimizeImagesEngines: optimizeImagesEngines,
readMetadataEngines: readMetadataEngines,
writeMetadataEngines: writeMetadataEngines,
passwordEngines: passwordEngines,
@@ -189,6 +192,17 @@ func (multi *multiPdfEngines) Flatten(ctx context.Context, logger *slog.Logger,
)
}
// OptimizeImages re-encodes the images of a PDF using the first available
// engine that supports image optimization.
func (multi *multiPdfEngines) OptimizeImages(ctx context.Context, logger *slog.Logger, imageQuality int, inputPath string) error {
return runWithFallbackVoid(ctx, "pdfengines.OptimizeImages", multi.optimizeImagesEngines,
func(ctx context.Context, engine gotenberg.PdfEngine) error {
return engine.OptimizeImages(ctx, logger, imageQuality, inputPath)
},
func(err error) error { return fmt.Errorf("optimize PDF images with multi PDF engines: %w", err) },
)
}
// Convert transforms the given PDF to a specific PDF format using the first
// available engine that supports PDF conversion.
func (multi *multiPdfEngines) Convert(ctx context.Context, logger *slog.Logger, formats gotenberg.PdfFormats, inputPath, outputPath string) error {

View File

@@ -32,6 +32,7 @@ type PdfEngines struct {
splitNames []string
flattenNames []string
convertNames []string
optimizeImagesNames []string
readMetadataNames []string
writeMetadataNames []string
encryptNames []string
@@ -44,6 +45,7 @@ type PdfEngines struct {
rotateNames []string
facturXNames []string
engines []gotenberg.PdfEngine
maxConcurrency int
disableRoutes bool
}
@@ -57,6 +59,7 @@ func (mod *PdfEngines) Descriptor() gotenberg.ModuleDescriptor {
fs.StringSlice("pdfengines-split-engines", []string{"pdfcpu", "qpdf", "pdftk"}, "Set the PDF engines and their order for the split feature - empty means all")
fs.StringSlice("pdfengines-flatten-engines", []string{"qpdf"}, "Set the PDF engines and their order for the flatten feature - empty means all")
fs.StringSlice("pdfengines-convert-engines", []string{"libreoffice-pdfengine"}, "Set the PDF engines and their order for the convert feature - empty means all")
fs.StringSlice("pdfengines-optimize-images-engines", []string{"pdfcpu"}, "Set the PDF engines and their order for the image optimization feature - empty means all")
fs.StringSlice("pdfengines-read-metadata-engines", []string{"exiftool"}, "Set the PDF engines and their order for the read metadata feature - empty means all")
fs.StringSlice("pdfengines-write-metadata-engines", []string{"exiftool"}, "Set the PDF engines and their order for the write metadata feature - empty means all")
fs.StringSlice("pdfengines-encrypt-engines", []string{"qpdf", "pdftk", "pdfcpu"}, "Set the PDF engines and their order for the password protection feature - empty means all")
@@ -68,6 +71,7 @@ func (mod *PdfEngines) Descriptor() gotenberg.ModuleDescriptor {
fs.StringSlice("pdfengines-stamp-engines", []string{"pdfcpu", "pdftk"}, "Set the PDF engines and their order for the stamp feature - empty means all")
fs.StringSlice("pdfengines-rotate-engines", []string{"pdfcpu", "pdftk"}, "Set the PDF engines and their order for the rotate feature - empty means all")
fs.StringSlice("pdfengines-factur-x-engines", []string{"qpdf"}, "Set the PDF engines and their order for the Factur-X XMP feature - empty means all")
fs.Int("pdfengines-max-concurrency", defaultMaxConcurrency, "Set the maximum number of PDF files a feature processes concurrently, across all requests - bounds how many qpdf, pdfcpu, pdftk and exiftool processes run at once, so raising it trades memory for speed. Does not apply to LibreOffice: scale Gotenberg containers instead")
fs.Bool("pdfengines-disable-routes", false, "Disable the routes")
// Deprecated flags.
@@ -91,6 +95,7 @@ func (mod *PdfEngines) Provision(ctx *gotenberg.Context) error {
splitNames := flags.MustStringSlice("pdfengines-split-engines")
flattenNames := flags.MustStringSlice("pdfengines-flatten-engines")
convertNames := flags.MustStringSlice("pdfengines-convert-engines")
optimizeImagesNames := flags.MustStringSlice("pdfengines-optimize-images-engines")
readMetadataNames := flags.MustStringSlice("pdfengines-read-metadata-engines")
writeMetadataNames := flags.MustStringSlice("pdfengines-write-metadata-engines")
encryptNames := flags.MustStringSlice("pdfengines-encrypt-engines")
@@ -102,8 +107,16 @@ func (mod *PdfEngines) Provision(ctx *gotenberg.Context) error {
stampNames := flags.MustStringSlice("pdfengines-stamp-engines")
rotateNames := flags.MustStringSlice("pdfengines-rotate-engines")
facturXNames := flags.MustStringSlice("pdfengines-factur-x-engines")
mod.maxConcurrency = flags.MustInt("pdfengines-max-concurrency")
mod.disableRoutes = flags.MustBool("pdfengines-disable-routes")
if mod.maxConcurrency > 0 {
maxFileConcurrency = mod.maxConcurrency
// One fewer than the ceiling: each request already reserves a unit of
// its own. See [engineExtraSlots].
engineExtraSlots = make(chan struct{}, mod.maxConcurrency-1)
}
engines, err := ctx.Modules(new(gotenberg.PdfEngine))
if err != nil {
return fmt.Errorf("get PDF engines: %w", err)
@@ -148,6 +161,11 @@ func (mod *PdfEngines) Provision(ctx *gotenberg.Context) error {
mod.convertNames = convertNames
}
mod.optimizeImagesNames = defaultNames
if len(optimizeImagesNames) > 0 {
mod.optimizeImagesNames = optimizeImagesNames
}
mod.readMetadataNames = defaultNames
if len(readMetadataNames) > 0 {
mod.readMetadataNames = readMetadataNames
@@ -214,6 +232,10 @@ func (mod *PdfEngines) Validate() error {
return errors.New("no PDF engine is available; enable at least one engine module (e.g. qpdf, pdfcpu, pdftk, libreoffice-pdfengine, exiftool)")
}
if mod.maxConcurrency < 1 {
return fmt.Errorf("PDF engines max concurrency must be at least 1, got %d; set --pdfengines-max-concurrency (env PDFENGINES_MAX_CONCURRENCY) to a positive value", mod.maxConcurrency)
}
availableEngines := make([]string, len(mod.engines))
for i, engine := range mod.engines {
@@ -247,6 +269,7 @@ func (mod *PdfEngines) Validate() error {
findNonExistingEngines(mod.mergeNames)
findNonExistingEngines(mod.splitNames)
findNonExistingEngines(mod.flattenNames)
findNonExistingEngines(mod.optimizeImagesNames)
findNonExistingEngines(mod.convertNames)
findNonExistingEngines(mod.readMetadataNames)
findNonExistingEngines(mod.writeMetadataNames)
@@ -275,6 +298,7 @@ func (mod *PdfEngines) SystemMessages() []string {
fmt.Sprintf("split engines - %s", strings.Join(mod.splitNames, " ")),
fmt.Sprintf("flatten engines - %s", strings.Join(mod.flattenNames, " ")),
fmt.Sprintf("convert engines - %s", strings.Join(mod.convertNames, " ")),
fmt.Sprintf("optimize images engines - %s", strings.Join(mod.optimizeImagesNames, " ")),
fmt.Sprintf("read metadata engines - %s", strings.Join(mod.readMetadataNames, " ")),
fmt.Sprintf("write metadata engines - %s", strings.Join(mod.writeMetadataNames, " ")),
fmt.Sprintf("encrypt engines - %s", strings.Join(mod.encryptNames, " ")),
@@ -286,6 +310,7 @@ func (mod *PdfEngines) SystemMessages() []string {
fmt.Sprintf("stamp engines - %s", strings.Join(mod.stampNames, " ")),
fmt.Sprintf("rotate engines - %s", strings.Join(mod.rotateNames, " ")),
fmt.Sprintf("factur-x engines - %s", strings.Join(mod.facturXNames, " ")),
fmt.Sprintf("max concurrency - %d", mod.maxConcurrency),
}
}
@@ -310,6 +335,7 @@ func (mod *PdfEngines) PdfEngine() (gotenberg.PdfEngine, error) {
engines(mod.splitNames),
engines(mod.flattenNames),
engines(mod.convertNames),
engines(mod.optimizeImagesNames),
engines(mod.readMetadataNames),
engines(mod.writeMetadataNames),
engines(mod.encryptNames),
@@ -341,6 +367,7 @@ func (mod *PdfEngines) Routes() ([]api.Route, error) {
mergeRoute(engine),
splitRoute(engine),
flattenRoute(engine),
optimizeRoute(engine),
convertRoute(engine),
readMetadataRoute(engine),
writeMetadataRoute(engine),

View File

@@ -208,14 +208,14 @@ func RotateStub(ctx *api.Context, engine gotenberg.PdfEngine, angle int, pages s
return nil
}
for _, inputPath := range inputPaths {
return forEachInputPath(ctx, inputPaths, func(inputPath string) error {
err := engine.Rotate(ctx, ctx.Log(), inputPath, angle, pages)
if err != nil {
return fmt.Errorf("rotate '%s': %w", inputPath, err)
}
}
return nil
return nil
})
}
// ValidatePdfFormatsCompat checks for incompatible combinations of PDF formats
@@ -334,19 +334,86 @@ func SplitPdfStub(ctx *api.Context, engine gotenberg.PdfEngine, mode gotenberg.S
// FlattenStub merges annotation appearances with page content for each given
// PDF, effectively deleting the original annotations.
func FlattenStub(ctx *api.Context, engine gotenberg.PdfEngine, inputPaths []string) error {
for _, inputPath := range inputPaths {
return forEachInputPath(ctx, inputPaths, func(inputPath string) error {
err := engine.Flatten(ctx, ctx.Log(), inputPath)
if err != nil {
return fmt.Errorf("flatten '%s': %w", inputPath, err)
}
return nil
})
}
// defaultImageQuality is the JPEG quality applied by the image optimization
// feature when the imageQuality form field is not set.
const defaultImageQuality = 80
// FormDataPdfOptimize extracts the image-optimization options from the form
// data: whether to optimize the images, and the JPEG quality (1 to 100) to
// apply to each re-encoded image.
func FormDataPdfOptimize(form *api.FormData) (bool, int) {
var (
optimizeImages bool
imageQuality int
)
form.
Bool("optimizeImages", &optimizeImages, false).
Custom("imageQuality", func(value string) error {
if value == "" {
imageQuality = defaultImageQuality
return nil
}
intValue, err := strconv.Atoi(value)
if err != nil {
return err
}
if intValue < 1 {
return errors.New("value is inferior to 1")
}
if intValue > 100 {
return errors.New("value is superior to 100")
}
imageQuality = intValue
return nil
})
return optimizeImages, imageQuality
}
// OptimizeStub re-encodes the images of each given PDF to shrink the file when
// optimizeImages is set, leaving text, vectors and structure untouched. It does
// nothing when optimizeImages is false.
func OptimizeStub(ctx *api.Context, engine gotenberg.PdfEngine, optimizeImages bool, imageQuality int, inputPaths []string) error {
if !optimizeImages {
return nil
}
return nil
return forEachInputPath(ctx, inputPaths, func(inputPath string) error {
err := engine.OptimizeImages(ctx, ctx.Log(), imageQuality, inputPath)
if err != nil {
return fmt.Errorf("optimize images of '%s': %w", inputPath, err)
}
return nil
})
}
// ConvertStub transforms a given PDF to the specified formats defined in
// [gotenberg.PdfFormats]. If no format, it does nothing and returns the input
// paths.
//
// This loop stays sequential on purpose. Convert is the one PDF engine method
// LibreOffice implements, and libreoffice-pdfengine is the default and only
// convert engine, so every iteration here drives the single soffice daemon. A
// LibreOffice instance is far too memory-hungry to run several of per
// container: the way to convert more documents at once is to scale Gotenberg
// containers, not to widen this loop. Do not route it through
// [forEachInputPath].
func ConvertStub(ctx *api.Context, engine gotenberg.PdfEngine, formats gotenberg.PdfFormats, inputPaths []string) ([]string, error) {
zeroValued := gotenberg.PdfFormats{}
if formats == zeroValued {
@@ -373,14 +440,35 @@ func WriteMetadataStub(ctx *api.Context, engine gotenberg.PdfEngine, metadata ma
return nil
}
for _, inputPath := range inputPaths {
return forEachInputPath(ctx, inputPaths, func(inputPath string) error {
err := engine.WriteMetadata(ctx, ctx.Log(), metadata, inputPath)
if err != nil {
return fmt.Errorf("write metadata into '%s': %w", inputPath, err)
}
return nil
})
}
// documentTitle returns the input PDF's Title metadata entry, falling back to
// the original filename without its extension when the entry is absent, blank,
// or cannot be read. It labels the per-document entries the merge route's
// titleBookmarks feature generates.
func documentTitle(ctx *api.Context, engine gotenberg.PdfEngine, inputPath, filename string) string {
fallback := strings.TrimSuffix(filename, filepath.Ext(filename))
metadata, err := engine.ReadMetadata(ctx, ctx.Log(), inputPath)
if err != nil {
ctx.Log().WarnContext(ctx, fmt.Sprintf("read metadata of '%s' for title bookmark, using filename: %s", filename, err))
return fallback
}
return nil
title, ok := metadata["Title"].(string)
if !ok || strings.TrimSpace(title) == "" {
return fallback
}
return title
}
func shiftBookmarks(bookmarks []gotenberg.Bookmark, offset int) []gotenberg.Bookmark {
@@ -411,28 +499,33 @@ func WriteBookmarksStub(ctx *api.Context, engine gotenberg.PdfEngine, bookmarks
return nil
}
for _, inputPath := range inputPaths {
return forEachInputPath(ctx, inputPaths, func(inputPath string) error {
err := engine.WriteBookmarks(ctx, ctx.Log(), inputPath, b)
if err != nil {
return fmt.Errorf("write bookmarks into '%s': %w", inputPath, err)
}
}
return nil
})
case map[string][]gotenberg.Bookmark:
for _, inputPath := range inputPaths {
return forEachInputPath(ctx, inputPaths, func(inputPath string) error {
filename := ctx.OriginalFilename(inputPath)
if specificBookmarks, ok := b[filename]; ok {
err := engine.WriteBookmarks(ctx, ctx.Log(), inputPath, specificBookmarks)
if err != nil {
return fmt.Errorf("write bookmarks into '%s': %w", inputPath, err)
}
specificBookmarks, ok := b[filename]
if !ok {
return nil
}
}
err := engine.WriteBookmarks(ctx, ctx.Log(), inputPath, specificBookmarks)
if err != nil {
return fmt.Errorf("write bookmarks into '%s': %w", inputPath, err)
}
return nil
})
default:
// Should not happen.
return fmt.Errorf("bookmarks type '%T' not supported", bookmarks)
}
return nil
}
// FormDataPdfEmbeds extracts embedded file paths from form data.
@@ -457,14 +550,14 @@ func EmbedFilesMetadataStub(ctx *api.Context, engine gotenberg.PdfEngine, metada
return nil
}
for _, inputPath := range inputPaths {
return forEachInputPath(ctx, inputPaths, func(inputPath string) error {
err := engine.EmbedFilesMetadata(ctx, ctx.Log(), metadata, inputPath)
if err != nil {
return fmt.Errorf("set embeds metadata on PDF '%s': %w", inputPath, err)
}
}
return nil
return nil
})
}
// FormDataPdfFacturX extracts the Factur-X parameters and the invoice XML path
@@ -657,14 +750,14 @@ func InjectFacturXXMPStub(ctx *api.Context, engine gotenberg.PdfEngine, facturX
return nil
}
for _, inputPath := range inputPaths {
return forEachInputPath(ctx, inputPaths, func(inputPath string) error {
err := engine.InjectFacturXXMP(ctx, ctx.Log(), facturX, inputPath)
if err != nil {
return fmt.Errorf("inject Factur-X XMP into PDF '%s': %w", inputPath, err)
}
}
return nil
return nil
})
}
// FormDataPdfEncrypt extracts the encryption parameters and permissions from
@@ -703,14 +796,14 @@ func EncryptPdfStub(ctx *api.Context, engine gotenberg.PdfEngine, opts gotenberg
return nil
}
for _, inputPath := range inputPaths {
return forEachInputPath(ctx, inputPaths, func(inputPath string) error {
err := engine.Encrypt(ctx, ctx.Log(), inputPath, opts)
if err != nil {
return fmt.Errorf("encrypt PDF '%s': %w", inputPath, err)
}
}
return nil
return nil
})
}
// EmbedFilesStub embeds files into PDF files.
@@ -739,170 +832,173 @@ func EmbedFilesStub(ctx *api.Context, engine gotenberg.PdfEngine, embedPaths []s
resolvedPaths[i] = resolvedPath
}
for _, inputPath := range inputPaths {
return forEachInputPath(ctx, inputPaths, func(inputPath string) error {
err := engine.EmbedFiles(ctx, ctx.Log(), resolvedPaths, inputPath)
if err != nil {
return fmt.Errorf("embed files into PDF '%s': %w", inputPath, err)
}
return nil
})
}
// FormDataPdfStamps builds the ordered list of stamps from the repeated stamp
// fields. See [formDataPdfStampsOrWatermarks].
func FormDataPdfStamps(form *api.FormData) ([]gotenberg.Stamp, error) {
return formDataPdfStampsOrWatermarks(form, "stamp")
}
// FormDataPdfWatermarks builds the ordered list of watermarks from the repeated
// watermark fields. See [formDataPdfStampsOrWatermarks].
func FormDataPdfWatermarks(form *api.FormData) ([]gotenberg.Stamp, error) {
return formDataPdfStampsOrWatermarks(form, "watermark")
}
// formDataPdfStampsOrWatermarks builds the ordered list of stamps or watermarks
// from the repeated {prefix}Source, {prefix}Expression, {prefix}Pages and
// {prefix}Options fields. The number of entries equals the number of
// {prefix}Source values, so a single occurrence of each field yields one entry,
// preserving the single-stamp/watermark behavior. Fields are aligned by
// position; a missing expression, pages or options entry defaults to empty.
// Image and pdf entries take their file from the uploaded files, in order (see
// [bindStampOrWatermarkFiles]).
func formDataPdfStampsOrWatermarks(form *api.FormData, prefix string) ([]gotenberg.Stamp, error) {
var sources, expressions, pages, options []string
form.
Strings(prefix+"Source", &sources).
Strings(prefix+"Expression", &expressions).
Strings(prefix+"Pages", &pages).
Strings(prefix+"Options", &options)
at := func(values []string, i int) string {
if i < len(values) {
return values[i]
}
return ""
}
stamps := make([]gotenberg.Stamp, 0, len(sources))
for i, source := range sources {
if source != gotenberg.StampSourceText && source != gotenberg.StampSourceImage && source != gotenberg.StampSourcePDF {
return nil, api.WrapError(
fmt.Errorf("wrong %sSource value '%s'", prefix, source),
api.NewSentinelHttpError(
http.StatusBadRequest,
fmt.Sprintf("Invalid form data: form field '%sSource' is invalid (got '%s', resulting to wrong value, expected either '%s', '%s' or '%s')", prefix, source, gotenberg.StampSourceText, gotenberg.StampSourceImage, gotenberg.StampSourcePDF),
),
)
}
var opts map[string]string
if raw := at(options, i); raw != "" {
err := json.Unmarshal([]byte(raw), &opts)
if err != nil {
return nil, api.WrapError(
fmt.Errorf("unmarshal %sOptions: %w", prefix, err),
api.NewSentinelHttpError(
http.StatusBadRequest,
fmt.Sprintf("Invalid form data: form field '%sOptions' is invalid", prefix),
),
)
}
}
stamps = append(stamps, gotenberg.Stamp{
Source: source,
Expression: at(expressions, i),
Pages: at(pages, i),
Options: opts,
})
}
return stamps, nil
}
// BindStampFiles assigns each image or pdf stamp its uploaded stamp file. See
// [bindStampOrWatermarkFiles].
func BindStampFiles(stamps []gotenberg.Stamp, stampFiles []string) error {
return bindStampOrWatermarkFiles(stamps, stampFiles, "stamp")
}
// BindWatermarkFiles assigns each image or pdf watermark its uploaded watermark
// file. See [bindStampOrWatermarkFiles].
func BindWatermarkFiles(watermarks []gotenberg.Stamp, watermarkFiles []string) error {
return bindStampOrWatermarkFiles(watermarks, watermarkFiles, "watermark")
}
// bindStampOrWatermarkFiles assigns each image or pdf entry its uploaded file,
// consuming files in order. Text entries take no file. It returns an [api] HTTP
// 400 error when an image or pdf entry has no file left to consume, which also
// prevents an anonymous caller from passing an arbitrary filesystem path via
// the expression field. kind is "stamp" or "watermark" and shapes the error.
func bindStampOrWatermarkFiles(stamps []gotenberg.Stamp, files []string, kind string) error {
fileIndex := 0
for i := range stamps {
if stamps[i].Source != gotenberg.StampSourceImage && stamps[i].Source != gotenberg.StampSourcePDF {
continue
}
if fileIndex >= len(files) {
return api.WrapError(
fmt.Errorf("not enough %s files for the image or pdf entries", kind),
api.NewSentinelHttpError(
http.StatusBadRequest,
fmt.Sprintf("Invalid form data: a %s file is required for image or pdf source", kind),
),
)
}
stamps[i].Expression = files[fileIndex]
fileIndex++
}
return nil
}
// FormDataPdfWatermark creates a [gotenberg.Stamp] for watermarking from the
// form data.
func FormDataPdfWatermark(form *api.FormData, mandatory bool) gotenberg.Stamp {
return formDataPdfStampOrWatermark(form, "watermark", mandatory)
}
// FormDataPdfStamp creates a [gotenberg.Stamp] for stamping from the form data.
func FormDataPdfStamp(form *api.FormData, mandatory bool) gotenberg.Stamp {
return formDataPdfStampOrWatermark(form, "stamp", mandatory)
}
func formDataPdfStampOrWatermark(form *api.FormData, prefix string, mandatory bool) gotenberg.Stamp {
var (
source string
expression string
pages string
options map[string]string
)
sourceFunc := func(value string) error {
if value != "" && value != gotenberg.StampSourceText && value != gotenberg.StampSourceImage && value != gotenberg.StampSourcePDF {
return fmt.Errorf("wrong value, expected either '%s', '%s' or '%s'", gotenberg.StampSourceText, gotenberg.StampSourceImage, gotenberg.StampSourcePDF)
// WatermarkStub applies each watermark to a list of PDF files, in order.
// Entries with no source are skipped, so an empty list does nothing.
func WatermarkStub(ctx *api.Context, engine gotenberg.PdfEngine, watermarks []gotenberg.Stamp, inputPaths []string) error {
// Watermarks stack on the same file, so the outer loop stays sequential;
// only the file dimension is parallel.
for _, watermark := range watermarks {
if watermark.Source == "" {
continue
}
source = value
return nil
}
optionsFunc := func(value string) error {
if value == "" {
err := forEachInputPath(ctx, inputPaths, func(inputPath string) error {
errWatermark := engine.Watermark(ctx, ctx.Log(), inputPath, watermark)
if errWatermark != nil {
return fmt.Errorf("watermark '%s': %w", inputPath, errWatermark)
}
return nil
}
err := json.Unmarshal([]byte(value), &options)
})
if err != nil {
return fmt.Errorf("unmarshal %s options: %w", prefix, err)
}
return nil
}
if mandatory {
form.
MandatoryCustom(prefix+"Source", func(value string) error {
return sourceFunc(value)
}).
String(prefix+"Expression", &expression, "").
String(prefix+"Pages", &pages, "").
Custom(prefix+"Options", func(value string) error {
return optionsFunc(value)
})
} else {
form.
Custom(prefix+"Source", func(value string) error {
return sourceFunc(value)
}).
String(prefix+"Expression", &expression, "").
String(prefix+"Pages", &pages, "").
Custom(prefix+"Options", func(value string) error {
return optionsFunc(value)
})
}
return gotenberg.Stamp{
Source: source,
Expression: expression,
Pages: pages,
Options: options,
}
}
// FormDataPdfWatermarkFile extracts the watermark file path from form data.
func FormDataPdfWatermarkFile(form *api.FormData) string {
var path string
form.Watermark(&path)
return path
}
// FormDataPdfStampFile extracts the stamp file path from form data.
func FormDataPdfStampFile(form *api.FormData) string {
var path string
form.Stamp(&path)
return path
}
// EnsureStampFile validates that, when stamp.Source is image or pdf, an
// uploaded stamp file was supplied, and replaces stamp.Expression with
// uploadedFile in that case. Returning an [api] HTTP 400 error prevents
// an anonymous caller from passing an arbitrary filesystem path via
// stampExpression and having pdfcpu read it. Source values of text or
// empty are passed through unchanged.
func EnsureStampFile(stamp *gotenberg.Stamp, uploadedFile string) error {
if stamp.Source != gotenberg.StampSourceImage && stamp.Source != gotenberg.StampSourcePDF {
return nil
}
if uploadedFile == "" {
return api.WrapError(
errors.New("no stamp file provided for image or pdf source"),
api.NewSentinelHttpError(
http.StatusBadRequest,
"Invalid form data: a stamp file is required for image or pdf source",
),
)
}
stamp.Expression = uploadedFile
return nil
}
// EnsureWatermarkFile mirrors [EnsureStampFile] for a watermark. The
// shape is identical: image or pdf sources must be accompanied by an
// uploaded file, and the file path replaces watermark.Expression to
// prevent pdfcpu from reading an attacker-controlled path.
func EnsureWatermarkFile(watermark *gotenberg.Stamp, uploadedFile string) error {
if watermark.Source != gotenberg.StampSourceImage && watermark.Source != gotenberg.StampSourcePDF {
return nil
}
if uploadedFile == "" {
return api.WrapError(
errors.New("no watermark file provided for image or pdf source"),
api.NewSentinelHttpError(
http.StatusBadRequest,
"Invalid form data: a watermark file is required for image or pdf source",
),
)
}
watermark.Expression = uploadedFile
return nil
}
// WatermarkStub applies a watermark to a list of PDF files. If the stamp has
// no source, it does nothing.
func WatermarkStub(ctx *api.Context, engine gotenberg.PdfEngine, stamp gotenberg.Stamp, inputPaths []string) error {
if stamp.Source == "" {
return nil
}
for _, inputPath := range inputPaths {
err := engine.Watermark(ctx, ctx.Log(), inputPath, stamp)
if err != nil {
return fmt.Errorf("watermark '%s': %w", inputPath, err)
return err
}
}
return nil
}
// StampStub applies a stamp to a list of PDF files. If the stamp has
// no source, it does nothing.
func StampStub(ctx *api.Context, engine gotenberg.PdfEngine, stamp gotenberg.Stamp, inputPaths []string) error {
if stamp.Source == "" {
return nil
}
// StampStub applies each stamp to a list of PDF files, in order. Entries with
// no source are skipped, so an empty list does nothing.
func StampStub(ctx *api.Context, engine gotenberg.PdfEngine, stamps []gotenberg.Stamp, inputPaths []string) error {
// Stamps stack on the same file, so the outer loop stays sequential; only
// the file dimension is parallel.
for _, stamp := range stamps {
if stamp.Source == "" {
continue
}
for _, inputPath := range inputPaths {
err := engine.Stamp(ctx, ctx.Log(), inputPath, stamp)
err := forEachInputPath(ctx, inputPaths, func(inputPath string) error {
errStamp := engine.Stamp(ctx, ctx.Log(), inputPath, stamp)
if errStamp != nil {
return fmt.Errorf("stamp '%s': %w", inputPath, errStamp)
}
return nil
})
if err != nil {
return fmt.Errorf("stamp '%s': %w", inputPath, err)
return err
}
}
@@ -924,33 +1020,42 @@ func mergeRoute(engine gotenberg.PdfEngine) api.Route {
bookmarks := FormDataPdfBookmarks(form, false)
encrypt := FormDataPdfEncrypt(form)
embedPaths := FormDataPdfEmbeds(form)
watermark := FormDataPdfWatermark(form, false)
watermarkFile := FormDataPdfWatermarkFile(form)
stamp := FormDataPdfStamp(form, false)
stampFile := FormDataPdfStampFile(form)
watermarks, wErr := FormDataPdfWatermarks(form)
if wErr != nil {
return fmt.Errorf("form data watermarks: %w", wErr)
}
stamps, sErr := FormDataPdfStamps(form)
if sErr != nil {
return fmt.Errorf("form data stamps: %w", sErr)
}
var watermarkFiles, stampFiles []string
form.Watermarks(&watermarkFiles).Stamps(&stampFiles)
angle, rotatePages := FormDataPdfRotate(form, false)
embedsMetadata := FormDataPdfEmbedsMetadata(form)
facturX, facturxXmlPath := FormDataPdfFacturX(form)
optimizeImages, imageQuality := FormDataPdfOptimize(form)
var inputPaths []string
var flatten bool
var autoIndexBookmarks bool
var titleBookmarks bool
err := form.
MandatoryPaths([]string{".pdf"}, &inputPaths).
Bool("flatten", &flatten, false).
Bool("autoIndexBookmarks", &autoIndexBookmarks, false).
Bool("titleBookmarks", &titleBookmarks, false).
Validate()
if err != nil {
return fmt.Errorf("validate form data: %w", err)
}
err = EnsureWatermarkFile(&watermark, watermarkFile)
err = BindWatermarkFiles(watermarks, watermarkFiles)
if err != nil {
return fmt.Errorf("validate watermark: %w", err)
return fmt.Errorf("bind watermark files: %w", err)
}
err = EnsureStampFile(&stamp, stampFile)
err = BindStampFiles(stamps, stampFiles)
if err != nil {
return fmt.Errorf("validate stamp: %w", err)
return fmt.Errorf("bind stamp files: %w", err)
}
err = ValidatePdfFormatsCompat(pdfFormats, encrypt.UserPassword, embedPaths)
@@ -976,12 +1081,12 @@ func mergeRoute(engine gotenberg.PdfEngine) api.Route {
outputPaths := []string{outputPath}
err = WatermarkStub(ctx, engine, watermark, outputPaths)
err = WatermarkStub(ctx, engine, watermarks, outputPaths)
if err != nil {
return fmt.Errorf("watermark PDFs: %w", err)
}
err = StampStub(ctx, engine, stamp, outputPaths)
err = StampStub(ctx, engine, stamps, outputPaths)
if err != nil {
return fmt.Errorf("stamp PDFs: %w", err)
}
@@ -998,6 +1103,11 @@ func mergeRoute(engine gotenberg.PdfEngine) api.Route {
}
}
err = OptimizeStub(ctx, engine, optimizeImages, imageQuality, outputPaths)
if err != nil {
return fmt.Errorf("optimize PDF images: %w", err)
}
pdfFormats = FacturXPdfFormats(ctx, engine, facturX, pdfFormats, false, outputPaths)
outputPaths, err = ConvertStub(ctx, engine, pdfFormats, outputPaths)
@@ -1013,7 +1123,7 @@ func mergeRoute(engine gotenberg.PdfEngine) api.Route {
finalBookmarks = b
} else {
bMap, _ := bookmarks.(map[string][]gotenberg.Bookmark)
if bMap != nil || autoIndexBookmarks {
if bMap != nil || autoIndexBookmarks || titleBookmarks {
offset := 0
for _, inputPath := range inputPaths {
filename := ctx.OriginalFilename(inputPath)
@@ -1023,7 +1133,9 @@ func mergeRoute(engine gotenberg.PdfEngine) api.Route {
fileBookmarks = bMap[filename]
}
if len(fileBookmarks) == 0 && autoIndexBookmarks {
// titleBookmarks nests each input's own outline under
// its title entry, so its outline is read here too.
if len(fileBookmarks) == 0 && (autoIndexBookmarks || titleBookmarks) {
fb, err := engine.ReadBookmarks(ctx, ctx.Log(), inputPath)
if err != nil {
return fmt.Errorf("read bookmarks of '%s': %w", filename, err)
@@ -1031,8 +1143,16 @@ func mergeRoute(engine gotenberg.PdfEngine) api.Route {
fileBookmarks = fb
}
if len(fileBookmarks) > 0 {
finalBookmarks = append(finalBookmarks, shiftBookmarks(fileBookmarks, offset)...)
fileBookmarks = shiftBookmarks(fileBookmarks, offset)
if titleBookmarks {
finalBookmarks = append(finalBookmarks, gotenberg.Bookmark{
Title: documentTitle(ctx, engine, inputPath, filename),
Page: offset + 1,
Children: fileBookmarks,
})
} else if len(fileBookmarks) > 0 {
finalBookmarks = append(finalBookmarks, fileBookmarks...)
}
pageCount, err := engine.PageCount(ctx, ctx.Log(), inputPath)
@@ -1101,13 +1221,20 @@ func splitRoute(engine gotenberg.PdfEngine) api.Route {
metadata := FormDataPdfMetadata(form, false)
encrypt := FormDataPdfEncrypt(form)
embedPaths := FormDataPdfEmbeds(form)
watermark := FormDataPdfWatermark(form, false)
watermarkFile := FormDataPdfWatermarkFile(form)
stamp := FormDataPdfStamp(form, false)
stampFile := FormDataPdfStampFile(form)
watermarks, wErr := FormDataPdfWatermarks(form)
if wErr != nil {
return fmt.Errorf("form data watermarks: %w", wErr)
}
stamps, sErr := FormDataPdfStamps(form)
if sErr != nil {
return fmt.Errorf("form data stamps: %w", sErr)
}
var watermarkFiles, stampFiles []string
form.Watermarks(&watermarkFiles).Stamps(&stampFiles)
angle, rotatePages := FormDataPdfRotate(form, false)
embedsMetadata := FormDataPdfEmbedsMetadata(form)
facturX, facturxXmlPath := FormDataPdfFacturX(form)
optimizeImages, imageQuality := FormDataPdfOptimize(form)
var inputPaths []string
var flatten bool
@@ -1119,13 +1246,13 @@ func splitRoute(engine gotenberg.PdfEngine) api.Route {
return fmt.Errorf("validate form data: %w", err)
}
err = EnsureWatermarkFile(&watermark, watermarkFile)
err = BindWatermarkFiles(watermarks, watermarkFiles)
if err != nil {
return fmt.Errorf("validate watermark: %w", err)
return fmt.Errorf("bind watermark files: %w", err)
}
err = EnsureStampFile(&stamp, stampFile)
err = BindStampFiles(stamps, stampFiles)
if err != nil {
return fmt.Errorf("validate stamp: %w", err)
return fmt.Errorf("bind stamp files: %w", err)
}
err = ValidatePdfFormatsCompat(pdfFormats, encrypt.UserPassword, embedPaths)
@@ -1148,12 +1275,12 @@ func splitRoute(engine gotenberg.PdfEngine) api.Route {
return fmt.Errorf("split PDFs: %w", err)
}
err = WatermarkStub(ctx, engine, watermark, outputPaths)
err = WatermarkStub(ctx, engine, watermarks, outputPaths)
if err != nil {
return fmt.Errorf("watermark PDFs: %w", err)
}
err = StampStub(ctx, engine, stamp, outputPaths)
err = StampStub(ctx, engine, stamps, outputPaths)
if err != nil {
return fmt.Errorf("stamp PDFs: %w", err)
}
@@ -1170,6 +1297,11 @@ func splitRoute(engine gotenberg.PdfEngine) api.Route {
}
}
err = OptimizeStub(ctx, engine, optimizeImages, imageQuality, outputPaths)
if err != nil {
return fmt.Errorf("optimize PDF images: %w", err)
}
pdfFormats = FacturXPdfFormats(ctx, engine, facturX, pdfFormats, false, outputPaths)
convertOutputPaths, err := ConvertStub(ctx, engine, pdfFormats, outputPaths)
@@ -1262,6 +1394,42 @@ func flattenRoute(engine gotenberg.PdfEngine) api.Route {
// convertRoute returns an [api.Route] which can convert PDFs to a specific ODF
// format.
func optimizeRoute(engine gotenberg.PdfEngine) api.Route {
return api.Route{
Method: http.MethodPost,
Path: "/forms/pdfengines/optimize",
IsMultipart: true,
Handler: func(c echo.Context) error {
ctx := c.Get("context").(*api.Context)
form := ctx.FormData()
// This route optimizes unconditionally, so the optimizeImages toggle
// is ignored; only the image quality is read.
_, imageQuality := FormDataPdfOptimize(form)
var inputPaths []string
err := form.
MandatoryPaths([]string{".pdf"}, &inputPaths).
Validate()
if err != nil {
return fmt.Errorf("validate form data: %w", err)
}
err = OptimizeStub(ctx, engine, true, imageQuality, inputPaths)
if err != nil {
return fmt.Errorf("optimize PDF images: %w", err)
}
err = ctx.AddOutputPaths(inputPaths...)
if err != nil {
return fmt.Errorf("add output paths: %w", err)
}
return nil
},
}
}
func convertRoute(engine gotenberg.PdfEngine) api.Route {
return api.Route{
Method: http.MethodPost,
@@ -1335,14 +1503,25 @@ func readMetadataRoute(engine gotenberg.PdfEngine) api.Route {
return fmt.Errorf("validate form data: %w", err)
}
res := make(map[string]map[string]any, len(inputPaths))
for _, inputPath := range inputPaths {
metadata, err := engine.ReadMetadata(ctx, ctx.Log(), inputPath)
if err != nil {
return fmt.Errorf("read metadata: %w", err)
// Collected per index, then folded into the map on this
// goroutine: a shared map cannot be written concurrently.
collected := make([]map[string]any, len(inputPaths))
err = forEachInputPathIndexed(ctx, inputPaths, func(i int, inputPath string) error {
metadata, errRead := engine.ReadMetadata(ctx, ctx.Log(), inputPath)
if errRead != nil {
return fmt.Errorf("read metadata: %w", errRead)
}
res[ctx.OriginalFilename(inputPath)] = metadata
collected[i] = metadata
return nil
})
if err != nil {
return err
}
res := make(map[string]map[string]any, len(inputPaths))
for i, inputPath := range inputPaths {
res[ctx.OriginalFilename(inputPath)] = collected[i]
}
err = c.JSON(http.StatusOK, res)
@@ -1413,14 +1592,25 @@ func readBookmarksRoute(engine gotenberg.PdfEngine) api.Route {
return fmt.Errorf("validate form data: %w", err)
}
res := make(map[string][]gotenberg.Bookmark, len(inputPaths))
for _, inputPath := range inputPaths {
bookmarks, err := engine.ReadBookmarks(ctx, ctx.Log(), inputPath)
if err != nil {
return fmt.Errorf("read bookmarks: %w", err)
// Collected per index, then folded into the map on this
// goroutine: a shared map cannot be written concurrently.
collected := make([][]gotenberg.Bookmark, len(inputPaths))
err = forEachInputPathIndexed(ctx, inputPaths, func(i int, inputPath string) error {
bookmarks, errRead := engine.ReadBookmarks(ctx, ctx.Log(), inputPath)
if errRead != nil {
return fmt.Errorf("read bookmarks: %w", errRead)
}
res[ctx.OriginalFilename(inputPath)] = bookmarks
collected[i] = bookmarks
return nil
})
if err != nil {
return err
}
res := make(map[string][]gotenberg.Bookmark, len(inputPaths))
for i, inputPath := range inputPaths {
res[ctx.OriginalFilename(inputPath)] = collected[i]
}
err = c.JSON(http.StatusOK, res)
@@ -1590,23 +1780,37 @@ func watermarkRoute(engine gotenberg.PdfEngine) api.Route {
ctx := c.Get("context").(*api.Context)
form := ctx.FormData()
stamp := FormDataPdfWatermark(form, true)
watermarkFile := FormDataPdfWatermarkFile(form)
watermarks, err := FormDataPdfWatermarks(form)
if err != nil {
return fmt.Errorf("form data watermarks: %w", err)
}
var inputPaths []string
err := form.
var watermarkFiles []string
err = form.
MandatoryPaths([]string{".pdf"}, &inputPaths).
Watermarks(&watermarkFiles).
Validate()
if err != nil {
return fmt.Errorf("validate form data: %w", err)
}
err = EnsureWatermarkFile(&stamp, watermarkFile)
if err != nil {
return fmt.Errorf("validate watermark: %w", err)
if len(watermarks) == 0 {
return api.WrapError(
errors.New("no watermark provided"),
api.NewSentinelHttpError(
http.StatusBadRequest,
"Invalid form data: form field 'watermarkSource' is required",
),
)
}
err = WatermarkStub(ctx, engine, stamp, inputPaths)
err = BindWatermarkFiles(watermarks, watermarkFiles)
if err != nil {
return fmt.Errorf("bind watermark files: %w", err)
}
err = WatermarkStub(ctx, engine, watermarks, inputPaths)
if err != nil {
return fmt.Errorf("watermark PDFs: %w", err)
}
@@ -1633,23 +1837,41 @@ func stampRoute(engine gotenberg.PdfEngine) api.Route {
ctx := c.Get("context").(*api.Context)
form := ctx.FormData()
stamp := FormDataPdfStamp(form, true)
stampFile := FormDataPdfStampFile(form)
// Reading the stamp fields as parallel arrays applies several
// stamps in one request. A single occurrence of each field is the
// existing single-stamp behavior.
stamps, err := FormDataPdfStamps(form)
if err != nil {
return fmt.Errorf("form data stamps: %w", err)
}
var inputPaths []string
err := form.
var stampFiles []string
err = form.
MandatoryPaths([]string{".pdf"}, &inputPaths).
Stamps(&stampFiles).
Validate()
if err != nil {
return fmt.Errorf("validate form data: %w", err)
}
err = EnsureStampFile(&stamp, stampFile)
if err != nil {
return fmt.Errorf("validate stamp: %w", err)
if len(stamps) == 0 {
return api.WrapError(
errors.New("no stamp provided"),
api.NewSentinelHttpError(
http.StatusBadRequest,
"Invalid form data: form field 'stampSource' is required",
),
)
}
err = StampStub(ctx, engine, stamp, inputPaths)
err = BindStampFiles(stamps, stampFiles)
if err != nil {
return fmt.Errorf("bind stamp files: %w", err)
}
err = StampStub(ctx, engine, stamps, inputPaths)
if err != nil {
return fmt.Errorf("stamp PDFs: %w", err)
}

View File

@@ -0,0 +1,181 @@
package pdfengines
import (
"errors"
"net/http"
"reflect"
"testing"
"github.com/gotenberg/gotenberg/v8/pkg/gotenberg"
"github.com/gotenberg/gotenberg/v8/pkg/modules/api"
)
func TestFormDataPdfStamps(t *testing.T) {
for _, tc := range []struct {
scenario string
values map[string][]string
expect []gotenberg.Stamp
expectErr bool
expectCode int
}{
{
scenario: "single text stamp (backward compatible)",
values: map[string][]string{
"stampSource": {"text"},
"stampExpression": {"CONFIDENTIAL"},
"stampOptions": {`{"rot":"45"}`},
},
expect: []gotenberg.Stamp{
{Source: "text", Expression: "CONFIDENTIAL", Options: map[string]string{"rot": "45"}},
},
},
{
scenario: "multiple stamps aligned by position",
values: map[string][]string{
"stampSource": {"text", "image"},
"stampExpression": {"ONE"},
"stampPages": {"1-2", "3"},
"stampOptions": {`{"pos":"tl"}`, `{"pos":"br"}`},
},
expect: []gotenberg.Stamp{
{Source: "text", Expression: "ONE", Pages: "1-2", Options: map[string]string{"pos": "tl"}},
{Source: "image", Expression: "", Pages: "3", Options: map[string]string{"pos": "br"}},
},
},
{
scenario: "no stamp fields",
values: map[string][]string{},
expect: []gotenberg.Stamp{},
},
{
scenario: "invalid source",
values: map[string][]string{"stampSource": {"text", "foo"}},
expectErr: true,
expectCode: http.StatusBadRequest,
},
{
scenario: "invalid options JSON",
values: map[string][]string{
"stampSource": {"text"},
"stampOptions": {"{"},
},
expectErr: true,
expectCode: http.StatusBadRequest,
},
} {
t.Run(tc.scenario, func(t *testing.T) {
ctx := &api.ContextMock{Context: &api.Context{}}
ctx.SetValues(tc.values)
form := ctx.FormData()
got, err := FormDataPdfStamps(form)
if tc.expectErr {
if err == nil {
t.Fatal("expected an error, got nil")
}
var httpErr api.HttpError
if !errors.As(err, &httpErr) {
t.Fatalf("expected an api.HttpError, got %T", err)
}
if status, _ := httpErr.HttpError(); status != tc.expectCode {
t.Fatalf("status = %d, want %d", status, tc.expectCode)
}
return
}
if err != nil {
t.Fatalf("expected no error, got %v", err)
}
if !reflect.DeepEqual(got, tc.expect) {
t.Fatalf("stamps = %#v, want %#v", got, tc.expect)
}
})
}
}
func TestBindStampFiles(t *testing.T) {
for _, tc := range []struct {
scenario string
stamps []gotenberg.Stamp
files []string
expect []gotenberg.Stamp
expectErr bool
expectCode int
}{
{
scenario: "text stamps consume no files",
stamps: []gotenberg.Stamp{{Source: "text", Expression: "FOO"}},
expect: []gotenberg.Stamp{{Source: "text", Expression: "FOO"}},
},
{
scenario: "image and pdf stamps consume files in order, overwriting expression",
stamps: []gotenberg.Stamp{
{Source: "image", Expression: "ignored"},
{Source: "text", Expression: "MIDDLE"},
{Source: "pdf"},
},
files: []string{"/a.png", "/b.pdf"},
expect: []gotenberg.Stamp{
{Source: "image", Expression: "/a.png"},
{Source: "text", Expression: "MIDDLE"},
{Source: "pdf", Expression: "/b.pdf"},
},
},
{
scenario: "not enough files for the image or pdf stamps",
stamps: []gotenberg.Stamp{{Source: "image"}, {Source: "image"}},
files: []string{"/a.png"},
expectErr: true,
expectCode: http.StatusBadRequest,
},
} {
t.Run(tc.scenario, func(t *testing.T) {
err := BindStampFiles(tc.stamps, tc.files)
if tc.expectErr {
if err == nil {
t.Fatal("expected an error, got nil")
}
var httpErr api.HttpError
if !errors.As(err, &httpErr) {
t.Fatalf("expected an api.HttpError, got %T", err)
}
if status, _ := httpErr.HttpError(); status != tc.expectCode {
t.Fatalf("status = %d, want %d", status, tc.expectCode)
}
return
}
if err != nil {
t.Fatalf("expected no error, got %v", err)
}
if !reflect.DeepEqual(tc.stamps, tc.expect) {
t.Fatalf("stamps = %#v, want %#v", tc.stamps, tc.expect)
}
})
}
}
func TestFormDataPdfWatermarks(t *testing.T) {
ctx := &api.ContextMock{Context: &api.Context{}}
ctx.SetValues(map[string][]string{
"watermarkSource": {"text", "image"},
"watermarkExpression": {"DRAFT"},
"watermarkOptions": {`{"opacity":"0.5"}`, ""},
})
form := ctx.FormData()
got, err := FormDataPdfWatermarks(form)
if err != nil {
t.Fatalf("expected no error, got %v", err)
}
want := []gotenberg.Stamp{
{Source: "text", Expression: "DRAFT", Options: map[string]string{"opacity": "0.5"}},
{Source: "image"},
}
if !reflect.DeepEqual(got, want) {
t.Fatalf("watermarks = %#v, want %#v", got, want)
}
}

View File

@@ -219,6 +219,20 @@ func (engine *PdfTk) Convert(ctx context.Context, logger *slog.Logger, formats g
return err
}
// OptimizeImages is not available in this implementation.
func (engine *PdfTk) OptimizeImages(ctx context.Context, logger *slog.Logger, imageQuality int, inputPath string) error {
_, span := gotenberg.Tracer().Start(ctx, "pdftk.OptimizeImages",
trace.WithSpanKind(trace.SpanKindClient),
trace.WithAttributes(engine.spanAttrs()...),
)
defer span.End()
err := fmt.Errorf("optimize PDF images with PDFtk: %w", gotenberg.ErrPdfEngineMethodNotSupported)
span.RecordError(err)
span.SetStatus(codes.Error, err.Error())
return err
}
// ReadMetadata is not available in this implementation.
func (engine *PdfTk) ReadMetadata(ctx context.Context, logger *slog.Logger, inputPath string) (map[string]any, error) {
_, span := gotenberg.Tracer().Start(ctx, "pdftk.ReadMetadata",

View File

@@ -120,6 +120,30 @@ func (engine *QPdf) spanAttrs(extra ...attribute.KeyValue) []attribute.KeyValue
return append(attrs, extra...)
}
// qpdfPageRange matches the page range syntax qpdf accepts, and nothing else.
//
// A term is a page number, "z" for the last page, or "rN" counting from the
// end, optionally prefixed with "x" to exclude it. Terms combine into ranges
// with "-", ranges join with ",", and the whole thing takes an optional ":odd"
// or ":even".
var qpdfPageRange = regexp.MustCompile(`^x?(?:z|r\d+|\d+)(?:-x?(?:z|r\d+|\d+))?(?:,x?(?:z|r\d+|\d+)(?:-x?(?:z|r\d+|\d+))?)*(?::odd|:even)?$`)
// validateSplitSpan returns span when it is a qpdf page range.
//
// qpdf reads the argument after "--pages ." as either a page range or another
// source file, so a span carrying a path makes qpdf append that file's pages
// to the output. Other engines in the split chain accept spellings qpdf does
// not, such as pdfcpu's "2-end", so a span this rejects is not necessarily
// invalid. Returning an error lets the chain move on to an engine that
// understands it.
func validateSplitSpan(span string) error {
if qpdfPageRange.MatchString(span) {
return nil
}
return fmt.Errorf("split span '%s' is not a QPDF page range: %w", span, gotenberg.ErrPdfSplitModeNotSupported)
}
// Split splits a given PDF file.
func (engine *QPdf) Split(ctx context.Context, logger *slog.Logger, mode gotenberg.SplitMode, inputPath, outputDirPath string) ([]string, error) {
ctx, span := gotenberg.Tracer().Start(ctx, "qpdf.Split",
@@ -139,6 +163,12 @@ func (engine *QPdf) Split(ctx context.Context, logger *slog.Logger, mode gotenbe
span.SetStatus(codes.Error, err.Error())
return nil, err
}
err := validateSplitSpan(mode.Span)
if err != nil {
span.RecordError(err)
span.SetStatus(codes.Error, err.Error())
return nil, err
}
args = append(args, inputPath)
args = append(args, engine.globalArgs...)
args = append(args, "--pages", ".", mode.Span)
@@ -255,6 +285,20 @@ func (engine *QPdf) Convert(ctx context.Context, logger *slog.Logger, formats go
return err
}
// OptimizeImages is not available in this implementation.
func (engine *QPdf) OptimizeImages(ctx context.Context, logger *slog.Logger, imageQuality int, inputPath string) error {
_, span := gotenberg.Tracer().Start(ctx, "qpdf.OptimizeImages",
trace.WithSpanKind(trace.SpanKindClient),
trace.WithAttributes(engine.spanAttrs()...),
)
defer span.End()
err := fmt.Errorf("optimize PDF images with QPDF: %w", gotenberg.ErrPdfEngineMethodNotSupported)
span.RecordError(err)
span.SetStatus(codes.Error, err.Error())
return err
}
// ReadMetadata is not available in this implementation.
func (engine *QPdf) ReadMetadata(ctx context.Context, logger *slog.Logger, inputPath string) (map[string]any, error) {
_, span := gotenberg.Tracer().Start(ctx, "qpdf.ReadMetadata",
@@ -454,7 +498,7 @@ func (engine *QPdf) EmbedFilesMetadata(ctx context.Context, logger *slog.Logger,
return err
}
catalogRef, catalogValue, filespecRefs, updateObjects := patchFilespecMetadata(logger, objects, metadata)
catalogRef, catalogValue, filespecRefs, updateObjects := patchFilespecMetadata(ctx, logger, objects, metadata)
if len(filespecRefs) == 0 {
span.SetStatus(codes.Ok, "")
return nil
@@ -506,7 +550,7 @@ func parsePdfObjects(output []byte) (map[string]json.RawMessage, error) {
// metadata keys. It sets /AFRelationship and /Subtype on matching objects
// and returns the catalog reference, catalog value, filespec references,
// and the update objects map.
func patchFilespecMetadata(logger *slog.Logger, objects map[string]json.RawMessage, metadata map[string]map[string]string) (string, map[string]any, []string, map[string]any) {
func patchFilespecMetadata(ctx context.Context, logger *slog.Logger, objects map[string]json.RawMessage, metadata map[string]map[string]string) (string, map[string]any, []string, map[string]any) {
updateObjects := make(map[string]any)
var catalogRef string
var catalogValue map[string]any
@@ -556,7 +600,7 @@ func patchFilespecMetadata(logger *slog.Logger, objects map[string]json.RawMessa
if ef, ok := value["/EF"].(map[string]any); ok {
efRef, _ := ef["/F"].(string)
if efRef != "" {
setStreamSubtype(logger, objects, updateObjects, efRef, mimeType)
setStreamSubtype(ctx, logger, objects, updateObjects, efRef, mimeType)
}
}
}
@@ -653,38 +697,38 @@ func (engine *QPdf) writeAndApplyUpdate(ctx context.Context, logger *slog.Logger
// setStreamSubtype finds a stream object by reference and sets the /Subtype
// key in its dict.
func setStreamSubtype(logger *slog.Logger, objects map[string]json.RawMessage, updateObjects map[string]any, ref, mimeType string) {
func setStreamSubtype(ctx context.Context, logger *slog.Logger, objects map[string]json.RawMessage, updateObjects map[string]any, ref, mimeType string) {
objKey := ref
if !strings.HasPrefix(objKey, "obj:") {
objKey = "obj:" + objKey
}
raw, ok := objects[objKey]
if !ok {
logger.Warn(fmt.Sprintf("set stream subtype on %s: object not found", ref))
logger.WarnContext(ctx, fmt.Sprintf("set stream subtype on %s: object not found", ref))
return
}
var obj map[string]json.RawMessage
if err := json.Unmarshal(raw, &obj); err != nil {
logger.Warn(fmt.Sprintf("set stream subtype on %s: unmarshal object: %s", ref, err))
logger.WarnContext(ctx, fmt.Sprintf("set stream subtype on %s: unmarshal object: %s", ref, err))
return
}
streamRaw, ok := obj["stream"]
if !ok {
logger.Warn(fmt.Sprintf("set stream subtype on %s: no stream key", ref))
logger.WarnContext(ctx, fmt.Sprintf("set stream subtype on %s: no stream key", ref))
return
}
var stream map[string]any
if err := json.Unmarshal(streamRaw, &stream); err != nil {
logger.Warn(fmt.Sprintf("set stream subtype on %s: unmarshal stream: %s", ref, err))
logger.WarnContext(ctx, fmt.Sprintf("set stream subtype on %s: unmarshal stream: %s", ref, err))
return
}
dict, ok := stream["dict"].(map[string]any)
if !ok {
logger.Warn(fmt.Sprintf("set stream subtype on %s: stream dict is not a map", ref))
logger.WarnContext(ctx, fmt.Sprintf("set stream subtype on %s: stream dict is not a map", ref))
return
}

View File

@@ -1,10 +1,14 @@
package qpdf
import (
"context"
"encoding/json"
"errors"
"log/slog"
"os"
"testing"
"github.com/gotenberg/gotenberg/v8/pkg/gotenberg"
)
func TestStripQpdfStringPrefix(t *testing.T) {
@@ -99,7 +103,7 @@ func TestPatchFilespecMetadata(t *testing.T) {
"factur-x.xml": {"relationship": "Data"},
}
catalogRef, _, filespecRefs, updateObjects := patchFilespecMetadata(logger, objects, metadata)
catalogRef, _, filespecRefs, updateObjects := patchFilespecMetadata(context.Background(), logger, objects, metadata)
if catalogRef != "obj:1 0 R" {
t.Errorf("catalogRef = %q, want %q", catalogRef, "obj:1 0 R")
@@ -125,7 +129,7 @@ func TestPatchFilespecMetadata(t *testing.T) {
"factur-x.xml": {"relationship": "Data"},
}
_, _, filespecRefs, _ := patchFilespecMetadata(logger, objects, metadata)
_, _, filespecRefs, _ := patchFilespecMetadata(context.Background(), logger, objects, metadata)
if len(filespecRefs) != 0 {
t.Errorf("filespecRefs = %v, want empty", filespecRefs)
}
@@ -139,7 +143,7 @@ func TestPatchFilespecMetadata(t *testing.T) {
"factur-x.xml": {"relationship": "Alternative"},
}
_, _, filespecRefs, updateObjects := patchFilespecMetadata(logger, objects, metadata)
_, _, filespecRefs, updateObjects := patchFilespecMetadata(context.Background(), logger, objects, metadata)
if len(filespecRefs) != 1 {
t.Fatalf("filespecRefs = %v, want 1 entry", filespecRefs)
}
@@ -158,7 +162,7 @@ func TestPatchFilespecMetadata(t *testing.T) {
"factur-x.xml": {"mimeType": "text/xml"},
}
_, _, _, updateObjects := patchFilespecMetadata(logger, objects, metadata)
_, _, _, updateObjects := patchFilespecMetadata(context.Background(), logger, objects, metadata)
streamObj, ok := updateObjects["obj:3 0 R"]
if !ok {
t.Fatal("expected obj:3 0 R in updateObjects")
@@ -223,7 +227,7 @@ func TestSetStreamSubtype(t *testing.T) {
}
updateObjects := make(map[string]any)
setStreamSubtype(logger, objects, updateObjects, "obj:3 0 R", "text/xml")
setStreamSubtype(context.Background(), logger, objects, updateObjects, "obj:3 0 R", "text/xml")
streamObj := updateObjects["obj:3 0 R"].(map[string]any)["stream"].(map[string]any)
dict := streamObj["dict"].(map[string]any)
@@ -238,7 +242,7 @@ func TestSetStreamSubtype(t *testing.T) {
}
updateObjects := make(map[string]any)
setStreamSubtype(logger, objects, updateObjects, "5 0 R", "application/pdf")
setStreamSubtype(context.Background(), logger, objects, updateObjects, "5 0 R", "application/pdf")
if _, ok := updateObjects["obj:5 0 R"]; !ok {
t.Error("expected obj:5 0 R in updateObjects")
@@ -249,7 +253,7 @@ func TestSetStreamSubtype(t *testing.T) {
objects := map[string]json.RawMessage{}
updateObjects := make(map[string]any)
setStreamSubtype(logger, objects, updateObjects, "obj:99 0 R", "text/xml")
setStreamSubtype(context.Background(), logger, objects, updateObjects, "obj:99 0 R", "text/xml")
if len(updateObjects) != 0 {
t.Error("expected no updates for missing object")
@@ -262,10 +266,61 @@ func TestSetStreamSubtype(t *testing.T) {
}
updateObjects := make(map[string]any)
setStreamSubtype(logger, objects, updateObjects, "obj:3 0 R", "text/xml")
setStreamSubtype(context.Background(), logger, objects, updateObjects, "obj:3 0 R", "text/xml")
if len(updateObjects) != 0 {
t.Error("expected no updates for non-stream object")
}
})
}
func TestValidateSplitSpan(t *testing.T) {
for _, tc := range []struct {
span string
valid bool
}{
// qpdf page ranges.
{"1", true},
{"12", true},
{"1-5", true},
{"2-z", true},
{"z", true},
{"r1", true},
{"r3-r1", true},
{"1,3,5-9", true},
{"1-5,x3", true},
{"1-z:odd", true},
{"1-z:even", true},
// Other engines' spellings. Not valid here, so the chain moves on.
{"2-end", false},
{"2-", false},
{"foo", false},
// A span qpdf would read as a source file.
{"/tmp/secret.pdf", false},
{"secret.pdf", false},
{"./secret.pdf", false},
{"../../etc/hosts", false},
{"1,/tmp/secret.pdf", false},
{"1 /tmp/secret.pdf", false},
{"", false},
{"--password=x", false},
} {
t.Run(tc.span, func(t *testing.T) {
err := validateSplitSpan(tc.span)
if tc.valid && err != nil {
t.Fatalf("validateSplitSpan(%q) = %v, want nil", tc.span, err)
}
if !tc.valid {
if err == nil {
t.Fatalf("validateSplitSpan(%q) = nil, want an error", tc.span)
}
// The chain must be able to try the next engine.
if !errors.Is(err, gotenberg.ErrPdfSplitModeNotSupported) {
t.Fatalf("error %v does not wrap ErrPdfSplitModeNotSupported", err)
}
}
})
}
}

View File

@@ -31,10 +31,33 @@ type client struct {
extraHttpHeaders map[string]string
startTime time.Time
// deliveryTimeout bounds one delivery including retries. See
// [Webhook.deliveryTimeout].
deliveryTimeout time.Duration
client *retryablehttp.Client
logger *slog.Logger
}
// deliveryContext returns the context one delivery runs on.
//
// It keeps the values of ctx, so trace propagation and logging correlation
// survive, and replaces its cancellation with a fresh budget. Threading the
// conversion context straight through does not work: a delivery starts after
// the handler returned, so that deadline may already be spent and the callback
// would fail without a single attempt.
func (c client) deliveryContext(ctx context.Context) (context.Context, context.CancelFunc) {
timeout := c.deliveryTimeout
if timeout <= 0 {
// An unset budget would expire the delivery before its first attempt.
// [Webhook.deliveryTimeout] never returns a non-positive value, so this
// only guards a caller that builds a client without one.
timeout = minDeliveryTimeout
}
return context.WithTimeout(context.WithoutCancel(ctx), timeout)
}
// send call the webhook either to send the success response or the error response.
func (c client) send(ctx context.Context, body io.Reader, headers map[string]string, errored bool) error {
url := c.url
@@ -57,6 +80,9 @@ func (c client) send(ctx context.Context, body io.Reader, headers map[string]str
spanName = fmt.Sprintf("%s Webhook Error", method)
}
ctx, cancel := c.deliveryContext(ctx)
defer cancel()
tracer := gotenberg.Tracer()
ctx, span := tracer.Start(ctx, spanName,
trace.WithSpanKind(trace.SpanKindClient),
@@ -64,7 +90,10 @@ func (c client) send(ctx context.Context, body io.Reader, headers map[string]str
)
defer span.End()
req, err := retryablehttp.NewRequest(method, url, body)
// The request must carry ctx: retryablehttp.NewRequest builds on
// [context.Background], and its wait between attempts is a select on the
// request context, so a contextless request cannot be interrupted.
req, err := retryablehttp.NewRequestWithContext(ctx, method, url, body)
if err != nil {
span.RecordError(err)
span.SetStatus(codes.Error, err.Error())
@@ -114,13 +143,12 @@ func (c client) send(ctx context.Context, body io.Reader, headers map[string]str
return fmt.Errorf("send '%s' request to '%s': %w", method, url, err)
}
if resp.StatusCode >= http.StatusBadRequest {
err := fmt.Errorf("send '%s' request to '%s': got status: '%s'", method, url, resp.Status)
span.RecordError(err)
span.SetStatus(codes.Error, err.Error())
return err
}
// Registered before the status check below. [retryablehttp.Client.Do] hands
// back a live body for a status it does not retry, which is every 4xx but
// 429, so returning early without closing it strands the connection and the
// transport goroutines that serve it for the lifetime of the process. The
// transport is built per delivery in [gotenberg.NewOutboundHttpClient], so
// nothing reclaims it later either.
defer func() {
err := resp.Body.Close()
if err != nil {
@@ -128,6 +156,13 @@ func (c client) send(ctx context.Context, body io.Reader, headers map[string]str
}
}()
if resp.StatusCode >= http.StatusBadRequest {
err := fmt.Errorf("send '%s' request to '%s': got status: '%s'", method, url, resp.Status)
span.RecordError(err)
span.SetStatus(codes.Error, err.Error())
return err
}
// Last piece for calculating the latency.
finishTime := time.Now()
@@ -165,6 +200,9 @@ func (c client) sendEvent(ctx context.Context, correlationIdHeader, correlationI
return
}
ctx, cancel := c.deliveryContext(ctx)
defer cancel()
tracer := gotenberg.Tracer()
ctx, span := tracer.Start(ctx, "POST Webhook Event",
trace.WithSpanKind(trace.SpanKindClient),
@@ -172,7 +210,7 @@ func (c client) sendEvent(ctx context.Context, correlationIdHeader, correlationI
)
defer span.End()
req, err := retryablehttp.NewRequest(http.MethodPost, c.eventsUrl, b)
req, err := retryablehttp.NewRequestWithContext(ctx, http.MethodPost, c.eventsUrl, b)
if err != nil {
span.RecordError(err)
span.SetStatus(codes.Error, err.Error())

View File

@@ -222,6 +222,7 @@ func webhookMiddleware(w *Webhook) api.Middleware {
eventsUrl: webhookEventsUrl,
extraHttpHeaders: extraHttpHeaders,
startTime: startTime,
deliveryTimeout: w.deliveryTimeout(),
client: &retryablehttp.Client{
HTTPClient: gotenberg.NewOutboundHttpClient(w.clientTimeout, w.allowList, w.denyList, w.enableEnvironmentProxy, ipOpts...),
@@ -230,7 +231,10 @@ func webhookMiddleware(w *Webhook) api.Middleware {
RetryWaitMax: w.retryMaxWait,
Logger: gotenberg.NewLeveledLogger(ctx.Log()),
CheckRetry: retryablehttp.DefaultRetryPolicy,
Backoff: retryablehttp.DefaultBackoff,
// Not DefaultBackoff: it returns a remote Retry-After
// verbatim, ignoring --webhook-retry-max-wait (env
// WEBHOOK_RETRY_MAX_WAIT).
Backoff: gotenberg.ClampedBackoff,
},
logger: ctx.Log(),
}

View File

@@ -42,7 +42,7 @@ func (w *Webhook) Descriptor() gotenberg.ModuleDescriptor {
FlagSet: func() *flag.FlagSet {
fs := flag.NewFlagSet("webhook", flag.ExitOnError)
fs.Bool("webhook-enable-sync-mode", false, "Enable synchronous mode for the webhook feature")
fs.StringSlice("webhook-allow-list", []string{}, "Set the allowed URLs for the webhook feature using regular expressions - supports multiple values")
fs.StringSlice("webhook-allow-list", []string{}, `Set the allowed URLs for the webhook feature using regular expressions - supports multiple values. A match bypasses --webhook-deny-private-ips (WEBHOOK_DENY_PRIVATE_IPS) and --webhook-deny-public-ips (WEBHOOK_DENY_PUBLIC_IPS), so terminate the host or the pattern also matches suffix hosts, for example ^https?://internal\.svc(:|/|$)`)
fs.StringSlice("webhook-deny-list", []string{}, "Set the denied URLs for the webhook feature using regular expressions - supports multiple values")
fs.Bool("webhook-deny-private-ips", false, "Reject webhook URLs whose host resolves to a non-public IP address (loopback, RFC1918, link-local, unique-local). Enable on deployments that accept untrusted webhook destinations to mitigate SSRF against internal services")
fs.Bool("webhook-deny-public-ips", false, "Reject webhook URLs whose host resolves to a public IP address. Enable on air-gapped or data-governed deployments to prevent callbacks from leaving a private network")
@@ -50,7 +50,7 @@ func (w *Webhook) Descriptor() gotenberg.ModuleDescriptor {
fs.Int("webhook-max-retry", 4, "Set the maximum number of retries for the webhook feature")
// Deprecated flags.
fs.StringSlice("webhook-error-allow-list", []string{}, "Set the allowed URLs in case of an error for the webhook feature using regular expressions - supports multiple values")
fs.StringSlice("webhook-error-allow-list", []string{}, `Set the allowed URLs in case of an error for the webhook feature using regular expressions - supports multiple values. A match bypasses --webhook-deny-private-ips (WEBHOOK_DENY_PRIVATE_IPS) and --webhook-deny-public-ips (WEBHOOK_DENY_PUBLIC_IPS), so terminate the host or the pattern also matches suffix hosts, for example ^https?://internal\.svc(:|/|$)`)
fs.StringSlice("webhook-error-deny-list", []string{}, "Set the denied URLs in case of an error for the webhook feature using regular expressions - supports multiple values")
err := fs.MarkDeprecated("webhook-error-allow-list", "use --webhook-allow-list instead")
if err != nil {
@@ -92,6 +92,27 @@ func (w *Webhook) Provision(ctx *gotenberg.Context) error {
return nil
}
// minDeliveryTimeout is the floor for [Webhook.deliveryTimeout], so that a
// deliberately tiny --webhook-client-timeout (env WEBHOOK_CLIENT_TIMEOUT) never
// leaves a delivery with no budget at all.
const minDeliveryTimeout = 1 * time.Second
// deliveryTimeout bounds one webhook delivery including its retries. It is the
// worst case a correctly behaving remote produces: one client timeout per
// attempt, plus the capped wait between attempts.
//
// A delivery runs after the handler returned, so it cannot borrow the
// conversion deadline. Without this budget it would have none, because
// [retryablehttp] builds its requests on [context.Background].
func (w *Webhook) deliveryTimeout() time.Duration {
timeout := w.clientTimeout*time.Duration(w.maxRetry+1) + w.retryMaxWait*time.Duration(w.maxRetry)
if timeout < minDeliveryTimeout {
return minDeliveryTimeout
}
return timeout
}
// Middlewares returns the middleware.
func (w *Webhook) Middlewares() ([]api.Middleware, error) {
if w.disable {

View File

@@ -0,0 +1,95 @@
package webhook
import (
"context"
"testing"
"time"
)
func TestWebhook_deliveryTimeout(t *testing.T) {
for _, tc := range []struct {
scenario string
clientTimeout time.Duration
maxRetry int
retryMaxWait time.Duration
want time.Duration
}{
{
scenario: "shipped defaults",
clientTimeout: 30 * time.Second,
maxRetry: 4,
retryMaxWait: 30 * time.Second,
want: 270 * time.Second,
},
{
scenario: "no retry is one client timeout",
clientTimeout: 30 * time.Second,
maxRetry: 0,
retryMaxWait: 30 * time.Second,
want: 30 * time.Second,
},
{
scenario: "a zero client timeout still gets a budget",
clientTimeout: 0,
maxRetry: 0,
retryMaxWait: 0,
want: minDeliveryTimeout,
},
} {
t.Run(tc.scenario, func(t *testing.T) {
w := &Webhook{
clientTimeout: tc.clientTimeout,
maxRetry: tc.maxRetry,
retryMaxWait: tc.retryMaxWait,
}
if got := w.deliveryTimeout(); got != tc.want {
t.Fatalf("deliveryTimeout() = %s, want %s", got, tc.want)
}
if w.deliveryTimeout() <= 0 {
t.Fatal("deliveryTimeout() must always be positive")
}
})
}
}
// A delivery must not inherit the conversion deadline. It runs after the
// handler returned, so that deadline is often already spent, which would fail
// the callback without a single attempt.
func TestClient_deliveryContext_IgnoresAnExpiredParentDeadline(t *testing.T) {
expired, cancelExpired := context.WithTimeout(context.Background(), -1*time.Second)
defer cancelExpired()
if expired.Err() == nil {
t.Fatal("expected the parent context to be expired")
}
c := client{deliveryTimeout: 30 * time.Second}
ctx, cancel := c.deliveryContext(expired)
defer cancel()
if ctx.Err() != nil {
t.Fatalf("delivery context inherited the expired parent: %v", ctx.Err())
}
deadline, ok := ctx.Deadline()
if !ok {
t.Fatal("delivery context has no deadline, so a delivery would be unbounded")
}
if remaining := time.Until(deadline); remaining <= 0 {
t.Fatalf("delivery budget = %s, want a positive value", remaining)
}
}
// The delivery context must still be bounded, so a hostile remote cannot hold
// the goroutine and its output file open indefinitely.
func TestClient_deliveryContext_IsBounded(t *testing.T) {
c := client{deliveryTimeout: 50 * time.Millisecond}
ctx, cancel := c.deliveryContext(context.Background())
defer cancel()
select {
case <-ctx.Done():
case <-time.After(5 * time.Second):
t.Fatal("delivery context never expired")
}
}

View File

@@ -23,12 +23,12 @@ make test-integration PLATFORM=linux/arm64 # force a specific platform
Available tags:
| Group | Tags |
| ----------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Chromium | `chromium`, `chromium-concurrent`, `chromium-convert-html`, `chromium-convert-markdown`, `chromium-convert-url`, `chromium-screenshot-html`, `chromium-screenshot-markdown`, `chromium-screenshot-url` |
| LibreOffice | `libreoffice`, `libreoffice-convert` |
| PDF Engines | `pdfengines`, `pdfengines-convert`, `pdfengines-merge`, `merge`, `pdfengines-split`, `split`, `pdfengines-flatten`, `flatten`, `pdfengines-rotate`, `rotate`, `pdfengines-embed`, `embed`, `pdfengines-encrypt`, `encrypt`, `pdfengines-watermark`, `watermark`, `pdfengines-stamp`, `stamp`, `pdfengines-metadata`, `metadata`, `pdfengines-bookmarks`, `bookmarks` |
| Infra | `health`, `debug`, `root`, `version`, `output-filename`, `prometheus-metrics`, `webhook`, `download-from` |
| Group | Tags |
| ----------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Chromium | `chromium`, `chromium-concurrent`, `chromium-convert-html`, `chromium-convert-markdown`, `chromium-convert-url`, `chromium-screenshot-html`, `chromium-screenshot-markdown`, `chromium-screenshot-url`, `chromium-ssrf` |
| LibreOffice | `libreoffice`, `libreoffice-convert`, `libreoffice-ssrf` |
| PDF Engines | `pdfengines`, `pdfengines-convert`, `pdfengines-merge`, `merge`, `pdfengines-split`, `split`, `pdfengines-flatten`, `flatten`, `pdfengines-optimize`, `optimize`, `pdfengines-rotate`, `rotate`, `pdfengines-embed`, `embed`, `pdfengines-encrypt`, `encrypt`, `pdfengines-watermark`, `watermark`, `pdfengines-stamp`, `stamp`, `pdfengines-metadata`, `metadata`, `pdfengines-bookmarks`, `bookmarks` |
| Infra | `health`, `debug`, `root`, `version`, `output-filename`, `prometheus-metrics`, `webhook`, `download-from` |
## Writing a new test

View File

@@ -62,6 +62,24 @@ Feature: /forms/chromium/convert/html
Page 12
"""
# A wide table on one landscape page: the page must expand its width to the
# content, so the rightmost column is not truncated and the page comes out
# landscape instead of the narrow-tall strip the height-only expansion used
# to produce. See https://github.com/gotenberg/gotenberg/issues/1390.
Scenario: POST /forms/chromium/convert/html (Single Page Landscape)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/html" endpoint with the following form data and header(s):
| files | testdata/wide-table-html/index.html | file |
| singlePage | true | field |
| landscape | true | field |
| Gotenberg-Output-Filename | foo | header |
Then the response status code should be 200
Then the response header "Content-Type" should be "application/pdf"
Then there should be 1 PDF(s) in the response
Then the "foo.pdf" PDF should have 1 page(s)
Then the "foo.pdf" PDF should be set to landscape orientation
Then the "foo.pdf" PDF should have content matching "Column 12" at page 1
Scenario: POST /forms/chromium/convert/html (Landscape)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/html" endpoint with the following form data and header(s):
@@ -195,6 +213,26 @@ Feature: /forms/chromium/convert/html
Wait delay > 2 seconds or expression window globalVar === 'ready' returns true.
"""
# A thenable (async) expression is awaited and its resolved value gates the
# print. Without awaiting, the Promise object never evaluates to true and the
# request would time out. See https://github.com/gotenberg/gotenberg/pull/1617.
Scenario: POST /forms/chromium/convert/html (Wait For Thenable Expression)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/html" endpoint with the following form data and header(s):
| files | testdata/feature-rich-html/index.html | file |
| waitForExpression | (async () => window.globalVar === 'ready')() | field |
| Gotenberg-Output-Filename | foo | header |
Then the response status code should be 200
Then the response header "Content-Type" should be "application/pdf"
Then there should be 1 PDF(s) in the response
Then there should be the following file(s) in the response:
| foo.pdf |
Then the "foo.pdf" PDF should have 1 page(s)
Then the "foo.pdf" PDF should have the following content at page 1:
"""
Wait delay > 2 seconds or expression window globalVar === 'ready' returns true.
"""
Scenario: POST /forms/chromium/convert/html (rAF / ResizeObserver / IntersectionObserver fire with waitForExpression)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/html" endpoint with the following form data and header(s):
@@ -416,6 +454,87 @@ Feature: /forms/chromium/convert/html
Then the Gotenberg container should log the following entries:
| 'file:///etc/passwd' matches the expression from the denied list |
# Control for the WebSocket scenario below. An ordinary fetch to a loopback
# address is surfaced as a Fetch.requestPaused event, so it is blocked by
# CHROMIUM_DENY_PRIVATE_IPS and the block is logged. The allow-list is
# cleared because a matching allow-list entry bypasses the IP-based check.
@chromium-ssrf
Scenario: POST /forms/chromium/convert/html (Fetch to a non-public address is filtered)
Given I have a Gotenberg container with the following environment variable(s):
| CHROMIUM_ALLOW_LIST | |
| CHROMIUM_DENY_PRIVATE_IPS | true |
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/html" endpoint with the following form data and header(s):
| files | testdata/ssrf-fetch-html/index.html | file |
| waitDelay | 1s | field |
Then the response status code should be 200
Then the Gotenberg container should log the following entries:
| 'http://127.0.0.1:9999/ssrf-fetch' targets a non-public address |
# A WebSocket handshake is never surfaced as a Fetch.requestPaused event, so
# it escapes the filter in listenForEventRequestPaused. The page opens
# WebSockets to two non-public addresses (loopback and the link-local cloud
# metadata IP). listenForEventWebSocketCreated logs each disallowed handshake
# with its full ws:// URL (detection), and the pinning proxy severs the
# connection now that the implicit loopback bypass is removed (enforcement).
@chromium-ssrf
Scenario: POST /forms/chromium/convert/html (WebSocket to a non-public address is filtered)
Given I have a Gotenberg container with the following environment variable(s):
| CHROMIUM_ALLOW_LIST | |
| CHROMIUM_DENY_PRIVATE_IPS | true |
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/html" endpoint with the following form data and header(s):
| files | testdata/ssrf-websocket-html/index.html | file |
| waitDelay | 1s | field |
Then the response status code should be 200
Then the Gotenberg container should log the following entries:
| 'ws://127.0.0.1:9999/ssrf-websocket' targets a non-public address |
| CONNECT blocked for '127.0.0.1:9999' |
# A Web Worker is a separate CDP target, so its WebSocket handshake is not
# observed by listenForEventWebSocketCreated. Enforcement must not depend on
# that listener: the pinning proxy sees the handshake and severs it whatever
# the originating context. Only the proxy's block is asserted, since no
# detection log is produced for the worker target.
@chromium-ssrf
Scenario: POST /forms/chromium/convert/html (WebSocket from a Web Worker is filtered)
Given I have a Gotenberg container with the following environment variable(s):
| CHROMIUM_ALLOW_LIST | |
| CHROMIUM_DENY_PRIVATE_IPS | true |
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/html" endpoint with the following form data and header(s):
| files | testdata/ssrf-websocket-worker-html/index.html | file |
| waitDelay | 1s | field |
Then the response status code should be 200
Then the Gotenberg container should log the following entries:
| CONNECT blocked for '127.0.0.1:9999' |
# wss:// (TLS) handshakes tunnel through the proxy via CONNECT, the same path
# as ws://, and must be filtered identically.
@chromium-ssrf
Scenario: POST /forms/chromium/convert/html (Secure WebSocket to a non-public address is filtered)
Given I have a Gotenberg container with the following environment variable(s):
| CHROMIUM_ALLOW_LIST | |
| CHROMIUM_DENY_PRIVATE_IPS | true |
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/html" endpoint with the following form data and header(s):
| files | testdata/ssrf-websocket-tls-html/index.html | file |
| waitDelay | 1s | field |
Then the response status code should be 200
Then the Gotenberg container should log the following entries:
| 'wss://127.0.0.1:9999/wss-test' targets a non-public address |
| CONNECT blocked for '127.0.0.1:9999' |
# EventSource issues an ordinary HTTP GET, so unlike a WebSocket it IS surfaced
# as a fetch.EventRequestPaused and blocked by listenForEventRequestPaused.
@chromium-ssrf
Scenario: POST /forms/chromium/convert/html (EventSource to a non-public address is filtered)
Given I have a Gotenberg container with the following environment variable(s):
| CHROMIUM_ALLOW_LIST | |
| CHROMIUM_DENY_PRIVATE_IPS | true |
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/html" endpoint with the following form data and header(s):
| files | testdata/ssrf-eventsource-html/index.html | file |
| waitDelay | 1s | field |
Then the response status code should be 200
Then the Gotenberg container should log the following entries:
| 'http://127.0.0.1:9999/sse' targets a non-public address |
Scenario: POST /forms/chromium/convert/html (Main URL does NOT match allowed list)
Given I have a Gotenberg container with the following environment variable(s):
| CHROMIUM_ALLOW_LIST | ^file:(?!//\\/tmp/).* |
@@ -1088,6 +1207,26 @@ Feature: /forms/chromium/convert/html
Then there should be 1 PDF(s) in the response
Then the response PDF(s) should be flatten
# Post-processing image optimization re-encodes the embedded lossless image to
# JPEG. The same page is ~700 KB without it and well under 300 KB with it.
# See https://github.com/gotenberg/gotenberg/issues/359.
Scenario: POST /forms/chromium/convert/html (Optimize Images)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/html" endpoint with the following form data and header(s):
| files | testdata/optimize-image-html/index.html | file |
| files | testdata/optimize-image-html/image.png | file |
| Gotenberg-Output-Filename | foo | header |
Then the response status code should be 200
Then the "foo.pdf" file size should be greater than 300 KB
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/html" endpoint with the following form data and header(s):
| files | testdata/optimize-image-html/index.html | file |
| files | testdata/optimize-image-html/image.png | file |
| optimizeImages | true | field |
| Gotenberg-Output-Filename | foo | header |
Then the response status code should be 200
Then the response header "Content-Type" should be "application/pdf"
Then the "foo.pdf" file size should be less than 300 KB
@encrypt
Scenario: POST /forms/chromium/convert/html (Encrypt - user password only)
Given I have a default Gotenberg container

View File

@@ -22,6 +22,41 @@ Feature: /forms/chromium/convert/url
Page 1
"""
# localStorage is per-origin and shared by the long-lived browser, so without
# clearing it accumulates across same-origin conversions.
# See https://github.com/gotenberg/gotenberg/issues/919.
Scenario: POST /forms/chromium/convert/url (localStorage leaks without clearing)
Given I have a default Gotenberg container
Given I have a static server
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/url" endpoint with the following form data and header(s):
| url | http://host.docker.internal:%d/html/testdata/local-storage-html/index.html | field |
| Gotenberg-Output-Filename | foo | header |
Then the response status code should be 200
Then the "foo.pdf" PDF should have content matching "localStorageCount=1" at page 1
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/url" endpoint with the following form data and header(s):
| url | http://host.docker.internal:%d/html/testdata/local-storage-html/index.html | field |
| Gotenberg-Output-Filename | foo | header |
Then the response status code should be 200
Then the "foo.pdf" PDF should have content matching "localStorageCount=2" at page 1
# With --chromium-clear-storage every conversion clears the origin's
# localStorage first, so the counter never carries over.
# See https://github.com/gotenberg/gotenberg/issues/919.
Scenario: POST /forms/chromium/convert/url (Clear Storage)
Given I have a Gotenberg container with the following environment variable(s):
| CHROMIUM_CLEAR_STORAGE | true |
Given I have a static server
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/url" endpoint with the following form data and header(s):
| url | http://host.docker.internal:%d/html/testdata/local-storage-html/index.html | field |
| Gotenberg-Output-Filename | foo | header |
Then the response status code should be 200
Then the "foo.pdf" PDF should have content matching "localStorageCount=1" at page 1
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/url" endpoint with the following form data and header(s):
| url | http://host.docker.internal:%d/html/testdata/local-storage-html/index.html | field |
| Gotenberg-Output-Filename | foo | header |
Then the response status code should be 200
Then the "foo.pdf" PDF should have content matching "localStorageCount=1" at page 1
Scenario: POST /forms/chromium/convert/url (Single Page)
Given I have a default Gotenberg container
Given I have a static server
@@ -499,6 +534,7 @@ Feature: /forms/chromium/convert/url
Then the response header "Content-Type" should be "application/pdf"
Then there should be 1 PDF(s) in the response
@chromium-ssrf
Scenario: POST /forms/chromium/convert/url (Main URL is a non-public IP literal, deny-private-ips on)
Given I have a Gotenberg container with the following environment variable(s):
| CHROMIUM_ALLOW_LIST | |
@@ -512,6 +548,55 @@ Feature: /forms/chromium/convert/url
Forbidden
"""
# IPv6 loopback literal is parsed as an IP and rejected by the IP-class check,
# like the IPv4 loopback literal above.
@chromium-ssrf
Scenario: POST /forms/chromium/convert/url (Main URL is an IPv6 loopback literal, deny-private-ips on)
Given I have a Gotenberg container with the following environment variable(s):
| CHROMIUM_ALLOW_LIST | |
| CHROMIUM_DENY_PRIVATE_IPS | true |
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/url" endpoint with the following form data and header(s):
| url | http://[::1]/ | field |
Then the response status code should be 403
Then the response header "Content-Type" should be "text/plain; charset=UTF-8"
Then the response body should match string:
"""
Forbidden
"""
# An alternate IP encoding (decimal for 127.0.0.1) that Chromium would read
# as loopback but the resolver rejects as a hostname. It must fail closed as
# filtered (a generic 403), not surface as a 500.
@chromium-ssrf
Scenario: POST /forms/chromium/convert/url (Main URL is a decimal-encoded loopback IP, deny-private-ips on)
Given I have a Gotenberg container with the following environment variable(s):
| CHROMIUM_ALLOW_LIST | |
| CHROMIUM_DENY_PRIVATE_IPS | true |
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/url" endpoint with the following form data and header(s):
| url | http://2130706433/ | field |
Then the response status code should be 403
Then the response header "Content-Type" should be "text/plain; charset=UTF-8"
Then the response body should match string:
"""
Forbidden
"""
# A classic SSRF vector: an allow-listed URL that redirects to an internal
# address. The redirected request must not inherit the initial URL's
# allow-list pass. listenForEventRequestPaused re-validates it; it does not
# match the allow-list, so it is blocked (Chromium reports ERR_ACCESS_DENIED
# and the conversion renders the resulting error page).
@chromium-ssrf
Scenario: POST /forms/chromium/convert/url (Redirect to a non-allow-listed address is re-filtered)
Given I have a Gotenberg container with the following environment variable(s):
| CHROMIUM_ALLOW_LIST | ^https?://host\\.docker\\.internal(:[0-9]+)?/ |
Given I have a static server
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/url" endpoint with the following form data and header(s):
| url | http://host.docker.internal:%d/redirect-to-private | field |
Then the response status code should be 200
Then the Gotenberg container should log the following entries:
| 'http://127.0.0.1:9999/redirected' does not match any expression from the allowed list |
Scenario: POST /forms/chromium/convert/url (Main URL resolves to a non-public IP, deny-private-ips on with allow-list bypass)
Given I have a Gotenberg container with the following environment variable(s):
| CHROMIUM_ALLOW_LIST | .+ |
@@ -1316,3 +1401,19 @@ Feature: /forms/chromium/convert/url
Then the response header "Content-Type" should be "application/pdf"
Then there should be 1 PDF(s) in the response
Then the "foo.pdf" PDF should have 1 page(s)
# chrome://crash makes the renderer crash deterministically, the same
# failure class as a renderer crash triggered by the page content. The
# request must fail fast with a 503 instead of hanging until the API
# timeout.
# See https://github.com/gotenberg/gotenberg/issues/1640.
Scenario: POST /forms/chromium/convert/url (Chromium crash fails fast with 503)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/chromium/convert/url" endpoint with the following form data and header(s):
| url | chrome://crash | field |
Then the response status code should be 503
Then the response header "Content-Type" should be "text/plain; charset=UTF-8"
Then the response body should contain string:
"""
Chromium crashed while processing the request
"""

View File

@@ -2,6 +2,22 @@
@chromium-screenshot-html
Feature: /forms/chromium/screenshot/html
# Route parity: the WebSocket outbound filter lives in the shared browser
# code path, so the screenshot route enforces it exactly like conversion.
@chromium-ssrf
Scenario: POST /forms/chromium/screenshot/html (WebSocket to a non-public address is filtered)
Given I have a Gotenberg container with the following environment variable(s):
| CHROMIUM_ALLOW_LIST | |
| CHROMIUM_DENY_PRIVATE_IPS | true |
When I make a "POST" request to Gotenberg at the "/forms/chromium/screenshot/html" endpoint with the following form data and header(s):
| files | testdata/ssrf-websocket-html/index.html | file |
| waitDelay | 1s | field |
Then the response status code should be 200
Then the response header "Content-Type" should be "image/png"
Then the Gotenberg container should log the following entries:
| 'ws://127.0.0.1:9999/ssrf-websocket' targets a non-public address |
| CONNECT blocked for '127.0.0.1:9999' |
Scenario: POST /forms/chromium/screenshot/html (Default)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/chromium/screenshot/html" endpoint with the following form data and header(s):
@@ -53,6 +69,32 @@ Feature: /forms/chromium/screenshot/html
Then the response status code should be 200
Then the response header "Content-Type" should be "image/png"
# The target element is 300x200 and sits below a spacer, so a correct clip
# proves both the element size and its page offset. See issue #947.
Scenario: POST /forms/chromium/screenshot/html (Selector)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/chromium/screenshot/html" endpoint with the following form data and header(s):
| files | testdata/screenshot-selector-html/index.html | file |
| selector | #target | field |
| Gotenberg-Output-Filename | foo | header |
Then the response status code should be 200
Then the response header "Content-Type" should be "image/png"
Then the "foo.png" image should be 300x200 pixels
# A centered red pixel proves the clip landed on the element, not on the
# white spacer above it, i.e. the page offset was applied.
Then the "foo.png" image pixel at 150,100 should be "#ff0000"
Scenario: POST /forms/chromium/screenshot/html (Selector Not Found)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/chromium/screenshot/html" endpoint with the following form data and header(s):
| files | testdata/screenshot-selector-html/index.html | file |
| selector | #does-not-exist | field |
Then the response status code should be 400
Then the response body should contain string:
"""
The selector '#does-not-exist' (selector) matched no element with a visible box
"""
Scenario: POST /forms/chromium/screenshot/html (Quality)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/chromium/screenshot/html" endpoint with the following form data and header(s):

View File

@@ -83,6 +83,7 @@ Feature: /debug
"chromium-auto-start": "false",
"chromium-clear-cache": "false",
"chromium-clear-cookies": "false",
"chromium-clear-storage": "false",
"chromium-deny-list": "[^file:(?!//\\/tmp/).*]",
"chromium-deny-private-ips": "false",
"chromium-deny-public-ips": "false",
@@ -221,6 +222,7 @@ Feature: /debug
"chromium-auto-start": "false",
"chromium-clear-cache": "false",
"chromium-clear-cookies": "false",
"chromium-clear-storage": "false",
"chromium-deny-list": "[^file:(?!//\\/tmp/).*]",
"chromium-deny-private-ips": "false",
"chromium-deny-public-ips": "false",

View File

@@ -1,6 +1,5 @@
# TODO:
# 1. Check if down for each module.
# 2. Restarting modules do not make health check fail.
@health
Feature: /health
@@ -106,7 +105,19 @@ Feature: /health
When I make a "HEAD" request to Gotenberg at the "/foo/health" endpoint
Then the response status code should be 200
# A planned restart, the eager cycle after LIBREOFFICE_RESTART_AFTER
# conversions, must not fail the health check: requests arriving during it
# are requeued, not rejected. Setting the limit to 1 restarts LibreOffice
# after every conversion, so each probe lands right on a restart.
# See https://github.com/gotenberg/gotenberg/issues/1648.
Scenario: GET /health (Planned LibreOffice Restart)
Given I have a Gotenberg container with the following environment variable(s):
| LIBREOFFICE_RESTART_AFTER | 1 |
When I make 5 sequential "POST" requests to Gotenberg at the "/forms/libreoffice/convert" endpoint, probing "/health" after each, with the following form data and header(s):
| files | testdata/page_1.docx | file |
Then all sequential response status codes should be 200
Then all probe response status codes should be 200
# TODO:
# 1. Check if down for each module.
# 2. Restarting modules do not make health check fail.
# 1. Check if down for each module.

View File

@@ -18,6 +18,33 @@ Feature: /forms/libreoffice/convert
Page 1
"""
# LibreOffice detects OOXML PowerPoint shows from content and converts them
# like .pptx; only Gotenberg's extension allow list gated them out.
# See https://github.com/gotenberg/gotenberg/pull/1626.
Scenario: POST /forms/libreoffice/convert (PowerPoint Show .ppsx)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/libreoffice/convert" endpoint with the following form data and header(s):
| files | testdata/slideshow.ppsx | file |
| Gotenberg-Output-Filename | foo | header |
Then the response status code should be 200
Then the response header "Content-Type" should be "application/pdf"
Then there should be 1 PDF(s) in the response
Then there should be the following file(s) in the response:
| foo.pdf |
Then the "foo.pdf" PDF should have 1 page(s)
Scenario: POST /forms/libreoffice/convert (PowerPoint Show .ppsm)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/libreoffice/convert" endpoint with the following form data and header(s):
| files | testdata/slideshow.ppsm | file |
| Gotenberg-Output-Filename | foo | header |
Then the response status code should be 200
Then the response header "Content-Type" should be "application/pdf"
Then there should be 1 PDF(s) in the response
Then there should be the following file(s) in the response:
| foo.pdf |
Then the "foo.pdf" PDF should have 1 page(s)
# A CSV becomes a single Calc sheet named after the input file, and Calc's
# default page style prints that sheet name as a centered header. Uploads are
# stored under a UUID-based filename, so the UUID must not leak into the PDF.
@@ -211,7 +238,7 @@ Feature: /forms/libreoffice/convert
Then the response header "Content-Type" should be "text/plain; charset=UTF-8"
Then the response body should match string:
"""
Invalid form data: no form file found for extensions: [.123 .602 .abw .bib .bmp .cdr .cgm .cmx .csv .cwk .dbf .dif .doc .docm .docx .dot .dotm .dotx .dxf .emf .eps .epub .fodg .fodp .fods .fodt .fopd .gif .htm .html .hwp .jpeg .jpg .key .ltx .lwp .mcw .met .mml .mw .numbers .odd .odg .odm .odp .ods .odt .otg .oth .otp .ots .ott .pages .pbm .pcd .pct .pcx .pdb .pdf .pgm .png .pot .potm .potx .ppm .pps .ppt .pptm .pptx .psd .psw .pub .pwp .pxl .ras .rtf .sda .sdc .sdd .sdp .sdw .sgl .slk .smf .stc .std .sti .stw .svg .svm .swf .sxc .sxd .sxg .sxi .sxm .sxw .tga .tif .tiff .txt .uof .uop .uos .uot .vdx .vor .vsd .vsdm .vsdx .wb2 .wk1 .wks .wmf .wpd .wpg .wps .xbm .xhtml .xls .xlsb .xlsm .xlsx .xlt .xltm .xltx .xlw .xml .xpm .zabw]
Invalid form data: no form file found for extensions: [.123 .602 .abw .bib .bmp .cdr .cgm .cmx .csv .cwk .dbf .dif .doc .docm .docx .dot .dotm .dotx .dxf .emf .eps .epub .fodg .fodp .fods .fodt .fopd .gif .htm .html .hwp .jpeg .jpg .key .ltx .lwp .mcw .met .mml .mw .numbers .odd .odg .odm .odp .ods .odt .otg .oth .otp .ots .ott .pages .pbm .pcd .pct .pcx .pdb .pdf .pgm .png .pot .potm .potx .ppm .pps .ppsm .ppsx .ppt .pptm .pptx .psd .psw .pub .pwp .pxl .ras .rtf .sda .sdc .sdd .sdp .sdw .sgl .slk .smf .stc .std .sti .stw .svg .svm .swf .sxc .sxd .sxg .sxi .sxm .sxw .tga .tif .tiff .txt .uof .uop .uos .uot .vdx .vor .vsd .vsdm .vsdx .wb2 .wk1 .wks .wmf .wpd .wpg .wps .xbm .xhtml .xls .xlsb .xlsm .xlsx .xlt .xltm .xltx .xlw .xml .xpm .zabw]
form field 'landscape' is invalid (got 'foo', resulting to strconv.ParseBool: parsing "foo": invalid syntax)
form field 'exportFormFields' is invalid (got 'foo', resulting to strconv.ParseBool: parsing "foo": invalid syntax)
form field 'allowDuplicateFieldNames' is invalid (got 'foo', resulting to strconv.ParseBool: parsing "foo": invalid syntax)
@@ -942,7 +969,7 @@ Feature: /forms/libreoffice/convert
# An embedded image is stored inside the document, not linked, so blocking
# untrusted linked content leaves it untouched. Guards against over-blocking.
@libreoffice-linked-content
@libreoffice-ssrf
Scenario: POST /forms/libreoffice/convert (Embedded Image Survives)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/libreoffice/convert" endpoint with the following form data and header(s):
@@ -955,7 +982,7 @@ Feature: /forms/libreoffice/convert
# An uploaded document always loads from an untrusted location, so soffice
# refuses to resolve any content it links (absolute file:// path or external
# URL). Closes the SSRF and local-file-read vector.
@libreoffice-linked-content
@libreoffice-ssrf
Scenario: POST /forms/libreoffice/convert (Linked External Resource Blocked)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/libreoffice/convert" endpoint with the following form data and header(s):
@@ -964,3 +991,17 @@ Feature: /forms/libreoffice/convert
Then the response status code should be 200
Then there should be 1 PDF(s) in the response
Then the "foo.pdf" PDF should have 0 image(s)
# The workbook was saved scrolled down to row 37. Without the topLeftCell
# reset, SinglePageSheets would start the page there and drop the header rows,
# so the "Meteor" column header only appears when the whole sheet is rendered.
# See https://github.com/gotenberg/gotenberg/issues/1222.
Scenario: POST /forms/libreoffice/convert (SinglePageSheets renders a scrolled workbook in full)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/libreoffice/convert" endpoint with the following form data and header(s):
| files | testdata/singlepagesheets-scrolled.xlsx | file |
| singlePageSheets | true | field |
| Gotenberg-Output-Filename | foo | header |
Then the response status code should be 200
Then there should be 1 PDF(s) in the response
Then the "foo.pdf" PDF should have content matching "Meteor" at page 1

View File

@@ -291,6 +291,49 @@ Feature: /forms/pdfengines/merge
}
"""
# titleBookmarks adds a top-level bookmark per merged document, labeled by its
# Title metadata (falling back to the filename) and pointing to its first page,
# with the document's own outline nested underneath.
# See https://github.com/gotenberg/gotenberg/issues/867.
@bookmarks
Scenario: POST /forms/pdfengines/merge (Title Bookmarks)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/pdfengines/merge" endpoint with the following form data and header(s):
| files | testdata/titled_alpha.pdf | file |
| files | testdata/titled_bravo.pdf | file |
| files | testdata/untitled_gamma.pdf | file |
| titleBookmarks | true | field |
| Gotenberg-Output-Filename | foo | header |
Then the response status code should be 200
Then the response header "Content-Type" should be "application/pdf"
When I make a "POST" request to Gotenberg at the "/forms/pdfengines/bookmarks/read" endpoint with the following form data and header(s):
| files | teststore/foo.pdf | file |
Then the response status code should be 200
Then the response header "Content-Type" should be "application/json"
Then the response body should match JSON:
"""
{
"foo.pdf": [
{
"title": "Alpha",
"page": 1,
"children": [
{ "title": "A1", "page": 1 },
{ "title": "A2", "page": 2 }
]
},
{
"title": "Bravo",
"page": 3
},
{
"title": "untitled_gamma",
"page": 4
}
]
}
"""
@bookmarks
Scenario: POST /forms/pdfengines/merge (Auto-index Bookmarks)
Given I have a default Gotenberg container

View File

@@ -0,0 +1,61 @@
@pdfengines
@pdfengines-optimize
@optimize
Feature: /forms/pdfengines/optimize
# image-heavy.pdf is a ~710 KB PDF whose single image Chromium embedded
# losslessly (FlateDecode). Re-encoding it to JPEG shrinks the file well
# below this threshold while leaving the structure intact.
# See https://github.com/gotenberg/gotenberg/issues/359.
Scenario: POST /forms/pdfengines/optimize (Image-heavy PDF)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/pdfengines/optimize" endpoint with the following form data and header(s):
| files | testdata/image-heavy.pdf | file |
| Gotenberg-Output-Filename | foo | header |
Then the response status code should be 200
Then the response header "Content-Type" should be "application/pdf"
Then there should be 1 PDF(s) in the response
Then the "foo.pdf" file size should be less than 300 KB
Scenario: POST /forms/pdfengines/optimize (Custom Image Quality)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/pdfengines/optimize" endpoint with the following form data and header(s):
| files | testdata/image-heavy.pdf | file |
| imageQuality | 40 | field |
| Gotenberg-Output-Filename | foo | header |
Then the response status code should be 200
Then the "foo.pdf" file size should be less than 300 KB
Scenario: POST /forms/pdfengines/optimize (PDF Without Images)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/pdfengines/optimize" endpoint with the following form data and header(s):
| files | testdata/page_1.pdf | file |
Then the response status code should be 200
Then the response header "Content-Type" should be "application/pdf"
Then there should be 1 PDF(s) in the response
Scenario: POST /forms/pdfengines/optimize (Invalid Image Quality)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/pdfengines/optimize" endpoint with the following form data and header(s):
| files | testdata/image-heavy.pdf | file |
| imageQuality | 200 | field |
Then the response status code should be 400
Then the response header "Content-Type" should be "text/plain; charset=UTF-8"
Scenario: POST /forms/pdfengines/optimize (Bad Request)
Given I have a default Gotenberg container
When I make a "POST" request to Gotenberg at the "/forms/pdfengines/optimize" endpoint with the following form data and header(s):
| Gotenberg-Output-Filename | foo | header |
Then the response status code should be 400
Then the response header "Content-Type" should be "text/plain; charset=UTF-8"
Then the response body should match string:
"""
Invalid form data: no form file found for extensions: [.pdf]
"""
Scenario: POST /forms/pdfengines/optimize (Routes Disabled)
Given I have a Gotenberg container with the following environment variable(s):
| PDFENGINES_DISABLE_ROUTES | true |
When I make a "POST" request to Gotenberg at the "/forms/pdfengines/optimize" endpoint with the following form data and header(s):
| files | testdata/image-heavy.pdf | file |
Then the response status code should be 404

View File

@@ -51,6 +51,39 @@ Feature: /forms/pdfengines/stamp
Then the response header "Content-Type" should be "application/pdf"
Then there should be 1 PDF(s) in the response
# Repeating the stamp fields applies several stamps in one request, in order.
# Image and pdf stamps consume the uploaded stamp files in order; text stamps
# take none. Both text stamps must land, so their content is asserted (the
# options keep them unrotated and apart so pdftotext reads them cleanly).
# See https://github.com/gotenberg/gotenberg/pull/1601.
Scenario: POST /forms/pdfengines/stamp (Multiple Stamps - pdfcpu)
Given I have a Gotenberg container with the following environment variable(s):
| PDFENGINES_STAMP_ENGINES | pdfcpu |
When I make a "POST" request to Gotenberg at the "/forms/pdfengines/stamp" endpoint with the following form data and header(s):
| files | testdata/page_1.pdf | file |
| stampSource | text | field |
| stampExpression | STAMPONE | field |
| stampOptions | {"rotation":"0","position":"tl","scale":"0.2 abs"} | field |
| stampSource | text | field |
| stampExpression | STAMPTWO | field |
| stampOptions | {"rotation":"0","position":"br","scale":"0.2 abs"} | field |
| stampSource | image | field |
| stamp | testdata/watermark.png | file |
| Gotenberg-Output-Filename | foo | header |
Then the response status code should be 200
Then the response header "Content-Type" should be "application/pdf"
Then there should be 1 PDF(s) in the response
Then there should be the following file(s) in the response:
| foo.pdf |
Then the "foo.pdf" PDF should have the following content at page 1:
"""
STAMPONE
"""
Then the "foo.pdf" PDF should have the following content at page 1:
"""
STAMPTWO
"""
Scenario: POST /forms/pdfengines/stamp (PDF - pdfcpu)
Given I have a Gotenberg container with the following environment variable(s):
| PDFENGINES_STAMP_ENGINES | pdfcpu |

View File

@@ -51,6 +51,39 @@ Feature: /forms/pdfengines/watermark
Then the response header "Content-Type" should be "application/pdf"
Then there should be 1 PDF(s) in the response
# Repeating the watermark fields applies several watermarks in one request, in
# order. Image and pdf watermarks consume the uploaded watermark files in
# order; text watermarks take none. Both text watermarks must land, so their
# content is asserted (the options keep them unrotated and apart so pdftotext
# reads them cleanly).
Scenario: POST /forms/pdfengines/watermark (Multiple Watermarks - pdfcpu)
Given I have a Gotenberg container with the following environment variable(s):
| PDFENGINES_WATERMARK_ENGINES | pdfcpu |
When I make a "POST" request to Gotenberg at the "/forms/pdfengines/watermark" endpoint with the following form data and header(s):
| files | testdata/page_1.pdf | file |
| watermarkSource | text | field |
| watermarkExpression | MARKONE | field |
| watermarkOptions | {"rotation":"0","position":"tl","scale":"0.2 abs"} | field |
| watermarkSource | text | field |
| watermarkExpression | MARKTWO | field |
| watermarkOptions | {"rotation":"0","position":"br","scale":"0.2 abs"} | field |
| watermarkSource | image | field |
| watermark | testdata/watermark.png | file |
| Gotenberg-Output-Filename | foo | header |
Then the response status code should be 200
Then the response header "Content-Type" should be "application/pdf"
Then there should be 1 PDF(s) in the response
Then there should be the following file(s) in the response:
| foo.pdf |
Then the "foo.pdf" PDF should have the following content at page 1:
"""
MARKONE
"""
Then the "foo.pdf" PDF should have the following content at page 1:
"""
MARKTWO
"""
Scenario: POST /forms/pdfengines/watermark (PDF - pdfcpu)
Given I have a Gotenberg container with the following environment variable(s):
| PDFENGINES_WATERMARK_ENGINES | pdfcpu |

View File

@@ -29,6 +29,17 @@ Feature: /
When I make a "GET" request to Gotenberg at the "/" endpoint
Then the response status code should be 401
# A request without a Bearer token is rejected. The JWKS URL is never fetched
# here, so no OIDC provider needs to be reachable.
Scenario: GET / (OIDC Auth)
Given I have a Gotenberg container with the following environment variable(s):
| API_ENABLE_OIDC_AUTH | true |
| API_OIDC_ISSUER | https://gotenberg.test/ |
| API_OIDC_AUDIENCE | gotenberg |
| API_OIDC_JWKS_URL | https://gotenberg.test/jwks.json |
When I make a "GET" request to Gotenberg at the "/" endpoint
Then the response status code should be 401
Scenario: GET /foo/ (Root Path)
Given I have a Gotenberg container with the following environment variable(s):
| API_ROOT_PATH | /foo/ |

View File

@@ -11,7 +11,6 @@ import (
"github.com/moby/moby/client"
"github.com/testcontainers/testcontainers-go"
"github.com/testcontainers/testcontainers-go/exec"
"github.com/testcontainers/testcontainers-go/network"
"github.com/testcontainers/testcontainers-go/wait"
)
@@ -21,11 +20,13 @@ import (
const testcontainersLabel = "org.testcontainers"
// PruneOrphanedNetworks removes dangling networks created by the test suite.
// Each scenario spins a dedicated network, and a failed container start can
// leak one before teardown records it. Leaked networks consume Docker's
// predefined address pools until none remain and every later scenario fails
// with "all predefined address pools have been fully subnetted". Call this
// before a run and between retries to reclaim the subnets.
// Scenarios no longer create one: the Gotenberg container is reached over its
// mapped port and the host-side helper over host.docker.internal, so the
// default bridge suffices. This stays as cheap insurance against networks
// leaked by an older suite version or an interrupted run, which consume
// Docker's predefined address pools until none remain and every later
// scenario fails with "all predefined address pools have been fully
// subnetted".
//
// Only unused networks bearing the testcontainers label are removed, so
// running containers and operator networks are never affected.
@@ -99,17 +100,18 @@ func applyDefaultEnv(env map[string]string) map[string]string {
return env
}
func startGotenbergContainer(ctx context.Context, env map[string]string) (*testcontainers.DockerNetwork, testcontainers.Container, error) {
// startGotenbergContainer starts a Gotenberg container on Docker's default
// bridge. No dedicated network is created: the suite addresses the container
// through container.Host plus its mapped port, and the container reaches the
// host-side webhook and static file server through the host.docker.internal
// alias below, so a per-scenario network would carry no traffic while still
// consuming one of Docker's predefined subnets.
func startGotenbergContainer(ctx context.Context, env map[string]string) (testcontainers.Container, error) {
ctx, cancel := context.WithTimeout(ctx, 2*time.Minute)
defer cancel()
env = applyDefaultEnv(env)
n, err := network.New(ctx)
if err != nil {
return nil, nil, fmt.Errorf("create Gotenberg container network: %w", err)
}
healthPath := "/health"
if env["API_ROOT_PATH"] != "" {
healthPath = fmt.Sprintf("%shealth", env["API_ROOT_PATH"])
@@ -122,7 +124,6 @@ func startGotenbergContainer(ctx context.Context, env map[string]string) (*testc
HostConfigModifier: func(hostConfig *container.HostConfig) {
hostConfig.ExtraHosts = []string{"host.docker.internal:host-gateway"}
},
Networks: []string{n.Name},
WaitingFor: wait.ForHTTP(healthPath),
Env: env,
}
@@ -148,19 +149,10 @@ func startGotenbergContainer(ctx context.Context, env map[string]string) (*testc
}
}
// The network is already created. The scenario teardown only
// removes networks it knows about, and the caller discards n on
// error, so remove it here to avoid leaking a subnet on every
// failed start. Leaked networks accumulate until Docker's address
// pools are fully subnetted and all later scenarios fail.
if errRemove := n.Remove(ctx); errRemove != nil {
err = fmt.Errorf("%w (also failed to remove network: %v)", err, errRemove)
}
return nil, nil, err
return nil, err
}
return n, c, nil
return c, nil
}
func execCommandInIntegrationToolsContainer(ctx context.Context, cmd []string, path string) (string, error) {

View File

@@ -29,14 +29,16 @@ func doRequest(method, url string, headers map[string]string, body io.Reader) (*
return resp, nil
}
func doFormDataRequest(method, url string, fields map[string]string, files map[string][]string, headers map[string]string) (*http.Response, error) {
func doFormDataRequest(method, url string, fields map[string][]string, files map[string][]string, headers map[string]string) (*http.Response, error) {
var b bytes.Buffer
writer := multipart.NewWriter(&b)
for name, value := range fields {
err := writer.WriteField(name, value)
if err != nil {
return nil, fmt.Errorf("write field %q: %w", name, err)
for name, values := range fields {
for _, value := range values {
err := writer.WriteField(name, value)
if err != nil {
return nil, fmt.Errorf("write field %q: %w", name, err)
}
}
}

View File

@@ -5,6 +5,8 @@ import (
"encoding/json"
"errors"
"fmt"
"image"
_ "image/png" // Register the PNG decoder for image.DecodeConfig.
"io"
"mime"
"net/http"
@@ -84,19 +86,22 @@ func findScenarioLine(filePath, name string) int {
}
type scenario struct {
resp *httptest.ResponseRecorder
concurrentResps []*httptest.ResponseRecorder
workdir string
teststoreDir string
gotenbergContainer testcontainers.Container
gotenbergContainerNetwork *testcontainers.DockerNetwork
server *server
hostPort int
resp *httptest.ResponseRecorder
concurrentResps []*httptest.ResponseRecorder
probeResps []*httptest.ResponseRecorder
sequentialResps []*httptest.ResponseRecorder
workdir string
teststoreDir string
gotenbergContainer testcontainers.Container
server *server
hostPort int
}
func (s *scenario) reset(ctx context.Context) error {
s.resp = httptest.NewRecorder()
s.concurrentResps = nil
s.probeResps = nil
s.sequentialResps = nil
err := os.RemoveAll(s.workdir)
if err != nil {
@@ -117,11 +122,10 @@ func (s *scenario) reset(ctx context.Context) error {
}
func (s *scenario) iHaveADefaultGotenbergContainer(ctx context.Context) error {
n, c, err := startGotenbergContainer(ctx, nil)
c, err := startGotenbergContainer(ctx, nil)
if err != nil {
return fmt.Errorf("create Gotenberg container: %s", err)
}
s.gotenbergContainerNetwork = n
s.gotenbergContainer = c
return nil
}
@@ -131,11 +135,10 @@ func (s *scenario) iHaveAGotenbergContainerWithTheFollowingEnvironmentVariables(
for _, row := range envTable.Rows {
env[row.Cells[0].Value] = row.Cells[1].Value
}
n, c, err := startGotenbergContainer(ctx, env)
c, err := startGotenbergContainer(ctx, env)
if err != nil {
return fmt.Errorf("create Gotenberg container: %s", err)
}
s.gotenbergContainerNetwork = n
s.gotenbergContainer = c
return nil
}
@@ -206,7 +209,7 @@ func (s *scenario) iMakeARequestToGotenbergWithTheFollowingFormDataAndHeaders(ct
return errors.New("no Gotenberg container")
}
fields := make(map[string]string)
fields := make(map[string][]string)
files := make(map[string][]string)
headers := make(map[string]string)
@@ -218,10 +221,10 @@ func (s *scenario) iMakeARequestToGotenbergWithTheFollowingFormDataAndHeaders(ct
switch kind {
case "field":
if name == "downloadFrom" || name == "url" || name == "cookies" {
fields[name] = strings.ReplaceAll(value, "%d", fmt.Sprintf("%d", s.hostPort))
fields[name] = append(fields[name], strings.ReplaceAll(value, "%d", fmt.Sprintf("%d", s.hostPort)))
continue
}
fields[name] = value
fields[name] = append(fields[name], value)
case "file":
if strings.Contains(value, "teststore") {
if s.teststoreDir == "" {
@@ -361,7 +364,7 @@ func (s *scenario) iMakeConcurrentRequestsToGotenberg(ctx context.Context, count
return errors.New("no Gotenberg container")
}
fields := make(map[string]string)
fields := make(map[string][]string)
files := make(map[string][]string)
headers := make(map[string]string)
@@ -372,7 +375,7 @@ func (s *scenario) iMakeConcurrentRequestsToGotenberg(ctx context.Context, count
switch kind {
case "field":
fields[name] = value
fields[name] = append(fields[name], value)
case "file":
wd, err := os.Getwd()
if err != nil {
@@ -468,6 +471,130 @@ func (s *scenario) iMakeConcurrentRequestsToGotenberg(ctx context.Context, count
return nil
}
// iMakeSequentialRequestsToGotenbergProbing mirrors the client loop from
// https://github.com/gotenberg/gotenberg/issues/1648: a conversion, then a
// probe, repeated. It records both the conversion and the probe responses so a
// scenario can assert that a planned process restart neither fails the probe
// nor breaks the conversions. Requests are sequential on purpose, since the bug
// only surfaces between two conversions.
func (s *scenario) iMakeSequentialRequestsToGotenbergProbing(ctx context.Context, count int, method, endpoint, probeEndpoint string, dataTable *godog.Table) error {
if s.gotenbergContainer == nil {
return errors.New("no Gotenberg container")
}
fields := make(map[string][]string)
files := make(map[string][]string)
headers := make(map[string]string)
for _, row := range dataTable.Rows {
name := row.Cells[0].Value
value := row.Cells[1].Value
kind := row.Cells[2].Value
switch kind {
case "field":
fields[name] = append(fields[name], value)
case "file":
wd, err := os.Getwd()
if err != nil {
return fmt.Errorf("get current directory: %w", err)
}
value = fmt.Sprintf("%s/%s", wd, value)
files[name] = append(files[name], value)
case "header":
headers[name] = value
default:
return fmt.Errorf("unexpected %q %q", kind, value)
}
}
base, err := containerHttpEndpoint(ctx, s.gotenbergContainer, "3000")
if err != nil {
return fmt.Errorf("get container HTTP endpoint: %w", err)
}
record := func(resp *http.Response) (*httptest.ResponseRecorder, error) {
defer resp.Body.Close()
body, readErr := io.ReadAll(resp.Body)
if readErr != nil {
return nil, fmt.Errorf("read response body: %w", readErr)
}
rec := httptest.NewRecorder()
rec.Code = resp.StatusCode
for key, values := range resp.Header {
for _, v := range values {
rec.Header().Add(key, v)
}
}
_, writeErr := rec.Body.Write(body)
if writeErr != nil {
return nil, fmt.Errorf("write response body: %w", writeErr)
}
return rec, nil
}
s.probeResps = make([]*httptest.ResponseRecorder, 0, count)
s.sequentialResps = make([]*httptest.ResponseRecorder, 0, count)
for i := range count {
resp, reqErr := doFormDataRequest(method, fmt.Sprintf("%s%s", base, endpoint), fields, files, headers)
if reqErr != nil {
return fmt.Errorf("request %d: do request: %w", i+1, reqErr)
}
rec, recErr := record(resp)
if recErr != nil {
return fmt.Errorf("request %d: %w", i+1, recErr)
}
s.resp = rec
s.sequentialResps = append(s.sequentialResps, rec)
probeResp, probeErr := doRequest(http.MethodGet, fmt.Sprintf("%s%s", base, probeEndpoint), nil, nil)
if probeErr != nil {
return fmt.Errorf("probe %d: do request: %w", i+1, probeErr)
}
probeRec, probeRecErr := record(probeResp)
if probeRecErr != nil {
return fmt.Errorf("probe %d: %w", i+1, probeRecErr)
}
s.probeResps = append(s.probeResps, probeRec)
}
return nil
}
func (s *scenario) allProbeResponseStatusCodesShouldBe(expected int) error {
if len(s.probeResps) == 0 {
return errors.New("no probe responses recorded")
}
for i, resp := range s.probeResps {
if resp.Code != expected {
return fmt.Errorf("probe %d: expected status %d, got %d %q", i+1, expected, resp.Code, resp.Body.String())
}
}
return nil
}
func (s *scenario) allSequentialResponseStatusCodesShouldBe(expected int) error {
if len(s.sequentialResps) == 0 {
return errors.New("no sequential responses recorded")
}
for i, resp := range s.sequentialResps {
if resp.Code != expected {
return fmt.Errorf("sequential response %d: expected status %d, got %d %q", i+1, expected, resp.Code, resp.Body.String())
}
}
return nil
}
func (s *scenario) allConcurrentResponseStatusCodesShouldBe(expected int) error {
if len(s.concurrentResps) == 0 {
return errors.New("no concurrent responses recorded")
@@ -1010,6 +1137,74 @@ func (s *scenario) thePdfShouldHaveImages(ctx context.Context, name string, imag
return nil
}
func (s *scenario) theImageShouldBePixels(_ context.Context, name string, width, height int) error {
path := fmt.Sprintf("%s/%s/%s", s.workdir, s.resp.Header().Get("Gotenberg-Trace"), name)
file, err := os.Open(path) //nolint:gosec // path is built from test-controlled values.
if err != nil {
return fmt.Errorf("open image %q: %w", path, err)
}
defer file.Close()
config, format, err := image.DecodeConfig(file)
if err != nil {
return fmt.Errorf("decode image %q: %w", path, err)
}
if config.Width != width || config.Height != height {
return fmt.Errorf("expected %s image %dx%d, but actual is %dx%d", format, width, height, config.Width, config.Height)
}
return nil
}
func (s *scenario) theImagePixelShouldBe(_ context.Context, name string, x, y int, want string) error {
path := fmt.Sprintf("%s/%s/%s", s.workdir, s.resp.Header().Get("Gotenberg-Trace"), name)
file, err := os.Open(path) //nolint:gosec // path is built from test-controlled values.
if err != nil {
return fmt.Errorf("open image %q: %w", path, err)
}
defer file.Close()
img, format, err := image.Decode(file)
if err != nil {
return fmt.Errorf("decode image %q: %w", path, err)
}
r, g, b, _ := img.At(x, y).RGBA()
// RGBA returns 16-bit channels; shift down to the 8-bit hex form.
got := fmt.Sprintf("#%02x%02x%02x", r>>8, g>>8, b>>8)
if !strings.EqualFold(got, want) {
return fmt.Errorf("expected %s pixel at %d,%d to be %s, but actual is %s", format, x, y, want, got)
}
return nil
}
func (s *scenario) theFileSizeShouldBe(_ context.Context, name, comparator string, sizeKB int) error {
path := fmt.Sprintf("%s/%s/%s", s.workdir, s.resp.Header().Get("Gotenberg-Trace"), name)
info, err := os.Stat(path)
if err != nil {
return fmt.Errorf("stat %q: %w", path, err)
}
limit := int64(sizeKB) * 1024
switch comparator {
case "less":
if info.Size() >= limit {
return fmt.Errorf("expected %s (%d bytes) to be smaller than %d KB", name, info.Size(), sizeKB)
}
case "greater":
if info.Size() <= limit {
return fmt.Errorf("expected %s (%d bytes) to be larger than %d KB", name, info.Size(), sizeKB)
}
}
return nil
}
func (s *scenario) thePdfShouldBeSetToLandscapeOrientation(ctx context.Context, name string, kind string) error {
var path string
if !strings.HasPrefix(name, "*_") {
@@ -1047,7 +1242,9 @@ func (s *scenario) thePdfShouldBeSetToLandscapeOrientation(ctx context.Context,
}
output = strings.ReplaceAll(output, " ", "")
re := regexp.MustCompile(`Pagesize:(\d+)x(\d+).*`)
// The dimensions can be fractional (e.g. a content-sized single page),
// so match floats, not just integers.
re := regexp.MustCompile(`Pagesize:([\d.]+)x([\d.]+)`)
matches := re.FindStringSubmatch(output)
if len(matches) < 3 {
@@ -1056,22 +1253,22 @@ func (s *scenario) thePdfShouldBeSetToLandscapeOrientation(ctx context.Context,
invert := kind == "should NOT"
width, err := strconv.Atoi(matches[1])
width, err := strconv.ParseFloat(matches[1], 64)
if err != nil {
return fmt.Errorf("convert width value %q to integer: %w", matches[1], err)
return fmt.Errorf("convert width value %q to float: %w", matches[1], err)
}
height, err := strconv.Atoi(matches[2])
height, err := strconv.ParseFloat(matches[2], 64)
if err != nil {
return fmt.Errorf("convert height value %q to integer: %w", matches[2], err)
return fmt.Errorf("convert height value %q to float: %w", matches[2], err)
}
if invert && height < width {
return fmt.Errorf("expected height %d to be greater than width %d", height, width)
return fmt.Errorf("expected height %g to be greater than width %g", height, width)
}
if !invert && width < height {
return fmt.Errorf("expected width %d to be greater than height %d", width, height)
return fmt.Errorf("expected width %g to be greater than height %g", width, height)
}
return nil
@@ -1572,10 +1769,13 @@ func InitializeScenario(ctx *godog.ScenarioContext) {
ctx.When(`^I make a "(GET|HEAD)" request to Gotenberg at the "([^"]*)" endpoint with the following header\(s\):$`, s.iMakeARequestToGotenbergWithTheFollowingHeaders)
ctx.When(`^I make a "(POST)" request to Gotenberg at the "([^"]*)" endpoint with the following form data and header\(s\):$`, s.iMakeARequestToGotenbergWithTheFollowingFormDataAndHeaders)
ctx.When(`^I make (\d+) concurrent "(POST)" requests to Gotenberg at the "([^"]*)" endpoint with the following form data and header\(s\):$`, s.iMakeConcurrentRequestsToGotenberg)
ctx.When(`^I make (\d+) sequential "(POST)" requests to Gotenberg at the "([^"]*)" endpoint, probing "([^"]*)" after each, with the following form data and header\(s\):$`, s.iMakeSequentialRequestsToGotenbergProbing)
ctx.When(`^I wait for the asynchronous request to the webhook$`, s.iWaitForTheAsynchronousRequestToWebhook)
ctx.Then(`^the Gotenberg container (should|should NOT) log the following entries:$`, s.theGotenbergContainerShouldLogTheFollowingEntries)
ctx.Then(`^the response status code should be (\d+)$`, s.theResponseStatusCodeShouldBe)
ctx.Then(`^all concurrent response status codes should be (\d+)$`, s.allConcurrentResponseStatusCodesShouldBe)
ctx.Then(`^all probe response status codes should be (\d+)$`, s.allProbeResponseStatusCodesShouldBe)
ctx.Then(`^all sequential response status codes should be (\d+)$`, s.allSequentialResponseStatusCodesShouldBe)
ctx.Then(`^all concurrent responses should have (\d+) PDF\(s\)$`, s.allConcurrentResponsesShouldHavePdfs)
ctx.Then(`^the (response|webhook request|file request|server request) header "([^"]*)" should be "([^"]*)"$`, s.theHeaderValueShouldBe)
ctx.Then(`^the webhook request header "([^"]*)" should carry trace id "([^"]*)"$`, s.theWebhookRequestHeaderShouldCarryTraceID)
@@ -1599,6 +1799,9 @@ func InitializeScenario(ctx *godog.ScenarioContext) {
ctx.Then(`^the "([^"]*)" PDF (should|should NOT) have the following content at page (\d+):$`, s.thePdfShouldHaveTheFollowingContentAtPage)
ctx.Then(`^the "([^"]*)" PDF (should|should NOT) have content matching "([^"]*)" at page (\d+)$`, s.thePdfShouldHaveContentMatchingAtPage)
ctx.Then(`^the "([^"]*)" PDF should have (\d+) image\(s\)$`, s.thePdfShouldHaveImages)
ctx.Then(`^the "([^"]*)" image should be (\d+)x(\d+) pixels$`, s.theImageShouldBePixels)
ctx.Then(`^the "([^"]*)" image pixel at (\d+),(\d+) should be "([^"]*)"$`, s.theImagePixelShouldBe)
ctx.Then(`^the "([^"]*)" file size should be (less|greater) than (\d+) KB$`, s.theFileSizeShouldBe)
ctx.After(func(ctx context.Context, sc *godog.Scenario, err error) (context.Context, error) {
if s.gotenbergContainer != nil {
errTerminate := s.gotenbergContainer.Terminate(ctx, testcontainers.StopTimeout(0))
@@ -1606,12 +1809,6 @@ func InitializeScenario(ctx *godog.ScenarioContext) {
return ctx, fmt.Errorf("terminate Gotenberg container: %w", errTerminate)
}
}
if s.gotenbergContainerNetwork != nil {
errRemove := s.gotenbergContainerNetwork.Remove(ctx)
if errRemove != nil {
return ctx, fmt.Errorf("remove Gotenberg container network: %w", errRemove)
}
}
return ctx, nil
})
ctx.After(func(ctx context.Context, sc *godog.Scenario, err error) (context.Context, error) {

View File

@@ -181,6 +181,12 @@ func newServer(ctx context.Context, workdir string) (*server, error) {
}
return c.HTML(http.StatusOK, string(b))
})
srv.GET("/redirect-to-private", func(c echo.Context) error {
s.req = c.Request()
// Redirect the browser to a non-public address so the outbound filter
// is exercised on the redirected request rather than on this URL.
return c.Redirect(http.StatusFound, "http://127.0.0.1:9999/redirected")
})
return s, nil
}

Binary file not shown.

View File

@@ -0,0 +1,18 @@
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8" />
<title>Local storage counter</title>
</head>
<body>
<!-- Increments a per-origin localStorage counter and renders it, so a value
above 1 means a previous conversion of the same origin leaked into this
one. See https://github.com/gotenberg/gotenberg/issues/919. -->
<pre id="out"></pre>
<script>
var n = (parseInt(localStorage.getItem("n"), 10) || 0) + 1;
localStorage.setItem("n", n);
document.getElementById("out").textContent = "localStorageCount=" + n;
</script>
</body>
</html>

Some files were not shown because too many files have changed in this diff Show More