Skip to content

chore(release): merge werf upstream into delivery-kit (3.6.2-dk.1) - #386

Merged
alexey-igrychev merged 84 commits into
mainfrom
chore/release/merge-werf-upstream-15
Oct 1, 2026
Merged

alexey-igrychev merged 84 commits into
mainfrom
chore/release/merge-werf-upstream-15

Conversation

@alexey-igrychev

Copy link
Copy Markdown
Collaborator

Summary

Sync werf main through 0ba870348 (3.6.2) into delivery-kit main, preserving the fork's commands, SBOM support and release configuration. Pin the next release to 3.6.2-dk.1 with Release-As.

What

Compatibility

  • Docker-backed image operations now require Docker Engine 19.03 / API 1.40 or newer; older daemons must be upgraded.
  • Registry builds again use shared synchronization by default: --repo commands contact https://synchronization.werf.io and fail if synchronization cannot be established. This merge retains the upstream default. Restricted-egress CI must configure WERF_SYNCHRONIZATION=kubernetes://NAMESPACE or a self-hosted synchronization service. WERF_SYNCHRONIZATION=:local avoids the public service but protects only builders on the same host, not concurrent builders on different hosts. Local builds use host locks. Synchronization also stores client-id-* / sync-server metadata tags in the primary repository.
  • Preserve delivery-kit commands, including verify and sbom get, alongside the restored synchronization command. SBOM builds use the same storage lock manager and synchronization/Kubernetes flags as other image builds.

Upstream changes

  • Import the build, Docker, cache/export, SSH, cleanup and deploy fixes documented in werf 3.6.2's versioned changelog, with the migration contract in docs/pages_en/resources/migration_from_v2_to_v3.md and its Russian counterpart.
  • Keep fork orphaned-artifact cleanup, including its dry-run preview, while preventing dry runs from updating last-cleanup metadata.
  • Preserve fork image-file reads with the upstream Moby client's archive-copy API.
  • Initialize operation statistics once for every sbom get mode so its shared flag continues to cover the whole command, including config rendering and failure paths.
  • Update Nelm from 6f628e8bd443 to 387f78685e1a: retain resources with a live skip-delete policy; recreate custom-validated immutable resources when requested; read release history as metadata; parse manifests once during rendering; support legacy programmatic patches and render-context jq variables.

Release and documentation

  • Leave CHANGELOG.md and .release-please/main/manifest.json unchanged; release-please owns their update after merge.
  • Set Release-As: v3.6.2-dk.1 in an empty commit; retain workflows deliberately removed by the fork.
  • Regenerate CLI reference documentation and retain the fork's test image sources when updating context-checksum expectations.

Why

Bring the fork up to the released werf 3.6.2 baseline without replacing delivery-kit customizations or importing upstream release history into the fork's changelog.

alexey-igrychev and others added 30 commits September 24, 2026 21:41
Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
…erf#7856)

## Summary

Operations statistics (`--build-report-operations` / `--log-debug`) now
cover the whole command run instead of only the conveyor: werf config
render, giterminism initialization and pre-build registry/git calls are
measured, and the summary is printed once before the command exits.
Previously a command spending most of its time before the build reported
a summary covering a fraction of it (observed: `build time: 0.14s` out
of an 11.8s command, with a 4.2s config render invisible).

## What

- The operations summary blocks are printed before the command exits on
the commands that save a build report (`build`, `bundle publish`,
`converge`, `lint`, `plan`, `render`, `stages copy`); the console line
is `command time:` instead of `build time:`. For `converge` the single
summary prints after the deploy.
- The accepted log level captured at command start is restored before
printing, so a nelm action lowering it (e.g. `render` sets Error) does
not drop the summary.
- Printing respects the effective log verbosity at command start: under
`--log-quiet` nothing is printed, and `helm get-autogenerated-values` is
quiet by default, so its summary stays hidden as before this PR.
- Two new operations appear in the report and the summary: `config
render` (werf.yaml templating) and `giterminism init` (git
worktree/submodules initialization).
- Pre-build registry, git and docker calls are now counted: the existing
hooks were no-ops until the collector entered the context.
- VERIFIED: smoke on macOS with the Docker backend for `build` and
`render` — the summary lists `config render` and `giterminism init`, the
JSON report `Operations` section contains them alongside build-phase
operations, and without the flag no summary is printed and no sections
appear.
- The console summary covers the whole command run; a saved JSON report
covers the operations recorded since the previous report of the same
command — with `--follow` that includes the polling between builds and
failed retry attempts. A failed report write does not lose observations:
they stay pending until a report is successfully written.
- VERIFIED: a 25s `build --follow --dev` session prints one summary with
`giterminism init: 13` while the last report carries `giterminism init:
1` and a single-iteration `StageCache`.
- Commands not saving a build report (`compose`, `run`, `kube-run`,
`export`) keep the previous conveyor-scoped behavior under
`--log-debug`.
- The JSON report schema does not change; `Operations`/`StageCache` stay
opt-in behind the same gate.
- SSH passphrase wait is deliberately not measured: it is interactive
input, absent in CI.

## Why

The collector was created inside `Conveyor.Build()`/`ShouldBeBuilt()`,
coupling the statistics scope to the conveyor lifetime: everything
before it — config render, giterminism init, sync clientID registry
calls — ran with no collector in the context, so the already-placed
measurement hooks recorded nothing and the printed total covered only
the build. Command-scoped installation with collector reuse in the
conveyor closes the gap without touching the hooks or the report schema;
the alternative of leaving per-conveyor scope and wrapping each
pre-build phase separately would keep the misleading `build time` total
and duplicate gate logic in every command. The report keeps per-build
semantics via flush-delta summaries because CI parses
`--build-report-path` per build, while the console block is a
human-facing whole-command view.

---------

Signed-off-by: Evgeniy Frolov <evgeniy.frolov@flant.com>
🤖 I have created a release *beep* *boop*
---


## [2.80.0](werf/werf@v2.79.2...v2.80.0)
(2026-09-25)


### Features

* **build:** collect operations statistics for the whole command run
([werf#7856](werf#7856))
([51c0d31](werf@51c0d31))

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).
…f#7936)

### What this PR does

Adds one bullet to `test-the-tests` — "Common ways a test looks strong
but isn't": an empirical probe must rebuild the exact wiring production
uses, or its conclusion describes the probe, not the system.

### Why do we need it?

During the werf#7928–werf#7934 series a cancellation probe on a bare context
"proved" that a canceled git command panics past the worktree heal
branch. Production wraps commands in `graceful.WithTermination`, where
`Terminate` returns normally, so the branch did execute — Ctrl-C wiped
healthy worktree caches. The false premise shipped in werf#7932 and needed
werf#7935 to fix. This is the session lesson recorded where mutation-testing
guidance already lives.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
Describe the v2-to-v3 migration in both English and Russian as a
comparison of major versions, without minor or patch release qualifiers.
Remove the intermediate-release history for `.helmignore`,
`stageDependencies`, and `cleanup --kube-scan-namespaces`, assuming
readers are moving directly from v2 to the current v3. Runtime behavior
is unchanged.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
…#7942)

## Summary

The English and Russian image configuration and template-engine examples
no longer use untagged Alpine or Ubuntu base images, which current werf
configuration validation rejects.

## What

- Replace previously untagged Alpine references with `alpine:3.21`,
including the templated `BaseImage` value used by `from`.
- Replace previously untagged Ubuntu references with `ubuntu:22.04` in
both languages.
- Leave internal image references, previously tagged examples, and
migration-guide examples of invalid syntax unchanged. No runtime
behavior changes.

## Why

External `from` references require an explicit tag or digest. These
examples retained the old implicit-latest syntax. Use version tags
already present in neighboring examples rather than following the latest
distribution release.

---------

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
…erf#7941)

## Summary

Keep provenance attestations disabled for each Dockerfile build,
including concurrent builds using the in-process Docker/buildx backend.
Previously, overlapping builds could restore a process-wide environment
variable before another build read it; unexpected index manifests were a
risk on attestation-capable drivers, not a reproduced failure.

## What

- Each Docker-backed Dockerfile build explicitly disables provenance
attestations, independently of other builds and
`BUILDX_NO_DEFAULT_ATTESTATIONS`.
- Builds no longer modify `BUILDX_NO_DEFAULT_ATTESTATIONS` in the
process environment.
- Metadata provenance remains disabled; no CLI flags or configuration
fields change.
- UNVERIFIED: prevention of unexpected index manifests has not been
exercised with concurrent builds on an attestation-capable driver.

## Why

Temporarily setting and restoring a process-wide variable cannot isolate
concurrent in-process buildx calls. Pass the equivalent of
`--provenance=false` in each build's attestation options instead of
relying on worker timing or serializing builds.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
## Summary

Reject invalid import sources even when a Stapel image uses `from:
scratch`. Previously, `werf config list` accepted both an untagged
external import such as `alpine` and an import from `scratch` when the
importing image had a scratch base.

Reproduce in a Git repository with this committed `werf.yaml`, then run
`werf config list`:

```yaml
project: smoke
configVersion: 1
---
image: app
from: scratch
import:
- from: alpine
  add: /src
  to: /app
  after: install
```

Replacing the import source with `scratch` reproduces the other bypass;
`alpine:3.20` remains valid.

## What

- With `from: scratch`, an external import without a tag or digest is
rejected with `must include a tag`.
- Importing from `scratch` is rejected with `scratch has no filesystem
to copy from`, regardless of the importing image's base.
- Tagged, digested, and internal import sources remain accepted;
`scratch` remains a valid base image.
- No flags, defaults, or error messages change.

## Why

An early `continue` exempted the entire image from validation when its
base was scratch or empty. Restrict that exemption to the base reference
so import validation cannot be bypassed by choosing a scratch base.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
## Summary

After a Stapel image has been built, renaming a mapped Git file without
changing its contents or traversal order can reuse an image containing
the old filename. Subsequent builds now produce the renamed file instead
of returning that stale content anchor.

## What

- Renaming a mapped file or directory changes the build checksum even
when file contents and modes are unchanged.
- Git-based checksums change on upgrade, requiring a one-time rebuild of
affected Stapel and Dockerfile images.
- Identical mapped paths, contents and modes remain independent of
commit history; no flags or configuration fields change.

## Why

The content-anchor checksum enumerates files and previously hashed only
blob hashes and modes. Unlike Git tree hashes, blob hashes do not encode
filenames. Include each full repository-relative path with a NUL
separator so pure renames cannot collide with the previous layout; using
commit IDs instead would lose reuse across commits with identical build
inputs.

---------

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
## Summary

Keep SSH aliases with different deploy keys from reusing the same
authenticated connection, without an extra SSH configuration probe for
each Git operation. This affects private Git repositories or includes
whose aliases resolve to the same host, user and port but specify
different `IdentityFile` values.

## What

- Different original SSH aliases or connection arguments use separate
control sockets. VERIFIED: real SSH servers on macOS and Linux
authenticated two aliases with their respective generated keys; the
previous `%C` implementation used the first key for both.
- Repositories using the same alias and connection arguments still share
a connection; the remote Git command is excluded from the socket key.
- Git recognizes the generated command as SSH without an extra
autodetection invocation. Explicit `GIT_SSH_VARIANT` and `ssh.variant`
selections remain unchanged.
- No executable wrapper is written to the temporary directory. Socket
names remain fixed-length inside a private per-process directory, with
the existing length limit and cleanup behavior.
- Automatic multiplexing tries an available `sha256sum`, then `openssl`,
then `shasum`, and is enabled only when the shell/SSH/hash probe
resolves a usable control path. Missing tools or failed probes leave
ordinary Git SSH behavior unchanged.
- A hash failure after initialization disables connection sharing for
that invocation. VERIFIED: real Git operations still succeeded with
fresh authentication and no control socket on both macOS and Linux.
- Relative temporary paths, shell-active characters and SSH expansion
tokens are rejected in favor of the existing `/tmp` fallback.
- Explicit `GIT_SSH_COMMAND`, `GIT_SSH` and `core.sshCommand` overrides
remain unchanged; automatic multiplexing remains disabled on Windows. No
new dependencies, CLI flags or configuration keys are introduced.

## Why

OpenSSH's `%C` hashes the resolved connection tuple, not the original
alias or its identity file. Reusing that socket bypasses the second
alias's authentication and can deny access to repositories requiring its
deploy key. Hash the original connection arguments inline rather than
disabling multiplexing or introducing a separate wrapper that Git must
autodetect on every invocation.

---------

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
…7946)

## Summary

The English Stapel-to-Dockerfile import example uses `from: golang`,
which werf v3 rejects as an external image reference without a tag or
digest.

## What

- Use `from: golang:1.23rc1-alpine3.20`, matching the existing Russian
example.
- Leave migration guidance and runtime behavior unchanged.

## Why

The example must follow v3's explicit-tag requirement. Reusing the
Russian example's tag keeps both language versions consistent without
changing the rest of the example.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
## Summary

Swapping content between two Git mappings no longer reuses a stale
content anchor. The failure requires an existing cached image and
mappings whose parameters stay unchanged while their content checksums
exchange places: after changing A=alpha, B=beta to A=beta, B=alpha, werf
previously returned the original image.

## What

- Swapping distinct contents between Git mappings changes the image's
content digest instead of reusing the old content anchor.
- Reordering complete mappings leaves their content digest unchanged;
new commits with unchanged mapped content also leave it unchanged.
- Changing mapped content or destination paths changes the content
digest.
- Content digests and content-based tags change on upgrade for images
with Git mappings, so previously published content anchors are not
reused under the new keys.
- Fetch behavior and ordinary Git archive stage dependency keys are
unchanged.

## Why

Mapping parameter hashes and content checksums were sorted together as
independent values. Exchanging two checksums preserved that sorted list
and therefore its digest, even though the expected image contents
differed. Hash each mapping's parameters together with its own checksum
before sorting to preserve the association.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
## Summary

Run documentation deployment, preview dismissal, registry cleanup and
release-image publishing with werf 3 dev instead of werf 2.

## What

- Install werf through `werf/trdl/actions/setup-app@v0.13.0` with
explicit group `3`, channel `dev`.
- Preserve an existing matching trdl repository with `force: false`; a
mismatched repository URL fails instead of being replaced.
- VERIFIED: the official action selects the published 3/dev release and
preserves repository state on a second invocation in an isolated macOS
smoke test.
- Remove the five `--virtual-merge=false` arguments rejected by werf 3.
- Leave published release versions, test binaries built from source,
local development helpers and user examples unchanged.
- The channel is intentionally moving, including production jobs; the
first v3 build can rebuild caches, and v3 cleanup can remove v2-era
images under retention policies.
- UNVERIFIED: live deployment/build and cleanup behavior against
existing infrastructure; validate in the target environment before
merging.

## Why

`werf/actions/install@v2` hardcodes major 2 and rejects major-3 version
inputs. The pinned official trdl action supports an explicit
group/channel without a custom installer or changing product version
defaults.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
## Summary

Skip ACR in the registry-cleanup workflow while its expired Azure
credentials keep ACR integration tests disabled.

## What

- Remove ACR from the cleanup matrix for scheduled and manually
triggered runs.
- Keep ECR cleanup, workflow triggers, and authentication configuration
unchanged.

## Why

ACR integration tests were disabled in werf#7920, but the independent
cleanup matrix still attempts Azure login and fails with
`AADSTS7000222`. Cleanup should follow the same temporary exclusion
until ACR credentials are renewed and testing is restored.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
## Summary

The `2` documentation config fails under werf 3 because it still uses
the removed Ansible builder. Reproduce with werf 3: `werf config list
--dir docs --env test` on the unmodified branch.

## What

- Write the web image's nginx configuration with the shell setup already
used on main, instead of `ansible.copy`.
- Preserve the nginx configuration text, including literal shell
variables, and retain the existing `fromImage` and `import.image`
references.
- VERIFIED: the updated config renders with werf 3.6.1 and werf 2.77.2
for both test and production.

## Why

CI must deploy versioned documentation with the same werf 3 driver as
current documentation. Port only the incompatible builder block so the
maintained branch remains usable by werf 2 as well.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
## Summary

Builds performing repeated reads and writes against a
bearer-authenticated registry exchange a new token for each operation,
even while an earlier token remains valid. Reuse valid bearer
credentials across matching operations to reduce requests to the
registry's authorization service.

## What

- Sequential and concurrent operations with the same registry, TLS
policy, token URL and request identity reuse an in-memory bearer
credential; concurrent cache misses share one exchange.
- Different credentials, registries and requested scopes remain
isolated, and a canceled waiter does not cancel another caller's
exchange.
- Cached credentials are reused only with more than 30 seconds
remaining; near-expiry credentials and credentials rejected by the
registry are acquired again, and failed exchanges are not retained.
- Registry rejection on an equivalent explicit default port invalidates
the cached token; different schemes and nondefault ports remain
isolated.
- Token-only responses use the Distribution protocol's 60-second default
lifetime, supporting GitLab responses without `expires_in`; explicit
lifetimes and earlier `issued_at` values bound reuse conservatively.
- The cache holds at most 128 credentials per registry API instance and
does not persist tokens or add configuration.
- OAuth exchanges, refresh-token responses and malformed token responses
retain the existing uncached behavior.
- Registry operations keep independent upload state and progress
completion, including concurrent image and index pushes, tagging and
deletion.

## Why

Each remote operation constructs a new go-containerregistry client and
bearer transport, whose initialization performs a token exchange.
Reusing the underlying HTTP connection transport does not retain that
authorization. Caching credentials at this transport boundary avoids
sharing the dependency's mutable bearer and upload state or retaining
per-operation progress options.

The fallback lifetime follows the [Distribution token authentication
specification](https://distribution.github.io/distribution/spec/auth/token/).
Tokens remain opaque; the cache does not interpret JWT claims or change
the registry's authorization decisions.

The expiry reserve limits reuse near expiration, but does not guarantee
token validity throughout a long upload. The existing
go-containerregistry behavior when retrying a consumed upload body after
an authorization challenge is outside this change.

---------

Signed-off-by: Anton Peretrukhin <anton.peretrukhin@flant.com>
Co-authored-by: Anton Peretrukhin <anton.peretrukhin@flant.com>
…erf#7953)

## Summary

Repeated Dockerfile builds with large shared contexts no longer create a
full temporary tar copy for each image when hard-linking is available.
With persistent local cache, eligible contexts are also reused across
commits when their selected Git files and checkout inputs are unchanged.
This applies to ordinary Git attributes as well as contexts without
attributes; unsupported environments retain commit-based caching.

Refs werf#7844. This reuses the Git base within a repository/context scope,
not every final overlaid or extracted context.

## What

### Archive reuse

- Dockerfile archive identities include selected paths, object hashes
and modes; applicable committed attribute-file identities; external
attribute-file contents; relevant Git configuration, Git version and OS.
- Parent `.gitattributes` files outside the context and nested attribute
files excluded from the tar still affect its identity; dirty tracked
attribute files do not replace committed rules.
- Standard `text`, `eol`, `crlf`, `ident`, `working-tree-encoding`,
`binary` and `-text` rules no longer disable reuse; changing a rule can
invalidate the archive even when resulting payload bytes would be
identical.
- Checkout filters declared for selected files retain commit keys when
filter drivers are configured; installed but unused drivers and rules
applying only to excluded files do not disable reuse.
- Declared submodules, conditional Git includes, worktree-specific
configuration, disabled symlink checkout and overridden attribute
sources retain commit keys; uncertain discovery, relative external
attribute paths and command failures do likewise.
- Git capability discovery runs before context traversal; environments
without attribute-path discovery fall back without calculating a content
key.
- Identical context calculations and conservative fallback decisions are
shared within a repository instance, assuming checkout settings stay
fixed during that build.
- VERIFIED: with Git 2.49.1, three images using a 10,000-file context
and root `* text=auto` retained one base tar across three
outside-context commits; the previous conservative guard retained four.
- VERIFIED: Git 2.34.1 retains commit caching without the previous
repeated content-key traversal penalty in the same many-file workload.
- On a content-cache miss, selected committed blobs are exported through
Git with committed attributes and a private index before archiving. This
avoids stale conversion bytes or staged files from the cached worktree,
without modifying its files or index; it adds one temporary
context-sized write on a miss, none on a hit.
- Failed private exports are removed before retry. Reuse across
invocations requires persistent cache and the same absolute repository
path/context scope; archive headers retain first-creation timestamps
rather than normalizing timestamps across independently populated
caches.

### Streaming and compatibility

- Without `contextAddFiles`, hard-link-capable cache and temporary
storage avoid full per-image tar copies; unavailable hard links and
additional files retain the materialized-copy path.
- VERIFIED: the streaming implementation eliminated three additional
archives for three Docker images sharing a 128 MiB context;
independently rebuilt Docker images retained layer diff IDs, normalized
metadata and Dockerfile permissions.
- Dockerfile overrides, including allowed uncommitted files, are read
for each context and preserve overlay precedence; archive pins remain
readable after GC removes the original cache path.
- Failed context opening or extraction removes incomplete extraction
directories before retry; pinned disk space remains allocated until all
cache and temporary links are removed.
- Image stage digest algorithms, CLI options, archive storage/GC formats
and non-Dockerfile archive keys are unchanged; older content-key
namespaces are not reused by the new identity scheme.
- Commit-cache limitations with externally changing filters remain
unchanged; the new key does not fingerprint arbitrary executable
behavior or the entire host environment.
- UNVERIFIED: native Buildah end-to-end execution; extraction
equivalence is tested, but a CGO-enabled native Linux build is still
needed for runtime confirmation.

## Why

Commit-only archive keys discard reuse when unrelated repository files
change, while appending a Dockerfile previously copied an entire tar.
Git object identities avoid hashing all payload bytes again, but omit
checkout rules; rejecting every attribute instead excludes common
repositories and repeats expensive checks even when reuse is
unavailable. Fingerprinting those inputs and exporting committed files
on cache misses preserves Git conversions without trusting cached
checkout bytes. Conservative filter fallbacks and per-build memoization
avoid claiming arbitrary programs are content-addressable or charging
every image for the same calculation.

---------

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
Follow up on werf#7955 with test-only changes. Add an HTTP request without
an explicit port so the authority test detects removal of scheme
isolation. Widen the expiry test’s cache-hit window from two to ten
seconds and its renewal timeout from six to fifteen seconds to tolerate
scheduling delays; production code is unchanged.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
## Summary

Preserve release resources whose live Kubernetes object carries
`helm.sh/resource-policy: keep` or `werf.io/resource-policy:
keep`/`skip-delete`, even when the chart has no retention annotation.
Previously, removing such a resource from the chart or uninstalling its
release could delete it despite the documented live-object policy.

## What

- VERIFIED: removing a ConfigMap from a chart or uninstalling its
release preserves its UID, data and live-only retention annotation; an
unprotected ConfigMap is still deleted.
- The dependency also restores live-only retention checks for
delete-on-failed and delete-on-succeeded operations; these paths are
covered by dependency-level tests, not a new werf cluster scenario.
- Render patches can use `$Values`, `$Release`, `$Chart` and
`$Capabilities`; chart-shipped subchart patches receive their own chart
and values. Diff patches still reject these render-only variables.
- UNVERIFIED: the newly available render variables have not been
exercised end-to-end through werf rollback/uninstall scenarios; a chart
using those variables would settle that integration coverage.
- The nelm update adds a programmatic `LegacyPatches` option; werf does
not populate it or add a CLI option for it.
- Update nelm from `6f628e8bd443` to `57f1772f2328`; keep the Go version
and other dependency versions unchanged.

## Why

The deletion planner consulted policies in stored or rendered manifests
but omitted the fetched live resource. The normal dependency update
includes [the retention fix](werf/nelm#729),
[render-context variables](werf/nelm#727) and
[programmatic patches](werf/nelm#720), rather
than carrying a local dependency patch. The full old/new endpoint delta
was checked because the revisions sit within an unreleased dependency
range.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
…erf#7959)

## Summary

Export images after a content-cache hit without a nil-pointer panic, and
publish/export images built with `--repo :local --final-repo REPO`
without the `expected non empty LegacyImage parameter` panic. A warm
local export no longer requires the final repository's image tag to
remain in the local backend.

## What

- A repeated single-platform `werf export` reuses the content anchor and
exports successfully even though no last-nonempty stage was recorded.
- VERIFIED: local-primary/final-repository cold and warm exports succeed
with both Docker and native Buildah.
- VERIFIED: with Docker, removing only the local final-repository alias
between exports still permits a warm export from the cached primary
image.
- VERIFIED: reusing the local primary cache with a new empty final
repository publishes and exports the same image on both backends.
- Exported images retain the same manifest across cold and warm runs and
omit werf service labels.
- Remote-primary/final-repository, multiplatform and ordinary
build-report exports retain their existing paths. The special
`--use-build-report` plus local/final-repository case after losing the
local final alias is not changed by this PR.

## Why

Content-anchor reuse bypasses stage traversal, so the
last-nonempty-stage pointer is not initialized. Final publication also
needs the initialized primary image when copying from a local backend,
and it replaces the image-level descriptor with the remote final
descriptor afterward. Use the initialized anchor for both the local copy
and the export source, without adding backend-specific state or changing
storage APIs.

---------

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
## Summary

Keep repository metadata unchanged during `werf cleanup --dry-run`.
Previously, a dry run with a separate `--meta-repo` created a persistent
safeguard marker, after which commands without `--meta-repo` failed and
required the flag or `werf meta-repo detach`.

## What

- `cleanup --dry-run` does not publish a last-cleanup record or create a
meta-repository safeguard as a side effect of that record.
- Dry runs still calculate and report the stages that cleanup would
remove.
- Actual cleanup still publishes its last-cleanup record and retains the
existing metadata safeguards.
- VERIFIED: against empty primary and metadata repositories, a
successful dry run leaves both repositories unchanged and a following
command without `--meta-repo` succeeds.

## Why

The final metadata-write block ran unconditionally after the deletion
preview. The meta-repository decorator turns that write into a
persistent repository contract change. Skip only that final write for
dry runs rather than returning before preview computation.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
## Summary

Make `werf release get`, `list` and `history` JSON/YAML output parseable
with their default logging settings, including in GitHub Actions and
GitLab CI. Previously, informational output could precede the document
and CI detection could inject ANSI escapes even when stdout was piped.

## What

- Without an explicit verbosity override, these three commands use their
command-specific logging defaults instead of forcing info-level progress
output.
- Automatic color selection respects parseable output instead of forcing
ANSI highlighting in GitHub Actions or GitLab CI.
- Explicit verbosity options keep their precedence: trace/debug, then
quiet, then verbose; other commands keep their existing logging
defaults.
- Explicit color or verbosity overrides remain user-controlled and are
not promised to produce a machine-only stream.
- VERIFIED: a real `release list --output-format=json` invocation with
GitHub Actions enabled parses successfully without `--log-quiet` or
`--log-color-mode=off`.

## Why

The shared log-level helper always returned `Info`, making the
command-specific fallback unreachable. Separately, logging setup did not
mark these commands as parseable. Fix both decisions: disabling color
alone leaves the informational prefix, while quiet alone leaves
CI-injected ANSI. Exercise command wiring rather than testing only the
helper, and stop masking the default behavior in the lifecycle
regression.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
## Summary

Make the English and Russian v3 installation instructions select werf
group 3 rather than group 2. Document two migration changes that would
otherwise break existing configurations or report consumers.

## What

- The GitHub Actions example uses `werf/trdl/actions/setup-app@v0.13.0`
with the werf preset, group `3`, channel `dev` and `force: false`; it
does not claim that an alpha channel already exists.
- Scheduled host-cleanup examples select `trdl use werf 3 dev` without
changing their schedule or execution user.
- The v2-to-v3 migration guide instructs users to remove
`disableGitAfterPatch`, explains that no equivalent switch exists, and
distinguishes build-stage dependencies from updating Git files
themselves.
- The guide documents the deploy-report discriminator change from
numeric `version: 3` to string `apiVersion: "v3"` and tells consumers to
update their checks.
- Apply the same guidance to both languages; do not change release
channels or generated release notes.

## Why

The old installation action hardcodes group 2, so following the v3
instructions installed the wrong major version. The removed
configuration field is rejected by the strict parser, and report
consumers need the actual new discriminator rather than an assumption
that the old field remains compatible. The migration guide remains the
authority for the rest of the v3 changes.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
alexey-igrychev and others added 25 commits September 30, 2026 22:23
…ket (werf#7981)

## Summary

A Docker socket which accepts connections but never answers — a stopped
Docker Desktop, a dead ssh tunnel — hung every werf command which
initializes the container backend, with no way out but a CI job timeout.
On top of that, commands which only talk to a container registry (`werf
cleanup`, `werf purge`, `werf dismiss`, `werf managed-images`, `werf
meta-repo`) failed at startup on a Docker daemon older than 19.03 with
`check Docker daemon API: API version 1.39 is not supported by this
client: the minimum supported API version is 1.40`, although they never
issue an Engine API call. The scope of both fixes is the commands which
initialize the container backend; `werf ci-env`, which reaches the
daemon through the docker CLI `docker login`, is untouched and listed as
a follow-up.

## What

- A command which needs the Docker daemon now fails after 10s against an
unresponsive daemon socket with `check Docker daemon API: Docker daemon
did not answer within 10s`, instead of waiting forever. VERIFIED: `werf
host purge --dry-run` against a socket which accepts and never answers
exits in 14s with that error; on `origin/main` the same command was
still running after 120s.
- A command which does not need the daemon treats a silent daemon like a
refused one and goes on: the registry mirror and insecure registry
lookups give up after 10s and return nothing. VERIFIED: `werf cleanup
--repo … --dry-run` against the same socket exits in 15s (failing on the
project's own missing git remote) instead of hanging; `werf
managed-images ls --repo …` reaches the registry client.
- Auto host cleanup deferred by such a command asks the daemon for its
storage path; that request has the same 10s deadline, so a silent daemon
makes werf skip the cleanup with `WARNING: unable to check if auto host
cleanup should be run: …` and does not affect the command's own result.
- An unresponsive daemon is no longer reported as an unavailable one:
`AllowDaemonUnavailable` still covers a refused connection or a missing
socket, never a silent one.
- UNVERIFIED: a busy-but-alive remote daemon which needs more than 10s
to answer `/_ping` or `/info` is now treated as unresponsive; only a
daemon behind a real slow link would settle it.
- The Docker daemon API version check runs only for commands which build
or run containers: `build`, `converge`, `render`, `plan`, `lint`,
`export`, `run`, `kube-run`, `compose`, `bundle publish`, `stage image`,
`stages copy`, `helm get-autogenerated-values`, `host cleanup`, `host
purge`.
- BREAKING (scope correction of fbb7e62): Docker 19.03 / Engine API
1.40 is required by those commands only. `cleanup`, `purge`, `dismiss`,
`managed-images` and `meta-repo` work again against an older daemon —
they take `--repo`, so their stage storage is always a registry and the
container backend they hold is never asked to do anything. UNVERIFIED:
if something does make one of them reach the daemon, the failure wording
is the moby client's, not werf's — the client refuses the sub-1.40
daemon at negotiation (`moby/client@v0.6.0 client.go:344-345`), but
`getAPIPath` ignores that error (`client.go:317`) and sends the request
at the client's own version, so the daemon's `client version … is too
new` answer can surface instead. No sub-1.40 daemon was available to
settle which one a user sees.
- `werf converge` and `werf render` check the daemon only on the path
that builds: `converge` skips the check with `--without-images` or
`--plan-artifact-path` (`cmd/werf/converge/converge.go:283,293-296`),
`render` with `--without-images`, `--stub-tags` or a local repository
address (`cmd/werf/render/render.go:228-235`). `converge` has no
`--stub-tags` flag. `lint`, `plan`, `bundle publish`, `stages copy` and
`helm get-autogenerated-values` initialize the backend before they know
whether they will build, so they require the daemon in every mode —
unchanged from `main`.
- The negotiated daemon API version is cached per api client, so
`ContainerCreate` no longer does a full `/_ping` round trip before every
Stapel stage. Only a verified version is cached: a daemon which is not
running at init is pinged again and can be used as soon as it comes up,
and a slow `/info` never makes the next check report an unresponsive
daemon.
- `werf build` fails with `unable to get import server container <name>
ip address: bridge network reports no ip address` (or `... container is
not attached to the bridge network`) instead of building the rsync url
`rsync://user@invalid IP:873/import/...`.
- Docs: `usage/build/backends.md` (en, ru) lists which commands require
Docker 19.03 / Engine API 1.40, which do not, and what happens to auto
host cleanup on an old or silent daemon.

## Why

Nothing in the Docker client path had a deadline: the connection check
pinged, the mirror lookup asked `/info`, and the host cleaner asked the
backend for its storage path, all on the bare command context, so a
socket which never answers turned into an unbounded wait, and the
`AllowDaemonUnavailable` fallback could only ever fire for a daemon
which actively refused the connection. The docker cli solves the same
problem with an init timeout (`cli/command/cli.go:32,373-382`); this
reuses that shape with more headroom, since unlike the cli's serverInfo
ping a timeout here can be fatal.

Caching is deliberately limited to the API version, which cannot change
for the lifetime of an api client. Availability can: a dind sidecar or a
Docker Desktop which starts between werf's init and the first Stapel
stage has to be picked up, exactly as on `main`, where every
`ContainerCreate` pinged anew.

`InitProcessDocker` is shared by every command which initializes the
container backend, and registry-only commands initialize it to get a
`ContainerBackend` for the storage manager, not to use the daemon.
Enforcing the daemon version there turned a Docker backend requirement
into a requirement of werf as a whole; before the moby client migration
the legacy SDK negotiated down to the daemon's API version, so these
commands worked on Docker 18.09. The moby client refuses anything below
API 1.40 at negotiation on its own (`client.go:344-345`), so the werf
check is there to word the error, not to enforce the floor — which is
why dropping it where the daemon is unused is safe.

Rejected alternative: moving the version check into the first real
daemon call (`ContainerCreate` already calls it). That leaves the other
daemon entry points — buildx builds, image and container listing,
pruning — reporting the client-worded error from deep inside an
operation, while the build-time check is the one users see first.

Fixes: fbb7e62 ("fix(build): require Docker 19.03+ for Docker backend
operations")

---------

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
…ds (werf#7977)

## Summary

A git mapping with `owner`/`group` left directories owned by `root:root`
when they first appeared in an incremental build, on both the Docker
(legacy stapel) and the Buildah backend. A directory added together with
a new file (`app-src/nested/new-file`) is `0:0` after a
`gitCache`/`gitLatestPatch` patch, but `1001:1002` when the same commit
is built from scratch — same content-based tag, two different images.

## What

- A directory that first appears in a git patch now gets the mapping's
`owner`/`group`, matching a from-scratch build of the same commit. Holds
for the Docker backend and for Buildah (native rootless included).
- Docker backend: every ancestor directory of an extracted entry, up to
the mapping's destination, is included in the `chown` list; entries are
deduplicated and written parent-first.
- Buildah backend: every ancestor directory created implicitly while
extracting the patch archive into the container rootfs is chowned, once
per directory.
- Stage digests and content-based tags do NOT change: the generated
commands and the extraction code were never digest inputs, so
already-built images with the wrong ownership are not rebuilt.
- `stapel.ChownBinPath` no longer takes a `context.Context` (it ignored
it); no user-visible effect.

## Why

The ownership is applied per tar entry of the archive being extracted. A
cold-build archive carries `tar.TypeDir` entries for intermediate
directories, so they were chowned. A patch archive is first filtered
through `util.CopyTar` with `IncludePaths`, and that filter matches
entry names by exact equality — the entry `nested` never equals the
include path `nested/new-file`, so the directory is dropped from the
archive and recreated implicitly as `root:root` by the extraction (`tar
-x` on Docker, `os.MkdirAll` on Buildah), with nothing left to name it
for the chown. Leaving it alone means the image an incremental build
produces differs from the image its own content-based tag promises.

Fixing the filter in `werf/common-go` was rejected: it is a versioned
dependency (a PR there plus a bump here), and reconstructing the
ancestors at extraction time covers every archive source, filtered or
not.

This completes werf#7972. That PR made the Docker backend apply the
mapping's `owner`/`group` to the files and symlinks a patch carries; the
directories a patch does not carry stayed `root:root` there, and the
Buildah backend had the same gap in `extractTarWithChown` independently
of it. No `Fixes:` line: the gap is not traceable to a single
introducing commit — on Docker it predates werf#7972 (extraction ran `tar
-xf` without `--no-same-owner`, and implicit directories were created as
root anyway), and the Buildah side is older still.

---------

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
…erf#7979)

## Summary

A stage rejected mid-build stayed invisible to the build that rejected
it, so every retry picked the same broken stage and the build ended with
`exhausted 3 retries on unexpected stages storage state` instead of
rebuilding. The same tags cache also silently gave up on reusing
registry tokens for any registry that redirects to another host (GHCR,
Quay, ACR and similar), and could abort a push with a panic.

## What

- A stage rejected during a build is seen as rejected by the retry that
follows it in the same process, and the build rebuilds it instead of
failing with `exhausted 3 retries on unexpected stages storage state`.
- Every tag werf publishes into the stages repo — rejected-stage
markers, managed images, image metadata, custom tag metadata and custom
tags, the cleanup record, posted manifests — is visible to the cached
lookups of the same process without waiting for the next tags listing.
- A push whose reference werf cannot parse as a tagged one no longer
aborts the command with a panic; the reference is skipped and lookups
fall back to a tags listing. A digest reference is no longer recorded as
the `latest` tag.
- Against a registry that redirects to another host, tokens for the
registry's own scopes keep being served from the cache, while a token
request is exchanged afresh as soon as any of its scope entries was
named by any bearer challenge the redirected host answered with — so a
token exchanged for that host is not reused.
- The exclusion is by scope entry and deliberately conservative: a
registry scope that a redirected host happens to name too stops being
cached for the rest of the operation, which costs a repeated exchange
and never a wrong token.
- Deleting tags from a GitLab registry reuses one bearer token per scope
instead of exchanging a new one for every tag.
- Registries without cross-host redirects, anonymous pulls and the
contents of the DELETE requests werf sends to GitLab are unchanged.

## Why

`werf#7965` (`2377615a8`) started recording published tags in the tags
cache, but wired it into `StoreImage` and `MutateAndPushImage` only.
`RejectStage` pushes its marker into the very repo whose listing is
cached, while the read side checks it with `WithCachedTags()`, i.e.
always from the cache; the retry backoff of
`RetryOnUnexpectedStagesStorageState` (2s/4s/5s) is shorter than the
one-minute listing freshness threshold, so the marker could not become
visible in time and every retry selected the rejected stage again.

`a84a30b1f` reacted to a bearer challenge from a redirected host by
disabling the token cache for the whole transport; the transport is
built per operation, so such a registry re-triggered it on every
operation and the caching added by `werf#7955` never applied. Invalidating
only the rejected entry is not enough on its own: the cross-host
challenge arrives on a request that carries no token of ours, and a
token obtained for the redirected host must not be reused at all.
Excluding only the scope entries the redirected host asked for keeps the
registry's own tokens cached and costs at most a repeated exchange when
a redirected host happens to demand the same scope as the registry.
Scope values are compared entry by entry on both sides, because
go-containerregistry sends one `scope` query parameter per entry of its
scope list and an entry copied from a challenge may itself hold several
space-delimited scopes.

Fixes: 2377615a8f0d ("fix(build): remember stage tags after publishing
them (werf#7965)")

---------

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
…erf#7978)

## Summary

`werf host purge` gave up on the first Docker daemon error: a single
stapel container held by a
concurrent build left the stapel containers of every other version and
platform, their volumes and
the stapel image on the host, and only the first failure was reported. A
stapel container that has
no volume at `/.werf/stapel` was also skipped entirely, so it kept the
stapel image referenced,
the final `docker rmi` failed and every later `werf host purge` failed
the same way. Both need a
host where a previous purge or a build left such a container behind;
there is no repro from a
clean checkout.

## What

- `werf host purge` removes every stapel container it owns and its
volume even when the removal of
another one fails, and then tries to remove the stapel image. The image
goes away only when no
container of that version is left: a kept container (a foreign holder,
or a 409 from the daemon)
makes `docker rmi` fail, and that failure is reported alongside the
others.
- A failed purge reports every failure it hit, not just the first one.
- A stapel container without a volume at `/.werf/stapel` is removed too,
so `werf host purge` stops
  failing forever on a host that holds one.
- A stapel container whose volume another container still mounts is
kept, and the failure is
reported as `keep container <id>: volume <name> is in use by <holder
ids>`. The next
  `werf host purge` after that holder is gone removes both.
- A volume removal that fails anyway — a holder attached between the
check and the removal — is
reported as `remove volume <name>: ...` and does not stop the purge of
the other containers. On
that genuine race the stapel container is already gone, so no later
`werf host purge` can find
the volume again: removing it by hand with `docker volume rm <name>`
(the name is in the
  reported error) is the only way out.
- `werf host purge` still sends `DELETE /containers/<id>` for stapel
containers with no `v` and no
`force`: a running stapel container and the volumes werf does not own
are left alone.
- Exit code and the purge log lines do not change.

## Why

The stapel container is the only record of its volume's name: `Purge`
discovers volumes through
`inspect.Mounts` of the containers it recognizes, and no code path
prunes volumes afterwards. So
removing the container first and then failing on the volume — the 409
the daemon answers while
another container still mounts it — leaves a volume no `werf` command
can ever find again. The
container removal is therefore preceded by a listing of the containers
that reference the volume
(`GET /containers/json?all=1&filters={"volume":...}`): when a foreign
holder exists, the stapel
container stays as the anchor and the failure is reported, so the disk
space is reclaimed by a
later purge instead of leaking. A 409 that still happens after that
check is a genuine race with a
container created in between; it is reported and the purge continues.
That case is not
recoverable by werf: the stapel container — the only record of the
volume's name — is already
removed, so the volume is invisible to every later `werf host purge` and
has to be removed by hand
with `docker volume rm <name>`, using the name from the reported error.
Closing that window
entirely would need a daemon-side lock the Engine API does not offer.

Skipping a volume-less container was a side effect of collecting the
volumes before deciding to
remove the container. The pre-werf#7969 code removed it unconditionally, and
it has to: `rmiIfExist`
runs `docker rmi` without `--force`, so one leftover container makes the
purge unrecoverable by
any werf command. Neither failure may stop the purge either, hence
`errors.Join` over all
containers and the image.

The second commit makes those guarantees testable at all. The fake
daemon parsed the id out of a
`DELETE` and ignored the query, so
`ContainerRemoveOptions{RemoveVolumes: true, Force: true}` on
the stapel container removal kept the whole suite green — the property
"werf does not delete the
volumes it does not own" could not fail. It now models `v` (anonymous
volumes of the container
only), `force` (a running container answers 409 otherwise) and rejects
the unmodeled `link`.

Fixes: 666afe129d6c ("fix(cleanup): purge platform Stapel containers
with their volumes")

---------

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
…f#7980)

## Summary

Building several Dockerfile images that share one git context copied the
whole context once per image on any host where the werf temp dir and
`$WERF_HOME` are on different filesystems — the default on Linux
distributions that mount `/tmp` as tmpfs. The reuse introduced in werf#7953
only worked when both happened to sit on one filesystem.

Repro (Linux, `/tmp` on tmpfs): a `werf.yaml` with three dockerfile
images over the same context, `WERF_TMP_DIR=/tmp werf build` — before,
`$WERF_TMP_DIR` gained one full context copy per image.

## What

- When the temp dir and `$WERF_HOME` are on different filesystems, a
build of N dockerfile images over an unchanged context no longer writes
N copies of that context; the cached archive is hard-linked instead.
UNVERIFIED: a host with `/tmp` on tmpfs is needed to observe the EXDEV
path; on macOS both roots are always one filesystem. The unit specs
assert the property that makes it true — the pin lives under
`$WERF_HOME`, never under the temp dir.
- The hard link is created at
`$WERF_HOME/service/tmp/context_pins/werf-<version>-context-pin-*/archive.tar`,
on the same filesystem as `$WERF_HOME/local_cache` by construction (both
are subdirectories of `$WERF_HOME`); a bind-mounted `local_cache` still
falls back to a copy per image. It is removed by `werf host cleanup`
together with the project temp dir of the same build, as before.
- A pin dir left behind by a process killed before it delegated its
cleanup is no longer reclaimed only by `werf host purge`: `werf host
cleanup` now also removes unregistered pin dirs whose mtime is older
than 24 hours. A build running longer than that loses its pin.
- `$WERF_HOME/local_cache` layout does not change: the pin is a service
dir, not a cache root, so gitdata GC neither collects it nor mistakes it
for a cache entry.
- Behavior when the link cannot be created is unchanged: werf falls back
to a materialized copy of the context with the Dockerfile overlay.

## Why

`os.Link` cannot cross filesystems. The pin was created under the
conveyor temp dir (`Image.TmpDir` ⊂ `werf.GetTmpDir()`, `os.TempDir()`
by default), while its target is the cached archive under
`$WERF_HOME/local_cache/git_archives`. With `/tmp` on a separate mount
every `os.Link` returned `EXDEV` and every image took the fallback copy,
so the benefit the feature claims never materialized on a large share of
Linux hosts.

Placing the pin under `$WERF_HOME/local_cache` instead would put a
non-cache directory into a tree whose collectors treat unknown entries
as invalid or as LRU entries (see `pkg/git_repo/gitdata/doc.go`), and a
new cache root would need its own version namespace and GC.
`$WERF_HOME/service/tmp` is the existing home for per-process temp state
(`git_data`, `buildah`, bundles), is registered with the tmp-manager GC
like the project temp dir, and is wiped by `werf host purge`.

An orphaned pin is worse than an orphaned temp dir: its hard link keeps
the git archive inode alive after the gitdata LRU evicted the archive,
so the LRU reports freed bytes that the filesystem never gets back.
Under `/tmp` the orphan was swept by the OS (tmpfs, reboot,
systemd-tmpfiles); under `$WERF_HOME` nothing swept it, hence the
age-based sweep.

Fixes: f2eab7e ("fix(build): reuse unchanged Dockerfile contexts
without extra copies")

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
…erf#7983)

## Summary

With `--final-repo`, a multi-platform image is published to the final
repo, but `werf export` read the manifest list back from the primary
repo, while `werf export --use-build-report` already read it from the
final one. No user-reachable failure follows from this today: `werf
export` builds in the same process and re-publishes the index into the
primary repo right before exporting, and multi-platform with a local
primary storage is rejected earlier, so this aligns the two export paths
rather than fixing an observable error.

## What

- With `--final-repo` set and the multi-platform index copied there,
`werf export` reads the index from the final repo; before, it read it
from the primary repo.
- Without `--final-repo`, or when the index was not copied there, the
export still reads from the primary repo.
- Nothing else about the export changes: the destination references, the
tags and the exported manifest content are the same, because both repos
hold the same index at that moment.
- No bug report is fixed here — with the current call order the primary
repo always still holds the index when the export runs.
- UNVERIFIED: the end-to-end behavior against a real registry; no
Docker/registry available in the working environment. A `--platform
linux/amd64,linux/arm64` entry in `test/e2e/export/final_repo_test.go`
("exports the same image before and after content-anchor reuse") would
settle it.
- The multi-platform + local primary storage rejection in `Exporter.Run`
is deliberately left in place: it guards manifest-list creation, not the
export.

## Why

`exportMultiplatformImage` always called
`GetStagesStorage().ExportStage()` with `img.GetStageDesc()`, although
`MultiplatformImage` also carries `GetFinalStageDesc()`, set exactly
when the index was copied into the final repo. Leaving it means the two
export paths disagree about which repo is the source of the exported
index, so anything that stops the primary repo from holding the index at
export time — cleanup between build and export, or multi-platform
support for a local primary storage — would break the non-report path
alone and silently. The report path picks its storage by comparing repo
addresses; that is not needed here, since the final descriptor is only
ever set by `publishMultiplatformFinalImage`.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
…werf#7984)

## Summary

On a host without `sha256sum`, `openssl` or `shasum` on `PATH` — a slim
CI image is the usual case — werf silently ran every git request over
ssh with a fresh handshake instead of reusing one multiplexed
connection, and nothing in the log said so. werf now warns and names the
utilities that would restore multiplexing. The same branch also drops a
dead helper.

## What

- When `true_git.Init` finds none of `sha256sum`, `openssl` or `shasum`
on `PATH`, werf prints `WARNING: ssh connection multiplexing is
disabled: none of sha256sum, openssl or shasum was found in PATH.`
followed by `Every git request over ssh pays a full handshake; install
one of them to let werf reuse connections.` The warning is emitted once
per werf process, at the default log level.
- When any one of the three is found, nothing is printed and
multiplexing is set up exactly as before.
- The other paths that disable multiplexing — Windows, a user-set
`GIT_SSH_COMMAND`/`GIT_SSH`/`core.sshCommand`, an ssh binary that
rejects `-o ControlMaster`, no directory able to hold the socket — stay
silent, unchanged.
- No user-visible change from the `chore(dev)` commit:
`common.GetManagedImagesNames` had no callers left and its logic lives
in `pkg/cleaning/cleanup.go`. The invariant a reviewer checks: `werf
cleanup` output is byte-identical.

## Why

The control socket name is derived by hashing the resolved ssh alias
options in a shell pipeline, because OpenSSH's own `%C` token folds
together aliases that differ only by `IdentityFile` and would hand one
build's authenticated connection to another. That hash needs an external
utility, and when the lookup failed the setup simply returned no
`GIT_SSH_COMMAND` — correct, but indistinguishable from a host where
multiplexing works, so a build slower than the same build elsewhere had
nothing to explain it. Falling back to `%C` was rejected: it
reintroduces the alias collision the hash exists to prevent.

Fixes: 6ba82d0 ("fix(build): keep deploy keys separate across SSH
aliases")

---------

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
`acr` is listed as `# - acr` in the registry cleanup matrix instead of
being absent, matching the other temporarily disabled registries in the
same workflow, so re-enabling ACR is one uncomment. The matrix result is
unchanged — `ecr` is still the only implementation that runs — and the
guarded `Setup (acr)` step stays in place, as the ACR login step was
kept in `_test_integration_per-container-registry.yml` while ACR is
disabled.

Fixes: 154df22 ("chore(ci): skip cleanup for the disabled ACR test
registry (werf#7954)")

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
## Summary

Reduce repetitive service output during image builds, especially with
debug logging enabled.

## What

- Hide successful `git show-ref` and `git ls-remote` stdout from general
debug output. With live Git output enabled, `WERF_DEBUG_TRUE_GIT=1`
restores the complete ref listing; command result buffers, stderr and
failure diagnostics remain available.
- Remove per-task worker-start messages and context addresses from
cancellation messages; keep worker IDs and cancellation diagnostics.
- Print stage-digest lock progress only when the mutex is contended,
retaining lock-wait statistics and the existing unlock lifecycle.
- Remove the repeated `Using git stages` detail line.
- Show each repository/commit pair once per image-tree calculation,
including the repository name. Keep distinct commits and distinct
repositories visible.
- Remove the `Initializing git mappings` wrapper while retaining
individual mappings and resolution errors.
- Under debug logging, print `Listed … tags` with duration once per
successful registry request. Replace per-call cache-hit and
shared-request noise with `registry tags cache hit` and `registry tags
shared result` counters in the existing operation statistics when a
collector is active.
- Omit empty `next stage dependencies` debug lines; retain non-empty
dependencies and all checksum input logging.

## Why

Large builds repeat ref listings, commit selections and scheduler
bookkeeping thousands of times, obscuring stage selection and build
failures. Tag listing messages previously also described cache hits and
each caller sharing a request. Explicit Git tracing retains the full ref
dump when investigating Git behavior.

---------

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
## Summary

Correct the v2-to-v3 migration instructions for custom Helm
configuration directories. Setting `XDG_CONFIG_HOME` to the previous
`HELM_CONFIG_HOME` does not make the old repository list visible: werf
looks in an additional `helm` subdirectory.

## What

- Both English and Russian migration guides show that
`XDG_CONFIG_HOME=/custom/config` selects
`/custom/config/helm/repositories.yaml`, not
`/custom/config/repositories.yaml`.
- The migration instructions tell users to copy the old repository list
into the new `helm` directory or add their repositories again.
- VERIFIED: `werf helm repo list` finds the repository fixture in the
`helm` subdirectory and does not find the same fixture directly under
`XDG_CONFIG_HOME`.
- Runtime behavior, configuration defaults and supported environment
variables do not change.

## Why

The Helm path resolver appends `helm` to the XDG base directory. The
previous wording treated the XDG base as the final Helm configuration
directory, leaving migrated users with an apparently empty repository
list.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
## Summary

Keep registry tag-cache requests out of the build's stage counters. With
operations statistics enabled, cached or shared tag listings previously
appeared as `stage(s)` and inflated the stage total even though they did
not represent built or reused stages.

## What

- `StageCache` and the `Stage cache summary` console block contain only
built stages and stages obtained from local, primary or secondary
storage, counted in `stage(s)`.
- The JSON report adds an optional `RegistryCache` map with the existing
`registry tags cache hit` and `registry tags shared result` keys.
- A separate `Registry cache summary` console block counts those events
in `request(s)`; empty cache sections are omitted.
- Registry counters follow the existing statistics opt-in:
`--build-report-operations`, `WERF_BUILD_REPORT_OPERATIONS`, or debug
logging. No new flag or default is introduced.
- Both cache sections cover the interval since the previous successful
report; a failed write retains pending counters, while the console
summary still covers the whole command run.
- The flag help, generated CLI reference and English/Russian
build-process documentation describe the separate sections. Shared
results count every caller, including the request initiator; they are
not a measure of avoided network requests.
- Existing event names, operation timings, image records and the ability
to read older reports do not change.

## Why

Registry tag-cache events and stage events share a collector, but the
report and logger treated every event as a stage. Separating the output
groups preserves the meaning of existing stage counters without
discarding registry diagnostics or changing collection and flush
behavior.

Fixes: a10017a ("fix(build): reduce repetitive service output
(werf#7985)")

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
…erf#7988)

## Summary

Chart-only `werf plan`, `werf lint`, `werf helm
get-autogenerated-values` and `werf bundle publish` no longer fail
because a local Docker daemon is too old or accepts connections without
answering. Commands such as `werf cleanup --repo … --dry-run` also stop
failing during optional registry-settings discovery on Docker API
versions below 1.40.

## What

### Chart-only commands

- Projects without images and `--without-images` skip container-backend
initialization in `plan`, `lint`, `helm get-autogenerated-values` and
`bundle publish`.
- `plan`, `lint` and `helm get-autogenerated-values` also skip
initialization with `--stub-tags` or a local repository address; `plan
--show-plan` does not need the backend.
- Custom Docker credentials remain available for private OCI chart
dependencies and chart-only bundle publishing.
- Image-processing paths retain the Docker 19.03 / API 1.40 requirement
and the response deadline, including `--require-built-images` and
`--use-build-report`.
- Automatic host cleanup is only requested by these lazy-initializing
commands after a container backend was actually initialized.

### Optional daemon settings

- Registry-only commands skip daemon-provided mirrors and
insecure-registry settings when a successful ping advertises an
unsupported API version, rather than issuing an incompatible `/info`
request.
- Missing sockets, refused connections, down or unreachable hosts and
networks (including Windows TCP), unresolvable hosts and unreachable SSH
endpoints do not block optional settings discovery; ping and info share
the existing 10-second deadline for an unresponsive daemon.
- Authentication, direct socket permission, TLS, malformed-response,
unexpected server and caller-cancellation errors reported by the Docker
client propagate instead of being treated as daemon absence.
- An inherited Docker-over-SSH limitation remains: some remote socket
permission failures are normalized by the pinned Docker client into the
same generic error as a stopped daemon. Optional settings discovery
cannot distinguish those errors and skips them; this change does not
reconstruct causes already discarded by the client.
- Only successful version observations are cached; a daemon unavailable
during initialization is checked again when it becomes available.
- The English and Russian backend documentation describes the lazy
paths, the retained image-processing checks and optional-settings
failure behavior.

## Why

Disabling the explicit daemon requirement was insufficient: the shared
initialization path still requested registry settings, and Moby refuses
to negotiate below API 1.40 before constructing an incompatible
versioned info URL. Chart commands also initialized their backend before
deciding whether any image work was needed. Keep credentials
initialization independent and defer the backend until that decision,
while preserving real daemon errors instead of suppressing every
connection failure.

Fixes: 082baeb ("fix(build, docker): stop container backend init
hanging on a mute socket (werf#7981)")

---------

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
## Summary

Stop staged Dockerfile builds with COPY --from on Buildah from retaining
temporary source containers and pinning their images. Repeated
multi-stage builds previously accumulated working containers because the
copy instruction created and mounted a source without releasing it.

## What

- Defer source-container cleanup immediately after allocation, including
partially successful creation, subsequent errors and panic unwinding.
Unmount first when mounting succeeded and attempt removal even if
unmounting fails.
- Log cleanup failures with COPY --from, the source image and source
container. Cleanup errors do not fail a successful copy or replace an
existing operation error; this matches temporary build-container cleanup
elsewhere in the backend.
- Detach cleanup from build cancellation with context.WithoutCancel. No
cleanup time bound is promised: the native storage operations do not
observe context deadlines.
- Leave source images, destination containers, copy options and ordinary
build-context COPY behavior unchanged.
- Cover returned errors, cancellation, panic cleanup and source-specific
error logging with focused regressions. Keep suite bootstrap and
concrete test helpers separate.

## Why

Copy.Apply calls FromCommand and Mount to use an image as its source
directory, but previously never unmounted or removed that container.
These leaked containers retain image references and prevent safe image
deletion. Release the resource in the instruction that owns it instead
of teaching post-test cleanup to force-delete containers it did not
create, while retaining the established non-fatal cleanup behavior.

---------

Signed-off-by: Evgeniy Frolov <evgeniy.frolov@flant.com>
…erf#7992)

## Summary

A committed ignore pattern ending in a backslash previously made `werf
build` panic with `syntax error in pattern`. Builds now return an error
identifying the selected ignore file and invalid pattern.

Reproduce with a reachable Docker backend:

```sh
mkdir dockerignore-repro && cd dockerignore-repro
git init -q
printf 'project: ignore-error\nconfigVersion: 1\n---\nimage: app\ndockerfile: Dockerfile\nstaged: false\n' > werf.yaml
printf 'FROM scratch\n' > Dockerfile
printf '%s\n' Dockerfile 'archive-old\' > .dockerignore
git add .
git -c user.name=Test -c user.email=test@example.com -c commit.gpgsign=false commit -qm repro
werf build --repo=:local
```

## What

- Invalid patterns in `.dockerignore`, `Dockerfile.dockerignore`,
`.containerignore`, and `Dockerfile.containerignore` produce a normal
build error naming the file and pattern.
- Patterns rejected during lazy compilation, including exclusions such
as `![z-a]`, fail during matcher creation instead of panicking during
path matching.
- VERIFIED: the reproducer exits with code 1 and reports `read ignore
file ".dockerignore": parse ignore pattern "archive-old\\": syntax error
in pattern` without a panic.
- Valid matching, ignore-file selection order, Dockerfile reinclusion,
and cache identities remain unchanged; `.gitignore` continues to be
interpreted by Git.

## Why

The matcher constructor panicked on user-supplied syntax errors, and
Moby defers some pattern compilation until matching. Validate every
pattern before exposing the matcher and propagate errors through the
existing build error path so users can locate and correct the rule.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
… immutable field change (werf#7994)

Signed-off-by: Dmitry Mordvinov <dmitry.mordvinov@flant.com>
)

## Summary

Concurrent builders can publish duplicate images for the same build
inputs, and a later retry can select a different result when those
outputs differ. Restore synchronization of publication in werf v3 for
every stage and built image.

## What

- Registry builds use the public synchronization service by default;
local builds use host file locks. Synchronization failures stop the
operation, including failures of the implicit public service.
- Restore `--synchronization` / `-S`, `WERF_SYNCHRONIZATION`, HTTP(S),
`:local`, direct `kubernetes://`, and `werf synchronization` with
Kubernetes connection flags and default kubeconfig discovery.
- Explicit synchronization settings print `Using sync server: ...` in
the normal log before initialization. HTTP credentials, query/fragment
and inline base64 kubeconfig are redacted; implicit defaults emit no
selection line.
- Publication from a build or secondary repository rechecks the primary
under the shared project/digest lock and adopts an existing suitable
image. The fresh listing cannot join a request started before lock
acquisition; final/cache copies retain the selected identity.
- Synchronization client-ID records remain in the primary repository
when a separate meta-repo is used.
- VERIFIED: the restored locking implementation passed synthetic
concurrent Docker/Buildah publication and retry checks, plus direct and
HTTP-backed Kubernetes checks. This does not establish registry-side
fencing after lease loss or server-state loss.
- English and Russian migration/build-process documentation retains
synchronization and describes HTTP daemon storage. Public-service
fallback, policy markers and explicit opt-out are outside this change.

## Why

An input digest does not make the existence check and publication
atomic. Builders with different stage chains can reach the same final
inputs and race to publish. Shared locks and an authoritative recheck
coordinate their selection while they share the backend and hold valid
leases.

Fixes: ce1ab9f ("refactor(build):
remove synchronization subsystem (werf#7603)")

---------

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
🤖 I have created a release *beep* *boop*
---


## [3.6.2](werf/werf@v3.6.1...v3.6.2)
(2026-10-01)


### Bug Fixes

* **build, dockerfile:** stop copying the build context per image
([werf#7980](werf#7980))
([ac62101](werf@ac62101))
* **build, docker:** keep chart-only commands independent of Docker
([werf#7988](werf#7988))
([1eadf7c](werf@1eadf7c))
* **build, docker:** stop container backend init hanging on a mute
socket ([werf#7981](werf#7981))
([082baeb](werf@082baeb))
* **build, stapel:** set owner/group on dirs added by incremental builds
([werf#7977](werf#7977))
([ae5a5be](werf@ae5a5be))
* **build:** export cached images with local and final repositories
([werf#7959](werf#7959))
([e19f8b1](werf@e19f8b1))
* **build:** export final-repository images from build reports
([870a276](werf@870a276))
* **build:** export final-repository images from build reports
([werf#7964](werf#7964))
([7aa2265](werf@7aa2265))
* **build:** export multi-platform images from the final repository
([werf#7983](werf#7983))
([a2efb5f](werf@a2efb5f))
* **build:** honor git mapping ownership with Docker
([werf#7972](werf#7972))
([152ee34](werf@152ee34))
* **build:** honor introspection on content-cache hits
([1191bf1](werf@1191bf1))
* **build:** honor introspection on content-cache hits
([werf#7967](werf#7967))
([d7da627](werf@d7da627))
* **build:** keep deploy keys separate across SSH aliases
([werf#7945](werf#7945))
([6ba82d0](werf@6ba82d0))
* **build:** keep provenance disabled in parallel Dockerfile builds
([werf#7941](werf#7941))
([48fbd9e](werf@48fbd9e))
* **build:** keep registry requests out of stage counts
([werf#7987](werf#7987))
([2c8339c](werf@2c8339c))
* **build:** prevent duplicate images from concurrent builders
([werf#7989](werf#7989))
([bfaad77](werf@bfaad77))
* **build:** rebuild a rejected stage instead of exhausting retries
([werf#7979](werf#7979))
([ee4f4ae](werf@ee4f4ae))
* **build:** rebuild images after renaming git-tracked files
([werf#7940](werf#7940))
([8412d4e](werf@8412d4e))
* **build:** rebuild images when git mappings swap content
([werf#7947](werf#7947))
([2b46672](werf@2b46672))
* **build:** reduce repeated registry token requests
([werf#7955](werf#7955))
([08dd45b](werf@08dd45b))
* **build:** reduce repetitive service output
([werf#7985](werf#7985))
([a10017a](werf@a10017a))
* **build:** refresh rejected tokens after registry redirects
([a84a30b](werf@a84a30b))
* **build:** reject invalid imports with a scratch base
([werf#7944](werf#7944))
([97367c9](werf@97367c9))
* **build:** release temporary containers after staged COPY
([werf#7938](werf#7938))
([4e27b5f](werf@4e27b5f))
* **build:** report invalid .dockerignore patterns without panicking
([werf#7992](werf#7992))
([29abab1](werf@29abab1))
* **build:** report when ssh connection multiplexing cannot be enabled
([werf#7984](werf#7984))
([851e641](werf@851e641))
* **build:** require Docker 19.03+ for Docker backend operations
([werf#7973](werf#7973))
([fbb7e62](werf@fbb7e62))
* **build:** respect remote Docker contexts and SSH hosts
([b9ca56d](werf@b9ca56d))
* **build:** respect remote Docker contexts and SSH hosts
([werf#7971](werf#7971))
([4df8c09](werf@4df8c09))
* **build:** reuse unchanged Dockerfile contexts without extra copies
([werf#7953](werf#7953))
([f2eab7e](werf@f2eab7e))
* **build:** use a tagged Go image in the documentation example
([werf#7946](werf#7946))
([8a9d42c](werf@8a9d42c))
* **build:** use versioned base images in documentation examples
([werf#7942](werf#7942))
([f591663](werf@f591663))
* **cleanup:** keep dry-run metadata unchanged
([werf#7960](werf#7960))
([189ff24](werf@189ff24))
* **cleanup:** preserve concurrently published final images
([werf#7975](werf#7975))
([31e31e3](werf@31e31e3))
* **cleanup:** purge platform Stapel containers with their volumes
([666afe1](werf@666afe1))
* **cleanup:** purge platform Stapel containers with their volumes
([werf#7969](werf#7969))
([83cab69](werf@83cab69))
* **deploy:** honor OCI credentials in chart-only converge
([f447c2c](werf@f447c2c))
* **deploy:** honor OCI credentials in chart-only converge
([werf#7974](werf#7974))
([52c9706](werf@52c9706))
* **deploy:** keep release output machine-readable by default
([werf#7961](werf#7961))
([d4cc45f](werf@d4cc45f))
* **deploy:** preserve resources with live retention policies
([werf#7958](werf#7958))
([81587cf](werf@81587cf))
* **deploy:** recreate StatefulSet and other custom-validated kinds on
immutable field change
([werf#7994](werf#7994))
([a0128c4](werf@a0128c4))
* **git:** preserve warm caches on cancellation
([4a27ec0](werf@4a27ec0))
* **git:** preserve warm caches on cancellation
([werf#7966](werf#7966))
([9d2ef8c](werf@9d2ef8c))
* **host-cleanup:** purge every stapel container and volume werf owns
([werf#7978](werf#7978))
([6104d14](werf@6104d14))
* **storage:** remember stage tags after publishing
([2377615](werf@2377615))
* **storage:** remember stage tags after publishing
([werf#7965](werf#7965))
([ec099be](werf@ec099be))

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

---------

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
Co-authored-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
Release-As: v3.6.2-dk.1
Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
…tream merge

Upstream now mixes file paths into git ls-tree checksums, so every context digest changed, and the fork's fixtures build from a custom base image, so upstream's recomputed values do not apply either. Replace the five expected digests with values observed on a linux docker run of the suite.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
Add a sbom get entry to the Kubernetes synchronization flags table so the
--synchronization and --kube-* flags the upstream merge added to the command
stay registered. Removing the setup calls from the command makes the new entry
fail.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
Regenerate the StagesStorage mock for the four methods upstream added, move
the docker image inspect test to the moby InspectResponse type, pass the spec
context into prepareWerfConfig, build the report-only BuildPhase as a literal
instead of NewBuildPhase(nil) that the fork dereferences for sbom, and expect
the fork's separate final content tag descriptor in the exporter test. Drop
the duplicate upstream GetFinalStagesStorage stub on fakeStorageManager, which
collided with the fork's final-repo-aware one.

Test-only changes; production code is unchanged.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
Initialize the upstream command-scoped collector before every SBOM get mode so pre-build operations and failures reach the summary. Extend the disabled-SBOM regression to assert config-render timing and command scope.

Signed-off-by: Aleksei Igrychev <aleksei.igrychev@palark.com>
@alexey-igrychev

Copy link
Copy Markdown
Collaborator Author

Validation for 7a9a7f58f:

  • Generated CLI documentation with WERF_* unset; preserved CHANGELOG and release manifest byte-for-byte.
  • macOS: task format, task build, task lint, task test:unit; generated mocks checked with task mock:check.
  • Linux compilation: task build:dev:linux:amd64:go; all test packages compiled using GOOS=linux GOARCH=amd64 task test:go-test paths='./pkg/... ./cmd/... ./schemas/...' -- -exec=/usr/bin/true. The latter is compile-only, not execution; off-host task test:unit stops at the first exec format error.
  • Native Linux/CGO binary built with task build. Scoped runtime checks: Dockerfile context checksums (5 specs), SBOM license/get path (1), SBOM statistics failure path (1), build copy-flags/git-ownership (5).
  • Regression checks rejected the five old context digests and deliberately broken SBOM flags/statistics, primary/final descriptor separation, config import validation, and digest inspection; restored variants passed.
  • Independent review covered conflict resolutions, fork/upstream interaction, generated mocks and final test adaptations.

Scope limits: not a complete e2e/integration run. The broader legacy build run was stopped without a suite result; only the scoped checks above count. Docker daemon was unavailable locally, so runtime checks used an isolated checkout on Linux. The public synchronization default is intentionally retained and explicitly documented in the PR's compatibility section.

@alexey-igrychev

Copy link
Copy Markdown
Collaborator Author

Maintainer explicitly authorized merging despite the current CI failures. The failed jobs report No space left on device (unit binary linking/runner logs; also the v2 documentation image build). Independent validation results are recorded above. This is an accepted CI exception, not a green CI result.

@alexey-igrychev
alexey-igrychev marked this pull request as ready for review October 1, 2026 18:30
@alexey-igrychev
alexey-igrychev merged commit 3e0c325 into main Oct 1, 2026
13 of 14 checks passed
@alexey-igrychev
alexey-igrychev deleted the chore/release/merge-werf-upstream-15 branch October 1, 2026 18:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants