Skip to content

Commit fa34711

Browse files
committed
docs(snapshot s3): document the metadata fingerprint source
Say what --fingerprint-source metadata changes and, as importantly, what it does not. The pipeline is shared, so keys, .kosli_ignore rules and the fingerprint itself are the same in both modes; only where each object's digest comes from differs. The two conditions a bucket must meet -- a stored full-object SHA256 on every contributing object, and no composite multipart checksums -- come with the aws command that fixes each. The obvious assumption is that reading metadata needs weaker permissions than downloading. AWS requires s3:GetObject for both, so the help says so plainly rather than leaving the reader to infer a benefit that is not there.
1 parent e728370 commit fa34711

1 file changed

Lines changed: 13 additions & 0 deletions

File tree

‎cmd/kosli/snapshotS3.go‎

Lines changed: 13 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -19,6 +19,11 @@ In all cases, the content is reported as one artifact. If you wish to report sep
1919
Object keys are never used as local file names: each object is downloaded to a temporary file, hashed and removed, and the fingerprint is computed from the keys and the content digests, so any key S3 accepts can be fingerprinted on any operating system.
2020
Keys that cannot form a directory tree are rejected and fail the snapshot, naming every key involved: a key containing a ^..^ segment, two keys that resolve to the same path (such as ^a//b^ and ^a/b^), or an object whose key is also a prefix of other objects (such as ^a^ beside ^a/b^). A legitimate key of that shape can be left out with ^--exclude-regex^ (anchor and escape it, since the pattern is a regular expression matched against the whole key); when ^--include^ or ^--include-regex^ is set, exclude filters are ignored, so narrow the include filter instead.
2121
22+
By default each object's SHA256 comes from downloading the object and hashing it. ^--fingerprint-source metadata^ reads the SHA256 checksum S3 stores for the object instead, which skips the download, the temporary disk and the hashing. Everything else -- the keys, the ^.kosli_ignore^ rules, the way digests combine into the fingerprint -- is the same in both modes, so the fingerprint is identical and a snapshot matches the artifact you attested either way. Two conditions apply:
23+
- Every contributing object must carry a full-object SHA256 checksum. S3 only stores one when the upload asked for it, for example ^aws s3api put-object --checksum-algorithm SHA256^. Objects without one fail the snapshot, all named in one run.
24+
- A multipart upload gets a composite SHA256, which hashes the checksums of the parts rather than the object content, so it cannot serve as the object's fingerprint. Such an object can be collapsed into a single part in place with ^aws s3api copy-object --checksum-algorithm SHA256 --copy-source yourBucket/yourKey --bucket yourBucket --key yourKey^.
25+
A root ^.kosli_ignore^ is still downloaded in this mode, because its rules decide which objects contribute; the objects it excludes are never fetched and need no checksum. Reading a checksum does not need fewer permissions than downloading: AWS requires ^s3:GetObject^ for both, and an SSE-KMS encrypted object additionally needs ^kms:GenerateDataKey^ and ^kms:Decrypt^ either way.
26+
2227
` + kosliIgnoreDescNoExclude
2328

2429
const snapshotS3Example = `
@@ -69,6 +74,14 @@ kosli snapshot s3 yourEnvironmentName \
6974
--exclude-regex '.*\.png$' \
7075
--api-token yourAPIToken \
7176
--org yourOrgName
77+
78+
# report contents of an AWS S3 bucket without downloading the objects,
79+
# using the SHA256 checksums S3 stores for them:
80+
kosli snapshot s3 yourEnvironmentName \
81+
--bucket yourBucketName \
82+
--fingerprint-source metadata \
83+
--api-token yourAPIToken \
84+
--org yourOrgName
7285
`
7386

7487
// fingerprint sources accepted by --fingerprint-source

0 commit comments

Comments
 (0)