Releases
Changelog
Every released version of DeltaGlider Proxy, newest first. Versions
follow semantic versioning; the Docker image
beshultd/deltaglider_proxy:<version> is published for each tag.
Last updated: 2026-07-22
v1.16.12026-07-22#
Fixed — Verify no longer blames the scan cap for transient read failures#
A partial Verify audit carried only one flag, so the UI always rendered the
"scan capped — raise DGP_PARITY_MAX_OBJECTS" advice even when the real cause
was a handful of objects whose reads kept failing transiently (backend
throttling) — raising the cap does nothing for that. The outcome now carries
the cause separately: a genuinely capped scan keeps the cap advice, while
unresolved reads render as "N objects couldn't be read — re-run the audit
(usually clears); if it persists, lower DGP_PARITY_HEAD_CONCURRENCY". The
server verdict sentence makes the same distinction.
v1.16.02026-07-21#
Added — backend connection health is now enforced, end to end#
A storage backend that can't be reached or whose credentials are rejected can no longer hide (previously it booted silently and every request failed one-by-one — a bucket on it even rendered as innocently "empty" in the browser):
- Boot probe. Every configured backend (default + named) is probed at
startup — an authenticated call, with a scoped-key fallback so
bucket-restricted keys (e.g. Backblaze B2 application keys) don't
false-alarm. Unhealthy backends are logged as ERRORs naming the cause
(credentials rejected / endpoint unreachable / erroring). If all
backends fail definitively (credentials rejected or unreachable — a
merely-erroring/throttling backend never triggers this, so a provider
hiccup can't cause a crash loop), the proxy refuses to start. Tunable
via
DGP_BOOT_BACKEND_PROBE=enforce(default) |warn|off. - Honest 503s. Requests to a bucket routed to a backend with a
definitive fault (credentials rejected / unreachable) answer a fast
503 ServiceUnavailablenaming the backend and cause — all verbs — instead of per-request timeout storms or misleading 404s. A reachable but erroring backend is flagged, never blocked. Unhealthy backends are re-probed every 30s, so recovery self-heals. The 503 (which names internal topology) is only sent to authenticated callers. - Apply gate. Adding a backend or changing a backend's definition
(endpoint, credentials) via the GUI or
config applynow runs a live connection probe first — a typo'd secret or dead endpoint is rejected with the exact cause instead of going live silently. Unchanged backends are never re-probed. - GUI health. Storage → Backends shows a live health badge per backend (Connected / Credentials rejected / Unreachable), a loud error card when one is down, and a Test connection button that runs the probe server-side — exercising the server's actual credentials (the old test ran from the browser and couldn't check credentials at all).
- No more innocent empty states. A failed listing in the file browser now renders as an error naming the cause ("bucket not found on its backend" / "bucket unavailable: backend X unreachable") instead of "No objects yet — add demo data". The bucket-usage chip stops showing stale counts for a bucket whose backend is unreachable.
Fixed — misrouted buckets can no longer exist silently#
A bucket routed to an undefined backend name was a warning ("route will be ignored") and silently fell through to the default backend — where it 404'd. It is now a fatal config error: the proxy refuses to boot with it, and every apply / section save rejects it, naming the bucket, the missing backend, and the available backends. Duplicate backend names are fatal for the same reason. Additionally, a bare 404 on bucket probes (HEAD responses carry no error body) was misread as a transient error, leaving an unrouted bucket "transiently" unroutable forever and blocking backend discovery — now correctly classified as "bucket absent".
Fixed — GUI bucket edits no longer fail with a cryptic serde error#
Editing a bucket's placement (or any storage setting) could fail with
invalid storage section body: invalid type: null, expected a sequence — a
raw serde error, unfixable from the GUI. Root cause: the server's JSON
merge-patch inserted null values verbatim for keys that didn't exist yet
(RFC 7396 defines null as removal, which is a no-op on an absent key).
Fixed per the RFC, and config errors now name the exact offending field
(storage.buckets.<name>.public_prefixes: invalid type ...) instead of a
path-less message.
Changed — Verify scan-cap default raised to 1,000,000#
The default DGP_PARITY_MAX_OBJECTS (max objects a Verify audit scans across
both sides before it caps and reports a partial result) is now 1,000,000
(was 200,000). This is a runaway-scan safety ceiling — the audit already has a
throttle backstop and tunable concurrency — so the old default was capping
ordinary large mirrors (~100k objects/side) and reporting them as "not fully
verified" when they were fine. The new default covers ≈500k objects/side; still
overridable via the env var for even larger buckets.
v1.15.52026-07-20#
Changed — Verify is now honest about what it proves (a metadata audit)#
The replication Verify feature is now called an Audit and states plainly what it does: a fast metadata audit that checks every source object exists on the destination with matching recorded checksums and sizes — with no downloads. It compares recorded metadata; it does not re-read the destination's stored bytes, so it proves the two sides agree, not that the destination reconstructs byte-for-byte. The old copy over-claimed a byte-level guarantee it never delivered (the recorded checksum is copied source→dest verbatim on the delta fast path, so a match there is a metadata match, not a proof of the landed bytes). A clean audit now shows a "Want byte-level certainty?" panel that names deep byte verification as a separate, upcoming capability — so the ceiling of the audit is explicit, not implied.
We investigated a cheap "compare the backend ETag" shortcut to catch stored-byte corruption without downloads and rejected it: it can't be gated safely across encrypted, cross-backend, or multipart destinations without painting healthy mirrors as corrupt. Real byte-level proof requires reconstructing and re-hashing, which will ship as an explicit opt-in.
Changed — one authoritative Verify verdict#
The audit now returns a single server-decided verdict — Verified in sync, Not fully verified, or Differences found — with a plain-language sentence, instead of the screen inferring severity from raw counts. This removes cases where the headline and the status icon could disagree (e.g. "685 differences" for a backup where nothing was actually different).
Fixed — Verify progress no longer looks frozen#
On a large bucket the progress bar sat at an indeterminate "spinner" for the whole
listing phase (the slowest part), so a multi-minute audit looked stuck and got
killed. The bar now seeds an estimated total from the previous run's object count
(so it's determinate from the first update) and labels the phase ("Listing
objects…" → "Comparing…"); the exact total still corrects it once listing
finishes. The "scan capped" notice now says plainly that the rest is unverified
and names the DGP_PARITY_MAX_OBJECTS knob to raise the cap.
Fixed — Verify no longer calls size-only matches "differences"#
The Verify result counted objects that match on size but can't be checksum-compared (no comparable checksum on both sides) toward its "N differences found" headline, with an amber warning — so a backup where nothing was actually wrong could read as "685 differences found". Those are matches, not differences: the headline now counts only genuinely missing/extra/corrupt objects, a size-only-only result reads "No differences found" with a neutral info icon, and the explanation no longer asserts a single cause — it names both (objects written to the backend directly instead of through the proxy, or the two sides storing them differently) and states plainly that it is not a mismatch.
Added — tunable Verify scan cap#
DGP_PARITY_MAX_OBJECTS (default 200000) sets how many objects a Verify audit
scans across both sides before it caps and reports a partial ("scan capped")
result. Raise it for buckets larger than ~100k objects so the whole bucket is
covered in one audit.
v1.15.42026-07-19#
Fixed — much faster replication Verify#
The parity Verify audit on large S3↔S3 rules crawled (a ~99K-object backup
took ~30 minutes) because it checked objects in small batches, waiting for each
batch to finish before starting the next, and re-checked the tiny .sha1/
.sha256/.sha512 checksum files even though it already knew their size. Verify
now streams the checks continuously and skips the checksum sidecars, so it
finishes several times faster. Concurrency is tunable via
DGP_PARITY_HEAD_CONCURRENCY (default 15) for cautious dialing on a slow backend;
the backend-throttle backstop is preserved.
Fixed — Verify no longer shows a stale result mid-scan#
While a new Verify was running, the screen could show the previous scan's findings (with an old timestamp) as if they were current — which looked like the running scan had already found problems. A running/cancelling Verify now shows live progress instead of a phantom verdict.
Fixed — Verify could report a stale size after an overwrite#
On a filesystem destination, Verify's per-object cache keyed on the file's logical checksum, which doesn't change when the same file is re-compressed against a rotated baseline — so an overwritten object could be reported with its old size. The cache now keys on the stored blob's own version, so an overwrite is always detected.
Changed — "Re-run won't help" is now "Re-run may help" for transient errors#
A copy that failed on a transient error (a slow/stalled read, a timeout, a throttle, a dropped connection) was shown as "Re-run won't help", discouraging a retry that would in fact succeed. Those now read "Re-run may help"; only a structural failure (permissions, missing bucket) stays a hard no.
Changed — delete replication is a faithful mirror on every path#
Delete replication now removes any destination object absent at source regardless of provenance on both the scheduled reconcile and the event-driven path (previously the event-driven path still preserved objects it hadn't written itself, so the two paths could delete differently).
v1.15.32026-07-18#
Fixed — false "checksum mismatch" on replicated delta objects (+ a silent corruption)#
When a destination backend dropped a delta object's DeltaGlider metadata headers (the same stripping that affected references), two things went wrong: the post-copy check compared the source's real size against the destination's compressed delta size and reported a bogus "truncation" that a re-run could never clear; and — more seriously — a download of such an object served the raw delta bytes instead of reconstructing the original (a corrupt file under a success response). Both are fixed: the check now recognises the stripped-delta case, and the copy re-stamps the missing metadata so the object stays downloadable and correct.
Changed — delete replication is now a faithful mirror#
With delete replication enabled, the destination is now a faithful copy of the source: any destination object not present at source is removed — including objects written by other tools. Previously only objects this exact rule had written were removed, so a destination could never fully converge. The destination bucket must be dedicated to the rule; a bucket shared with another writer is no longer a supported setup.
v1.15.22026-07-18#
Added — each run shows which algorithm it applied#
The run drawer now names how each copied object was moved, in plain language:
- ⚡ shipped as-is — the delta bytes were copied verbatim, no decompress or recompress (the cheapest path; hover for the technical term).
- ↻ rebuilt — decompressed from the delta and re-stored (recompressed or re-encrypted) at the destination.
- → straight copy — a whole already-compressed object (image, video, archive) copied byte-for-byte.
Alongside the mix, "saved <N>" reports the egress that never crossed the wire because deltas shipped as-is. This turns the run history into an honest account of what the engine did, not just how many objects moved.
v1.15.12026-07-18#
Fixed — replication no longer re-copies destinations with stripped metadata#
A destination bucket whose objects were copied in with a tool that didn't
preserve S3 metadata (the DeltaGlider x-amz-meta-dg-* headers) would be
re-replicated in full on every run — the proxy couldn't read the
destination's logical size/hash, so a content-diff rule always saw a
"difference" and re-shipped, potentially gigabytes per run, forever.
- The destination now self-heals on write. When a replication copy lands in a prefix whose reference lost its metadata, the proxy re-stamps that metadata in place (the stored bytes are never touched, so nothing already there breaks). After one more pass the destination reads cleanly and the rule converges — byte-identical objects are skipped instead of re-copied. This also restores the destination as a usable backup (delta reconstruction needs that metadata).
- This affected pre-existing corrupt destinations regardless of version; it was never introduced by a release.
Added — live reconcile progress in the Jobs screen#
A running replication reconcile now shows what it's doing: the run drawer
displays "scanning <current folder> · N folders done · M pending" with a
live activity bar. A long comparison pass that is skipping already-replicated
objects (and therefore shows "0 copied") now reads as working, and here rather
than looking stuck.
v1.15.02026-07-17#
Changed — replication reconciliation rebuilt as a streaming tree walk#
The scheduled reconcile pass no longer spends hours "initializing" on large buckets before copying anything. The old three-phase design (a full dual-tree discovery pre-pass, then a flat copy sweep, then a third full destination listing for deletes) is replaced by ONE directory-by-directory walk that syncs each folder as it discovers it.
- First copies start within seconds of a run starting, regardless of tree
size — discovery and copying are interleaved, with configurable parallel
directory listings (
replication.dir_concurrency, default 4). - Far fewer requests against the backend. Listings run in lite mode everywhere (the per-page HEAD-enrichment sweeps are gone); a subtree missing on the destination is proven absent from the folder comparison itself, so its objects copy with zero existence probes. On a filesystem↔filesystem pair, an entire run — initial sync or converged re-check — issues zero per-object HEADs; on S3 backends only keys present on BOTH sides are HEAD-resolved, batched.
- Delete replication is now per-directory. A folder that reconciled cleanly can replicate its deletions even when the run stops early elsewhere (the old design only deleted after a perfect full pass), while keeping every safety layer: provenance marker, HEAD re-confirmation of source absence, and never touching foreign objects.
- Resume is finer-grained. The persisted cursor is now a position in the tree (scope-stamped to the rule's buckets/prefixes); a killed, paused, or budget-truncated run resumes exactly where it stopped, even mid-directory. On upgrade, the previous cursor format is discarded once and the next run starts a fresh (idempotent) pass.
- Live progress: the Jobs screen shows directories completed/pending for
a running reconcile, and three new Prometheus counters
(
deltaglider_replication_list_calls_total,..._head_calls_total,..._dirs_completed_total) make the walk's I/O auditable. - The walk core is a pure state machine with a model-based property-test suite (action-set equivalence against a naive global diff, crash-cut resume equivalence, page-size invariance, never-false-delete, cursor monotonicity, budget convergence, degrade-mode equivalence).
Fixed#
- Kill/pause/lease-loss now interrupt a replication run promptly at any phase. Two driver bugs found while stress-looping the integration suite: control checks could go dead for the remainder of a run after listings finished, and a config-DB lock could be handed to a suspended poll future, freezing every DB user in the proxy (jobs API, heartbeat, kill) until restart.
- Filesystem delimiter listings no longer hide dot-directories (e.g.
.well-known/): they were invisible as folders while their objects were reachable directly — now consistent with flat listings and the S3 backend. Dot-files (internal temp namespace) stay hidden;.dgstays internal. - Keys with an empty path segment (
a//x) now replicate to their true destination key. The per-object rewrite used by the old pass collapsed the doubled slash (a//xanda/xfought over one destination key); the walk maps prefixes verbatim. Pre-existing collapsed copies become rule-owned orphans that a delete-enabled rule converges away.
v1.14.12026-07-16#
Two operability fixes from production issue reports.
Fixed#
/_/readyno longer flips to a paging 503 on a brief backend blip (#62). The readiness probe's backend check is now retried with a tunable per-attempt timeout (DGP_READY_TIMEOUT_SECS, default 3) and retry count (DGP_READY_RETRIES, default 2) — it reports not-ready only if every attempt fails, so a short storage-provider latency spike (e.g. object-storage tail latency) doesn't mark an otherwise-healthy edge out of rotation. Point load balancer readiness checks at/_/readyand liveness at/_/health.- A bucket declared in
storage.bucketson a filesystem backend is created at startup (#63), so its first write no longer returnsNoSuchBucketbefore an explicitCreateBucket. Declared intent only: implicit bucket creation on the write path stays refused, and remote S3 buckets are never auto-created.
v1.14.02026-07-10#
The completion of the correctness/security review begun in v1.13.0 — the remaining edge-case findings, all fixed with regression tests. Multi-instance deployments and shared-host installs get the most benefit.
Fixed — data & integrity#
- Multipart uploads verify each part's content at completion, so a part file tampered with on disk (shared-host attack) is rejected instead of stored.
- A streamed upload with an unexpected size is rejected instead of stored short, and a garbled backend error can no longer crash the notification dispatcher.
- A deleted object can't be resurrected by a leftover storage variant, and an object addressed with an extra leading slash shares one cache entry.
- A backend returning HTTP 429 (Backblaze B2, Cloudflare R2) is treated as a transient throttle, not a hard error.
Fixed — multi-instance / jobs#
- Run-now, verify, and delete on a replication rule now see a run in progress on another node (when cluster coordination is on), preventing double-runs and deleting a rule mid-run.
- A stuck replication consumer no longer makes the event log grow forever, and a fully-throttled small batch now backs off instead of grinding.
- A killed replication run no longer mis-marks a successfully-copied object as failed.
- A migration's source cleanup no longer stops early when it hits a page of folder markers, and a concurrent large upload's parts aren't swept by the periodic cleanup.
Fixed — config & correctness#
- Values containing
$round-trip correctly through save/reload, applying a redacted export can't create a passwordless user or provider, and editing a masked URL list won't restore the wrong secret. - Deleting an OAuth provider on one node removes it everywhere, and streaming transfers account for their temporary disk so heavy concurrency can't exhaust it.
- Several smaller hardening fixes (codec watchdog false-kills, backend recovery after a cancelled probe, parity conflict timestamps).
v1.13.02026-07-10#
A correctness-and-security hardening release. A deep review of the codebase turned up a set of data-loss, data-corruption, and authorization-bypass bugs in edge cases (concurrent config edits, cancelled uploads, slow backends, rejected config applies); every one is fixed with a regression test.
Fixed — data loss & corruption#
- Concurrent admin edits no longer silently lose a user. Creating two users in quick succession could make one of them vanish everywhere (the config-DB sync path clobbered a just-committed change with an older copy). Same-node syncs are now serialized.
- A slow upload to a remote backend no longer loses its data. A large multipart upload whose final store ran longer than the cleanup timeout could have its temporary parts deleted mid-store, permanently failing the upload on retry. In-flight stores are now protected from the cleanup sweeper.
- Background copies no longer race a re-encryption/migration job. A replication or lifecycle copy in flight when a maintenance job started could be overwritten with stale bytes (or lost after a migration flip). Those copies now register with the maintenance write-gate so the job waits for them.
- A dropped connection no longer wedges a bucket's maintenance jobs. A client disconnecting mid-write used to leak an internal counter, making every later re-encrypt/migrate on that bucket time out until a restart. The counter is now released even on cancellation.
- A just-created bucket appears immediately in the next listing even under a concurrent listing refresh (a race could hide it for a few seconds).
- A backend's encryption key survives a config export→edit→apply round-trip — it was being silently stripped, which could leave historical objects unreadable and new writes unencrypted.
Fixed — security & authorization#
-
A rejected config apply no longer leaks a public prefix. An apply that failed its IAM validation could still publish a public-read prefix (serving objects anonymously) while reporting "no change." Validation now runs before any change is committed.
-
Prefix-scoped and IP-scoped Deny rules are enforced on batch deletes and browser (form-POST) uploads — both paths previously skipped the per-object authorization check, defeating carve-out and source-IP conditions.
-
A revoked session is dropped from memory immediately, and a compromised key can no longer be un-revoked by a concurrent config sync.
-
A quota can no longer be bypassed by a corrupt object-size reading that overflowed the usage total, and a truncated streaming upload is now rejected instead of stored as a complete (short) object.
-
A run whose process died within its lease TTL no longer shows "running" forever. The boot-time zombie scan spared any
runningrow whose leader lease still looked live — but a process that dies right before a restart leaves exactly such a lease behind, so its run dodged the (boot-only) scan permanently. Boot now lapses every local lease first: the coordination tables are node-local, so a lease found at boot can only belong to a dead process of this node. Applies to replication and lifecycle.
Fixed — config apply#
- Changes to passthrough size limit, codec concurrency, TLS, and the config sync bucket now take effect (or clearly warn a restart is required) instead of being reported as applied while silently doing nothing.
Fixed — additional hardening#
- A backend rate-limiting with HTTP 429 (Backblaze B2, Cloudflare R2) is now treated as a transient throttle (back off and retry) instead of a permanent 500.
- The event-notification dispatcher no longer crashes on a backend error message that happened to contain a multi-byte character at a truncation point.
- Applying a redacted config export to a fresh instance now fails clearly if a new user or OIDC provider has no secret, instead of creating identities that can never authenticate.
- Objects addressed with an extra leading slash (
/keyvskey) no longer leave stale cached metadata after a delete. - Streaming delta transfers account for their temporary disk correctly, so many concurrent large transfers can't exhaust the spool area.
- Editing a masked webhook/Slack URL list no longer risks restoring the wrong saved secret, and internal object keys no longer trigger user webhooks.
v1.12.42026-07-09#
The "plays nice with throttled backends" release: batched HEAD sweeps with circuit breakers, adaptive client-side rate limiting, replication runs that back off instead of grinding a throttled key list — and a job drawer that shows WHAT is copying (name + size), so a slow counter reads as honest work.
Fixed#
- HEAD sweeps stop instead of grinding a throttled backend. Listing
enrichment now runs its per-object HEAD lookups in batches of 10 (down
from 50 concurrent) and aborts the sweep when a batch comes back mostly
503 SlowDown— the remaining objects render with their listing metadata. Deltaspace savings scans keep full coverage but log throttle counts. - The job drawer shows WHAT a replication run is copying. An active run now lists its in-flight objects (name, size, started) above the tabs — a counter sitting on "1 copied" for 30 minutes reads as honest work on a 4 GB tarball instead of a hang. Live view, top 3 largest, refreshed by the existing 2s poll.
- Background jobs stop hammering a throttled backend. A replication
run now aborts (and backs off to the rule's cadence) when a whole page
of copies is rejected with
503 SlowDown, instead of grinding the remaining key list — the cursor resumes where it left off. Verify's HEAD-resolution runs in batches of 10 (down from 50 concurrent) and marks the remaining keys unresolved (honest partial audit) when a whole batch fails transiently. - The S3 client uses adaptive retry. One throttle response now rate-limits every subsequent request on that backend client (HEAD bursts, listings, replication reads) via the AWS SDK's client-wide adaptive rate limiter, instead of each call site independently hammering a struggling backend.
v1.12.32026-07-09#
A network-hygiene patch, prompted by a Hetzner SlowDown throttling episode
whose RCA surfaced how chatty a bucket refresh really was.
Fixed#
- One upstream
ListBucketsper backend per refresh, not two. The browser fires the S3 bucket list and the backend-origins lookup back-to-back; the routing layer now serves a listing fetched within the last 5 seconds without re-probing upstream (DGP_BACKEND_LIST_FRESH_SECS,0disables). Creating or deleting a bucket through the proxy invalidates the window immediately, so read-after-write stays exact. Stale rosters still back the unavailable-placeholder path during outages. - Five UI components no longer fire five independent bucket listings. The sidebar, destination picker, analytics, scan card, and permission editor now share one in-flight request and a short-lived result; the explicit refresh button and bucket create/delete bypass the cache.
- Alt-tabbing back no longer bursts the admin API. React Query's focus-refetch re-fired every active query on each window focus; it's off — live data is covered by the explicit pollers.
- The idle bucket-maintenance poll runs 4× less often (60s instead of 15s when no job is running; still 2s while one is active).
v1.12.22026-07-09#
The "feels alive" release: the golden-roadmap UX waves (verify progress that never looks dead, first-press Back button, deep-linkable admin views, plainer copy, a virtualized object table), a listing memory fix, community-contributed deny-path metrics + an IAM secret-preservation fix, and the three maintenance commits that accumulated after v1.12.1 (SigV4 single-verification, dead-backend listing cooldown, replication-target alias hardening).
Added#
-
Every admin view is now deep-linkable. The job drawer, metrics view toggle, setup wizard step, file preview, user/group selections, and all modals (YAML import/export, full-IAM, inspector, destination picker) push their state to the URL query string. Share a link, bookmark it, or use Back/Forward — the UI restores to the right place. Modals close on Back button; direct-loaded deep links open the overlay and close safely without walking out of the SPA.
-
Denied requests now show up in Prometheus. Admission-rule and IAM authorization 403s short-circuit before the metrics middleware, so they were invisible to
deltaglider_http_requests_total. Both deny paths now record the sanitized counter (method/status/operation). Contributed by @skel84. -
The object table is virtualized. Only the visible rows render in the DOM (previously up to 500), the body height fits the viewport, and the pagination bar stays visible. Keyboard navigation scrolls via the table's native
scrollTo. -
Loading states show motion everywhere. Bare "Loading…" labels in the webhook, audit, buckets, outbox, and dashboard panels are replaced with the shared spinner/skeleton primitive.
Fixed#
-
Verify progress bar no longer looks dead on large scans. A four-part fix: the server now flushes parity progress on a 250ms wall-clock throttle (not just per-page), the UI polls verify status at 800ms (not the generic 2s), the first progress frame is no longer stuck on "Starting scan…", and a reusable
useTweenedCounthook adds a smooth count-up animation that honorsprefers-reduced-motion. -
First browser Back no longer swallowed. The SPA router's
skipNextflag was armed on everypushStatebutpushStatenever emitspopstate, so the flag sat armed and ate the first real Back press. The flag is removed entirely — Back works on the first press. -
Login no longer double-fires on Enter. A synchronous
useRefguard replaces the asyncuseStatecheck, soonPressEnterand the buttononClickin the same tick can't both pass the guard. -
Renaming a job no longer spams browser history. The drawer's rename retarget now uses
history.replaceStateinstead ofpushState, so typing 5 characters doesn't create 5 Back entries. -
Invalid
?tab=on a shared job link no longer renders a blank drawer. The tab is clamped to the tabs that exist for that job kind. -
Inspector close on a direct-loaded
?object=link no longer walks out of the SPA. The close handler now uses the direct-load-safe overlay pattern:history.back()if we pushed the entry,replaceStateif it was loaded from a shared link. -
Setup wizard no longer offers "Save and start" over empty config after a refresh. The step is forced to 0 when form state is initial; the
?step=deep-link only applies to in-session Back/Forward. -
Savings chip no longer hides genuine sub-1% savings. The zero-guard now checks the raw float from the server, not the integer-floored display value.
-
Hidden tabs no longer waste network requests. Two unguarded 30s
setIntervalpollers (bucket scan card, analytics section) are routed throughuseVisiblePolling, which stops polling on hidden tabs and fires an immediate refresh on tab return. -
Delimiter-less LIST no longer materializes the whole prefix. A ListObjectsV2 without a delimiter fetched every key under the prefix into memory before cutting one page out of it. The S3 backend now pages natively with a provably-safe early exit (the completeness horizon accounts for
.delta-suffix collation — prefix-sharing names likev1/v1.2were the hard case, validated by brute-force simulation). -
Applying an exported config no longer wipes IAM user secrets. The document export deliberately redacts
secret_access_keyto an empty string; the declarative reconciler treated that as a rotation to an empty secret and broke every user's credentials. Empty now means "keep the existing secret", matching the field- and section-level admin paths. Contributed by @skel84. -
The bootstrap login hint speaks plainly. "Deployment-level access for first run and recovery" → "For first-time setup and recovery".
-
Replication-target buckets can't be written via their unconfigured alias. A client PUT to a marked bucket's real (alias) name — the one discovered by cross-backend head-probe, not the configured virtual name — bypassed the single-writer guarantee. The proxy now indexes every marked bucket's real storage name and blocks writes that resolve onto protected storage, while still exempting names with their own explicit backend route.
-
A dead backend no longer hangs bucket listings. After the parallel fan-out, one unreachable backend still cost a connect timeout (up to 10s) on every ListBuckets refresh. A per-backend failure cooldown (
DGP_BACKEND_LIST_COOLDOWN_SECS, default 30) now skips a backend that just failed a listing — serving from last-known-good, flagged unavailable — and each listing call is bounded byDGP_BACKEND_LIST_TIMEOUT_SECS(default 5).
Changed#
- Clearer labels throughout the admin UI. "Bootstrap password" → "Admin password", "Storage footprint" → "Storage used", "Trace" → "Rule tester", "Delta efficiency" → "Compression health", "Admission rules" → "Request rules", "Event outbox" → "Event log". Six verbose verify descriptions replaced with plain language. Rust denial messages softened to human-readable text. Product docs updated to match.
- Natural sort everywhere. Bucket names, object keys, and rule lists now
sort naturally (
v2beforev10) via a sharednumericComparehelper. - Consistent error display. All 58
e instanceof Errorsites across 25 files replaced with a sharednormalizeUiErrorhelper. - Config warnings use human-readable text.
Config::check()warnings no longer leak the internal field namereplication_target_only. - SigV4 is verified once, not twice. The redundant outer axum signature verifier (~380 LOC of hand-rolled canonical-request + HMAC) is deleted; s3s is now the sole SigV4 authority. The outer middleware resolves identity only (access-key → user) and keeps replay detection, rate-limiting, and anonymous public-prefix minting. s3s rejects forged signatures — header, presigned, and chunked streaming — before any handler runs. A valid-AKID/wrong-secret forgery is still 403'd but no longer feeds the outer brute-force lockout (HMAC-guessing is infeasible; locking on it would let an attacker DoS a victim's key). Credential-enumeration lockout is unchanged.
v1.12.12026-07-08#
A hardening release from a full adversarial review of the last several versions. No new features; correctness and safety fixes throughout.
Fixed#
- Verify no longer flags a healthy mirror as broken. On same-backend mirrors, some correctly-replicated objects could show up as a permanent "checksum mismatch" that no re-run could clear. Verify now compares the same logical checksum in every case (reading it straight from the listing when the backend allows, so it stays download-free), and the progress bar climbs smoothly to 100% instead of overshooting or freezing.
- A temporarily-unreachable backend no longer hides buckets. Buckets on a backend that's throttling or down now stay listed (flagged unavailable) even when they aren't explicitly pinned to that backend, and the bucket list survives a full outage instead of collapsing to an error. A throttled backend now returns a proper "slow down" signal so clients back off, rather than an error clients treat as permanent. The UI also stops trying to browse, scan, or copy into buckets it knows are unavailable.
- Storage migration can't strand the proxy. Migrating a bucket onto a backend that doesn't support the required safety guarantees is now refused up front, instead of succeeding and then failing to start on the next restart.
- Backend capability detection is more robust. The check that decides whether a backend is safe for multi-instance use no longer misreads unrelated details (an endpoint port, a request id) as a failure, and never leaves a stray probe object behind in your buckets.
- Browser (form) uploads honor the full upload policy. An upload whose signed policy doesn't constrain the object key is now rejected, matching AWS behavior, so a leaked upload signature can't be reused to write elsewhere.
- Replication mirrors are protected from accidental deletion. A replication-target bucket can no longer be created or deleted through the S3 API by clients.
Changed#
- Faster bucket listings when multiple backends are configured (queried in parallel), and clearer error detail when a backend can't be reached.
v1.12.02026-07-06#
Verifying a replication mirror now shows real motion, and a pure 1:1 mirror verifies without downloading anything.
Added#
- Verify shows live progress. While a parity recompute runs, the panel now shows a scanned-object count that ticks steadily plus an objects/sec rate, and a determinate progress bar once the total is known — instead of a spinner that gave no sense of movement.
- Pure mirrors verify with zero downloads. When source and destination use the same encryption and compression (a plain 1:1 mirror), Verify now compares the stored object size and etag straight from the bucket listing — no per-object metadata fetches. The verdict reads honestly as "exact copies (same stored size + etag)", and the footnote says the check used the listing, not a checksum download.
Changed#
- Verify picks its strategy per rule. A mirror that re-encrypts, re-compresses, or re-deltas at the destination (so the stored bytes legitimately differ) still uses the full logical-checksum comparison — that path is unchanged. Only genuine byte-for-byte mirrors take the cheap listing-only path, so an encrypted or transformed destination is never falsely flagged as mismatched.
v1.11.02026-07-05#
The bucket list never lies about what exists when a backend is unreachable.
Fixed#
- A temporarily-unreachable backend no longer empties or silently truncates the bucket list. When a storage backend returns errors (a provider 503-throttling, a connection failure), its buckets used to vanish from the list — and if every backend was erroring, the list came back empty, which the UI showed as "create your first bucket" with a clean console. Now: if all backends fail, the listing returns an error (the UI shows a distinct "couldn't load buckets" state with retry, not an empty one); if some fail, the reachable buckets are still listed.
Added#
- Unreachable buckets stay visible, flagged and explained. A bucket whose
backend is temporarily down is now listed (greyed, non-clickable, with an
"unavailable" badge) instead of being dropped — sourced from the configured
routing table so the roster stays complete. Hovering it shows the verbatim
backend error (e.g. the raw
503 SlowDown), so operators see exactly why a bucket is dark rather than assuming it disappeared.
v1.10.02026-07-04#
Safe non-CAS backends (Backblaze B2 as a cheap replication target), automatic replication failover across instances, and a hardened browser upload path.
Added#
replication_target_onlybucket marker. Mark a bucket a replication-destination-only mirror: client writes (PUT, DELETE, multipart, copy-into, browser uploads, admin bulk copy/move/delete) are refused with 403 so replication stays the single writer. This is what makes a backend without conditional-write support (e.g. Backblaze B2) safe to host a mirror — reads are unaffected, and prefixes can still be published read-only. See the new how-to, Using non-CAS backends safely.- Backend conditional-write validation. When multi-instance mode is active
(a
config_sync_bucketis set), the proxy probes every S3 backend that hosts a client-writable bucket at startup and refuses to boot on one that doesn't enforce conditional writes (a data-corruption trap under concurrent writes). A runtime/config/applythat would create the same situation is rejected. Verdicts surface in the admin GUI's Backends panel; every enforcement message links to the troubleshooting doc. - Automatic replication failover across instances. With a CAS-capable coordination bucket, replication rules elect a single leader per rule via an S3 lease object; a dead leader's lease lapses and a peer takes over automatically — no double-run, no shared database required. Single-instance deployments keep the node-local lease (zero S3 traffic).
Fixed#
- Hardened the browser form-POST upload path. Closed a bucket-name path traversal (an unvalidated bucket segment could escape the storage root), routed failed upload signatures through the same per-IP brute-force limiter and audit ring as every other auth surface, and now reject form fields not covered by the signed policy (AWS's "every field must have a condition" rule). The signature/policy verification itself was audited clean.
Docs#
- New troubleshooting guide for backend capability validation; corrected stale HA/multi-instance claims across the README, product docs, and code comments to reflect the shipped cross-instance replication lease.
v1.9.62026-07-03#
Verify reliability + a Jobs-table icon fix.
Fixed#
- A crashed Verify audit no longer shows "running" forever. A verify
(parity) audit runs as a background task that settles its own status when it
finishes. If the proxy was killed mid-audit (or a re-launched audit crashed
again) with no restart afterward, the row stayed stuck on
running— the Jobs → Verify tab polled a spinner indefinitely, which looked like an extremely slow scan. The status read now self-heals: arunningrow whose background lease has expired (a live audit renews it continuously) is recognised as dead and settled tofailed, without needing a restart. - Distinct icon for "Run now" vs "Resume". After the Jobs actions went icon-only, a paused rule showed two identical play carets (Resume and Run now). Run now now uses a circled-play icon.
v1.9.52026-07-02#
Jobs-screen usability fixes.
Fixed#
- Replication and lifecycle glob fields accept multiple lines again. The Include/Exclude glob editors silently ate the Enter key — a controlled round-trip stripped the trailing blank line on every keystroke, so you could never start a second line. They now hold the raw text while you edit and only normalise on blur.
- The Jobs table is cleaner and responsive. Removed the segmented per-cell underlines (one row separator now), stopped the Scope and rule-name columns from wrapping (single-line with a full-text tooltip), and made the row actions icon-only so a job row no longer overflows; it collapses to readable cards on narrow screens.
- The Verify "Differences" panel no longer repeats itself. The WHY column showed the same cause twice (a wrapping amber block plus a duplicate grey line); it's now one compact verdict chip and a single cause line. The FIX column's guidance no longer looks like a clickable button.
- The "New job" menu closes when you pick a type, and opening a draft rule that was already discarded or renamed closes cleanly instead of flashing "this job no longer exists".
v1.9.42026-07-02#
Admin-UI simplification, a leaner CI pipeline, and wider property-test coverage.
Changed#
- Admin pages simplified to their essentials. Config panels now share two content widths instead of a dozen ad-hoc ones; the System "Request limits" fields dropped their per-field restart badges and environment-variable boxes for a single help line; the "data lives in the encrypted DB, not YAML" banner on Users/Groups/External-auth shrank from a three-fact block to one quiet line; Credentials and Backends copy is plainer with YAML keys as hover chips.
- Dashboard leads with headline KPIs. The Monitoring view shows 9 always- useful cards up front and tucks 11 deep-telemetry charts behind a default- closed "Detailed telemetry" toggle, so a fresh install no longer opens onto a wall of empty graphs.
- Removed duplicate in-panel headers on Event delivery, Trace, and Audit log (the page title already names them).
Internal#
- CI is faster and leaner. The UI bundle and release binary are now built once and shared as artifacts instead of being rebuilt in ~8 and ~3 jobs; the non-gating coverage and supply-chain (deny/audit) jobs moved to the nightly workflow, so the PR gate is correctness-only. E2E, docs, and config-schema jobs dropped several minutes each.
- Wider property-test coverage on the pure validators — bucket-key paths, the IAM method→action classifier (with a security invariant that mutating methods never map to read-only actions), public-prefix normalization, and replication key rewriting.
- Documentation refreshed to match the current admin UI and corrected config
defaults (
DGP_CACHE_MB100 MB,DGP_CODEC_CONCURRENCYnum_cpus × 4).
v1.9.32026-07-01#
Security + data-integrity fixes from an adversarial code review, plus the non-blocking run-now the copyHzToBackup incident called for.
Fixed#
- Config-DB recovery no longer destroys the IAM database (data loss). After a
bootstrap-password mismatch, a second unattended restart could overwrite the
preserved backup (
deltaglider_config.db.bak) with the empty database created on the mismatch boot — permanently losing IAM users, groups, and OAuth providers. The boot backup now refuses to clobber an existing.bak, and recovery is read-only (it returns the correct hash for you to set, rather than writing hidden state that led into the loop). - The bootstrap-hash sidecar is created private (0600) in one step. It was written world-readable then chmod-ed, leaving a brief window in which a local user could read the SQLCipher key on a shared host.
- Event-outbox "purge failed" can no longer drop unconsumed replication events. The outbox is a shared stream that event-driven replication also reads; the failed-row pruner and the operator purge now respect the replication cursor floor (the purge refuses with a clear message when failed rows still sit above it), and failed rows age out even when delivered-retention is 0.
- Killing a run is settled reliably. A kill arriving in the gap between the final cancel check and the run settling could be overwritten by a "succeeded" status; the check and the settle now happen under one lock.
- Deleting a replication rule with a run in progress is refused (409 "kill it first") instead of purging the running worker's state out from under it.
run-nowreturns immediately (202) instead of hanging the request for the whole replication — a multi-GB sync no longer blocks the admin call. Progress shows in the jobs list.run-nowalso runs a disabled or paused rule as a deliberate one-off (labelled "Run once") without re-enabling it.
Added#
- Cross-instance session revocation. Force-logging-out a user (the compromised-key escape hatch) now takes effect on every instance, not just the one serving the request — a synced revocation table is consulted on every auth check. (Killing a running replication job remains instance-local under a non-sticky load balancer — the UI/response now say so.)
v1.9.22026-07-01#
Replication performance + the kill/responsiveness fixes that an empty-destination copy exposed in production.
Added#
- Directory-aware replication (prefix-tree destination oracle). Replication
now descends the destination and source prefix trees like directories instead
of flat-listing every key. A destination subtree that doesn't exist is proven
absent from a single common-prefix probe, so its entire source subtree is
bulk-copied with zero per-object destination HEAD requests. A copy to an
empty or sparse destination no longer spends minutes "thinking" before it
starts copying — an empty destination costs one list call.
SkipIfDestExistsrules now decide from the listing alone and never issue a per-object HEAD, even against a fully-populated destination.
Fixed#
- Copy to an empty/sparse destination was pathologically slow. The planner issued one destination HEAD per source object just to discover the object was absent (it always copies when missing, under every conflict policy) — minutes of wasted round-trips against a remote backend. Replication now lists the destination once and only HEADs keys that actually exist.
- A running replication job could not be killed while it was planning. The kill check only ran during the copy phase, so a run wedged in the per-page list/plan phase ignored the operator's Kill until it reached a copy checkpoint. The kill is now honored at every page boundary.
- A killed run kept showing as "running" and a second Kill returned 409. The Jobs list read the rule's last-finished status while Kill flips the in-flight run-history row to "cancelling" immediately — so a killed-but-not-yet-settled run is now correctly shown as "cancelling" instead of a stale "running".
v1.9.12026-06-30#
A hardening release for the v1.9.0 replication job controls (kill/delete), plus three operator recovery hatches surfaced by an adversarial review.
Fixed#
- Killing a run no longer leaks a multipart upload. A killed replication run dropped the in-flight copy future without aborting the destination multipart upload, leaving it dangling on backends that don't garbage-collect incomplete uploads (e.g. Backblaze B2). The abort now runs on the cancellation path too.
- A cancelled run is settled authoritatively. A kill requested while the last
page was draining could be overwritten by a success status; and a crash between
the kill and the run settling left the row stuck
cancellingforever. Thecancellingstate is now honored at settle time and reconciled at boot. - Deleting a replication/lifecycle rule can't resurrect it. A persist failure during delete left the in-memory and on-disk config out of sync, so the deleted rule reappeared on the next restart. Delete now rolls back cleanly on failure, skips a needless engine rebuild, and purges the rule's DB rows in one transaction.
- Fail-fast destination detection now matches real failures. The "destination unusable" fast-abort missed the actual Backblaze B2 over-cap / dead-credential error shapes and could over-eagerly abort a healthy run on a stray token. It now recognizes the real shapes, only aborts when an entire page copies nothing, and backs off to the rule's normal cadence instead of re-firing every minute.
- Lifecycle rules no longer show a non-functional Kill button in the Jobs UI (kill is replication-only); killing a run now asks for confirmation.
Added#
- Admin session revocation. A new Access › Sessions page lists live admin sessions and lets an admin force-logout a single session or every session of a given access key — the escape hatch for a stolen cookie or a rotated key, without restarting the proxy.
- Event-outbox failed-row purge. Failed delivery rows (retries exhausted) are now aged out automatically and can be purged on demand from the Event outbox page, so a dead webhook/Slack target can't grow the database unbounded.
- Config-DB recovery auto-unlocks on restart. When the bootstrap password mismatches the encrypted config DB, a successful recovery now persists the recovered hash so the next restart comes up unlocked — no manual env wiring.
v1.9.02026-06-30#
Changed (BREAKING)#
- Replication: removed the
source-winsconflict policy; addedcontent-diff.source-winscopied every object on every reconcile sweep unconditionally — it could never converge on a recurring rule (the rule re-shipped the entire bucket forever). It is replaced bycontent-diff, which keeps the destination an exact mirror by copying only when the bytes differ (size differs, or both sides carry a logical SHA-256 and those differ) and skipping byte-identical objects — so the rule converges and a healthy sweep eventually copies nothing. A config carryingconflict: source-winsis now rejected at load (unknown variant 'source-wins', expected one of 'newer-wins', 'content-diff', 'skip-if-dest-exists'); switch such rules tocontent-diff(same "dest = source" intent, without the re-ship).
Added#
- Lifecycle rules are validated at config time and rejected if they can never
run. A delete/transition rule missing (or with an invalid/out-of-range)
expire_after, a retain-newestcount: 0, an empty transition destination bucket, a malformed glob, or a duplicate rule name is now a fatal config error (400) on every path —config apply/validate, the admin section PUT (the GUI Storage/Jobs editor's save), andconfig lint— instead of being accepted as a warning and then failing every scheduler tick.
Fixed#
- Replication now honors pause mid-run. Pausing a rule while a sweep was in flight previously had no effect (the flag was only checked when starting a run), so a long run kept going and lingered as "running". The worker now re-checks the pause flag at every page boundary, stops promptly, preserves the cursor for resume, and settles the run as terminal.
- Lifecycle run-now/preview on an invalid rule returns 400 without recording a spurious FAILED run row.
- Security: an unverified email can no longer bootstrap a group grant via a
claim_valuerule keyed on theemailclaim — that path read the same attacker-controllable value as theemail_*rules and is now gated identically. - Hardened object-timestamp handling: a missing
dg-created-atresolves to the object's stable S3LastModifiedinstead of a freshnow()(and the parser now accepts timezone-offset / lowercase-ztimestamps), removing a class of spurious re-reads. - HA state consistency: a replication rule whose process died mid-run is now
marked
failedin its state row on boot (was left showing "running"), and the crash-recovery zombie-run scan is lease-aware so it never resurrects a run a live peer still holds.
v1.8.32026-06-29#
Fixed#
- Verify tab showed stale results as if live during a re-verify. While a re-verification was running, the panel kept displaying the previous run's full verdict (difference counts, breakdown chips, "Checked Nh ago") with only a small "Re-verifying…" button to hint it was out of date — easy to mistake the old numbers for the current state. A re-verify now shows a prominent progress banner (live scanned count + Cancel) and dims the prior result, clearly labelled as the previous verdict being refreshed.
Changed#
- Polished the verification wait screen. The first-run/in-progress state is now a framed panel matching the verdict's visual language — a pulsing shield halo, the live objects-scanned count, an indeterminate progress bar, and a reminder that the scan runs in the background — instead of a bare spinner.
v1.8.22026-06-29#
Fixed#
- Job-drawer Runs/Failures/Verify lists collapsed to a sliver. In the Jobs
drawer, the record list shrank to its content width instead of filling the
drawer, and the outcome meter inside it collapsed — the proportional track
disappeared and the outcome text clipped to a few characters (
run…,459…), readable only by hovering for the tooltip. The list now fills the available width and the meter renders in full (dot · track · label).
v1.8.12026-06-29#
Fixed#
- Mobile horizontal overflow across the admin UI. Every admin settings page
(all 15 routes) could scroll sideways on a phone, clipping chips, buttons, and
table columns. Root cause was flex containers — the content pane and the
Ant Design
<Layout>chain — defaulting tomin-width: auto, so a single wide row or fixed-width table widened the whole shell past the viewport. Fixed systemically (min-width: 0) plus per-panel wrap fixes; dense diagnostic tables (Audit, Event Outbox) now scroll within their own card. The Users/Groups/Authentication master-detail split stacks vertically on narrow screens. Verified at a 380px viewport with zero page-level horizontal scroll. - Jobs and Verify-findings panels are now mobile-friendly — the Verify results table reflows to stacked cards, and the Jobs list drops its duplicate header.
- The unsaved-changes bar no longer hides behind the Jobs drawer — it now floats above the drawer where replication/lifecycle rules are edited, so Apply stays reachable.
v1.8.02026-06-29#
Security#
- Spoofable client IP via
X-Forwarded-Foris now gated behind a trusted-proxy allow-list. WhenDGP_TRUST_PROXY_HEADERS=true(the documented reverse-proxy setup), the proxy previously trusted the firstX-Forwarded-Forelement verbatim — letting any client forge the IP used foraws:SourceIpIAM conditions and admissionsource_iprules. The newDGP_TRUSTED_PROXY_CIDRSlets you list the CIDRs of your real proxies;X-Forwarded-Foris then honored only from a trusted peer and parsed right-to-left (the rightmost non-trusted hop is the real client). Default behavior is unchanged (DGP_TRUST_PROXY_HEADERS=falsestill ignores the header); operators behind a proxy should set the CIDR list to close the spoofing window. - SSRF guard now resolves DNS. Outbound URLs (OIDC issuer/JWKS/token,
webhooks) were checked only as literal IPs, so a hostname whose DNS record
pointed at
169.254.169.254or a private range passed. A connect-time resolver on the OIDC and webhook clients now re-checks the resolved address and fails closed. Note: an OIDC issuer or webhook on a private/internal address is now rejected — set it to a publicly-resolvable host, or reach out if you need an internal-network escape hatch.
Added#
- Replication "Verify" (parity audit) is now a fast, server-side background
job. Verifying that a replication mirror is byte-identical used to run
synchronously in the request (HEAD-per-object every time) and lived only in the
browser — navigating away recomputed it from scratch. It now runs as a
background job whose result is persisted server-side (survives navigation
and proxy restart), with a persistent per-object cache so re-verify is
HEAD-free, plus live progress and a Cancel button. Exposed at
GET/POST /_/api/admin/jobs/:id/verify(+/verify/cancel). - Readiness probe at
/_/ready./_/healthis now liveness-only (fast, no I/O) and honest about it; the new/_/readyactually probes the storage backend (bounded by a 3s timeout) and the config DB, returning503when not ready. Point your load balancer's readiness check at/_/ready. - Multi-instance / HA contract is documented.
CLAUDE.mdnow states plainly which planes are shared across instances (IAM/OAuth via the S3-synced DB, background-job leases) versus instance-local (sessions, multipart uploads, metadata cache, rate limiter), plus the hard prerequisites for running behind a load balancer. See alsodocs/plan/architecture-ha-audit-2026-06-28.md.
Fixed#
- Multi-instance config sync no longer clobbers in-flight job state. The encrypted config DB sync replaced the whole SQLCipher file on every download — wiping this node's per-node coordination state (replication / lifecycle / maintenance job cursors and leases, the event outbox, listener cursors) along with the IAM tables it was meant to share. The sync now merges only the IAM tables out of a peer's DB, leaving coordination state intact, so background jobs and event-driven replication stay consistent across a synced fleet. A schema-version guard skips the merge during a rolling upgrade rather than risk a column mismatch, and the sync ETag only advances on a successful merge (a failed merge retries on the next poll).
CompleteMultipartUploadon the wrong instance now explains itself. Multipart upload state is per-instance; behind a non-sticky load balancer a part/complete landing on a different node than the create returned a bareNoSuchUpload. The error message now names the per-instance cause and points at sticky sessions (the S3 error code is unchanged for client compatibility).- Startup banner no longer disagrees with the runtime on proxy-header trust.
The banner parsed
DGP_TRUST_PROXY_HEADERSas a narrower set than the runtime (yes/onlogged "untrusted" while the runtime trusted them); both now use the same parser.
Changed#
- The storage backend's streaming memory bound is now enforced by the type
system.
get_passthrough_stream/get_passthrough_stream_rangehad buffering fallbacks that a backend could silently inherit (defeating the memory bound for large ranged GETs); they are now required methods. No behavior change for the built-in filesystem and S3 backends. - Internal hardening + test-determinism: typed config-apply pipeline shared
by the backup-restore path, replication run/event version counters that replace
sleep-based test waits,
parse_copy_rangeunit + property tests, and a sharedRunLeasetype. The Jobs admin tables also render consistently across mobile and desktop.
v1.7.02026-06-28#
Fixed#
- Batch uploads under one presigned POST signature no longer 403. A
presigned browser/CI form-POST policy with
starts-with $keyis designed to upload many files under a single signature (e.g. a release's.zip+.sha512+.sha1). The replay guard treated the second and later files as a replay attack and intermittently returned403 SignatureDoesNotMatch— after the body had fully uploaded — whenever two files were signed in the same second. The guard now keys on the signature and the file fingerprint, so a batch of distinct files passes while an exact resend stays an idempotent overwrite. The policy's own conditions (starts-with $key,content-length-range,expiration) remain enforced on every request and are the real bound on what a captured signature can write.
Added#
- Streaming delta compression for unbounded object sizes (opt-in, dormant by
default). New code paths reconstruct (GET) and encode (PUT/POST/copy) delta
objects through bounded-memory spool files instead of buffering the whole
object in RAM — so delta dedup can scale past the previous in-memory ceiling.
This is off by default:
DGP_SPOOL_THRESHOLD_BYTESdefaults tomax_object_size, so the streaming paths are unreachable until an operator lowers the threshold. New tunables:DGP_SPOOL_DIR,DGP_SPOOL_MAX_BYTES,DGP_SPOOL_THRESHOLD_BYTES,DGP_SPOOL_ACQUIRE_TIMEOUT_SECS,DGP_CODEC_STALL_SECS,DGP_CODEC_ABSOLUTE_SECS. The xdelta3 codec gained a stall-based watchdog and streaming entry points; the storage backends gainedget_reference_to_file/put_reference_from_file(filesystem hardlinks, S3 streams). Seedocs/plan/streaming-delta-any-size.md.
v1.6.02026-06-27#
Added#
-
Instant per-bucket size — no more O(n) listing sweeps. S3 has no protocol call for "how big is this bucket?" — the only primitive is an O(n)
ListObjectsV2sweep, and the old stats scan capped out at 1000 objects per bucket (slow, and simply wrong for large buckets). The proxy now keeps a running per-bucket counter (object count + logical pre-delta bytes + stored bytes), updated on every write and delete, the way Ceph or Backblaze B2 surface an instant number./_/statsreads it in O(1) — the 1000-object cap andtruncatedflag are gone — and a bucket size chip in the browser top bar shows size · object count at a glance (admin session only). The counter is per-instance and best-effort (it never blocks or fails an S3 request); a Refresh button runs an uncapped full reconciling scan on demand.GET /_/api/admin/usage/bucket/:bucketexposes the O(1) read;POST /_/api/admin/usage/refreshtriggers the reconcile. -
Save-time config advisories. The admin Apply dialog (and
config lint) now surface cross-field "this combination is suspicious" warnings before you save, instead of leaving you to discover the misconfiguration in production. Seed rules catch a rate limit running with proxy-header trust disabled (which collapses every client onto the proxy's own IP and one shared rate-limit bucket), a stale${username}IAM permission template that silently denies access, a frozen bucket quota, and a public-prefix rule that is redundant when auth is already open. -
In-GUI operational logs — live tail + filter, no SSH required. A new System logs view in the admin UI tails the proxy's operational log stream live (server-sent events) with a Follow toggle, and filters a bounded in-memory backlog by level, target, and free-text search — so diagnosing an incident no longer means SSH-ing in to grep stdout. Backed by
GET /_/api/admin/logs(filtered backlog) andGET /_/api/admin/logs/stream(live SSE), both admin-session-gated. The ring is per-instance and in-memory (size viaDGP_LOG_RING_SIZE, default 2000; floor viaDGP_LOG_RING_LEVEL, default INFO) — it supplements, and does not replace, shipping the full stdout stream to your log aggregator. -
DGP_LOG_FORMAT=jsonstructured logs. Opt-in JSON log output so prod logs are field-greppable withjq(e.g. by client IP, bucket, or action) instead of being parsed out of free-text lines.
Changed#
-
Slimmer admin sidebar. The settings sidebar was consolidated from eight groups down to five — single- and two-leaf groups that cost more in headers than they earned are folded together (Dashboard, Trace, Delta efficiency, Audit log, and System logs now sit under one Observability group; the single Jobs screen moves under Storage). Every page URL is unchanged.
-
More readable audit & auth logs. Audit lines now resolve the client IP the same proxy-aware way the rate limiter does (no more
ip=unknownsplitting one client across two values), and brute-force / lockout log lines now include the rate-limit bucket key and proxy-trust state — the fields that make "all clients collapsed onto one bucket" obvious at a glance.
Fixed#
- Delta uploads work on newer xdelta3 builds (3.1+). A newer xdelta3
enables stream "armor" by default, which requires a seekable target the
proxy doesn't provide when piping — making every delta-eligible
PUTfail witharmor requires a seekable targeton those builds. The codec now probes its xdelta3 at startup and disables armor when supported, while staying compatible with the older 3.0.x builds that lack the flag. Deltas remain format-identical across versions (plain RFC-3284 VCDIFF), so a delta encoded on one xdelta3 version still decodes on any other. The exact xdelta3 version and armor state are now logged at boot.
v1.5.42026-06-27#
Changed#
- Jobs UI polish. Run/failure tables now show relative times ("3h ago") with the full timestamp on hover. The Runs table replaces the scanned/copied/skipped/errors number columns with a single proportional progress bar (green = copied, red = errors, blank = skipped/already-in-sync; the copied count is overlaid, full breakdown on hover) and a clock icon for scheduler-triggered runs — freeing horizontal space. Resuming or running a rule now immediately refreshes its runs + failures tables so the new run shows without reopening the drawer.
Fixed#
- Clearer "scanned vs processed" reporting. A run showing
scanned=600, processed=0previously hid the skipped count, making it look stalled when in fact all 600 objects were already in sync (nothing to copy). The skipped count is now surfaced (in the progress bar + on hover), so an incremental run's "nothing to do" outcome is transparent.
v1.5.32026-06-26#
Added#
- Delta-passthrough replication fast path. Replicating a delta-compressed
object between two compressed buckets no longer reconstructs the full object,
ships it whole, and re-compresses it at the destination. When the destination
already holds the byte-identical reference baseline (or has none yet — it's
seeded), the
.deltablob is shipped verbatim: no xdelta3 on either end, and only the delta's bytes cross the wire instead of the full logical object. For versioned-artifact mirrors this is a large egress + CPU saving. The fast path is gated hard for correctness — it only fires when the destination reference's checksum matches and both sides are plaintext, and always falls back to the proven reconstruct path otherwise. Each run reports how many objects took the fast path and the egress bytes saved (adeltaglider_replication_delta_passthrough_bytes_saved_totalmetric).
v1.5.22026-06-26#
Fixed#
-
Reject malformed
//keys at ingest. The proxy is the S3 gateway, but it used to store a client's path-join bug verbatim — keys with an empty path segment (a//b, distinct froma/bin S3). NowPUT, browserPOSTform-upload, and multipart uploads all refuse internal//keys. Reads and deletes stay permissive so any pre-existing//objects remain reachable for cleanup; a single trailing/(folder marker / listing prefix) is unaffected. -
Parity Verify survives transient listing throttle. A parity audit over a large bucket lists both sides page-by-page; a single Hetzner
503 SlowDownmid-scan used to fail the whole verify. The page-list now retries transient throttle errors with backoff, so a momentary 503 no longer aborts the audit. -
Live job Runs tab. A running job's Runs tab now updates its scanned/processed counters live, sourced from the jobs-list poll (no second poller) — and refetches run history once when the run finishes.
v1.5.12026-06-26#
Added#
-
Replication parity audit — "is my mirror verified identical?" A new Verify tab on each replication job runs an on-demand source↔destination parity check and returns an explicit verdict instead of inferring sync from
status=succeeded. It compares logical SHA-256 + size from metadata (no downloads — works for delta-stored objects too) and classifies every object Match / checksum-mismatch / missing-on-destination / extra-on-destination. The green "Verified in sync" verdict is the headline; a foreign-object three-tier verifier (sha → etag+size → size-only) avoids false alarms on objects not written through the proxy. -
Closed-loop remediation per finding. Every difference now explains its cause, says policy-awarely whether re-running the rule will fix it, and guides the operator to the right action. Crucially honest: a checksum mismatch under a
skip-if-dest-existsrule is reported as "re-run won't help — overwrite manually or change the policy", not a false "just re-run". A persistently failing copy surfaces its last error. "Run now" is offered only where it actually helps.
Fixed#
- Replication retries transient source 5xx. Transient Hetzner-phrased
throttle/gateway errors (
throttled (status=503),failed (status=502/504)) are now recognised as retryable, so small objects no longer land in the failure ring on a momentary source blip — they retry in-line and succeed within the same run.
v1.5.02026-06-25#
Added#
-
Gigantic-file replication: streaming multipart copy with rclone-class concurrency. Replicating a large passthrough object (e.g. a 30 GB backup tarball) used to buffer the whole object in memory twice and fail when the single monolithic GET dropped mid-stream. The transfer path now streams large passthrough objects source→destination via multipart upload with bounded memory (only
upload_concurrency × part_sizeresident, not the whole object), per-part range-resume (a dropped connection retries just that part, not the entire object), and concurrency on two axes — parts within an object and objects within a run:replication: transfers: 4 # objects copied concurrently per run upload_concurrency: 4 # multipart parts in flight per objectDelta objects still reconstruct transparently; proxy-side AES-encrypted backends fall back to the buffered path (native multipart for them is deferred). New
max_passthrough_object_size(default 64 GiB) decouples the streaming ceiling from the delta-reconstruction limit. -
Replication run resilience. A single slow or poison object can no longer kill a whole run. Lease defaults raised (TTL 60s→300s, heartbeat 20s→60s) so a multi-minute copy can't lapse the run lease; new
replication.object_timeout(default30m) bounds a stalled copy;replication.object_skip_after_failures(default 5) skips an object that fails N consecutive runs so it stops blocking the queue head (stays visible in the job's failures). -
Replication observability. New Prometheus metrics expose the streaming characteristics for Grafana:
deltaglider_replication_part_bytes_resident(+_peak),_parts_inflight(+_peak),_objects_inflight(+_peak),_multipart_parts_total,_part_retries_total,_bytes_streamed_total.
v1.4.32026-06-18#
Added#
-
Count-based lifecycle retention (
retain-newest). A new lifecycle action that keeps the newest N objects in a prefix and deletes the rest — selection by count, not age — for the canonical "keep the last N backups" cleanup that native S3 lifecycle never shipped:action: type: retain-newest count: 2 qualify: # eligibility filter (NOT a delete guard) min_size_bytes: 1048576 # ignore truncated/empty junk min_age: "1h" # ignore in-flight uploads protect_younger_than: "7d" # optional delete-side guardqualifyis an eligibility gate — an object failing it is invisible (never counted toward N, never deleted) — so an accidental empty or half-written file can't anchor the keep set or displace a real backup.protect_younger_thanis a separate delete-side guard (spares an object this run, never promotes it into the keep set). "Keep newest N" is set-relative, so it runs on a dedicated collect→rank→act worker path; the age-based path stays per-object/streaming and unchanged. -
Savings dashboard reads as a before/after bill. The metrics hero dollar line now tells the bill story directly: the regular monthly bill (what you'd pay without DeltaGlider) struck through, the compressed bill as the green headline (what you actually pay), and the saving kept quiet underneath ($/mo · $/yr). Previously the green headline showed the saved amount, which read like "what you pay" when it was the opposite.
Fixed#
- Browser form-POST upload returned 403 on an idempotent retry. The
presigned form-POST replay cache was keyed only on the SigV4 signature
and rejected any second request bearing an already-seen signature for
up to 24h (the policy-expiry-capped TTL). A presigned POST's signature
is deterministic for a given policy + second-resolution
x-amz-date, so a CI retry or pipeline re-run of the same artifact re-sent the identical signed request and got a spurious403 SignatureDoesNotMatch(observed uploading the deterministic.sha1/.sha512siblings right after a.zip). The replay entry now also fingerprints the(key, body)the signature first wrote: re-sending the same object is an idempotent retry and is allowed, while reusing a captured signature to write a different key or body is still blocked (form-POSTkeyisstarts-with "", so one signature could otherwise authorise writing anywhere). Replay protection is preserved; legitimate retries succeed.
Docs & site#
- Documentation site overhaul. A Warp-style collapsible sidebar tree
with a calmer editorial type scale and focus-on-current-section
navigation; clickable heading anchors; a readable search input with a
correctly-centred Clear button that closes on result selection; the
/docslanding reworked into 3-up feature cards with readable descriptors; and an accessibility/readability pass (type scale, code panels, link accent, inline-code chip sizing). - Accurate, more complete docs content. Acted on an expert docs-site
review (search, intro, S3-compatibility, FAQ) and onboarding/
architecture/versioning gaps; corrected the
authentication: nonetutorial; embedded all product-doc screenshots (not a stale subset); and replaced the ASCII data-path diagram with a rendered image. - Marketing site polish. Split-bleed hero with an upbeat captioned demo video, a lighter page background, and a branded glow on the demo video frame.
v1.4.22026-06-13#
Added#
- Inline video & audio preview in the object browser. Double-clicking
(or the inspector Preview button on) an
mp4/webm/mov/m4v/ogvobject now plays it in a native<video>player, andmp3/wav/ogg/m4a/aac/flac/opusin an<audio>player — streamed from a presigned URL so the proxy's range support (206 Partial Content) handles seeking without buffering through the browser tab. A codec the browser can't decode falls back to the existing Download affordance. Previously these fell through to "Preview not available".
v1.4.12026-06-12#
Removed (breaking)#
- TOML config support removed entirely. YAML is the only config
format. Gone: TOML loading (
Config::from_toml_file/from_toml_strand the.tomlboot path), TOML persisting, theconfig migratesubcommand, the--show-tomlflag, thedeltaglider_proxy.toml.examplefile, and theDGP_SILENCE_TOML_DEPRECATIONenv var. A.tomlconfig — pointed at viaDGP_CONFIG/--configor found on the default search path — now fails startup loudly with: "TOML configs are no longer supported (removed in v1.4.1). Convert withdeltaglider_proxy config migrateon v1.4.0, then point the server at the YAML file." Upgrade path: runconfig migrateon v1.4.0 (the last release that ships it) BEFORE upgrading.
Added#
${env:NAME}references round-trip. The proxy now records which config values were resolved from${env:...}references and re-emits the references — instead of materialized secrets — when persisting the config to disk (GUI changes) and inGET /config/export. Redactors keep references intact (a reference is not a secret). The admin/config/applydocument path now also expands${env:NAME}against the server environment. Together these make the IaC loop lossless: provision a secret-free template, tweak via the GUI, export, and commit the export straight back into IaC.
Fixed#
- Apply dialog showed "No changes detected" for credential
rotations. The section diff redacted secrets on both sides before
comparing, so rotating a secret access key, encryption key, OAuth
client secret, or webhook header produced an empty diff. Secrets are
now projected to deterministic one-way fingerprints (
fp:xxxxxxxx) for the comparison: unchanged secrets compare equal, rotations surface as a readable fingerprint swap, and no key material appears in the dialog.${env:NAME}references stay readable verbatim.
v1.4.02026-06-12#
Added#
- One Jobs surface. Replication rules, lifecycle rules, and one-off
maintenance jobs (bucket migration, re-encryption) now live on a single
admin screen backed by a unified API (
GET /jobs, per-kind pause / resume / run-now / preview / cancel,POST /jobs/reencrypt,POST /buckets/:bucket/migrate). The old per-subsystem/replication*,/lifecycle*,/maintenance*admin routes are gone. - Bucket migration between backends as a resumable background job
(stage → copy → verify → flip → cleanup) with a per-bucket write gate:
writes to a bucket under maintenance get
503 Slow Down, reads pass. Re-encrypt jobs use the same machinery, survive restarts, and never resurrect a peer instance's live job after an IAM-DB sync. - Redesigned Analytics dashboard. Percent-led savings hero ("270% smaller") with a before/after proof bar, one unified bucket list (ratio + footprint per bucket, per-row scan), an "est. $ left on the table" panel for compression-off buckets, a scan-coverage strip, and a designed empty state. The old gauge / facts-grid / duplicate top-buckets panels are gone.
- App-wide keyboard shortcuts: ⌘K command palette over the whole admin
IA, ⌘S apply-current-section,
?shortcuts help, arrow-key navigation in the object browser. - Multi-backend bucket management: create-a-bucket-on-a-named-backend admin API + UI, backend origin badges across the browser and admin.
- Documentation rewritten to Diátaxis — 3 executed tutorials, 25 goal-named how-to guides, 13 reference pages, 5 explanations; 14 new product screenshots; the marketing site renders the same corpus.
- Prod-config regression tests: a sanitized structure-true snapshot of
the production config validated in the CI gate (parse, declarative-IAM
reconciler validation, export round-trip) plus a local harness
(
scripts/test-prod-config.sh) that boots the current branch against the real prod backup with dynamically derived assertions.
Fixed#
- Creating a bucket on a non-default backend could land it on the wrong backend on case-mismatched names.
- Lifecycle crash-resume could replay a stale cursor after a same-named rule was redefined with a different scope (cursor is now scope-stamped).
- Maintenance requeue-on-boot is lease-aware — a synced config DB carrying a peer's live job is no longer resurrected locally.
- Phase machines (migrate / re-encrypt) now fail loudly when a pagination page budget truncates a phase instead of silently advancing.
- Admin bulk copy/move/delete now participates in the per-bucket write gate instead of bypassing it.
- Jobs table no longer renders job names one character per line in narrow columns.
v1.3.02026-06-05#
Changed (breaking) — IAM permission templates are now ${iam:...}#
- IAM identity templates in
resources/ string condition values are now${iam:username}and${iam:access_key_id}(previously bare${username}/${access_key_id}). Theiam:prefix makes them an explicit namespace, symmetric with the new${env:NAME}config expansion — every${ns:name}declares when it resolves (iam = request time, env = load time), and a bare${...}is now a literal. - Breaking: a bare
${username}no longer substitutes — it fails policy validation (user/group API + declarative apply). Update existing policies to the${iam:...}form. No back-compat shim.
Added — in-process ${env:...} config expansion#
- Config files now expand
${env:NAME}/${env:NAME:-default}against the environment when loaded, removing the need for an externalenvsubststep in deployments. Ship a secret-free config with${env:...}placeholders and inject the values as env vars; an unset reference (with no default) fails loudly at load instead of silently leaving a hole.- The
env:prefix is mandatory: it keeps env placeholders distinct from DGP's runtime IAM permission templates (${username},${access_key_id},${email},${filename}), which are left untouched for the auth layer. $$is a literal$escape (e.g. for a literal${...}in a comment).- Applies on the disk-load paths — server startup,
config lint, andconfig apply— but NOTconfig migrate(templates are preserved) nor the admin API's in-memory doc-apply (no surprise expansion against the server env).
- The
Added — Slack notifications connector#
- Object events can now post to Slack. Setting
event_delivery.format = slackrenders each outbox event as a Slack message (Block Kit + plain-text fallback) instead of the raw{schema,event}JSON envelope. Two delivery modes:- Incoming Webhook URL — point
webhook_url/webhook_urlsat ahooks.slack.comURL. Each URL is bound by Slack to a single channel; the cosmeticslack_username/slack_icon_emojioverrides apply here. - Bot token (
slack_bot_token,xoxb-…) — delivery uses the Slack Web API (chat.postMessage), posting toslack_channelwith support for@-mentions and any channel viachat:write.public.
- Incoming Webhook URL — point
- Per-bucket / per-prefix channel routing (bot-token mode only). When
slack_routesis non-empty, an eligible event posts to EVERY route it matches — so different buckets/prefixes can fan out to different channels (and one object can hit several).slack_notify_kinds(default["ObjectCreated"]) plusslack_include_globs/slack_exclude_globsare a global pre-filter for what's eligible; routes then pick the channels. - Editable from the admin UI at Configuration → Advanced → Webhook delivery, alongside the raw-webhook config. The bot token is a secret: it's masked in the GUI and in every export, and leaving the masked value untouched preserves the real token on save (both the section-PUT and full-document export→apply paths), exactly like webhook auth headers.
v1.2.02026-06-02#
Added — event-driven replication#
- Replication now reacts to object mutations in near-real time. Previously
replication was pure-poll: a timer did a full
list + HEADdiff every tick, scaling with bucket size rather than change rate. Object mutations (PUT / DELETE / COPY / CompleteMultipartUpload) are now appended to the durableevent_outboxand drained by a per-process consumer that fans each key out to the matching replication rules. The per-ruleintervalis repurposed as a slow (default 24h) full-reconcile safety net — events are the primary trigger.- Pub/sub via a per-listener cursor over the append-only outbox: webhook delivery and replication keep independent cursors and never contend.
- Per-key compaction collapses a burst of events for one key into a single liveness verdict; idempotency stays in one place (the planner + a dest HEAD), shared with reconcile — no new per-key sync table.
- At-least-once delivery: the cursor advances only to the highest contiguous fully-handled event id; reconcile backstops the rest.
- The webhook pruner clamps its delete floor to the smallest active listener cursor, so a stuck or disabled consumer can no longer pin the outbox and let it grow without bound.
Added — Webhook delivery GUI#
- Webhook (event-delivery) config is now editable from the admin UI at Configuration → Advanced → Webhook delivery — the last config surface that was YAML/env-only. Enable switch, endpoint list, auth-header editor, retry / retention / batching tuning, and a live delivery-status strip linking to the Event Outbox viewer.
- Header secrets are masked and preserved. Header values are shown masked in the GUI and in every export; leaving a masked value untouched preserves the real token on save (both the section-PUT and the full-document export→apply paths), while a removed header is deleted and a renamed-but-not-retyped secret is blocked rather than silently dropped.
Fixed — CI#
- Self-hosted runner "Too many open files". The CI integration job runs ~35
test binaries in parallel inside Ryzen LXC runners whose nested dockerd
defaulted to a 1024 nofile limit; the aggregate fd count overflowed it. Raised
the limit at the dockerd, workflow (
--ulimit), and test-harness (seed concurrency cap) layers.
v1.1.22026-06-02#
Fixed — hardening (from an adversarial Rust audit)#
- Form-POST uploads can no longer OOM the proxy. The
multipart/form-dataPOST interceptor buffered the request body with no ceiling (to_bytes(usize::MAX)); because it ran aboveDefaultBodyLimit— which only does an eagerContent-Lengthcheck — a chunked or oversized POST slipped past the limit. Now bounded bymax_object_size(oversized →413). This also fixes a real upload failure with aws-cli's chunked transfer encoding. - CopyObject on a concurrently-deleted source now returns
404, not500. The generic S3 error classifier didn't recogniseNoSuchKey, so a benign race surfaced as an internal error. - Form-POST signature length is no longer leaked via timing. The HMAC compare now rejects any signature that isn't exactly 64 hex characters before any length-dependent work.
v1.1.12026-06-01#
Changed — IAM permission editor#
- One-line grant editor. Each permission rule is now a single row —
WHERE(bucket / prefix) +CAN DO(the five atomic actions as a horizontal strip of on/off toggle-chips) + Conditions + Remove. The chips are an independent multi-select with a checkbox indicator (solid when on, dashed when off), so "write without delete" is expressible; the Admin chip auto-disables on prefix-scoped grants (a bucket-level op is meaningless on a sub-prefix), and a live plain-language caption summarises each grant. "Clear all" appears inline beside the actions when more than one is selected.
Fixed — IAM permission editor correctness#
- Resource rows no longer drop or re-key on blur. The on-blur normalizer re-split the whole comma-joined resource string, discarding in-progress empty rows and reassigning React keys to surviving rows; it now normalizes only the blurred row in place.
- Narrowing a grant's scope no longer silently loses the rule. Reducing a
resource from a bucket to a sub-prefix strips the now-invalid
adminaction, and any incomplete rule (missing actions or resource) is flagged inline ("…or this rule is dropped on save") instead of vanishing on Apply. - Invalid resource patterns are caught inline. Patterns the server rejects
(a
*that isn't trailing, embedded whitespace) now show an inline error before Apply, instead of an opaque HTTP 400.
v1.1.02026-05-31#
Added#
- Full IAM export/import as YAML. The admin GUI Account menu gains
"Export full IAM (YAML)" and "Import full IAM (YAML)", round-tripping the
entire IAM state (users, groups, OAuth/auth providers, group-mapping rules)
as declarative
access:-shaped YAML — distinct from the runtime-config YAML, which excludes the encrypted IAM DB. Export includes real secrets for a lossless round-trip (with a prominent live-credentials warning); import is a dry-run change preview → confirm → atomic single-transaction reconcile. Backed by new admin endpointsconfig/declarative-iam-{export,validate,apply}(mode-agnostic, admin-GUI-session gated).
Fixed — IAM permission editor (GUI)#
- Bucket-root listing ("" prefix) is now expressible. The
s3:prefixcondition editor round-tripped through a comma-joined string that silently dropped the empty-string entry, so "list from the bucket root" could not be saved. The editor now uses a string-array contract end-to-end with a dedicated "List bucket root (empty prefix)" toggle; the empty string is preserved through save and reload. - Prefix conditions no longer auto-append a trailing slash on blur. For an
s3:prefix StringLikecondition,ror/libsandror/libs/are not equivalent (the slash-less form also matchesror/libs-internal/…); the blur normalizer now preserves the operator's trailing-slash choice. - Resource rows keyed by stable ids (were array-indexed), removing a stale-key class of focus/row-confusion bugs when adding or deleting rows.
- Inline unknown-bucket warning. A resource pattern targeting a bucket that
doesn't exist (e.g.
ror/lib/*when the data is inbeshu/ror/libs/*) silently grants nothing; the rule editor now flags it inline with a near-miss suggestion.
Changed — bucket browser & dashboard#
- URL-routed bucket browser. Browser navigation (bucket, folder prefix) is reflected in the URL, so back/forward, deep-links, and reload now restore the exact view instead of resetting to the root.
- Dashboard scan state survives navigation. "Scan all buckets" results are server-side state the dashboard re-attaches to on mount (including in-flight scans), and cached results carry a derived staleness nudge after 6 hours rather than silently disappearing when you navigate away and back.
v1.0.1Hardening, cleanup & reliability2026-05-31#
A large quality release: the admin UI was hardened against a class of state-management bugs and then substantially refactored for less code and cleaner structure; the S3 bucket browser got the same treatment; and several correctness/reliability gaps were closed across the engine and S3 layer. No wire-format or breaking changes — drop-in over v1.0.0.
Fixed — data integrity & correctness#
- Same-location move no longer deletes data. The server bulk-move handler
did copy-then-delete-source for every item; a "move" whose destination
resolved to the same bucket+key as the source was a self-copy no-op, so the
unconditional source delete destroyed the only copy. A server-side guard
(
is_same_location_move) now skips the source delete for any such item, regardless of what the client sends. - Range-GET out-of-bounds panic on objects whose stored
file_sizemetadata exceeded the reconstructed byte length now returns400 InvalidRangeinstead of crashing the worker. - Doubled-quote ETag on passthrough HEAD/GET (
""abc"") fixed to a single quoted form; strict S3 clients rejected the doubled form. - Admin config editors: eliminated a class of row-drop / stale-closure / lost-edit / index-key bugs (prefix list, permissions, buckets, advanced, credentials, auth-rules, setup wizard); a failed Apply no longer discards in-flight edits.
- Bucket browser: inspector async-race (download/share landing in the wrong object), pagination desync on search, bulk-copy not clearing selection, destination-prefix join bug, 0-byte upload handling, and a perpetual "Loading compression stats…" spinner for passthrough objects.
Changed — performance & logging#
- LIST
metadata=trueskips the per-object HEAD for passthrough, non-delta files (checksum sidecars, images, …): they're stored verbatim so the LIST size is authoritative. Cuts the upstream HEAD burst on build-artifact listings (which was also a throttle/500 trigger). - Tamed the
PATHOLOGICALlog flood: missing DG metadata on a normal passthrough file is benign and now logs at DEBUG; the loud WARN is reserved for.delta/reference.binfiles where it actually breaks reconstruction. - Catch-all 500s now log the upstream cause before mapping, so production 500s are debuggable.
Changed — admin & browser UI refactor (no behavior change)#
- Admin UI clean-code + architecture pass: migrated Buckets/Lifecycle/
Replication onto the shared
useSectionEditor; extractedMasterDetailPanel,RuleListEditor, and a shared overlay-dropdown primitive; split the four largest panels into per-concern files; splitadminApi.tsinto a barrel over domain modules; ~1k LOC of duplication removed. - Bucket-browser clean-code + correctness pass: shared
getFileName/downloadBlobAsFile/pluralizehelpers,StorageTypeTag/FolderSizeCellextraction, s3client pure helpers.
Added — testing#
- Codec encode→reconstruct round-trip property test (the core data-integrity invariant), a criterion benchmark harness for the codec hot paths, and a panic-surface audit of the request hot paths. Plus ~15 new Node regression scripts covering the extracted pure UI helpers.
v1.0.0Project-shape milestone2026-05-22#
v1.0.0 marks the official single canonical implementation: one Rust
binary (deltaglider_proxy) ships the S3-compatible proxy, the
AWS-CLI-shaped s3 command group, and the web UI. The Python
deltaglider tool is
deprecated as of its parallel v6.2.0 release; PyPI installs continue
to work indefinitely but no further updates will land there. The wire
format remains byte-identical across both tools.
This release also closes the four post-v0.12.0 regression fixes
(HeadBucket 501, multipart x-amz-storage-type header drop, s3 ls
doctest, all from the v0.12.0 axum-handler retirement) plus the
small UX gaps the v1.0.0 plan flagged.
Changed#
s3 migratelearned--source-endpoint-urlfor cross-provider migrations (e.g. Hetzner → AWS). When unset, behaviour is identical to v0.12.0: a single engine handles both sides. When set, builds two engines (source + destination); credentials are shared across both. Per-side credential flags can land later if operators ask.s3 cp/s3 sync/s3 migratelearned--max-object-size-mbto override the engine's per-invocation 100 MiB defensive ceiling (memory-safe limit for xdelta3 — reference + delta + result all held in RAM simultaneously). Operators uploading large artifacts (release ZIPs, disk images) used to hit a "Object too large" error with no actionable hint; now there's a flag and the error text points at it.
Fixed#
- s3s: implement
HeadBucket(HEAD /<bucket>). The v0.12.0 axum-handler retirement deleted the legacy handler but the s3s trait impl never had one of its own — s3s's default returned 501. Real AWS / MinIO return404 NoSuchBucketwhen the bucket is missing; we now mirror that. Restoreserror_test::test_nosuchbucket_xml_responseandtest_entitytoolarge_response. - s3s multipart: emit
x-amz-storage-typeon CompleteMultipartUpload. The legacy axum handler set this header fromstore_result.metadata.storage_info.label(); the s3s impl dropped it during the v0.12.0 retirement. Restored to matchhead_object/get_object/put_objectparity, fixingtest_multipart_delta_compressionandtest_multipart_large_zip_forces_passthrough_on_s3_backend. - cli docs:
s3 lsshell example no longer breakscargo test --doc. The four-space indentation made rustdoc try to compile the example as Rust. Wrapped in ```text so it stays prose. - Stronger
--max-object-size-mberror path: DetectEngineError::TooLargeincp/sync/migrateand render an actionable message naming both the per-invocation flag and the server-sidemax_object_sizeYAML knob.
Deprecated (downstream — not in this repo)#
- The Python
deltagliderpackage is deprecated as of v6.2.0. Migration:brew install beshu-tech/tap/deltaglider_proxy(or download the binary), thenalias dg='deltaglider_proxy s3'. Every Python subcommand has a 1:1 Rust equivalent. The Python repo will be archived approximately one week after v6.2.0 hits PyPI.
v0.12.02026-05-19#
Removed (BREAKING)#
- Legacy axum-handler S3 adapter retired. The
s3scrate adapter has been the production S3 protocol path for several releases; the hand-rolled axum-handler implementation it was meant to replace is now deleted (~3500 LOC ofsrc/api/handlers/{object,bucket,multipart}.rsplus the supporting XML response builders, the parity test suite, andsrc/api/xml.rs). With it goes:- the
s3s-adapterCargo feature flag (now an unconditional dep), - the
DGP_S3_ADAPTERenv var (no more selector), - the
test-s3s-adapterCI job and the nightlytest-all-s3s-adapterjob (every test they ran is in the regular matrix already), - the
x-deltaglider-s3-adapter: s3sdiagnostic response header (only useful when both adapters coexisted), docs/dev/s3-adapter-decision.mdanddocs/plan/s3s-adapter-migration.md(superseded by this release). Operators previously rolling back viaDGP_S3_ADAPTER=axumhave no fallback path. If a behavior regression turns up, file a bug against the s3s adapter directly — there is no longer a second implementation to fall back on.
- the
Changed (BREAKING)#
- CLI client verbs moved under
s3subgroup. The 10 AWS-CLI-shaped client commands (cp,ls,rm,sync,migrate,stats,verify,purge,get-bucket-acl,put-bucket-acl) are now nested underdeltaglider_proxy s3 <verb>instead of being top-level. Top-leveldeltaglider_proxy --helpis now uncluttered, listing onlyconfig,admission, ands3as subcommand groups; future top-level proxy flags (init,set-bootstrap-password, etc.) no longer collide. Migration: replacedeltaglider_proxy cp ...withdeltaglider_proxy s3 cp ...(likewise for the other nine verbs). Operators preferring the Python-style invocation can aliasdg='deltaglider_proxy s3'. Library callers (deltaglider_proxy::cli::cp::runetc.) are unchanged — only the CLI surface moved.
Fixed#
-
Eliminated
unsafe { std::env::set_var("DGP_BACKEND_ALLOW_LOCAL", ...) }from CLI startup paths. The SSRF allow-local opt-in now flows through the typedBackendConfig::S3.allow_localfield instead of via process-env mutation. The legacyDGP_BACKEND_ALLOW_LOCALenv var is still honoured (backward-compat for existing deployments and the proxy's env-driven config layer), so no operator action is required. Eliminates a Rust-2024-unsafeblock from three CLI modules (engine_factory.rs,bucket_acl.rs,purge.rs) and makes the engine constructible without mutating process state. NewBackendConfig::S3.allow_local(defaults tofalse, serde-skipped when unset for clean YAML diffs) is the preferred path going forward; legacy YAMLs parse unchanged. Added regression testbuild_client_allows_dev_local_when_config_field_setpinning the typed-field path independent of env. -
CI workflow
DGP_BACKEND_ALLOW_LOCAL=truejob-level env (2fbcc5c + 577edd7 + e643ba3) — wave-1 security (3ff9edc) rejectedhttp://localhost:9000MinIO endpoints by default, breaking the[ci-integration]-gated Integration Tests batch (hidden on main pushes by the gate). Set the opt-in at job level inci.ymlandtest-all-nightly.yml; threaded throughinterop_test.rs'senv_clear()barrier; updated 4 cli_s3 test fixtures that forgot to create the source bucket before seeding. -
DGP_REPLAY_WINDOW_SECS=0in CI (577edd7) to disable the wave-3 SigV4 replay cache for tests that legitimately re-issue identical signed requests inside the 2s window. Production keeps the cache on; the two auth tests validating the wave-3 replay contract override the env back to2viaTestServerBuilder::env. -
DGP_USAGE_CACHE_TTL_SECSmade env-configurable (2c5ea89) so quota tests can shorten the 5-minute scan-cache TTL. Production default unchanged. Fixed two deterministic-race test failures inquota_test.rswhere the spawned scan finished BEFORE the seed PUT body landed on disk, caching a "0 bytes" stale read for 5 minutes. -
CI builder image now ships
clang+lld(95cdeac) — the compile-time-optimization PR addedlinker = "clang"+-fuse-ld=lldto.cargo/config.tomlbut the GHCR builder image only hadgcc. Cascaded every build-gated job intoerror: linker 'clang' not found. Two-lineapt-get installaddition plus verification step (clang --version,ld.lld --version).
v0.11.02026-05-16#
Added#
-
Bucket-wide scan endpoint with persistent on-disk cache, backing the dashboard's headline totals. Replaces the old
/_/statspath that capped at 1,000 objects per scan — useless on any real bucket (a 1.4 TB Hetzner bucket with 79k objects reported "1,000+" and guessed savings off the first kilobyte). NewGET /_/api/admin/diagnostics/scan/status[?bucket=X]returns the cached map (or single bucket);POST …/scan/startkicks a paginated walk;POST …/scan/stopcancels;DELETE …/scanforgets a cached result;GET …/scan/stream?bucket=Xopens an SSE feed of per-page progress (objects,original_bytes,stored_bytes,pages_done,has_more). Results persist to.deltaglider_scans/<bucket>.jsonwith a schema-version field — results survive proxy restart, no TTL, regenerate on deserialise failure. Per-bucket scan handle backed bytokio::sync::watch+CancellationToken; the scan continues to completion even after every SSE subscriber disconnects (useful for long PB-scale walks left running overnight). 6 unit tests cover persistence, corrupted- file recovery, version-mismatch rejection, and bucket-name sanitisation against..//traversal. -
"Money shot" Analytics dashboard redesign. The four flat KPI cards (Total storage · Space saved · Savings % · Est. monthly savings) collapse into one 12-col
HeroSavingsPanelshowing storage saved at billboard scale with a count-up animation on first visit per session (sessionStoragegate, honoursprefers-reduced-motion), a scale-accurate kept-vs-saved bar echoing the hero ratio, a dollar savings line with cost-rate preset popover, and a cache-age footer line ("Newest scan 47 s ago · Oldest 3 h 12 m ago · 2 buckets excluded"). Single new dependency:motion@^12for the orchestrated count-up + bar-grow animation (~5 KB initial, ~30 KB lazy). Below the hero, the old "Storage by bucket" recharts stacked bar — which collapsed to a single line when one bucket dwarfed the rest — is replaced by a Bucket fleet view where every bucket gets a full-width ratio bar plus a thin scale-relative footprint bar underneath, so even tiny buckets are legible. The dead "Scan progress" + "Compression opportunities" panels (empty 90% of the time) become a By the numbers insights grid (largest bucket / biggest single saving / best & worst ratio / total objects / avg ratio) and a Compression effectiveness radial gauge with explicit tier-breakdown (Excellent ≥ 50 %, Good 20–49 %, Low < 20 %, None 0 %, N/A < 2 objects). The gauge value is bytes-weighted — a 1.4 TB bucket at 93 % dominates a 88 KB bucket at 0 % — so the number matches the operator's intuition of "what's my storage bill doing", not the naive unweighted mean that penalises trivial buckets. Per-bucket Scan / Re-scan / Stop affordances surface directly in the Top buckets table; a backend chip identifies which named backend each bucket lives on; the table is sortable by original size / bytes saved / ratio / object count / most recent scan, with the operator's choice persisted tolocalStorage. -
Delta-efficiency per-prefix "verify" deep dive. New
POST /_/api/admin/diagnostics/delta-efficiency/verifydoes the HEAD-based scan path (vs the cheap lite scan that powers the bucket-wide overview) for a single prefix the operator opted into, returning true per-file savings + a sortedper_deltaarray suitable for percentile / distribution rendering. Cost: one prefix-scoped LIST + one HEAD per delta, bounded-parallel via the backend'sbounded_head_calls. Frontend exposes this as a "Verify savings" row affordance inDeltaEfficiencyPanel— the lite scan remains the default so the bucket overview stays fast, and the verify path is only run when the operator wants the exact numbers for one prefix. 5 unit tests pin the response shape and the sorting invariant.
Changed#
- Monitoring-tab refresh controls hidden on Analytics. The
cadence selector (
Off / 5s / 30s), time-range buttons (5m / 15m / 1h), manual-refresh icon, and "live" pulse dot were Monitoring-only chrome but rendered on every tab. They suggested Analytics also polled on a cadence — it doesn't, it backs onto the persistent disk-cached scan engine.DashboardToolbarnow gates that row onview === 'monitoring'.
Added#
- Delta efficiency diagnostics panel at
/_/admin/diagnostics/delta-efficiency. Scans every deltaspace in the chosen bucket and surfaces prefixes whose reference baseline is producing too-large deltas — the v0.9.17 incident shape wheres3://beshu/ror/builds/1.70.0-pre5/had 22 GB of effectively- uncompressed deltas because the first uploaded file (a Kibana ZIP in store-mode) became the reference for 30+ unrelated ES plugins (deflate-mode). Per-prefix verdicts: Excellent (median delta ≤ 200 KB AND ≤ 5 % of reference), Good (median ≤ 1 MB OR ≤ 20 % of reference), Fair (20–50 %; structurally bounded by multi-variant prefix mixing), Poor (≥ 50 %; wrong reference, re-upload), NoReference (deltas exist but baseline is missing — anomalous). Read-only surface; classifications are advisory and the operator decides what to re-upload. Pure-function core (classify_deltaspace) with truth-table unit tests covering each prod scenario from the audit; serde-stable JSON contract pinned by test. Background scan + cache + dedup mirroringUsageScanner: GET/_/api/admin/diagnostics/delta-efficiencyreturns 200 from a fresh 5-min cache or 202 + enqueued background scan (with panic-safe RAII dedup-key cleanup). POST…/scanforces a re-scan. The frontend panel polls every 2 s on 202 and shows a "cached" tag when the result came from cache so the operator can tell stale from fresh.
Documented#
- Reverse-proxy read-timeout requirement for large uploads.
Symptom: large objects (>~50 MB) fail with
502 Bad Gatewaysurfacing as "Upload failed" in the embedded UI. Root cause: Traefik 3.x defaultsrespondingTimeouts.readTimeoutto 60 s, killing the body-read mid-upload. The DeltaGlider Proxy itself then logs a400 BadRequestbecause hyper's body channel is closed beneath axum'sBytesextractor (BytesRejection:: FailedToBufferBody::UnknownBodyError, statically defined as#[status = BAD_REQUEST]in axum-core). Reproduced locally: default Traefik 3.6 in front of the proxy → 80 MB at 1 MB/s → Traefik returns504 Gateway Timeoutat exactly 60.7 s. Fix: setentryPoints.<name>.transport.respondingTimeouts.readTimeoutto30m(or0for no limit). For Coolify users: edit/data/coolify/proxy/docker-compose.yml, add the corresponding CLI flag to thetraefikservicecommand:block, restart withdocker compose -f /data/coolify/proxy/docker-compose.yml up -d. Cross-reverse-proxy reference table now indocs/product/20-production-deployment.md(Traefik, Caddy, nginx, AWS ALB, HAProxy).
Mitigations#
- Embedded uploader: cap concurrent files at 1. Pre-fix, the
upload queue ran
floor(DEFAULT_UPLOAD_QUEUE_SIZE / 2) = 2files in parallel, each driving 4 in-flight 16 MB UploadPart requests via@aws-sdk/lib-storage. Two large files queued together produced 8 × 16 MB = 128 MB of concurrent body data through the reverse proxy, with each part's slice of the upload pipe spread across more streams — making it more likely an individual part exceeds the 60 s default Traefik read-timeout. Solo uploads of equivalent or larger files succeeded — proven by a 160 MB upload that completed all 10 parts at ~43 s each in the same window where 294 MB and 358 MB files queued in parallel failed every part at exactly 60 000 ms (Coolify logs, 2026-05-08 ~20:34 UTC). SettingmaxConcurrentFiles = 1keeps the per-file 4-part queue alone governing ingress pressure, so individual parts get fair allocation of the upload pipe and complete inside the 60 s default window even without the operator-side timeout fix. Belt-and-braces with the docs change above.
Fixed#
- Form-POST upload rejected
acl=private. Browser presigned-POST builders (boto3, minio-js, aws-amplify) auto-includeacl=privatewhen the user constructs a presigned-POST without explicitly opting out. Pre-fix, the proxy returned501 NotImplemented: POST form upload ACL overrides are not supported. Now the field is silently accepted when the value is compatible with the proxy's owner-only default (private,bucket-owner-full-control,bucket-owner-read) — same semantics asx-amz-acl: privateon regularPUT objectoperations (header silently ignored). Public-grant variants (public-read,public-read-write,authenticated-read) still return 501 because silently accepting them would lie about object visibility. Same logic applies toaclpolicy conditions inside the signed POST policy. Newis_compatible_canned_aclpure function with truth-table unit tests; integration tests cover both the accept and the reject path.
v0.9.162026-05-07#
Correctness x-ray fix programme#
A focused audit by five specialised investigators (concurrency, auth, storage, HTTP/wire, error handling) surfaced ten distinct correctness bugs ranging from durable on-disk drift to silent panics on attacker-reachable input. All ten are now fixed; tests pin each contract.
- C-P0-1 / C-P1-1: DeleteBucket race with in-flight CompleteMultipartUpload.
purge_uploads_for_bucketno longer removes uploads inCompletingstate (returnsErr(count)for the operator to retry); filesystemensure_dirwalks bucket-relative paths with non-recursivemkdirsocreate_dir_allcan no longer silently resurrect a just-deleted bucket. Closes the data-resurrection class entirely. - E-P0-1: OAuth callback panic from
?next=header injection. Pre-fix, control bytes (CR/LF/NUL/0x01–0x1F/0x7F–0xFF) survived thenextvalidator and crashed the OAuth callback when handed toResponse::builder().header(LOCATION, ...).unwrap(). Now rejected up-front; defence-in-depth on the response build avoids panic even if the validator is later weakened. - S-P1-1: Anchor poisoning. Removed the
!has_existing_referencegate from themax_delta_ratiocheck. A dissimilar follow-up to a tiny anchor now stores as passthrough; the reference is preserved for delta siblings. - S-P1-2: Orphan reference on encode failure. When
set_reference_baselinesucceeds andencode_and_storesubsequently fails (codec saturated, codec panic, size cap), the freshly-minted reference is rolled back. Pre-fix it stayed durably on disk and poisoned every later PUT to the prefix. - S-P1-3: Filesystem fallback ETag. Unmanaged files (no DG
xattr) now get a synthetic hex-32 ETag derived from
(size, mtime)instead of the literal"". Empty files use the canonicald41d8cd9.... Compare-and-swap clients now work after rsync/tar migrations that strip xattrs. - H-P1-1: RFC 7232 §6 conditional precedence.
If-Matchnow suppressesIf-Unmodified-Since(and theIf-None-Match/If-Modified-Sincepair) on GET/HEAD per spec. Pre-fix, requests combining both headers got spurious 412s. - H-P1-3: Multi-range header. Pinned the existing fall-back-to- full-GET behaviour with regression tests so the contract doesn't regress.
- E-P1-1: Backend 503 SlowDown surfaces as 503, not 500. New
StorageError::Throttledvariant routes through toS3Error::SlowDown, restoring the AWS-SDK retry/backoff contract. - E-P1-2: Usage scanner panic leak. RAII guard ensures the
dedup key is removed from
scanningeven on panic unwind. Pre- fix, any panic indo_scanpermanently disabled scanning of that prefix until process restart.
A-P1-1 (form-POST $key policy variable parity drift) is documented
in source but kept at current semantics pending behavioural
verification against real AWS S3.
cargo test --lib 769 (was 746); +23 regression tests across nine
test modules. Clippy -D warnings clean on default features and
--features s3s-adapter. No env-var, schema, or wire-protocol
changes; backwards compatible.
v0.9.152026-05-07#
CI and compatibility follow-up#
- Stabilized
s3scompatibility runs by skipping form POST compatibility assertions on the experimentals3sadapter path (feature remains validated on the primary adapter). - Hardened full-flow Playwright smoke by asserting uploaded object visibility in the browse view, avoiding a flaky dependency on transient upload-queue status labels.
v0.9.142026-05-07#
Multipart complete memory bound fix#
- Fixed the CI-blocking
memory_testregression by changing MPU relay policy to threshold-triggered promotion instead of always-relay for passthrough uploads. - This preserves bounded-memory completes for normal passthrough MPU sizes while still enabling disk relay for large payloads.
- Added a relay-part passthrough path (
RelayedParts) so complete can hand ordered part files to storage without forcing a monolithic assembled temp object. - Filesystem storage now supports writing passthrough objects directly from ordered
relay part files (
put_passthrough_parts) in a single atomic temp+rename flow.
v0.9.132026-05-07#
Multipart passthrough relay and cleanup hardening#
- CompleteMultipartUpload now supports relay-backed passthrough completes for large payloads, avoiding monolithic in-memory assembly when delta reconstruction would be too expensive.
- Added passthrough file-store path in engine/storage (filesystem + S3) so relayed multipart payloads stream from disk into backend writes while preserving multipart ETag semantics.
- Multipart sweeper now reclaims stuck
Completinguploads after a configurable timeout and removes orphan relay artifacts on startup + periodic sweeps. - Added multipart sweep metrics and env knobs:
DGP_MULTIPART_SWEEP_INTERVAL_SECS,DGP_MULTIPART_SWEEP_MAX_AGE_SECS,DGP_MULTIPART_COMPLETING_TIMEOUT_SECS,DGP_MPU_DELTA_RECONSTRUCT_MAX_BYTES.
S3 form POST compatibility#
- Added minimal SigV4 policy-based
multipart/form-dataobject POST support on bucket endpoints forcreate_presigned_postflows. - Implemented safe constraints with explicit
NotImplementederrors for unsupported form features (e.g. ACL overrides,success_action_*, session-token policy forms). - Added compatibility tests for successful presigned form POST upload and unsupported form-option rejection behavior.
Coverage for relay and reclaim behavior#
- Added S3-backend integration coverage for large multipart passthrough completes.
- Added startup orphan relay cleanup integration coverage.
- Added completing-timeout reclaim integration coverage to verify stuck uploads are reclaimed and removed.
v0.9.122026-05-07#
Upload reliability and UX#
- Browser uploads now use managed self-signed multipart upload (
@aws-sdk/lib-storage) with bounded concurrency and per-part retries, replacing single-request large PUTs. - UI error normalization now maps gateway/plaintext failures during long uploads into explicit 502/503/504 guidance instead of parser/deserialization noise.
Multi-backend bucket creation#
- Added admin API support to create a bucket pinned to a selected backend
(
POST /_/api/admin/buckets) and wired the browser create-bucket flow to use it. - Create-bucket backend selector in the sidebar now uses
SimpleSelect(no Ant popup dependency), matching the rest of the embedded UI dropdown policy.
v0.9.112026-05-06#
Browser UI and admin#
- Open-access mode: admin bootstrap login now seeds anonymous S3 session credentials so a hard refresh does not strand the embedded file browser without a working SDK client.
- Connect / session flow refinements (capabilities, reconnect, config-DB mismatch handling).
- Replaced the browser-lift banner with a lighter session tip; small copy and layout tweaks across browse chrome, inspector, and bulk actions.
- Admin navigation: Recovery is renamed to Backup (route unchanged:
/_/admin/configuration/recovery).
Testing#
- Added Playwright full-flow E2E (bucket create, upload, admin login, sign-out, reconnect,
object still visible);
e2e-smoke.shruns the fulle2e/suite.
v0.9.82026-05-05#
Admin IAM and storage-path authoring#
- Admin UI: grouped autocomplete for IAM resource patterns and LIST condition prefixes (listed prefixes vs per-user placeholders), with clearer section layout and plain-language help.
- Admin UI: access-section YAML modal explains empty
access: {}responses when IAM state lives in the encrypted config database. storagePathhelpers, prefix listing for suggestions, and a small Node regression script for path formatting rules.
IAM policy variables (proxy)#
- Policy variables such as
${username}and${access_key_id}in resource patterns ands3:prefixconditions, with safe cloning (no glob corruption). - Admin read APIs and config DB plumbing for extended permission rows used by the UI.
- Documentation updates for SigV4, IAM conditions, and declarative IAM.
Reliability#
- Test harness unsets
DGP_BOOTSTRAP_PASSWORD_HASHandDGP_ADMIN_PASSWORD_HASHin the child process so developer shells cannot override the test config and break admin login tests.
v0.9.72026-05-04#
Lifecycle policies, event outbox, and replication operations#
- Added lifecycle v2 support for expiring objects, transitioning objects between storage classes/backends, previewing planned actions, and surfacing lifecycle state through the admin UI/API.
- Added a durable object-mutation event outbox with delivery/requeue flows, diagnostics, internal-only failed-count handling, and admin UI controls for inspecting and retrying stuck events.
- Added scheduled replication coordination with leases/progress tracking so multi-instance deployments can run replication without duplicate workers.
- Shipped Helm chart and Kubernetes deployment docs for running DeltaGlider Proxy with persistent config, ingress, and operational defaults.
Routing backend and admin UI polish#
- Fixed routed-bucket listing semantics so virtual buckets preserve their backend origin and duplicate real bucket names follow request-routing priority: explicit route, default backend, then stable backend order.
- Added an admin bucket-origin endpoint and browser sidebar backend badges for local, Hetzner, AWS, and generic S3 backends. The icons appear beside bucket names in the left bucket sidebar after login.
- Tightened admin/browser spacing in section headers, tab headers, and section overview pages so dense configuration pages waste less vertical space.
v0.9.12026-04-30#
Marketing site + release docs refresh#
- Added a GitHub Pages marketing minisite under
marketing/with SSG, sitemap/robots generation, SEO smoke checks, strict screenshot validation, and a dedicatedmarketing-pagesworkflow. - Reworked product positioning around the current release stories: repeated binary storage, regulated use of cheap/untrusted S3-compatible backends via proxy-side encryption, lower-cost S3 SaaS with a portable enterprise control plane, and MinIO migrations where storage is available but IAM/policy/operations are missing.
- Added marketing pages for About, Privacy, Terms, artifact storage, regulated workloads, cheaper S3 SaaS control-plane use, and MinIO migration.
- Refreshed product screenshots and added required screenshot checks for
filebrowser,analytics,iam,advanced_security,bucket-policies, andobject-replication. - Updated release-facing markdown to reflect shipped soft bucket quotas, object replication with delete replication, and the current proxy-side encryption story.
Fourth-wave correctness fixes (SigV4 integrity + replication provenance + stubs)#
Eight findings from the latest review. Two highs are silent correctness/security failures; the rest tighten honesty in stub handlers and replication-control endpoints.
H1 — SigV4 didn't verify the signed payload hash against the body
A credentialed client could sign hash A and ship body B; the
signature was computed over the canonical-request which only sees
the header value. SigV4's integrity contract requires the receiver
to verify the body downstream — we documented that as a future
guarantee but never implemented it. Fix: the auth middleware
stashes the signed x-amz-content-sha256 value in a
SignedPayloadHash request extension; put_object_inner and
upload_part recompute the body's SHA-256 and constant-time-compare.
Mismatch → 400 BadDigest. UNSIGNED-PAYLOAD and STREAMING-* sentinels
are recognised and skip verification (they have their own contracts;
see aws_chunked.rs for the streaming variant we don't yet validate
end-to-end). Three integration tests cover signed-mismatch reject,
unsigned-payload pass-through, and the signed-match happy path.
H2 — replication delete pass deleted unrelated destination objects
Pre-fix, replicate_deletes=true listed every key under
destination.prefix, mapped each back to source via prefix-rewrite,
and deleted on source-NoSuchKey. With no provenance marker, an
operator-placed object or a sibling rule's data could be wiped.
Fix: every replicated object carries a dg-replication-rule = <name>
user-metadata key. The delete pass only considers candidates whose
metadata carries THIS rule's marker; HEAD-fallback is used when the
listing didn't surface user-metadata; HEAD failure preserves
(false-delete is much worse than a leftover). Integration test seeds
both replicated AND manual objects on destination, runs replication,
then deletes a source key; verifies the replicated key gets deleted
on the next run while the manual one survives.
M1 — pause/resume created ghost rule rows
Both handlers called replication_ensure_state BEFORE looking up
the rule in config, so a request for a non-existent rule inserted
a state row even though the response was 404. Fix: new
rule_in_config() helper runs first; ensure_state only fires
when the rule actually exists. Test verifies that 404 on a ghost
rule leaves the overview clean (no orphan row).
M2 — run-now bypassed enabled flags
replication.enabled=false and rule.enabled=false were honoured
by the (future) scheduler but admin-triggered run-now ignored
both. Fix: explicit gates that 409 with a descriptive message.
Two tests cover the global and per-rule cases.
M4 — ACL/versioning PUT stubs returned fake 200
PUT /bucket?acl, PUT /bucket?versioning, and PUT /bucket/key?acl
returned 200 OK while silently discarding the request body. Clients
believed grants/versioning had been applied. Fix: all three return
501 NotImplemented (preceded by a bucket/object existence check so
404 wins when the target is missing). Five integration tests cover
both 501-on-existing and 404-on-missing.
L1 — tagging PUT/DELETE returned 501 even on missing target The previous wave fixed GET ?tagging precedence; PUT/DELETE were still 501-first. Fix: each calls head/head_bucket before returning 501 so NoSuchKey/NoSuchBucket wins. Three tests cover PUT-on-missing, DELETE-on-missing, and DELETE-on-existing-object (still 501).
M3 verified intact — bucket-subresource existence checks (?location, ?versioning, ?uploads) added in the previous wave still in place.
L2 verified intact — size_is_known = file_size > 0 || !md5.is_empty()
discriminator in build_object_headers correctly emits
Content-Length: 0 for known-zero managed objects.
Tests: 7 new integration in replication_test.rs and s3_correctness_test.rs,
3 new auth integration tests in auth_integration_test.rs. Full
suite: 659 lib + 28 s3_correctness + 8 replication + 28 auth +
existing suites green. Clippy clean.
Third-wave correctness fixes (replication + S3 conditionals + headers)#
Nine findings from a follow-up review of the replication v1 commits plus a couple of latent bugs in PUT / bucket subresource paths.
H1 — replication only copied the first page
run_rule listed one page with batch_size, copied it, finished
succeeded, and never resumed. A bucket with 50k keys and
batch_size=100 was 99.8% unreplicated while the dashboard reported
success. Fix: full pagination loop. Per-page continuation_token
persisted via replication_set_continuation_token so a crash mid-
run resumes from the last batch. Cleared on a clean complete pass.
Bound MAX_PAGES_PER_RUN = 10_000 defends against pagination
loops in pathological backends. Integration test seeds 17 objects
with batch_size=5 (≥4 pages).
H2 — replicate_deletes was config-only, no implementation
The flag was validated, documented, and ignored. Fix: new
run_delete_pass runs after the forward copy when
rule.replicate_deletes. Paginates the destination prefix, HEADs
each key on source, deletes destination on NoSuchKey source.
Other source-HEAD errors preserve the destination key (false-
delete is worse than a leftover). Skips the delete pass when the
forward pass hit a fatal error (otherwise a transient list
failure could trigger a destination-wide wipe). Integration test
seeds an orphan and verifies it's deleted while legitimate keys
survive.
H3 — replication lost multipart-ETag identity
copy_one always called plain engine.store, so a destination
HEAD returned a fresh full-body MD5 instead of the source's
"abc-N" multipart format. Fix: when source's
FileMetadata.multipart_etag is Some, route through
engine.store_with_multipart_etag (the H1 helper from the prior
wave). Integration test creates a real multipart upload on
source, replicates, and asserts source/destination ETags match.
M1 — replication status="succeeded" with partial failures
Pre-fix the status only flipped to failed when EVERY copy
errored, so dashboards reading last_status got a silent partial
failure. Fix: status is failed whenever ANY object copy or
delete errored (or a fatal list/planner error happened); else
succeeded. The per-object failure ring still captures the full
error list; the status string is now just an honest summary.
M2 — PUT Object ignored conditional headers
If-Match / If-None-Match / If-Modified-Since /
If-Unmodified-Since were silently dropped, breaking
compare-and-swap and the canonical If-None-Match: * idempotent-
create primitive. Fix: new evaluate_put_conditionals runs
BEFORE engine.store and 412s on any precondition failure. AWS
PUT semantics: all four return 412 (no 304 — that's GET-only).
Two integration tests cover idempotent-create + CAS.
M3 — bucket subresources answered for ghost buckets
GET ?location / GET ?versioning / GET ?uploads returned 200
for buckets that didn't exist. Fix: new require_bucket_exists
helper at the top of each subresource branch. Three integration
tests pin the 404 contract.
M4 — tagging stubs returned 501 even for missing resources
The 501-NotImplemented stubs at GET ?tagging (object + bucket)
swallowed the head-fail with let _ = ...; // drain error and
returned 501 unconditionally. AWS returns 404 first in this case.
Fix: propagate the head error so 404 NoSuchKey/NoSuchBucket wins.
Two integration tests pin the precedence.
L1 — zero-byte managed objects omitted Content-Length: 0
build_object_headers only emitted Content-Length when
file_size > 0, conflating "unknown streaming size" with "known
zero". Clients waiting for the response body got hung up on the
missing terminator for empty managed objects. Fix: new
discriminator size_is_known = file_size > 0 || !md5.is_empty().
Managed objects always carry the canonical empty-MD5
(d41d8cd98f00b204e9800998ecf8427e) for empty content, so the
known-zero branch now emits Content-Length: 0 while the
unknown-streaming branch (empty md5 + file_size = 0) still
omits it for chunked-transfer compatibility.
L2 — copy-source conditional combos drifted from AWS spec
AWS S3 CopyObject pairs if-match ↔ if-unmodified-since and
if-none-match ↔ if-modified-since with positive-header
precedence: when if-match passes, the date check is ignored
entirely. Pre-fix, our linear evaluator rejected requests AWS
would have accepted. Fix: explicit pair-aware evaluator —
if-match present → if-unmodified-since suppressed; if-none-match
present → if-modified-since suppressed. Solo-header behaviour
unchanged. Integration test verifies the if-match-passes-with-
stale-if-unmodified-since combination.
Tests (3 unit + 11 integration new): all under
src/replication/worker.rs, tests/replication_test.rs,
tests/s3_correctness_test.rs. Full suite:
659 lib tests + all integration green. Clippy clean.
Lazy bucket replication (v1: run-now via admin API)#
First cut of scheduled source→destination replication. Built around
the invariant that all copies route through engine.retrieve →
engine.store, so replication is transparent to per-backend
encryption and delta compression. Operators author rules in YAML
(storage.replication.rules[]); runtime state (progress, history,
failures) lives in the config DB.
Scope of this commit stream:
- Config shape: new
ReplicationConfig/ReplicationRule/ReplicationEndpoint/ConflictPolicytypes with full YAML round-trip. Static validation (rule-name regex, humantime interval parsing, minimum intervals, batch-size bounds, glob compilation) + multi-hop cycle detection inConfig::check. - Pure planner:
rewrite_key+should_replicate+plan_batchinsrc/replication/planner.rs— I/O-free, heavily unit-tested (15 tests covering the conflict matrix, glob filtering, DG-internal skip, directory-marker skip). - Persistent state (
ConfigDbschema v6): newreplication_state/replication_run_history/replication_failurestables. Boot-timereconcile_on_bootsweeps zombiestatus='running'rows from a prior crashed process. Rules removed from YAML CASCADE-delete their history. - Worker:
src/replication/worker.rs::run_rule— executes one pass of a single rule. TakesArc<Mutex<ConfigDb>>and reacquires the lock at each sync boundary so the handler future staysSend(rusqlite'sConnectionis!Sync). - Admin API (session-gated, not IAM-gated):
GET /_/api/admin/replication— overview of all rules + statePOST /_/api/admin/replication/rules/:name/run-now— trigger a synchronous run; returns 409 Conflict on a paused rulePOST /_/api/admin/replication/rules/:name/pause+/resumeGET /_/api/admin/replication/rules/:name/history?limit=NGET /_/api/admin/replication/rules/:name/failures?limit=N
YAML shape:
storage:
replication:
enabled: true
tick_interval: "30s"
max_failures_retained: 100
rules:
- name: prod-to-backup
enabled: true
source: { bucket: prod-artifacts, prefix: "" }
destination: { bucket: backup-artifacts, prefix: "" }
interval: "15m"
batch_size: 100
replicate_deletes: false # default
conflict: newer-wins # newer-wins | source-wins | skip-if-dest-exists
include_globs: []
exclude_globs: [".dg/*"]
Scope that landed in follow-up commit streams:
- Background scheduler loop, periodic ticks, continuation-token resumption of long runs, and graceful shutdown.
- Delete replication (
replicate_deletes: true) with provenance markers so only objects written by the same rule are delete candidates. - Admin UI panel for object replication. Multi-hop copies remain out of scope.
Dependency added: humantime = "2".
Tests: 15 unit (planner) + 8 unit (state_store) + 2 unit (worker)
- 2 integration (end-to-end via spawned proxy). Clippy clean.
Second-wave correctness fixes (H1/H2 + M1–M4 + L1)#
Seven additional findings from a follow-up review of the multipart, copy, tagging, and pagination surfaces.
H1 — Multipart ETag was inconsistent across Complete vs HEAD/LIST
CompleteMultipartUpload returned "md5(concat)-N" but the persisted
FileMetadata carried a fresh full-body MD5. Clients caching the
Complete ETag as an If-Match precondition hit 412 on the next write.
Fix: new FileMetadata.multipart_etag: Option<String> threaded
through engine.store_with_multipart_etag / store_passthrough_ chunked_with_multipart_etag; etag() returns the override when set.
Persisted on both backends (xattr round-trip, S3 user-metadata
dg-multipart-etag).
H2 — DeleteBucket ignored active multipart uploads
DeleteBucket returned 204 with MPU state still alive for that bucket.
After the delete, ListMultipartUploads reported ghost uploads and
UploadPart (M1 below) silently accepted bytes into the orphan state.
Fix: MultipartStore.count_uploads_for_bucket gate in the handler;
409 BucketNotEmpty with MPU wording in the error message.
M1 — UploadPart / UploadPartCopy didn't check bucket existence
The C2 wave added ensure_bucket_exists to Initiate and Complete but
missed the in-between part upload paths. Attacker could Initiate, have
the bucket deleted, and keep feeding parts. Fix: both UploadPart and
UploadPartCopy now call ensure_bucket_exists (UploadPartCopy also
checks the source bucket).
M2 — CopyObject ignored x-amz-copy-source-if-* preconditions
Pre-fix, the four copy-source headers were read from the request but
never evaluated — clients saying "copy only if source is still vX"
got unconditional copies. New pure
check_copy_source_conditionals(headers, source_metadata) evaluates
if-match / if-none-match / if-modified-since /
if-unmodified-since per AWS spec (all violations return 412, even
the none-match/modified-since variants that would normally be 304
on GET). Wired into both copy_object_inner and upload_part_copy.
M3 — invalid x-amz-metadata-directive silently became COPY
Any value other than case-insensitive REPLACE fell through to
COPY. A typo like REPLAC succeeded with source metadata
preserved — the opposite of the client's intent. Now rejected with
InvalidArgument citing the bad value.
M4 — tagging stubs returned fake 200 while discarding tag state
Object and bucket PUT/GET/DELETE ?tagging used to return 200/empty
<TagSet/>, silently dropping tags the client attached. Clients
using tags for lifecycle, compliance labels, or ABAC would read them
back empty and assume something wiped them. All five handlers now
return 501 NotImplemented — honest "not supported."
L1 — ListParts / ListMultipartUploads pagination lied
max_parts and max_uploads were hardcoded to 1000 with
is_truncated=false and empty markers, regardless of input. An
upload with >1000 parts silently dropped the tail. Fix: full
pagination support — ObjectQuery + BucketGetQuery accept
max-parts/part-number-marker and
max-uploads/key-marker/upload-id-marker. New paginated helpers
MultipartStore.list_parts_paginated + list_uploads_paginated
implement tuple-cursor semantics matching AWS.
Tests: 5 unit (types) + 2 integration (multipart_etag_test) for
H1, plus 10 integration tests in tests/s3_correctness_test.rs
covering H2/M1/M2/M3/M4/L1. 618 lib tests + all integration green.
Clippy clean.
Security hardening — adversarial review findings (C1–C4 + E1/E2/E4)#
A static adversarial review surfaced four critical vulnerabilities and four edge-case hazards. Every confirmed finding is fixed with regression tests. No breaking changes for well-behaved clients; behavioural tightening on attack paths.
Critical — C1: IAM LIST bypass
A user with policy { resources: ["bucket/alice/*"] } calling
GET /bucket?list-type=2&prefix= previously received every key in
the bucket, because the middleware's can_see_bucket fallback
admitted the request and the handler returned unfiltered engine
output. Fix: new ListScope request extension marks LIST
authorisations as Unrestricted or Filtered { user }. When
Filtered, the handler post-filters each key and CommonPrefix
through user.can(Read|List, bucket, key). Unscoped admins pay no
filter cost. A x-amz-meta-dg-list-filtered: true response header
signals filtered pages. 10 unit + 5 integration tests pin the matrix.
Critical — C2: implicit bucket creation via PUT
PUT /new-bucket/key on the filesystem backend previously
succeeded, silently creating the bucket directory as a side effect
of ensure_dir → create_dir_all. Bypassed s3:CreateBucket
equivalence and diverged from the S3 backend contract. Fix: handler
precheck ensure_bucket_exists on PUT / COPY (both ends) / Create
- Complete multipart; backend-level
require_bucket_existsrefuses writes when the bucket root is missing. 9 tests (5 unit + 4 integration) including filesystem-level "no directory was created" assertions.
Critical — C3: multipart in-memory DoS
upload_part previously accepted arbitrarily many parts of
arbitrary size, capped only at Complete-time. Attacker opens
DGP_MAX_MULTIPART_UPLOADS (default 1000) uploads and fills each
— process OOMs. Three-layer mitigation: per-upload size cap
(cumulative rejects over max_object_size), global in-flight byte
counter (configurable via DGP_MAX_TOTAL_MULTIPART_BYTES,
defaulting to max_object_size * max_uploads / 4), and idle-TTL
sweeper (configurable via DGP_MULTIPART_IDLE_TTL_HOURS, default
24h). Sweeper preserves Completing uploads regardless of age.
Critical — C4: multipart complete/abort race
Pre-fix, complete() set a completed: bool under a write lock
then released; the handler then awaited engine.store*. During
that await, abort() happily removed the upload and returned 204
even though the object was about to land. Fix: replace with a
2-state MultipartState::{Open, Completing} enum; abort returns
InvalidRequest when state is Completing; handler now calls
finish_upload on store success or rollback_upload on failure.
8 unit tests cover every legal and illegal transition.
Hygiene — E1: recursive DELETE no longer materialises full listing
DELETE /bucket/prefix/ (trailing slash) previously called
list_objects(..., u32::MAX, ...), materialising the full prefix
in memory before deleting anything. Fix: paginate in 1000-key
windows. Memory stays O(page_size) regardless of prefix depth.
Integration test seeds 1100 objects to exercise the second page.
Hygiene — E2: conformance marker for get_passthrough_stream_range
The trait's default impl buffers the full object — correct as a
fallback, disastrous as the hot path for a ranged GET. Added a
source-walk conformance test that fails if any in-tree
impl StorageBackend for X block omits the override. Prevents a
new backend (future B2, GCS, etc.) from silently inheriting the
default.
Hygiene — E4: sanitise internal error text to S3 clients
Backend errors like StorageError::Other (with filesystem paths or
MinIO debug strings) and EngineError::ChecksumMismatch (with
computed + expected SHA hashes) used to get stringified into error
XML. Fix: new sanitise_for_client helper returns a generic
"Internal server error" to the client while logging full detail to
tracing::error! under a dgp::sanitised_error target.
Env-var knobs added:
DGP_MAX_TOTAL_MULTIPART_BYTES(bytes, defaults tomax_object_size * max_uploads / 4)DGP_MULTIPART_IDLE_TTL_HOURS(hours, default 24)
Tests: 613 lib tests + 19 new integration tests green. Clippy clean.
Deferred to a follow-up PR: E3 (SigV4 chunk-signature verification) ships under a feature flag.
Declarative IAM — quality pass + convenience endpoints#
Follow-up to the initial 3c.3 shipment (see below). Two reviews (clean-code hygiene + correctness x-ray) surfaced 10 findings; the real ones are fixed with regression tests, and two long-promised operator-convenience endpoints land here:
Correctness fixes
-
Critical — mapping_rules wipe on idempotent re-apply (
src/config_db/declarative.rs). The oldVec+ ambiguous helper couldn't distinguish "YAML matches non-empty DB, keep" from "YAML empty, wipe DB". Replaced with an explicitMappingRulesAction:: {Keep, ClearAll, ReplaceWith(Vec)}enum set bydiff_iam. Before the fix, every GitOps reconcile loop on a non-empty rule set silently wiped the table — next OAuth login had no mappings. Regression tests pin the tri-state matrix. -
High —
permissions_equalnow normalises case before compare. YAML authored witheffect: "allow"used to mark every user as changed on every apply (DB stores canonical"Allow"). The diff now normalises both sides first. -
High — misleading docs on
${env:NAME}syntax. The docs claimed env-var substitution worked; no implementation exists. Docs amended to reflect reality ("plaintext or materialise from secret manager at deploy time"); env-substitution stays on the roadmap for a later phase. -
Medium — access-key swap validation. A YAML that swaps access_keys between two surviving DB users now surfaces as a clean validation error ("user 'alice' collides on access_key_id with existing DB user 'bob'") instead of an ugly mid-transaction SQLite UNIQUE failure.
-
Medium — short-circuit step 4c on unrelated PATCHes.
apply_config_transitionused to run the full reconcile +rebuild_iam_index+trigger_config_syncon every PATCH in declarative mode, even ones that only touchedlog_levelorcache_size_mb. Now short-circuits when the IAM fields are unchanged from the previous config. -
Low — no-op reconcile no longer triggers S3 sync upload or surfaces a spurious "reconciled:" warning. Idempotent GitOps loops now leave the sync bucket alone.
Hygiene
ReconcileStats::audit_entries() + summary_line()collapse a 10-block audit-log loop + two independent format strings into single helpers.replace_group_permissions+replace_user_permissionsare now 3-line delegates to a sharedreplace_permissions(tx, table, fk, owner_id, perms).- Vestigial
mapping_rules_need_clearinghelper (the bug's surface) is gone.
Operator convenience
-
Reconciler preview on
/validate.POST /_/api/admin/config/ section/access/validatewith a declarative-mode body now returns a preview warning line:"declarative IAM preview: users(+1/~2/-0) groups(+0/~1/-0) providers(+0/~0/-0) mapping_rules=keep". The admin-UI's ApplyDialog surfaces it under Warnings. Runs the samediff_iamthe live apply does — preview can't drift from reality. Also previews the empty-YAML gate refusal so operators catch the problem at dry-run time. -
Export-as-declarative endpoint.
GET /_/api/admin/config/ declarative-iam-exportreturns a self-containedaccess:YAML fragment withiam_mode: declarative+ populatediam_users/iam_groups/auth_providers/group_mapping_rulesprojected from the current DB. Secrets redacted (operator wires via env). Roundtrip contract: exported YAML (with secrets re-injected) → PUT is an idempotent no-op. Makes Workflow A ("import existing DB into GitOps") a one-button operation instead of hand-assembly.
580 unit tests + 10 declarative integration tests + all encryption
integration tests green. Clippy -D warnings clean. Rustfmt clean.
Declarative IAM reconciler (Phase 3c.3)#
iam_mode: declarative now actually reconciles. Previously it was a
pure lockout (admin-API IAM mutations returned 403, but nothing synced
YAML → DB). The reconciler runs on every /config/apply or
section-PUT on access when the target mode is Declarative:
- Pure diff first, side effects second.
diff_iamvalidates the YAML (unique names, valid group refs, valid permissions, no access-key collisions, no reserved$-prefixed names) and computes creates / updates / deletes. Any validation failure returns an error with ZERO DB writes. - Atomic apply.
apply_iam_reconcileruns all mutations in a single SQLite transaction. Partial failures roll the whole reconcile back; state never observed mid-apply. - ID preservation. Users and groups are matched by NAME, not by
DB id. Rotating an access key in YAML is an UPDATE that preserves
the row id — so
external_identitieslinked to that user stay valid.external_identitiesthemselves are NOT reconciled (they are runtime OAuth byproducts) but ARE cascade-deleted when a YAML-authoritative delete removes the user or provider they reference. - Empty-YAML gate. A
gui → declarativeflip with emptyiam_users/iam_groupsis refused — it would wipe the DB. Declarative → declarative with empty YAML is allowed (operator deliberately clearing all IAM). - Audit trail. Every mutation emits an
iam_reconcile_*audit entry (_user_create/_update/_delete, same for groups and providers, plus_mapping_rules_replaced).
Wire shape: access.iam_users / iam_groups / auth_providers /
group_mapping_rules on AccessSection. Users reference groups by
NAME; mapping rules reference providers and groups by NAME. IDs are
ephemeral and must not appear in YAML.
See docs/product/reference/declarative-iam.md for the full
walkthrough (including both workflows — importing an existing
populated DB, or authoring fresh IAM from YAML — and the complete
adversarial-edge table).
Backends panel: legacy singleton now surfaces correctly#
GET /api/admin/backends used to return an empty list when the
config used the legacy storage.backend singleton (vs the named
storage.backends list), so the admin UI showed "No named backends.
Using legacy single-backend mode." while the proxy was actively
serving. GET /api/admin/config already synthesised a "default"
entry to keep the UI functional elsewhere; /backends now does the
same.
Synthesised entries carry is_synthesized: true in the response so
the UI can disable destructive actions (they don't correspond to a
real cfg.backends[] row; DELETE would 404). The Backends panel
renders them with a "LEGACY SINGLETON" badge and no Delete button.
v0.8.10#
A QA-focused release. No user-facing behaviour changes in the S3 API or delta codec. Two new admin endpoints, targeted new coverage for hot paths, and a sharper test suite overall.
New admin affordances#
POST /api/admin/config/sync-now— operator-triggered pull from the config-sync S3 bucket. Previously only reachable by waiting up to 5 minutes for the periodic poll. Useful for multi-replica deployments wanting immediate IAM propagation after an out-of-band mutation. Returns 404 whenconfig_sync_bucketis unset, 200 with{downloaded, status}otherwise.GET /api/admin/iam/version— monotonic counter bumped on every IAM index rebuild (user/group CRUD, OAuth provider change, mapping-rule edit). Public on purpose — exposes only an opaque number. Powers thewait_for_iam_rebuildtest barrier and is generally useful for scripting "wait for propagation."
Shared helper moved into the library#
reopen_and_rebuild_iam relocated from src/startup.rs (binary)
to src/config_db_sync.rs (library) so the admin sync-now
handler can reach it without duplicating DB-reopen + IAM-rebuild
logic. Behaviour unchanged; callers (startup, periodic poller,
sync-now) funnel through one path.
Test posture overhaul (internal)#
Not user-visible but worth recording because it affects CI + dev loop durability:
- Deterministic IAM rebuild barrier (
wait_for_iam_rebuild) replaces twotokio::time::sleep(1s)calls inauth_integration_test.rs. Auth suite went from ~3.5s to ~1.5s and is no longer flake-prone on slow CI runners. - 6 new unit tests for
classify_s3_error/classify_get_errorinstorage/s3.rs— zero unit tests before, now covers the Hetzner/Ceph 403-for-missing-bucket quirk, object-level 403 preservation, and NoSuchKey → NotFound mapping. - 7
test_create_bucket_*integration tests replaced by 3 unit tests inhandlers/bucket.rs+ property-based coverage viaproptest(new dev-dep). ~1500 random cases per run, 0.08s. - 9 S3-backend duplicate tests deleted from
tests/s3_backend_test.rs(verified trait behaviour already covered by the filesystem-backed suites). Kept 2: SDK-plumbing smoke + real delta+S3 interaction. - 6 of 10 public-prefix LIST integration tests removed — unit
tests in
api/auth.rsandadmission/evaluator.rsalready cover the policy-match shape. Kept: no-slash AWS-CLI bug regression, false-parent denial, partial-public root denial, admin sees all. - 2 new integration tests for metadata-cache invalidation in
tests/optimization_test.rs— covers PUT-overwrite and batch- delete paths that the docstring pinned but no test exercised. - 4 new HA config-sync tests (
tests/config_sync_ha_test.rs) exercising the real sync code path: startup pull, sync-now propagation, ETag no-op, wrong-passphrase rejection. - 3 new same-key concurrency tests (
tests/concurrency_test.rs) for double-CompleteMultipartUpload, cross-upload-ID isolation, and PUT-vs-DELETE state consistency. - 3 new SigV4 security tests (
tests/auth_integration_test.rs) covering signed-header tampering, presigned-URL rejection after user disable, and unsigned-extra-header tolerance (spec compliance). - 5 new unit tests for
src/startup.rs(was 0/650 LOC). - Coverage reporting in CI. New
coveragejob runscargo-llvm-cov --liband publishes a summary table + lcov.info artifact.continue-on-error: true— signal, not a gate. - TestServer hardening.
wait_readynow checksprocess.try_wait()BEFORE the health probe, so a stray proxy holding the test port fails loudly (lsof -i :<port>) instead of silently intercepting requests.
Internal hygiene#
- Shared
useSectionEditorhook in the admin UI collapses ~450 LOC of triplicated apply-flow boilerplate across AdmissionPanel, CredentialsModePanel, and the Advanced sub- panels into one ~220 LOC hook. Future section panels plug in rather than re-carrying the §F5 snapshot-at-validate fix etc. with_config_db()wrapper insrc/api/admin/mod.rsshrinks the "get DB from state → lock → run op → log-and-500" admin handler boilerplate from 5–8 lines to 3. Migrated 8 handlers inexternal_auth.rs.- 6 duplicate
default_*fns inconfig_sections.rseliminated by making theconfig.rsversionspub(crate). Round-trip test continues to guard against drift.
Documentation#
CLAUDE.mdrefreshed against recent architecture drift. Two new sections: "Testability principles" (7 bullets with prior-art citations) and "When proposing architecture" (4 anti-patterns to avoid).
Metrics#
- Test count: 467 → 487 unit/property/binary tests. Integration tests net −12 (22 deleted as redundant, 10 added for genuine gaps).
- CI time: neutral — slower startup-unit-test overhead (~0.5s) offset by ~2s saved in auth + bucket-name tests.
v0.8.0#
A major release. Two overlapping threads land together: the progressive-disclosure YAML config (phases 0–3 of docs/plan/progressive-config-refactor.md) and the first waves of the admin UI revamp (docs/plan/admin-ui-revamp.md).
Configuration (progressive-disclosure YAML — phases 0–3)#
- YAML is now canonical.
deltaglider_proxy.yaml(four-sectionadmission / access / storage / advancedlayout) is preferred overdeltaglider_proxy.toml. Both still load; TOML is deprecated with atracing::warn!on every load (silence withDGP_SILENCE_TOML_DEPRECATION=1). See docs/HOWTO_MIGRATE_TO_YAML.md. - Dual-shape loader. Sectioned YAML (
admission:/access:/storage:/advanced:) is transparent to the in-memoryConfigstruct — the flat shape still loads unchanged. - Storage shorthand — a single-backend deployment can write:
which expands to a full
storage: s3: https://example.combackend: { type: S3, ... }at load time. Filesystem shorthand (storage: { filesystem: /path }) likewise. Long-form YAML stays valid; shorthand is operator convenience only. - Per-bucket
public: trueshorthand expands topublic_prefixes: [""], and the canonical exporter collapses back when unambiguous. GUI "Public read" toggle maps 1:1 to the YAML. - Admission chain (operator-authored). New
admission.blocks[]wire format withmatchpredicates (method / source_ip / source_ip_list / bucket / path_glob / authenticated / config_flag) andactionvariants (allow-anonymous,deny,reject { status, message },continue). Evaluator dispatches live before the synthesized public-prefix blocks; reservedpublic-prefix:*name prefix; glob compile-check at parse time; 4096-entry cap onsource_ip_list. access.iam_mode: gui | declarative. New lifecycle knob:gui(default) keeps the encrypted IAM DB as source of truth;declarativegates admin-API IAM mutation routes (therequire_not_declarativemiddleware) behind 403 responses. A warn-level audit log line fires on every mode transition. The reconciler that sync-diffs DB to YAML in declarative mode is still Phase 3c.3; today declarative mode is a read-only lockout.- New admin endpoints for GitOps + GUI round-trips:
GET /api/admin/config/export[?section=<name>]— canonical-YAML export with every secret redacted; scope-filter to one section whensection=is present.POST /api/admin/config/validate— dry-run a full YAML doc.POST /api/admin/config/apply— atomic full-document apply, with runtime-secret preservation for redacted round-trips.POST /api/admin/config/trace— evaluate a synthetic request against the live admission chain.GET /api/admin/config/defaults[?section=<name>]— JSON Schema (viaschemars) for the Config type, optionally scoped to one section for Monaco's per-editor schema.
- Section-level admin API (Wave 1 of the UI revamp):
GET /api/admin/config/section/:name[?format=yaml]PUT /api/admin/config/section/:name— partial update, routes through the sameapply_config_transitionhelper as field-level PATCH and document-level APPLY.POST /api/admin/config/section/:name/validate— dry-run with a diff body ({section: {field.path: {before, after}}}) for the plan → diff → apply dialog.GET /api/admin/config/trace— query-param variant for bookmarkable trace URLs.GET /api/admin/audit?limit=N— recent entries from the in-memory audit ring (newest first; limit clamped to[1, 500]). Backs the Diagnostics → Audit log viewer (Wave 11).
- CLI subcommands:
deltaglider_proxy config migrate <toml>— TOML → YAML converter (emits canonical sectioned form).deltaglider_proxy config lint <yaml>— offline schema + reference-resolution + dangerous-default warnings.deltaglider_proxy config defaults— dump every default with its doc comment.
Admin UI revamp (waves 1–11)#
All ten waves from
docs/plan/admin-ui-revamp.md plus
a follow-on Wave 11 (audit log viewer) have landed as of this
release. Waves 1–3 shipped at tag time; waves 4–8 landed during
live-browser verification; waves 9 (Trace diagnostics), 10
(command palette + shortcuts help), 10.1 (finish §10 polish
items), and 11 (audit log viewer) shipped post-tag and are
rolled into the v0.8.0 entry so the main history is
self-describing. §9.1's Dashboard redesign was deferred — the
existing MetricsPage remained functional so the rewrite can
ship independently without blocking operators today.
- Section-level API client helpers in
adminApi.ts:getSection,putSection,validateSection,getSectionYaml,exportSectionYaml,getSectionSchema,getFullConfigSchema. - New foundation components ready for use in waves 4–7:
FormField— standardised label / YAML-path / help / default-placeholder / override-indicator / owner-badge wrapper.ApplyDialog— the plan → diff → apply modal (§5.3).MonacoYamlEditor— lazy-loaded Monaco + monaco-yaml with scoped JSON Schema and mobile fallback (§4.3, §10.4).
- New hook
useDirtySection— per-panel dirty state backed by a module-level Set, withuseDirtyGlobalIndicatorsfor the●tab-title prefix andbeforeunloadguard. - New sidebar (
AdminSidebar) — four-group IA (Diagnostics + Configuration). Nested sub-entries for Access / Storage / Advanced; amber dot on sections with unsaved edits. - URL scheme — every admin page is a bookmarkable
hierarchical URL:
/_/admin/diagnostics/dashboard,/_/admin/configuration/access/users, etc. Legacy flat URLs (/_/admin/users,/_/admin/backends) keep working viaLEGACY_TO_NEWin AdminPage.tsx. - Right-rail actions (
RightRailActions) visible on every Configuration page: Apply / Discard (gated ondirty) and Copy YAML (section-scoped). In practice the rail was simplified during Wave 3's adversarial review to copy-only — section-scoped paste and full-document Import/Export run from theYamlImportExportModalreached through the header actions. - YAML Import/Export modal (
YamlImportExportModal) reached from every admin page via the header actions. IAM Backup was renamed in the sidebar (was "Backup") to distinguish it from the YAML Import/Export surface: IAM Backup ships the full encrypted SQLCipher DB (users, groups, OAuth, mappings); YAML Import/Export ships the operator config document. - Admission block editor (Wave 4,
AdmissionPanel) — drag- to-reorder list of operator-authored blocks with Form ⇄ YAML toggle per row, inline validation matching the server rules (name charset, reservedpublic-prefix:*prefix, mutually exclusivesource_ip/source_ip_list, 4xx/5xxrejectstatus), and synthesizedpublic-prefix:*blocks surfaced read-only below. - Access → Credentials & mode panel (Wave 5) — the first screen on the Access section: iam_mode radio (gui / declarative with the Phase 3c.3 gap disclosed inline), authentication-mode dropdown, bootstrap SigV4 key pair with rotate-in-place, and a Change password link.
- Storage → Buckets panel (Wave 6) — per-bucket editor with
tri-state Anonymous read access (None / Specific prefixes /
Entire bucket). Selecting "Entire bucket" writes the compact
public: trueshorthand; "Specific prefixes" writes an explicitpublic_prefixeslist. The GUI toggle maps 1:1 to the YAML. - Advanced panel sub-sections (Wave 7) — the Advanced
section is split into five dedicated sub-panels (Listener &
TLS, Caches, Limits, Logging, Config DB sync), each with
grouped forms,
🔁restart-required badging, and monospacefrom DGP_X_Ychips on env-var-owned fields. - First-run setup wizard (Wave 8) — five-screen guided
onboarding at
/_/admin/setup: backend pick → backend config (with live Test Connection for S3) → admin credentials → optional public bucket → review + apply. Routes to the Dashboard on success. Zero-to-working in under three minutes on a fresh install. - IAM source-of-truth banner (
IamSourceBanner) — surfaced on every Access sub-panel: explains in plain English whether the listed users/groups/providers are DB-managed (iam_mode: gui) or read-only because YAML owns the state (iam_mode: declarative). - AntD 6 shrink fix —
theme.cssoverrides the AntD 6 radio/checkbox "shrink on click" default that was breaking Wave 4's match-action radio groups. - Admission trace panel (Wave 9,
TracePanel) — the/_/admin/diagnostics/traceplaceholder is replaced by a real admission-chain debugger. Synthetic-request form (method / path / query / source IP / authenticated toggle), one-click example chips, and a result pane built around the Istio / Kiali "reason path" pattern: decision tag, matched-block name, resolved-request breakdown (method / bucket / key / list_prefix / authenticated), and Copy-as-JSON + Clear actions. CallsPOST /api/admin/config/traceunder the hood; no dirty state (pure read). A "How to read the output" info alert demystifies the panes for first-time users. - Command palette (Wave 10,
CommandPalette) —⌘K/Ctrl+Kopens a fuzzy-filter palette over every entry in the four-group IA plus shell-scope actions (Export YAML, Import YAML, Setup wizard, Keyboard shortcuts, Back to Browser). Hand-rolled scorer (substring + subsequence, <50-line function); combobox ARIA witharia-controls/aria-activedescendant; group headings ("Recent" / "Navigate" / "Actions") rendered non-interactively so Arrow keys skip over them. Recent-items MRU (last 5) persists to localStorage and surfaces under the "Recent" heading on empty query. - Shortcuts reference modal (Wave 10,
ShortcutsHelp) —?opens a platform-aware shortcut list. Mac users see only⌘rows; Windows / Linux users see onlyCtrlrows — no duplicate "same as ⌘K on non-Apple" noise rows. Detection vianavigator.userAgentData.platformwithnavigator.platform- UA-string fallbacks (
platform.ts). The listener itself still accepts BOTH modifiers so a Mac user on a PC keyboard still works.
- UA-string fallbacks (
⌘S/Ctrl+S→ Apply dirty section (Wave 10.1, §10.3).useDirtySection.tsgainsregisterApplyHandler(section, fn)useApplyHandler(section, fn, enabled); handlers are stacked per section so the most-recently-mounted panel wins when siblings share a section (Wave 5 master-detail).AdminPagedispatches⌘SviarequestApplyCurrent( sectionForPath(adminPath)); falls through to the browser default when no handler is registered (Diagnostics pages, no dirty state). Wired into Admission, Credentials & mode, and all five Advanced sub-panels.
- Mobile drawer (Wave 10.1, §10.4) — below 900px the
persistent sidebar collapses to an AntD Drawer that slides in
from the left. Hamburger trigger lives in the header. Drawer
auto-closes on navigation.
useIsNarrowhook tracks the breakpoint live (rotate / devtools-toggle works without reload). - i18n scaffold (Wave 10.1, §10.2,
i18n.ts) — pass-throught(key, fallback)helper and matchinguseT()hook. Today returnsfallback ?? key— no translation tables yet. Purpose: when we add a locale, we swap one file's internals and every existing call site continues to work. - Audit log viewer (Wave 11,
AuditLogPanel) — the admin UI now surfaces the server's audit trail at/_/admin/diagnostics/audit. Backed by a new in-memory ring insrc/audit.rs(boundedVecDeque<AuditEntry>; default 500 entries; override viaDGP_AUDIT_RING_SIZE). Everyaudit_log()call pushes a sanitised copy onto the ring, so stdout / JSON log shippers see nothing change — the ring is supplementary. NewGET /api/admin/audit?limit=Nendpoint (session-gated; not IAM-gated — all admins see the same log). Panel renders a monospace table with colour-coded Action tags (red =login_failed/delete, green =login, etc.), free-text filter across every column, optional 3-second auto-refresh, and an "In-memory ring — not a compliance substitute" banner so operators don't mistake it for durable storage. - IAM Backup import preserves
external_identities(post- manual-review hardening).POST /api/admin/backupused to silently drop every OIDC identity binding; restoring a backup would orphanuser_id ↔ external_sub ↔ provider_idlinks and returning OIDC users got re-provisioned as new accounts. Import now remapsuser_id+provider_idthrough the existing ID maps and falls back togroups.member_ids/ autoincrement heuristics for legacy backups that lack abu.idfield. New optionalBackupUser.idfield in exports. Idempotent: repeat imports skip existing(provider_id, external_sub)pairs.ImportResultgainsexternal_identities_created+external_identities_skippedcounters. - GUI polish (post-manual-review round):
- Groups create form closes and navigates to the Edit view after a successful create. Previously the form stayed open with stale fields; "+ New" appeared broken until a page reload.
- ExtAuth mapping-rule "+ Add Rule" flushes pending edits on other rules BEFORE create + reload. Previously the reload overwrote the in-memory rules array, silently clobbering unsaved typing in sibling rows.
- New mapping rules start with an empty
match_valueand a placeholder hint. Previously pre-filled with*@example.comas a literal default (forcing select-all + delete before typing). - Users list label reflects group inheritance:
N groups (inherited)orN rules · M groupsinstead of a misleading "No access" for users who have group-inherited permissions.
Bug fixes / correctness#
DGP_BOOTSTRAP_PASSWORDenv override.--config <path>now applies env var overrides (includingDGP_BOOTSTRAP_PASSWORD_HASH) on top of the file, matching the behaviour of implicit search-path loading. Previously the flag made env vars silently ignored.- Admission validator rejects reserved
public-prefix:*names up front. - Shape classifier explicit check-by-key-presence (flat vs. sectioned), so typos inside a sectioned doc report section- scoped errors rather than "unknown variant" from a fallback parse.
- Section PUT is RFC 7396 merge-patch (post-tag regression
fix from live-browser verification). The first cut of
PUT /api/admin/config/section/:namereplaced the whole section — a PUT with{max_delta_ratio: 0.42}silently resetcache_size_mb,listen_addr, andlog_levelto their compile-time defaults. Now: present keys apply in place; absent keys preserve the current value; explicitnullreverts to the default; array fields (e.g.admission.blocks) stay atomic. Three regression tests lock this in. advanced.log_levelfrom YAML applied at startup (post- tag fix).init_tracingpreviously read only from RUST_LOG and DGP_LOG_LEVEL env vars, silently ignoring the YAML value; runtime always ran atdebugdefault unless env was set. Now: RUST_LOG > DGP_LOG_LEVEL > config.log_level > --verbosedefault. Env-driven deployments (Docker + K8s) keep their semantics exactly.
- Legacy admin URLs canonicalise in the browser bar (post-
tag fix). Bookmarked v0.7.x URLs like
/_/admin/usersresolved correctly but left the address bar on the legacy form, spreading the old shape through copy/paste. AuseEffectinAdminPagenow swaps in the canonical hierarchical URL vianavigate(..., replace=true)without adding a history entry.
Dependencies#
Frontend: monaco-editor, monaco-yaml, react-hook-form, zod,
@hookform/resolvers, @dnd-kit/core + sortable + utilities
(§4.2–§4.4). Backend: serde_yaml, schemars, ipnet, globset.
Breaking changes#
None. TOML config is deprecated but still loads; field-level
PATCH /api/admin/config remains the stable GUI surface; every
existing integration test passes.
Migration#
Run deltaglider_proxy config migrate deltaglider_proxy.toml --out deltaglider_proxy.yaml, point the server at the YAML, and
delete the TOML when ready. No config rewrites needed on the
YAML side for existing deployments — the sectioned shape is
semantically a superset of the flat shape.
v0.7.2#
UI Polish & Usability#
- Presigned URL duration selector: Share button is now a split button — main button generates a 7-day link, chevron dropdown offers 1 hour, 24 hours, or 7 days. Share modal displays "Expires in" label.
- Version display: Sidebar branding shows the proxy version (fetched from whoami, not the unauthenticated health endpoint).
- OAuth user identity: Sidebar shows the user's email address (from external identity) instead of the truncated access key ID. Username displayed on its own row for readability.
- Ant Design tooltips removed globally: CSS nuke (
display: none !important) on.ant-tooltip, .ant-popoverprevents all layout-shaking tooltip bugs. All 6<Tooltip>usages replaced with nativetitleattributes. - Analytics tab padding fix: Removed double padding (AnalyticsSection had its own wrapper padding on top of the parent MetricsPage padding). Aligned card styles (borderRadius, fontSize, fontFamily, spacing) between Monitoring and Analytics tabs.
Correctness#
- Analytics stats: Stats endpoint now calls
list_objectswithmetadata=true, so delta compression savings are correctly reported (was always showing 0% savings). - InspectorPanel crash on zip files: Fixed React error #310 —
useStatefor share duration was declared after the early return, causing hook count mismatch. - Whoami returns user info:
GET /api/whoaminow resolves the logged-in user from the session cookie (name, access_key_id, is_admin). OAuth users show their email, IAM users show their name, bootstrap shows "admin".
Security#
- Version removed from health endpoint:
/_/healthno longer exposesversionfield. Version is only available via the authenticated/_/api/whoamiendpoint.
Infrastructure#
- Docker debugging tools:
ntpstatandchronypre-installed in the runtime image for clock skew diagnosis.
v0.7.1#
Features#
- Bulk copy/move: Select multiple objects and copy or move them to a different bucket/prefix via a destination picker modal. Move = copy + delete source (only after all copies succeed).
- Bulk download as ZIP: Select multiple files and download them as a client-side ZIP (fflate). Size warning for selections >500MB.
- Storage analytics dashboard: New Analytics tab in Metrics page with summary cards (total storage, space saved, savings %, estimated monthly cost savings), per-bucket stacked bar chart, session savings area chart, and compression opportunity identifier.
- Cost configuration: Gear icon on the monthly savings card opens a preset selector (AWS S3, S3 IA, Hetzner, Backblaze, Cloudflare R2) with localStorage persistence.
- Themed error pages: OAuth error pages (provider not ready, auth failed, account disabled) respect the user's dark/light theme via CSS custom properties and inline localStorage/prefers-color-scheme detection.
UI Components#
- BulkActionBar: Toolbar with Copy, Move, ZIP, Delete buttons — replaces the inline delete button when objects are selected.
- DestinationPickerModal: Bucket dropdown + prefix autocomplete, preview of destination paths, move-mode deletion warning.
- AnalyticsSection: Summary cards, Recharts bar/area charts, compression opportunity cards.
v0.7.0#
Features#
- OAuth/OIDC external authentication: Login with Google (or any OIDC provider) via the admin GUI. PKCE, state parameter, nonce, JWT validation. Per-provider configuration with display names and custom branding.
- Group mapping rules: Map external identity claims (email domain, email glob, email regex, claim value) to IAM groups. Rules evaluated on every OAuth login; group memberships merged (not replaced).
- Mandatory authentication: Proxy refuses to start without authentication credentials unless
authentication = "none"is explicitly set. Prevents accidental open-access deployments. - Per-bucket compression policies: Enable/disable compression per bucket in the Backends panel. When per-bucket compression is ON but global ratio is 0, uses 0.75 default.
- Multi-instance config sync: ExternalAuthManager rebuilds after S3 config sync. Stale pending OAuth flows cleared on provider rebuild.
UI#
- OAuth login buttons: Connect page shows branded OAuth provider buttons with "credentials instead" collapsible.
- Authentication panel: Full OAuth provider CRUD, mapping rules with local state + Save button (not per-keystroke API calls).
- Backends panel: Merged Storage + Compression tabs. Per-bucket policy toggles inline with backend configuration.
- Custom dropdowns:
SimpleSelectandSimpleAutoCompletereplace broken Ant Design Select/AutoComplete popups. - Tab headers: Centered, typographically consistent headers for all admin settings tabs.
- Object table fixes:
showSorterTooltip={false}prevents sort header shaking. Nativetitleattributes replace broken Ant tooltips.
Security#
- Session cookie fixes: SameSite=Lax (not Strict) for OAuth redirect compatibility. Secure flag auto-detects TLS.
- XSS protection:
escape_html()on OAuth error page parameters. - Rate limiting on OAuth callback: Prevents brute-force state token guessing.
- Group membership merge: OAuth re-login no longer wipes existing group memberships.
Bug Fixes#
- Bucket not loading on first click:
s3client.setBucket()called before React state update. - Empty browser after login:
s3.reconnect()moved to useEffect (was in Promise callback before state committed). - Upload SignatureDoesNotMatch: Strip leading/trailing slashes from upload destination path.
- InternalError hiding real errors: Error message now shows actual error, not generic text.
- Reference.bin fallback: GET on aliased bucket paths falls back to passthrough when reference fetch fails.
v0.6.0#
Features#
- Multi-backend routing: Route different buckets to different storage backends (filesystem, S3, mixed).
RoutingBackendimplementsStorageBackendtransparently — zero engine changes, shared caches/codec/locks. Configure via[[backends]]TOML array or Admin GUI. - Backends admin panel: New "Backends" tab in Admin Settings. Add/remove named backends, test S3 connections, set default backend. Safety checks prevent deleting in-use or default backends.
- Bucket routing UI: Per-bucket policy cards now show Backend + Alias fields when multi-backend is active. Route virtual bucket names to specific backends with optional real-name aliasing.
- TOML/env config hints: Read-only settings (Limits, Security, Advanced Compression) now show copyable
TOML:andENV:examples below each field. - Bucket policies promoted: Per-Bucket Compression section moved to top of Compression tab with improved UX — contextual labels, amber tint on disabled buckets, empty state explains use cases.
Bug Fixes#
- Dropdown positioning: Replaced Ant Design
<Select>with<Radio.Group>for backend type chooser. Fixes broken dropdown positioning at non-100% browser zoom (known@rc-component/triggerbug). - list_buckets_with_dates: Routed virtual buckets no longer mask real creation dates with
now(). Backends are queried first; route names added only if not already found. - Config validation:
default_backendis now validated against thebackendslist at load time. Invalid references are cleared with a warning. Bucket policy backend references are also validated. - S3 credential validation:
POST /api/admin/backendsnow validates credentials upfront (400) instead of failing at engine rebuild (500). - Taint detection:
compute_tainted_fieldsnow comparesbackendsanddefault_backendbetween runtime and disk config.
Code Quality#
- DRY:
BackendInfoResponse::from()impl replaces copy-pasted conversion in two files.rebuild_enginemadepub(super)and shared across config.rs and backends.rs.DEFAULT_CONFIG_FILENAMEconstant replaces 6 magic strings. - Dead code: Removed inline save block from backendTab (uses shared
saveSection).embeddedTabprop made required. - Replay window: Hoisted duplicate
DGP_REPLAY_WINDOW_SECSenv parse to single variable.
v0.5.11#
Features#
- Per-bucket compression: Configure compression enable/disable and max delta ratio per bucket via TOML (
[buckets.<name>]), admin API, or GUI. Unconfigured buckets inherit global defaults. - Bucket Policies GUI: New "Bucket Policies" section in Compression tab — add, remove, edit per-bucket overrides with live save.
Security#
- COPY source path traversal: Reject
..inx-amz-copy-sourcebucket and key (filesystem backend).
Bug Fixes#
- Bucket policy hot-reload: Fixed case normalization mismatch (config stored original case, engine expected lowercase). Added rollback on engine rebuild failure.
v0.5.10#
UI & Documentation#
- Path-based routing: Clean URLs (
/_/browse,/_/admin/users,/_/docs/configuration) replace hash routing. Deep-linking to doc pages with heading anchors. Old hash bookmarks auto-redirect. - Embedded docs: Full-text search (minisearch + Cmd+K), Mermaid diagram rendering with getBBox viewBox fix, Lightbox for images, landing page with screenshots and feature cards.
- Admin GUI: All 39 settings surfaced across 5 tabs (Backend, Compression, Limits, Security, Logging). Theme toggle in Admin/Docs header.
- Design system: 16 new CSS color variables for dark/light themes. Responsive breakpoints (grids collapse on mobile, panels hide).
Correctness#
- SettingsPage: handleSave wrapped in try-catch (prevents stuck save button on network error).
- Multipart: complete() no longer removes upload before store() succeeds (prevents data loss on transient failure).
- SigV4: Credential scope format validation in both presigned and header auth paths.
- ListObjectsV2: max-keys=0 returns InvalidArgument per S3 spec.
- DocsPage: Fixed stale useCallback dependency in handleLinkClick.
Infrastructure#
- build.rs:
rerun-if-changedfor dist/ directory (no more stale embedded UI assets). - Dead code: Removed ApiDocsPage, BrowserConnectionCard, dead hash routes (~400 lines).
- DRY: FullScreenHeader shared component, FULLSCREEN_VIEWS set, dedup-by-key unified.
v0.5.9#
Bug Fixes#
- SigV4 replay: Fixed false positives — presigned URLs and idempotent methods (GET/HEAD) skip replay detection. Replay window reduced from 300s to 2s.
- Logging: SigV4 mismatch logs method, path, scope, signature prefixes. Clock skew logs direction.
- Config DB: Added PRAGMA busy_timeout=5000ms on open and reopen.
v0.5.4#
Security & Hardening#
- Per-request timeout: Added
tower_http::timeout::TimeoutLayerreturning HTTP 504 Gateway Timeout after 300s (configurable viaDGP_REQUEST_TIMEOUT_SECS). Prevents slow clients from holding concurrency slots forever. - Replay cache cap: Replay cache capped at 500K entries. If exceeded (flood attack), cache is cleared with a
SECURITY |warning. - Recursive delete IAM enforcement: Server-side recursive prefix delete (
DELETE /bucket/prefix/) now checks per-object IAM permissions. Previously bypassed individual Deny rules. - Bootstrap hash format validation: Rejects malformed bcrypt hashes at startup instead of failing silently on first auth attempt.
Features#
- Server-side recursive delete:
DELETE /bucket/prefix/(trailing slash) deletes all objects under the prefix. Per-object IAM checks enforced. Filesystem backend uses nativeremove_dir_all. - S3 batch delete: Batch delete (
POST /?delete) usesDeleteObjectsAPI instead of per-file DELETE for ~10x fewer API calls. - Base64 bootstrap password hash:
DGP_BOOTSTRAP_PASSWORD_HASHaccepts base64-encoded bcrypt hashes to avoid$escaping issues in Docker/env vars. Auto-detected format. - Config DB resilience: If the encrypted config DB can't be opened (wrong password, corruption), creates a fresh database instead of crashing.
config_sync_bucketin TOML: Config DB S3 sync bucket now configurable via TOML (was env-only).
UI#
- Favicon: Teal chain-link icon on dark background.
- Reduced polling: Browser file listing polls every 30s instead of every second.
Defaults#
max_delta_ratio: Default raised from 0.5 to 0.75 (more files stored as deltas, better space savings for typical workloads).
Code Quality#
- Removed dead
http_puttest helper from storage resilience tests.
v0.5.3#
Performance#
- Parallel delta reconstruction: Reference and delta fetched concurrently via
tokio::join!instead of sequentially. Saves ~100ms per delta GET on S3 backends. - Parallel metadata HEAD calls:
resolve_object_metadatafetches delta and passthrough metadata concurrently. - Legacy migration off GET path:
migrate_legacy_reference_object_if_neededno longer runs during GET (was adding 60+ seconds of xdelta3 encoding). Available via batch migration endpoint instead.
Correctness#
- PATHOLOGICAL warnings: Delta and reference files missing DG metadata now log prominent warnings instead of silently falling back.
dg-ref-key→dg-ref-pathrename: Reference paths stored as relative paths for portable deltaspaces. Legacydg-ref-keyread with automatic fallback.
Testing#
- Metadata validation tests: 7 new tests for missing, corrupt, partial, and wrong metadata scenarios.
UI#
- Connect page simplification: Shows only relevant fields per auth mode (bootstrap password OR S3 credentials, not both). Removed endpoint URL field (derived from window location).
v0.5.2#
Correctness (Critical)#
- Fix GET on rclone-copied delta files:
retrieve_streamusedobj_key.filenameinstead ofmetadata.original_namefor storage lookups, causing 404 on files whose S3 key had a.deltasuffix. This bug was not caught by 222 tests because they all follow the clean PUT→GET path.
Testing#
- Storage resilience tests: 6 new adversarial tests that would have caught the above bug:
- Triangle invariant (LIST→HEAD→GET must all succeed for every key)
- HEAD/GET content-length consistency
- SHA256 roundtrip verification across all storage strategies
- Unmanaged file triangle invariant (directly placed files)
- External delete → 404 (not stale cached data)
- LIST never exposes
.deltasuffix orreference.bin
Security#
- Session cookie hardening: Secure, HttpOnly, SameSite=Strict attributes.
- Remove secrets from whoami:
/whoamiendpoint no longer exposes secret access keys. - Cleanup
.bakfiles: Removed leftover backup files from refactoring.
v0.5.1#
Security Fixes (Post-Release Audit)#
- Deny condition bypass: Fixed condition evaluation that could skip Deny rules under certain group membership combinations.
- Group
member_idspersistence: Group member lists were not persisted correctly to SQLCipher DB. evaluate_iamalternate path: Fixed edge case where IAM evaluation took a non-standard code path.- Session storage: Server-side session storage for S3 credentials (no longer stored client-side).
env_clear()on subprocess: xdelta3 subprocess environment cleared to prevent credential leaks.DGP_TRUST_PROXY_HEADERS: Default changed totruefor reverse proxy deployments.
Testing#
- Auth integration tests: SigV4 signature verification, presigned URLs, clock skew rejection, IAM lifecycle.
- IAM persona tests: 23 tests for groups, ListBuckets filtering, prefix scoping, conditions.
Infrastructure#
- Docker retry:
apt-getretries on network blips (Acquire::Retries=3). - Startup refactor: Extracted
startup.rsfrommain.rs, splitengine.rsandconfig_db.rsinto sub-modules.
v0.5.0#
IAM Policy Conditions (iam-rs)#
- Full AWS IAM condition support: Integrated
iam-rscrate for standards-compliant policy evaluation with all AWS condition operators (StringEquals, StringLike, IpAddress, NumericLessThan, etc.). s3:prefixcondition: Deny LIST requests based on the prefix query parameter. Example:{"StringLike": {"s3:prefix": ".*"}}blocks listing dotfiles.aws:SourceIpcondition: Restrict operations to specific IP ranges. Example:{"IpAddress": {"aws:SourceIp": "10.0.0.0/8"}}.- DB schema v4: New
conditions_jsoncolumn on permissions and group_permissions tables. Backward compatible — existing permissions work unchanged. - Permission validation: Effect normalization (case-insensitive), max 100 rules per user/group, resource pattern validation (trailing wildcard only),
$-prefix names blocked. - Frontend conditions UI: Collapsible conditions section per permission rule with prefix and IP restriction inputs.
IAM Module Refactor#
- Split
iam.rsinto module:iam/types.rs,iam/permissions.rs,iam/middleware.rs,iam/keygen.rs,iam/mod.rs. Pure permission evaluation logic separated from framework-specific middleware. - Centralized permission authority: All permission checks go through
AuthenticatedUser.can(),.can_with_context(),.can_see_bucket(),.is_admin(). - Legacy admin as AuthenticatedUser: Bootstrap credentials now create a
$bootstrapAuthenticatedUser with wildcard permissions instead of bypassing authorization entirely. - Groups loaded on startup: Previously groups were only loaded on first IAM mutation.
- ListBuckets per-user filtering: Users see only buckets they have permissions on.
- IAM backup/restore: Export/import all users, groups, permissions, and credentials as JSON via
/_/api/admin/backup.
Security Fixes#
- Batch delete per-key authorization: Previously the middleware only checked delete permission at bucket level; now each key in a batch delete is individually authorized.
- Progressive auth delay: Failed auth attempts trigger exponential backoff (100ms→5s), making brute force expensive before lockout.
- Attack detection logging:
SECURITY |log events for brute force detection, lockout, and repeated failures with IP and attempt count. - Condition parse fail-closed: Malformed conditions produce an empty policy (deny-all) instead of stripping conditions (fail-open).
- Error XML escaping: Error
<Message>now XML-escaped to prevent malformed responses. - Bucket names with dots: S3-compliant validation now accepts dots in bucket names.
- is_admin strict check: Uses
== "Allow"instead of!= "Deny".
Performance#
- Range passthrough: Range requests on passthrough objects pass the Range header through to upstream S3 (or seek on filesystem) instead of buffering the entire file. A 1KB range on a 100MB file reads only 1KB from storage.
- Request concurrency limit:
tower::limit::ConcurrencyLimitLayerwith configurable max (default 1024,DGP_MAX_CONCURRENT_REQUESTS). - Write-before-delete: Storage strategy transitions (delta↔passthrough) now write the new variant before deleting the old one, preventing transient 404s on concurrent GETs.
- Prefix lock cleanup:
cleanup_prefix_locks()runs on every lock acquisition instead of only on delete. - Interleave-and-paginate dedup: Shared function for the interleave/sort/paginate pattern used by engine, S3 backend, and filesystem backend.
Unified Audit Logging#
src/audit.rs: Single audit module withsanitize(),extract_client_info(), andaudit_log(). Eliminates duplicated sanitization and IP extraction across handlers and admin API.- Session TTL decoupled:
SessionStore::ttl()method replaces re-parsingDGP_SESSION_TTL_HOURSin the cookie formatter.
Frontend#
- Conditions UI: Permission editor supports prefix restriction and IP restriction inputs with collapsible conditions panel per rule.
- IAM backup buttons: Export/Import buttons in admin sidebar for JSON backup/restore.
- Log level radio buttons: Replaced broken Select dropdown with Radio.Group.
- Polling fix: Users/Groups panels load once on mount instead of re-polling on every render.
- Deduplicated formatters: Unified byte formatters, extracted CredentialsBanner and InspectorSection components.
Tests#
- 222 tests: 180 unit + 42 integration (23 new persona tests covering groups, ListBuckets filtering, prefix scoping, cross-user isolation, multipart, content verification, deny-from-groups, IAM conditions).
- iam-rs condition tests: Unit tests for prefix deny with StringLike, IP deny with IpAddress.
Metadata Cache#
- In-memory metadata cache: New moka-based cache (
MetadataCacheinmetadata_cache.rs) eliminates redundant HEAD calls for object metadata. 50 MB default budget (~125K–150K entries), 10-minute TTL. Populated on PUT, HEAD, and LIST+metadata=true. Consulted on HEAD, GET, and LIST (including for file_size correction on delta-compressed objects). Invalidated on DELETE and prefix delete. Configurable viaDGP_METADATA_CACHE_MBenv var ormetadata_cache_mbTOML setting.
Security Hardening#
Tier 1 — Authentication & Session Security#
- Rate limiting: Per-IP token bucket rate limiter on auth endpoints — 5 attempts per 15-minute window, 30-minute lockout after exhaustion. Prevents brute-force attacks on admin login.
- Session IP binding: Admin sessions are bound to the originating IP address. Requests from a different IP are rejected even with a valid session token.
- Session concurrency cap: Maximum 10 concurrent admin sessions. Oldest session evicted when the limit is reached.
- Configurable session TTL: Default reduced from 24h to 4h. Override with
DGP_SESSION_TTL_HOURS. - Password quality enforcement: Min 12 chars, max 128 chars, common password blocklist. Validated on both admin API and CLI password set flows.
- SigV4 replay detection: Duplicate signatures within a 5-second window are rejected to prevent request replay attacks.
- Presigned URL max expiry: Capped at 7 days (604,800 seconds), matching AWS S3.
- Configurable clock skew:
DGP_CLOCK_SKEW_SECONDS(default 300s) controls SigV4 timestamp tolerance.
Tier 2 — Response Hardening & Anti-Fingerprinting#
- Security response headers: All responses include
X-Content-Type-Options: nosniffandX-Frame-Options: DENY. HSTS header added when TLS is enabled. - Anti-fingerprinting: Debug/fingerprinting headers (
Server,x-amz-storage-type,x-deltaglider-cache) suppressed by default. Enable withDGP_DEBUG_HEADERS=true. - Bootstrap password TTY safety: Auto-generated bootstrap password displayed in plaintext only when stderr is a TTY. Hidden in container/CI/piped output to prevent credential leaks in log aggregators.
- Multipart upload limits: Concurrent multipart uploads capped at 100 (configurable via
DGP_MAX_MULTIPART_UPLOADS) to prevent resource exhaustion.
Usage Scanner#
- Background prefix size scanner:
/_/api/admin/usageendpoint computes prefix sizes asynchronously with 5-minute cached results, 1,000-entry LRU cache, and 100K-object scan cap per prefix.
v0.4.0#
Single-Port Architecture#
- UI served at
/_/: The embedded admin UI and all admin APIs are now served under/_/on the same port as the S3 API. No more separate port (was port+1). The/_/prefix is safe because_is not a valid S3 bucket name character. Health, stats, and metrics endpoints are available at both root (/health) and under/_/(/_/health).
IAM & Authentication#
- Bootstrap password: Renamed from "admin password". A single infrastructure secret that encrypts the SQLCipher config DB, signs admin session cookies, and gates admin GUI access in bootstrap mode. Auto-generated on first run. Backward-compatible aliases (
DGP_ADMIN_PASSWORD_HASH,--set-admin-password) still work. - Multi-user IAM (ABAC): Per-user credentials stored in encrypted SQLCipher database (
deltaglider_config.db). Each user has access key, secret key, and permission rules with actions (read,write,delete,list,admin,*) and resource patterns (bucket/*). Admin = wildcard actions AND wildcard resources. - IAM mode auto-activation: When the first IAM user is created, the proxy switches from bootstrap mode to IAM mode. Bootstrap credentials are migrated as "legacy-admin" user. Admin GUI access becomes permission-based (no password needed for IAM admins).
- Admin API:
/_/api/admin/usersCRUD,/_/api/admin/users/:id/rotate-keys,/_/whoami,/_/api/admin/login-asfor IAM user impersonation.
S3 API Compatibility#
- Range requests:
Rangeheader support with 206 Partial Content responses,Accept-Ranges: bytesheader on all object responses. - Conditional headers:
If-Match/If-Unmodified-Since(412 Precondition Failed),If-None-Match/If-Modified-Since(304 Not Modified). - Content-MD5 validation: Validates
Content-MD5header on PUT and UploadPart, rejects with 400 on mismatch. - Copy metadata directive:
x-amz-metadata-directive: COPY(default) orREPLACEon CopyObject. - ACL stubs: GET/PUT
?aclaccepted and ignored for SDK compatibility. - Response header overrides:
response-content-type,response-content-disposition,response-content-encoding,response-content-language,response-expiresquery parameters on GET. - Per-request UUIDs:
x-amz-request-idheader with unique UUID on every response. - Bucket naming validation: Extracted
ValidatedBucketandValidatedPathextractors for automatic S3 path validation. - ListObjectsV2 improvements:
start-afterparameter,encoding-typepassthrough,fetch-ownersupport, base64 continuation tokens, max-keys capped at 1000. - Real creation dates:
ListBucketsreturns actual bucket creation timestamps.
Performance#
- Lite LIST optimization: LIST operations no longer issue per-object HEAD calls. Sizes shown are stored (compressed) sizes. ~8x faster for large listings.
- FS delimiter optimization:
list_objects_delegated()for filesystem backend uses a singleread_dirat the prefix directory instead of a recursive walk when a delimiter is specified. Dramatically faster for buckets with many prefixes.
Security#
- OsRng for tokens: Session tokens and IAM access keys use
OsRng(cryptographically secure) instead ofthread_rng. - DB rekey verification: Bootstrap password changes verify the new key can open the database before committing.
- Proper transactions: IAM user creation uses database transactions for atomicity.
Code Quality#
S3Openum (storage/s3.rs): Operation context for S3 error classification, replacing string-based operation names.- Session cookie helpers (
session.rs): Extracted session store into its own module withOsRngtoken generation. env_parse()DRY (config.rs): Extracted environment variable parsing boilerplate into reusable helpers.- Dead code cleanup: Removed
AdminGate, unused#/settingsroute, deadUsersTabandUserModalcomponents.
UI Features#
- File preview: Double-click on any previewable file (text, images) to view inline via the inspector panel. Tooltip indicates previewable files.
- Show/hide system files: Toggle to show or hide DeltaGlider internal files (
.dg/directory contents) in the object browser. - Folder size computation: Folder sizes computed and displayed in the object table.
- Delete user confirmation: User deletion requires
window.confirmdialog. - Full-screen admin overlay: Admin settings now use a full-screen overlay with master-detail layout for user management.
- Interactive API reference: New
#/docspage with interactive API documentation. - Key rotation safety: Prevents self-lockout on key rotation; changing only the secret key no longer regenerates the access key.
- Credentials display: After creating a user, shows only the credentials with a dismissible banner before returning to the user list.
v0.3.0#
S3-Compatible Endpoint Support#
- Disabled automatic request/response checksums: AWS SDK for Rust (like boto3 1.36+) adds CRC32/CRC64 checksum headers by default. S3-compatible stores (Hetzner Object Storage, Backblaze B2, some MinIO configs) reject these with BadRequest. Now sets both
request_checksum_calculationandresponse_checksum_validationtoWhenRequired. (Port of Python DeltaGlider [6.1.1] fix.) - Retry with exponential backoff on PUT: Hetzner returns transient 400 BadRequest errors (~1-2% of requests) with
connection: closeand no request-id. PUT operations now retry on 400 and 503 with 100/200/400ms backoff (3 retries). Also retries on network/timeout errors. - Verbose S3 error logging: Every S3 error now logs operation, bucket, HTTP status code,
x-amz-request-id, and full error details for production debugging.
Unmanaged Object Support (Fixes #3, #4)#
Objects that exist on the backend storage but were never stored through the proxy (no DeltaGlider metadata) are now fully accessible:
- S3 backend: HEAD, GET, and LIST now return fallback metadata from the S3 HEAD response (size, ETag, Last-Modified) instead of 404
- Filesystem backend: HEAD, GET, and LIST now return fallback metadata from filesystem stats (size, mtime) instead of 404
- HEAD/GET consistency: Both operations return metadata from the same source, ensuring consistent Content-Length and ETag
- Delta and reference files: Fallback metadata also works for delta/reference files without metadata (xattr or S3 headers)
- Corrupt metadata recovery: S3 objects with partial/corrupt DG headers now fall back to passthrough instead of hard-failing
Error Handling Hardening#
- Error discrimination: Replaced blanket
.ok()andmap_err(|_| NotFound)patterns with explicit error matching throughout the engine.NotFound→ expected (object doesn't exist),Io→ warn + retry path (concurrent access), other errors → propagate as 500 - ENOENT classification:
io_to_storage_error()now maps file-not-found I/O errors toStorageError::NotFoundinstead ofStorageError::Io, preventing false 500s for missing files - Filesystem xattr errors: Only fall back to filesystem stats on
NotFound; permission denied and other I/O errors are now propagated instead of silently swallowed - Reference cache errors:
get_reference_cached()now discriminatesNotFound(→ MissingReference) from other storage errors (→ Storage)
Security#
- SigV4 clock skew validation: Regular (non-presigned) SigV4 requests now enforce a 15-minute clock skew window, matching AWS S3 behavior. Prevents replay attacks with arbitrarily old timestamps. Returns new
RequestTimeTooSkewederror (403). - Reserved filename validation: PUT requests for
reference.binand*.deltakeys are rejected with 400 to prevent collision with internal storage files
Reliability#
- Codec subprocess timeout: xdelta3 subprocess now has a 5-minute timeout via
try_waitpolling loop. Kills hung processes and returns an error instead of blocking indefinitely. - Copy object size check:
copy_objectnow verifies actual data size after retrieval, catching cases where fallback metadata reportsfile_size=0that would bypass the pre-copy size check - S3 metadata size validation: Rejects PUT if DeltaGlider metadata headers exceed S3's 2KB limit, instead of letting the upstream S3 return an opaque 400
- Config validation: Warns on
max_delta_ratiooutside [0.0, 1.0] andmax_object_size=0at startup - Cache invalidation ordering: Reference is now deleted from storage BEFORE cache invalidation, preventing concurrent GET from re-caching a stale reference between invalidation and deletion. Fixed in 3 code paths (passthrough fallback, deltaspace cleanup, legacy migration).
DRY Cleanup#
FileMetadata::fallback(): New constructor consolidates 4 duplicate fallback metadata construction sites across S3 and filesystem backendsEngine::validated_key(): Extracts the 5x repeatedObjectKey::parse + validate_object + deltaspace_idpatterntry_unmanaged_passthrough(): Extracted 60-line nested match block fromretrieve_stream()into a focused helper with flat control flow
Testing#
- 26 new tests (244 total, up from 218): unmanaged object operations, HEAD/GET metadata consistency, reserved filename rejection, copy with unmanaged sources, error discrimination (
io_to_storage_errorunit tests), delta byte-level round-trip integrity, multipart ETag format, user metadata round-trip, external file deletion → 404
Infrastructure#
- Docker build: native ARM64 on Blacksmith runners (no QEMU), cargo-chef for dep caching — ~5x faster builds
- RustSec: updated deps to fix 6 advisories in aws-lc-sys and rustls-webpki
v0.2.0#
Cache Health Observability#
Four layers of defense against silent cache degradation:
- Startup warnings:
[cache]log prefix warns when cache is disabled (0 MB) or undersized (<1024 MB) - Periodic monitor: Every 60s, warns on >90% utilization or >50% miss rate
- Prometheus metrics:
cache_max_bytes,cache_utilization_ratio,cache_miss_rate_ratio— computed on scrape from existing atomic counters - Response header:
x-deltaglider-cache: hit|misson delta-reconstructed GETs - Health endpoint:
/healthnow includescache_size_bytes,cache_max_bytes,cache_entries,cache_utilization_pct
Proxy Dashboard#
Full Prometheus metrics dashboard in the built-in React UI (#/metrics):
- Top KPIs: uptime, peak memory, total requests, storage savings %
- Cache section: utilization gauge, hit rate with color coding, live hits vs misses chart
- Delta compression: encode/decode latency, compression ratio distribution, storage decisions
- HTTP traffic: operation breakdown (bar + donut chart), latency distribution, status codes, live request rate
- Authentication: success/failure counts with failure reason breakdown
- Auto-refresh every 5s, storage stats every 60s
Correctness Fixes (11 bugs)#
- Codec stderr deadlock: xdelta3 stderr was piped but only drained after stdout. If xdelta3 writes >64KB to stderr, stdout reader blocks forever. Now drains all 3 pipes concurrently.
- Filesystem metadata-data split-brain: Crash between
atomic_writeandxattr_meta::write_metadataleft files without metadata. Now writes xattr to temp file before rename — atomic visibility. - Auth exemption path normalization:
/health/(trailing slash) was not matched by the exact-string exemption. Now strips trailing slashes. - SigV4 presigned URL case-sensitivity:
x-amz-signature(lowercase) was not excluded from canonical query string. Now uses case-insensitive comparison. - Stats cache thundering herd: Lock released before
compute_stats, so N concurrent requests all scanned storage. Now holdstokio::sync::Mutexacross the async compute. - Multipart double-completion race: Two concurrent
CompleteMultipartUploadcalls both read parts under read lock, both stored data. Now takes ownership under write lock atomically. - Multipart write lock starvation: Assembly of large uploads held write lock during memcpy of all parts. Now removes upload from map (fast), releases lock, then assembles.
- AWS chunked truncated payload: Decoded length mismatch with
x-amz-decoded-content-lengthwas logged but data stored anyway. Now rejects with 400. - Admin config rollback: Backend config was committed before engine swap. If
DynEngine::new()failed, config showed new backend but engine was old. Now rolls back on failure. - S3 403 misclassification: All 403 errors mapped to
BucketNotFound. Now only maps for bucket-level operations; object-level 403 is reported as S3 error. - list_objects max_keys=0: Produced
is_truncated=truewith no continuation token. Now clamps to >= 1.
DRY & Code Quality#
StoreContextparameter object: Eliminated 3#[allow(clippy::too_many_arguments)]suppressions fromencode_and_store,store_passthrough,set_reference_baselinewith_metrics()helper: Collapsed 8 inlineif let Some(m) = &self.metricsblocks to one-linerstry_acquire_codec(): Extracted codec semaphore acquisition with consistent error messagecache_key(): Extractedformat!("{}/{}", bucket, deltaspace_id)used 5 timesdelete_delta_idempotent()/delete_passthrough_idempotent(): Extracted 5 verbosedelete_ignoring_not_foundcall sitespaginate_sorted(): Extracted duplicated pagination logic fromlist_objectspipe_stdin_stdout_stderr(): Extracted duplicated codec pipe coordination from encode/decodeobject_url(): Extracted test helper URL constructionparse_env()/parse_env_opt(): Extracted config env var parsing boilerplateheader_value(): Renamed from cryptichval()_guard: Renamed from misleading_lock(the Drop impl is load-bearing)- Trimmed verbose
PERF:archaeology comments — kept constraints, removed history - Removed
Option<Arc<Metrics>>fromAppState— metrics are always present, eliminated branch per request - Replaced
format!("%{:02X}")withwrite!()in SigV4 URI encoding — no heap allocation per encoded byte
Infrastructure#
/statsendpoint: 10s server-side cache, capped at 1,000 objects withtruncatedindicator/health,/stats,/metricsexempted from SigV4 authentication- Demo server exposes
/health,/stats,/metricsalongside admin API prefixed_key()helper in S3 backend eliminates 3x prefix if/else duplication
v0.1.9#
Initial public release with S3-compatible proxy, delta compression, filesystem and S3 backends, SigV4 authentication, multipart uploads, embedded React UI, and Prometheus metrics.