Flip both rows to Implemented with precise Notes covering the tail-resume
model, the prefix-verify semantics (and the plain-append rsync-parity safety
statement), the new wire frames, and the PROTOCOL_VERSION 2.9.0 -> 2.10.0 bump.
Add a Phase-3 append-wave implementation-notes block. The Summary count line
is deliberately left untouched.
- CLI: --append/--append-verify acceptance (parse + imply --incremental,
validate) and incompatibility rejection with -s and --whole-file; both
removed from the unimplemented reject list.
- Config: on-the-wire append/append_verify round-trip.
- Unit: append_resume_eligible / append_tail_length pure resume math.
- Integration (test_append.py): matching-prefix resume is byte-identical and
tail-only (wire bytes << source size); --append with a wrong prefix keeps
prefix+tail (rsync parity) while --append-verify detects the mismatch and
falls back to a byte-exact full transfer; --append with --inplace and -m.
Implement rsync's append modes: when an existing destination file is SHORTER
than the source, the receiver negotiates a resume offset and only the tail is
transferred; the full file (retained prefix + tail) is rebuilt and installed
through the normal atomic store path, so the result is byte-identical to the
source whenever the prefix matches.
- --append: sends the tail without content-verifying the retained prefix
(rsync parity; the documented prefix-trust risk).
- --append-verify: verifies the retained prefix against the source's prefix
xxHash64 before appending and, on a mismatch, falls back to a clean full
transfer (never a corrupt prefix+tail blend).
New wire frames STATUS_APPEND / STATUS_APPEND_SIG / STATUS_APPEND_OK /
STATUS_APPEND_DATA; PROTOCOL_VERSION bumped 2.9.0 -> 2.10.0 (peers must match).
Both flags imply --incremental and are incompatible with -s (chunk
serialization) and --whole-file (rejected up front). Respects --inplace,
--partial/--partial-dir and --delay-updates via the shared store engine.
The suffix-trim loop using computed end offsets tripped cppcheck's
knownConditionTrueFalse value-range analysis (it unsoundly concluded the trims
always consume the whole middle). Rewrite it with explicit moving end indices
and add an inline suppression with a rationale for the residual false
positive; cppcheck --error-exitcode=1 is clean again. The trimming logic is
unchanged and was verified against a full DP reference over 200k random name
pairs.
The walker's all-or-nothing guarantee holds only while the destination is not
concurrently modified (rehearsal and delete are separate walks). The
protected-prefix list shares the 16 MB MAX_MANIFEST_BYTES budget with the
keep-set and each section is capped at MAX_MANIFEST_ENTRIES; an over-budget
frame is rejected on the receiver with STATUS_ERROR rather than truncated.
- ignore-errors scan-error test now runs single-threaded and under -m (the
exact scan_directory_multithreaded path that had the use-after-free);
- new integration test: --delete/--delete-before --ignore-errors with an
unreadable SOURCE ROOT (sequential and -m) must fail and delete NOTHING;
- new integration test: a destination-only file that merely matches an exclude
rule is deleted under plain --delete (protection is sender-derived), while a
source-excluded mirror is protected;
- delete-protects/delete-excluded coverage extended to --delete-after and
--delete-delay;
- new unit test: the 100000-entry server hard bound is all-or-nothing (more
extras than the bound -> nothing removed);
- new integration test: --force is inert under --delay-updates (documented);
- fixed the ignore-errors assertion message that stated the opposite of what it
asserted.
scan_directory_multithreaded read directory_scanner_had_io_error() /
parallel_scanner_had_io_error() AFTER destroying the scanner object (a
heap-use-after-free on every successful -m run that recorded an io_error); the
flag is now captured before the destroy. As a second line of defense against
an empty-keep-set wipe, all four manifest send sites (sequential and -m, early
pre-scan and commit data pass) now refuse to transmit a keep-set manifest when
the scan that built it recorded an io_error and produced no keep entries: a
source that merely LOOKS empty because part of it was unreadable must never
delete the whole destination. A genuinely empty source (no io_error) still
sends its empty keep-set and prunes extras.
A sequential scanner records an opendir failure of its seed/root directory as a
skippable io_error and would complete an EMPTY scan, whose keep-set manifest
would then delete every destination entry. The seed directory that maps to the
transfer root (relative path "") is now fatal regardless of --ignore-errors;
only subdirectories discovered during an otherwise-successful root scan are
skippable. The -m path never had this hole (its root open failure aborts
scanner creation), so sequential and -m now agree.
Unit (CLI): --fuzzy --no-incremental (either order) stays a valid plain-mode
config -- no forced handshake, no delta implication; --fuzzy --no-delta stays
covered.
Integration (TestFuzzy):
- worthless basis: a sibling that passes the name+size gates but shares no
blocks makes the sender reply STATUS_NEXT; the whole file is consumed inside
the delta handshake with byte-exact output (no protocol desync).
- non-displacement: an existing exact-path destination file INSIDE the delta
size bounds (same size, different content, older mtime) is used as the delta
basis instead of a byte-identical similar sibling (whole-file wire cost).
- asymmetric bases: basis larger than source (prefix reuse) and basis smaller
than source (appended tail as literals) both reconstruct byte-exactly with a
small delta.
- --no-fuzzy end-to-end equals the no-flag whole-file behavior.
CountingProxy closes its listener socket (fd hygiene).
Review follow-ups on the --fuzzy row: note that the receiver's block
signature is derived from a sibling file it may not otherwise send, exposing
destination sibling files to the sender at block granularity (the same
known-plaintext information class as the ordinary delta path); and document
that --fuzzy honors an explicit --no-incremental (unlike the basis-dir
options, which force it), with the delta implication suppressed accordingly.
The --fuzzy implication previously forced use_incremental back on even when
the user passed --no-incremental, while --no-delta and -W were honored.
Track a --no-incremental latch (like the no_delta latch): with it set, do not
force the handshake on, and because delta needs the handshake, also suppress
the delta implication so no invalid '--delta requires --incremental' config
results. A --fuzzy --no-incremental run is therefore a plain default-mode
transfer (fuzzy inert), consistent with -W/--no-delta. Documented in the
usage text (this deliberately differs from the basis-dir options, which still
force incremental unconditionally).
Review follow-ups on the --fuzzy candidate scan:
- Allocate the two DP rows once per directory scan instead of once per
candidate (4096-entry directories no longer do thousands of malloc pairs).
- Pre-prune before the DP with two cheap lower bounds on the edit distance:
the name length gap and the count of characters of one basename absent from
the other; a candidate whose gate (distance*2 <= longer) already fails on
the max of those bounds is skipped without running the DP.
- Trim the common prefix and non-overlapping common suffix before the DP so
it only runs over the differing middles.
- Note that the 4096 readdir cap bounds iterations, not per-entry DP cost,
and that the seen set is filesystem-order dependent (winner stays
deterministic via the total comparator).
- Open the chosen candidate with O_NONBLOCK so a name raced to a FIFO cannot
block the receive thread forever in open(2); the existing fstat S_ISREG gate
still rejects non-regular files. (The pre-existing basis_open_regular has
the same latent FIFO pattern and is intentionally left unchanged.)
- Free old_data in receive_delta_file's defensive NULL guard.
file_remove_tree_secure() opens the final component O_NOFOLLOW below the
authorized root and recursively wipes it with an fd-relative walk (symlinks are
removed by name, never followed); --force uses it to clear a destination
directory that blocks an incoming regular file.
--delete-excluded/--max-delete/--ignore-errors/--force/--prune-empty-dirs now
✅ with precise notes: the rsync-parity default (plain --delete protects
filter-excluded destination mirrors), the -m divergence (FastSync -m stays
multithreading, so --prune-empty-dirs is long-only), the all-or-nothing
max-delete/hard-bound error model, the 2.8.0 -> 2.9.0 protocol bump, and the
new two-section STATUS_MANIFEST frame. Summary counts left for the orchestrator
to recount after the wave.
Unit: walker all-or-nothing bounds (exceeded -> nothing removed + distinct
result; exact bound -> deletes), protected-prefix skipping, config wire
round-trip for force_delete/delete_excluded/prune_empty_dirs/max_delete, CLI
parse/validation for the new flags. Integration (TestDeletePolicy): default
delete-excluded protection and --delete-excluded opt-out (single-thread, -m,
early --delete-before, excluded-dir subtrees), --max-delete all-or-nothing over
and at the limit, --force file-over-nonempty-dir replacement, --prune-empty-dirs
(--dirs mode + recursion-mode parity), and --ignore-errors keeping deletion
active across a genuine scan I/O error (run as an unprivileged user).
The scanners now record every entry pruned by user-selection rules
(--filter/-C/per-dir, --exclude/--include, --max-size/--min-size) as a
destination-relative protected path on a caller-supplied sink (thread-safe in
the parallel scanner); --files-from subset pruning and -R relative wire paths
are never recorded. The sender transmits these as manifest protected prefixes,
giving rsync's default --delete behavior (excluded mirrors survive) with
--delete-excluded opting back into deleting them. --ignore-errors makes an
unreadable source directory a recorded, non-fatal scan error: the run continues,
the deletion still runs, and the exit code reports the ignored error.
--prune-empty-dirs omits an empty source directory's explicit --dirs entry.
Empty directories were never transferred by recursive scans (rsync -m parity).
The STATUS_MANIFEST frame now carries two count-delimited sections: the kept
paths and a protected-prefix list (excluded-on-source paths the walker must not
delete unless --delete-excluded opted out). The receiver's DeleteManifest is
passed through the commit/early paths unchanged. manifest_delete_extras
honors a client --max-delete (all-or-nothing) and produces a distinct error for
it versus the 100000-entry server bound. --force clears a non-empty directory
that blocks an incoming regular file (confined, symlink-safe) via a new
file_remove_tree_secure helper.
delete_extras_limited now rehearses a finite-capped deletion before unlinking
anything (an identical fd-relative walk that counts files and directories) and
returns DELETE_WALK_LIMIT_EXCEEDED with nothing removed when the run would
exceed the cap, so --max-delete is enforced per run instead of truncating the
deletion. A directory that still holds entries the walker leaves in place
(protected excluded prefix, manifest-kept file, symlink) is left behind rather
than failing the whole deletion, matching rsync's leave-non-empty-dirs
behavior. Rehearsal/delete each open an independent file description so a
prior pass cannot drain the directory stream.
Adds ignore_errors (client-only) and force_delete (wire) booleans, makes
max_delete default -1 (no client limit), and parses --delete-excluded,
--max-delete=NUM, --ignore-errors, --force, --prune-empty-dirs. Bumps
PROTOCOL_VERSION 2.8.0 -> 2.9.0 for the new on-the-wire force_delete field.
Document the receiver-side decision location, the exact deterministic
similarity heuristic, the byte-exactness argument, when fuzzy applies (and
when it deliberately does not), the 2.8.0 -> 2.9.0 protocol bump, and the
divergences from rsync. Update the basis-dir rows' protocol references to
the now-current 2.9.0.
Unit: parse -y/--fuzzy and --no-fuzzy negation (order-independent), -W and
--no-delta leave fuzzy inert, --fuzzy implies --incremental/--delta, and
validate_config rejects fuzzy with -s (chunk serialization) and -f
(sendfile). Config: fuzzy survives the socketpair wire round-trip.
Integration (TestFuzzy, byte counts via a new CountingProxy helper): the
rename case transfers a 2 MiB file in a few percent of its size with byte-
exact output under --fuzzy, while the no-fuzzy, no-candidate, dissimilar-
sibling and --whole-file runs send the whole file; -y alias; -m parity;
dest-holds-unsuitable-file falls back to the sibling; --delay-updates and
--remove-source-files keep their semantics on fuzzy transfers.
When a file must be transferred and the destination holds no usable content
at the exact path (file absent, or outside the delta engine's size bounds),
the receiver now searches the target's own destination directory for an
existing regular file with a similar basename and uses it as the delta basis
via the existing receiver-driven STATUS_DELTA_SIGNATURE handshake. The
sender never learns the basis was another file, so no wire change beyond the
new config flag was required.
Heuristic (deterministic, simpler than rsync's, documented): candidates are
sibling entries confined below the root and opened O_NOFOLLOW (symlinks are
never followed, nothing outside the destination root is read); dotfiles,
directories, the target's own name and stage/temp names are excluded; size
gate is delta_should_attempt; name gate is a Levenshtein distance <= half
the longer basename; the closest candidate (size tie-break, then lexical) is
loaded; the scan is capped at 4096 entries.
Byte-exactness is independent of the basis: block matches are verified by
Adler-32 + xxHash32, delta_apply validates every reference, and a basis that
shares nothing makes the sender reply with a whole-file transfer. When no
candidate qualifies the normal whole-file transfer runs unchanged.
Register --fuzzy (alias -y) as a boolean option and fuzzy as negatable so
--no-fuzzy works through the generic negation machinery. FastSync's delta
machinery is off by default (unlike rsync, where --fuzzy implies nothing
because delta is the default), so --fuzzy implies --incremental and --delta
unless --whole-file or an explicit --no-delta switched delta off (leaving
fuzzy inert, matching rsync where -W makes fuzzy irrelevant). Add help text.
-y/--fuzzy is receiver-side similar-file basis selection, so the receiver
must learn the flag: add a bool to Config, serialize it as a trailing field
on the config frame, and validate the received value. Because the config
frame layout changed, PROTOCOL_VERSION moves to 2.9.0 (peers must match).
The seed run did not preserve timestamps, so the destination copy's mtime was
the write time; the incremental --remove-source-files rerun only skipped the
file when both writes happened to land in the same whole second, making the
test flaky (observed intermittently in local full-suite runs and on CI). Seed
with -M so the destination stores the source's exact mtime.
- --link-dest destination entries share the basis inode: a later --inplace run
against such a path mutates the basis snapshot through the shared inode
(recommend --copy-dest when the destination must stay independently writable).
- basis-hit files take mode/uid/gid and mtime from the basis file, not the
sender's metadata (with --size-only the mtime can differ from the source).
- --remove-source-files sources satisfied by a basis dir are retained.
- basis runs refuse files above the 256 MiB whole-file limit up front (FastSync
caps every whole-file payload path at 256 MiB; rsync supports arbitrary sizes);
the config frame always carries a basis-count field (protocol 2.8.0).
- unit: config basis-path normalization (trailing slash, a//b, ./x/./y collapse;
degenerate inputs rejected).
- integration: same-size/same-mtime/different-content fixtures prove the xxHash
gate -- the basis changed.txt now has the SAME byte size as the source so the
size short-circuit can no longer mask the hash comparison, plus a dedicated
parametrized same-size mismatch test asserting link/copy never use a
content-mismatched basis and compare-dest transfers.
- basis-dir priority is first-match-wins: two link-dest dirs (inode of the
first), and compare-dest before link-dest stays sparse while the reverse
order hard-links.
- --size-only links a same-content basis file with a different mtime;
--ignore-times never links even an exact match.
- --delay-updates + --delete removes extras inside a nested .fastsync-stage
dir (regression guard) while keeping the real staging dir and basis tree.
- a basis run containing a file above the whole-file limit fails up front with
a clear error and transfers nothing.
Basis dirs imply the per-file incremental check, which (like every whole-file
payload path in FastSync) is bounded by MAX_RECEIVE_WHOLE_FILE_SIZE. A source
tree with a larger file used to abort the whole run mid-stream on the receiver
with no client-side diagnostic. With basis dirs configured the client now
preflights the scan (respecting filters/size rules) and fails up front with a
clear error naming the offending file before connecting, matching the
documented 256 MiB whole-file limit instead of aborting silently.
- copy-dest basis hits leaked the heap-allocated basis path: BASIS_DEST_COPY
did not transfer it (only LINK does) and returned before the basis cleanup.
basis_match_free is now called on every materialization return path (success
and send-failure) after content/link ownership is transferred.
- the --delete walker regression: the delay-updates staging name must be
protected only as a DIRECT child of the receive root, while basis dirs may
be skipped at any depth. delete_extras_limited now takes DeleteSkipEntry
entries carrying a top_level_only flag instead of a flat prefix list, so a
nested destination directory named .fastsync-stage is ordinary content again
(its extras are deleted) and a basis tree is still never removed.
config_basis_path_valid and config_basis_append now share one normalizer
(basis_path_normalize): interior empty components (a//b) collapse, '.'
components and trailing slashes are dropped, and the stored form is exactly
the canonical relative path used by validation, the delete-walker prefix match
and the receiver's basis lookup. Degenerate inputs (empty, absolute, '..',
'.' that normalizes to nothing) stay rejected.
Usage help now gives --delete-during a complete description with --del on its
own line, and notes that timing flags imply --delete while timing+--no-delete
is rejected regardless of argument order. RSYNC_COMPAT.md documents: the
receiver's MAX_MANIFEST_ENTRIES/MAX_MANIFEST_BYTES caps now abort an early-mode
run before any data (previously only the deletion step failed), the extended
early-delete ACK deadline, and the order-independent flag-conflict policy.
- Unit (leak guards): drive receiver_process_pending() past a parked keep-set
into STATUS_ABORT, EOF, and a second manifest frame; each must return -1 with
no manifest handed out. Verified leak-free under ASan.
- Unit: receive_status_timed reads a status and fails cleanly on EOF.
- Integration: test_late_flags_commit_only_after_success now parametrizes the
-m path, proving a failed -m late-timing run preserves every extra and that
the deferred manifest is dropped (never applied) when the writer fails.
The receiver performs the whole bounded deletion walk (up to
MAX_SERVER_DELETE_COUNT unlinks) before answering the delete-before/during
manifest, so its STATUS_OK reply can take far longer than the default 60 s
per-message receive window. Waiting with the default would make the sender
abort AFTER the deletion had already committed on the receiver. Add a timed
receive variant (receive_status_timed / protocol_receive_n_data_timed) and use
it for the early-manifest ACK with a 1 h explicit deadline; connection errors
and EOF still abort immediately.
The late/commit path keeps the received keep-set in a local list until
STATUS_FINISHED. Error exits after it was parked (STATUS_ABORT, a failing
receive_status / non-FINISHED status, a second manifest frame, or a later
file/chunk/store failure) previously dropped the only reference and leaked up
to ~16 MB of path strings + pointer array per connection. Both failure labels
now discard the parked list exactly once; the successful FINISHED path still
hands ownership to *pending_manifest (the -m caller) without freeing it.
RSYNC_COMPAT.md: flip --delete-before, --del/--delete-during, --delete-delay
and --delete-after to Implemented with precise notes (default-under--delete,
safety model, 2.7.0 -> 2.8.0 protocol bump, exact divergences from rsync).
README option tables list the new flags and the delete-after default.
- CLI: each timing flag (+ --del alias) accepted and implies --delete;
conflicting timings and a timing with --no-delete are rejected.
- Config: delete_during/delete_delay survive config_send/config_receive;
two simultaneous timings are rejected by the receiver-side validation;
config_delete_timing_early() mapping is unit-tested.
- Integration (TCP, single- and multithreaded): every flag removes extras on
a successful transfer; early modes (--delete-before/--delete-during/--del)
delete before data is applied so a destination file blocking a nested
write is removed and the transfer succeeds, while plain --delete /
--delete-after / --delete-delay keep it and fail with every extra intact
(commit-style). Early timing also completes (without deleting) when the
server refuses deletion.
Deletion timing is now real and selected by the four rsync flags plus the
plain --delete default. Wire protocol bumps to 2.8.0: two new config
booleans (delete_during, delete_delay) are serialized and validated, joining
the existing delete_before/delete_after.
- Early modes (--delete-before, --delete-during/--del): the sender pre-scans
the whole tree (paths only), transmits the keep-set manifest BEFORE any
file data, and the receiver removes extras and acks STATUS_OK; the sender
only streams data after the deletion committed. Deletion is thus performed
even if a later transfer phase fails (rsync delete-before/during are
destructive by definition). FastSync streams in a single scan so it cannot
interleave per-directory like rsync delete-during; --delete-during selects
the same engine mode as --delete-before (documented divergence).
- Late/commit modes (plain --delete, --delete-after, --delete-delay): the
manifest closes the data stream and deletion is committed only after
STATUS_FINISHED proves the whole transfer succeeded, preserving FastSync's
commit-style safety. --delete-delay converges with --delete-after because
FastSync never snapshots the destination during data flow (documented).
- The STATUS_MANIFEST frame is now self-delimiting and position-independent.
Single-threaded receivers delete before the success frame; the -m receiver
hands the keep-set to server.c, which commits the deletion only after the
disk writer thread has drained (fixes a delete-vs-in-flight-temp race).
- Every timing flag implies --delete; at most one timing flag is allowed.
- Each timing flag implies --delete, matching rsync; conflicts are rejected.
Document receiver-side basis semantics, relative-to-destination-root
confinement, content-verified matching, hard-link vs copy vs compare-only
behavior, cross-filesystem copy fallback, repetition/priority, --delete
exclusion, the implied --incremental and -s incompatibility, and each exact
divergence from rsync (no dest deletion on compare-dest stale entries, no
attribute re-application on basis hits, no re-linking of already-up-to-date
dest files, sources matched from basis dirs kept under --remove-source-files).
- compare-dest skips an exact basis match (leaving a sparse destination) and
still transfers files the basis cannot satisfy; a content mismatch forces a
normal transfer.
- copy-dest materializes the unchanged file as a real local copy (distinct
inode) and transfers content mismatches.
- link-dest hard-links (asserted same inode/nlink to the DIR file) and falls
back to a normal transfer on content mismatch.
- a missing basis dir is a clean full-transfer no-op for all three flags.
- link-dest works through -m multithreading and --delay-updates (staged and
published as a real link, staging cleaned up).
- --delete removes genuine extras while leaving the basis dir untouched.
- parse_args accepts each flag in both forms, keeps repetition order/types,
implies --incremental + metadata, and rejects absolute/escaping/degenerate
paths; validate_config rejects basis dirs combined with -s.
- config wire round-trips a mixed basis list and rejects escaping/absolute
paths on the receiver side.
Generalize the delete walker's protected-root-child skip into a prefix list.
The receiver now passes both the --delay-updates staging directory and every
basis-dir path, so a --delete run can never treat a basis snapshot (which a
--link-dest run just linked from) as destination content to remove.
The receiver's per-file incremental check now consults the ordered basis-dir
list whenever the destination is not already up to date. An exact basis match
requires equal size, equal mtime (unless --size-only; --ignore-times disables
basis matching like rsync), and an equal content xxHash64 -- the sender sends
its xxHash for every file whenever basis dirs are configured (not only under
--checksum), so a hard link or local copy is only ever made from byte-identical
content.
On a match:
- compare-dest: reply STATUS_OK and skip data only when the destination does
not already hold the file (sparse, rsync parity). A destination that holds
a DIFFERENT version falls back to a normal transfer instead of rsync's
delete, keeping the mirror complete.
- copy-dest: reply STATUS_OK and hand a synthetic File (bytes read from the
basis file, basis metadata) to the normal store sink, so the file is
installed as a real local copy through the existing atomic temp+rename
engine and honors --existing/--ignore-existing/--update/--backup/
--delay-updates/--partial-dir unchanged.
- link-dest: same, but File.basis_link records the basis path and the store
engine calls the new file_to_disk_secure_link(): an atomic temp hard link +
rename. Cross-filesystem/refused links fall back to a byte-identical local
copy (never a corrupt or partial file); the copy fallback applies metadata,
while a successful link keeps the basis inode's own attributes so the basis
file is never mutated.
Basis-materialized files carry File.skip so they are not acknowledged to a
--remove-source-files sender (the sender already saw STATUS_OK and keeps the
source). The no-match path is byte-for-byte identical to the existing delta /
full-data transfer.
Replace the vestigial single compare_dest/copy_dest/link_dest Config fields with
an ordered BasisDest list (type + path per entry) that is serialized to the
receiver and interpreted relative to the destination root. Paths must be
relative with no '.'/'..' components (confined like --backup-dir); trailing
slashes are normalized. Wire layout changes, so PROTOCOL_VERSION -> 2.8.0.
CLI: each flag is parsed in both --flag=DIR and --flag DIR forms, is
repeatable, and keeps command-line order as basis priority. Supplying any
basis dir implies --incremental (and therefore metadata) on the sender because
the unchanged decision is receiver-side; combining basis dirs with -s chunk
serialization is rejected in validate_config. Usage text updated.
After the writer releases bytes, the blocked enqueuer may already have
admitted the next payload, so the transient queued_bytes==0 read is
scheduling-dependent (lost the race under ASan). The post-join state
(covers unblocking + budget re-limit) is deterministic.
scan_root_entry str_dup'd the entry name for every root-level file but only
consumed it on the -R + --files-from path, leaking 1 string per root file in
normal parallel scans (found by CI ASan/valgrind).
B1: destination-root existence/creation mishandled two legitimate forms.
file_directory_exists_secure/file_ensure_directory_secure now normalize a
trailing-slash destination (so the last component is never an empty leaf)
and treat a destination equal to the already-open authorized root as present
(no spurious <root>/<basename> nested dir with --mkpath). Confinement and
O_NOFOLLOW probing are unchanged. Regression integration tests: existing
dest with trailing slash works with and without --mkpath; dest == authorized
root works and creates no stray nested dir.
W1: chunk serialize/deserialize round-trip unit test for a directory entry
mixed with a regular file, with and without metadata; --dirs coverage under
-s and -s -m.
W2: --list-only and --dry-run now print file_wire_path() so all outputs show
the -R transformed relative name, matching -i/--out-format.
W3: RSYNC_COMPAT -d/--dirs note: dirs are created immediately under
--delay-updates (only regular files are staged).
W4: --dirs generator flushes chunks by element count too, so a long
--files-from list of empty directories cannot exceed the per-chunk file cap.
N1: STATUS_MKDIR added to status_to_string.
N2: removed dead TestMkpath._transfer.
N3: removed redundant --no-implied-dirs OPTION_TABLE row.
N4: integration tests: --dirs --delete keeps the just-created empty dir (and
deletes extras); a listed dir colliding with a regular file at the dest fails
cleanly.
RSYNC_COMPAT Phase-2 row M: path-list construction and destination
directory creation while preserving traversal safety.
- -R/--relative with --files-from: transmit each listed entry under its
bare relative destination path (no source-root mirror). Files keep an
absolute local read path plus a separate wire/dest path (File.send_path);
manifest, incremental quick-check and change output follow the wire path,
so --delete and --remove-source-files stay consistent. -R without
--files-from is unchanged (full mirror).
- --no-implied-dirs: client-only, only meaningful with -R + --files-from.
A listed file whose parent dir is not itself (or via an ancestor)
explicitly listed cannot be placed; the run fails up front with a clear
error. No effect otherwise.
- --dirs/-d + --old-dirs/--old-d aliases: -d <dir> transmits the source
root as an explicit empty directory entry (STATUS_MKDIR frame); with
--files-from listed dirs are created empty and listed files transferred,
never descending. Works single-threaded, -m (sequential scanner in the
-m scan thread) and chunk-serialization (per-file type marker).
Directory entries appear in the delete manifest.
- --mkpath: new wire bool; server creates the destination root (and missing
leading components under its authorized root) at connection start. A
missing destination root is now rejected by default.
- Protocol bumped to 2.7.0 (mkpath wire field + STATUS_MKDIR + chunk type
marker). All receive paths funnel through file_save_to_disk_full which
creates directories via the secure confined mkdir engine; dir entries are
excluded from --remove-source-files outcome acknowledgements on both ends.
- Unit coverage: CLI parse (relative/dirs aliases/mkpath/no-implied-dirs),
config round-trip (relative + mkpath), scanner --dirs non-recursion and
-R send_path (sequential + parallel), receiver dir-entry save.
- Integration coverage: TestRelativeFilesFrom, TestNoImpliedDirs, TestDirs,
TestMkpath (single and -m).
- RSYNC_COMPAT: 4 rows move to Implemented (Summary 67/3/5/1/71 = 147).
Receiver stages writes and publishes only after the full transfer succeeds
(single and -m, before the outcomes frame). delete walker skips staging;
backup-dir name reserved; flock serializes concurrent delayed sessions;
protocol bump to 2.6.0. c-review REQUEST CHANGES -> blockers fixed; PR #264.
Review fixes for --delay-updates:
- --delete no longer deletes the staged files: the delete walker gains a
skip_root_child parameter and receive_manifest passes DELAY_UPDATES_STAGING_DIR
when delay_updates is active, so deletion removes genuine extras while the
staging dir (a direct child of the receive root) is left for publication in
both single and -m modes.
- --backup-dir is rejected when it collides with the reserved internal staging
name .fastsync-stage (trailing slash normalized), in client validation and in
the received-config wire validation, preventing old backups from being
silently installed as new files.
- Staging dir is now held under an exclusive advisory flock for the whole
transfer (context lifetime): two simultaneous delayed transfers to one
destination root no longer share/destroy each other's staged data - the
second fails cleanly. Cleanup only touches the staging dir when this context
owns the lock, so a lock-contention failure cannot wipe a live session.
- Post-publish staging cleanup now returns/logs instead of discarding failures
(warning when the staging dir cannot be fully removed).
- Reworked the publish-failure integration test to exercise real mid-publish
semantics (top-level file published, nested rename fails, no rollback,
sources retained under --remove-source-files) and added integration tests for
--delete + --delay-updates ordering and reserved --backup-dir rejection.
- RSYNC_COMPAT note documents delete ordering, the reserved-name hazard, and
the concurrency guard.
Address an independent c-review of the files-from/filter feature:
- .rsync-filter precedence now matches rsync: evaluate the innermost
(current) directory's rules first, then ancestors, then the command-line
base (--filter/-C), so a deeper file's '+' can re-include what a shallower
'-' excluded (regression tests in both scan modes; single-thread and -m).
- --files-from: a listed entry missing on disk and an empty list are now hard
errors surfaced pre-transfer in send_files, send_files_multithreaded,
dry-run and --list-only; '.' (whole tree) and empty listed dirs stay valid.
- scanner_path_relative now handles a transfer root of / (previously the
scanner aborted on children of /).
- Reject unsupported rsync filter syntax explicitly (no silent no-ops):
+/- modifiers other than '/' (! C s r p x) and rules beginning with ':'/'.'
/'!' (merge/dir-merge/list-clear shorthands). Docs updated.
- -0/--from0 NUL mode preserves entry bytes (no CR/LF trimming); only newline
mode trims. Absolute-entry error message no longer includes the newline.
- --no-from0/--no-cvs-exclude registered as negatable booleans.
- RSYNC_COMPAT rows updated for the precedence, rejection list, NUL-mode
detail and the documented O(entries x files) scalability bound of the
allow-set (Summary unchanged: 62/3/5/1/76 = 147).
Implement the RSYNC_COMPAT Phase-2 filter/parser feature group:
- --files-from=FILE (repeatable) plus -0/--from0 NUL delimiters: parse the
source file list relative to the source root into a shared read-only
allow-set; the scanner transfers listed files and the whole subtree of
listed directories and prunes everything else in single- and
multithreaded mode. Absolute/'..' entries and missing files are hard
CLI errors.
- --filter=RULE: rsync-style +/- rules (anchored '/', dir-only trailing '/',
word include/exclude forms) evaluated first-match-wins with a default of
include, as an independent layer from legacy --exclude/--include.
Unsupported directives (merge/hide/... ) are rejected explicitly. -f stays
sendfile.
- -C/--cvs-exclude: well-known rsync CVS default exclude set.
- -F: per-directory .rsync-filter files read during traversal and applied to
the owning directory's subtree (single + parallel), never transferred.
- Delete manifest still derives from what was actually sent.
Client-only config fields; no wire/protocol change. Adds unit coverage
(CLI parse, allow-set and filter scanning single+parallel) and integration
tests (TestFilesFrom, TestFilters). RSYNC_COMPAT matrix rows updated:
5 rows move to Implemented (Summary 62/3/5/1/76 = 147).
Stage every successfully written file under a private 0700 .fastsync-stage
directory inside the receive root and atomically publish all staged files
only after the whole protocol stream (manifest/delete handling included)
has completed, immediately before the success/outcome frame. On any
abort/error before publication nothing is installed and staging is removed;
a publish failure aborts the transfer with best-effort cleanup of the
remainder (already-published files are not rolled back). Crash leftovers
are wiped when the next delayed transfer starts.
Wire: new delay_updates config flag (selection-options block), protocol
version bumped to 2.6.0, client/server validation rejects --inplace.
CLI/usage/validation updated. Works in single-threaded and -m modes
(exactly one write_thread stages files; the staged-file registry is
mutex-protected; publication runs once after both threads join).
--existing/--ignore-existing/--update decide against the final destination
at stage time; --backup is deferred to publication. remove_source_files
outcomes are only sent after publication so skipped/unpublished sources are
never deleted. Default (no flag) behavior is unchanged.
Tests: config wire round-trip, CLI parse, --inplace rejection, new
test_delay_updates unit suite (27 suites total), and integration
TestDelayUpdates covering single/-m parity, incremental reruns, remove
source files, receiver-skip ordering, and a deterministic publish-failure
abort path.
Receiver writes temps into a confined scratch dir and atomically renames
into place; EXDEV aborts; inplace/partial bypass; thread-safe temp names.
Uses existing wire field; no protocol bump. Reviewed (c-review APPROVE
WITH NITS, all fixed); PR #262.
Sender-side scanner stays within the source filesystem (-x), single and
multithreaded; default unchanged; client-only, no wire change. Reviewed
(c-review APPROVE WITH NITS, all fixed); PR #261.
Address c-review nits on the itemize/output feature:
- RSYNC_COMPAT.md: state that %b is the source length (always == %l) because
no wire-byte counter exists; keep Summary equal to the matrix (recounted:
54 implemented / 84 not-implemented, 147 rows total - four rows flipped).
- change_list.h/.c: document bytes_sent == size; note itemize/out-format lines
never interleave with each other but may interleave with legacy log
messages sharing the stream; mark the %M stat() path best-effort.
- change_render_format scan in format_uses_mtime now mirrors the tokenizer
(skips '%%' and unknown '%X' pairs) so a literal '%%M' no longer triggers
the stat() fallback.
- Integration tests: --list-only under -m; a changed file on a second
--incremental run emits exactly one '>f' line while unchanged files print
nothing; --log-file + --log-file-format under -m.
Add a rootless unit test that reaches the actual st_dev skip branch in both
the sequential and parallel (-m) scanners: a symlink nested under the scan
root points at a directory on /dev/shm (a different device than the build
fs) and, under --copy-links semantics, -x must drop that subtree while a
plain scan includes it. Skips only when no cross-device target exists.
Integration OneFileSystem test now cleans both dest dirs up front and
reports a busy test mountpoint instead of ignoring the umount result.
RSYNC_COMPAT.md notes that cross-filesystem mount-point subdirectories are
dropped entirely (rsync parity).
Implement rsync-style itemized output backed by one shared change-event
engine (src/client/change_list.c):
- -i/--itemize-changes prints ">f+++++++++ <path>" for files actually sent
(single-threaded and -m); unchanged files print nothing.
- --list-only prints an ls-style listing of files that would be transferred
without contacting the server or writing anything.
- --out-format=FORMAT prints a printf-style template per changed file
(tokens %%f %%n %%l %%b %%M %%%%; unknown escapes preserved).
- --log-file-format=FMT logs each transferred file when --log-file is set.
Events are emitted from the per-file sender path shared by both transfer
modes, so the single sender thread is the only reporter (no races).
Capture the transfer root's device (st_dev) at scanner creation and skip
descending into any subdirectory on a different device (a mount point).
Implemented sender/client-side only: sequential BFS and parallel (-m) root
scan apply the same scanner_same_filesystem decision; no wire/protocol change
and default behavior is unchanged. Unit tests cover the pure decision, same
device scanning in both modes, and CLI parsing; integration tests prove -x
leaves a single-filesystem tree byte-identical and, when root can mount a
tmpfs, skips a genuine cross-device subtree.
Reconcile the security-fixes branch's RSYNC_COMPAT.md updates onto dev:
- Keep dev's implemented statuses as authoritative and port the previously
undocumented rsync 3.4.1 rows (one-file-system, -F, temp-dir/-T, identity
mapping options, remote-option/-M, max-alloc) with code-verified statuses.
- Add the Implementation Difficulty Plan (Phases 1-6 + delivery order) as the
sequencing roadmap; note that it is a superset snapshot and the Summary
matrix is authoritative for shipped features.
- Recomputed the Summary counts from the table (was stale on dev): 51
implemented / 3 alt-arg / 5 partial / 1 no-op / 87 not implemented.
Review nits from independent review of the four fix branches:
- receive_delta_file STATUS_NEXT oversize branch now sets *failed=true
- receive_incremental_check oversize branch returns NULL (receiver sends the
single STATUS_ERROR) instead of double-sending
- add multithreaded -m --ignore-existing --remove-source-files integration
coverage so the writer-thread outcome path is exercised
- Correct misleading doc comments on set_string_option /
set_positive_int_option / set_nonneg_int_option (they return 0/-1,
not true/false).
- find_table_option_with_equals() now matches every OPTION_TABLE value
option (OPT_STRING/OPT_POS_INT/OPT_NONNEG_INT/OPT_ULL), so forms such
as --max-size=2G, --min-size=1K, --suffix=.bak, --timeout=30,
--max-depth=5, --backup-dir=X parse instead of dying as 'Unknown option'.
--max-size/--min-size now accept rsync-style binary suffixes (0 remains a
valid 'no limit' byte count). Existing special handling for
--compress-choice, --compress-level, --modify-window=, --chmod=,
--skip-compress=, --compress-threads= is preserved.
- Options that require a separate value (-p, --exclude, --include,
--delta-block, --delta-max, --server-port, --bwlimit, --chunk-size,
--log-file, --exclude-from, --include-from, -T, --skip-compress,
--compress-threads) now emit an explicit 'missing argument' diagnostic
instead of falling through to the generic 'Unknown option' branch when
given as the final argv entry.
- Add unit tests covering the = forms (--max-size=2G, --min-size=1K,
--suffix=.bak, --timeout=30, --max-depth=5, --backup-dir=X) and a clean
'missing argument' (not 'Unknown option') diagnostic for trailing
--exclude/--server-port/--skip-compress/-T.
receive_files() in src/server/server.c was a non-static, unprototyped,
zero-caller duplicate of receiver_process()/receiver_receive_files() in
src/server/receiver.c. Delete it along with the includes it uniquely pulled
(chunk.h, metadata.h, sys/stat.h, duplicate quoted unistd.h); remaining code
still uses file.h/config.h/protocol.h (via multiprocessing.h) directly.
mkdir_r() in src/shared/utils.c had zero callers in src/ and tests/; remove
the function and its declaration in utils.h, plus the unused libgen.h include.
- remove-source-files keeps sources skipped by --existing/--ignore-existing/
--update (rsync reference behavior)
- --backup keeps <file>~, --suffix .bak, and --backup-dir backups
- --partial --partial-dir installs completed files in the destination
- a 100 MB file transfers end to end (>64 MiB whole-file cap regression)
#251 --remove-source-files deletes sources that were skipped receiver-side
(--existing/--ignore-existing/--update). The receiver now reports a
per-file outcome for every processed data file when the sender requests
removal; the client only unlinks sources the receiver actually wrote.
add remove_source_files to the wire config and bump the protocol to 2.5.0.
#252 --backup/--suffix/--backup-dir broken by NULL-vs-empty wire loss. Receivers
canonicalize the empty wire string back to NULL for backup_dir, temp_dir,
partial_dir and suffix, and --suffix is received unconditionally.
#253 --partial --partial-dir never installed completed files. file_save_to_disk
now renames a fully written partial-dir file into the real destination.
#255 STATUS_CHECK read the entire old file before the size/mtime quick check.
Old contents are only read when a checksum compare or delta needs them.
#256 receive_delta_file failure paths did not set *failed, so the caller sent
STATUS_NEXT and waited for a body that never came. Every NULL return now
marks the transfer failed.
#257 files >64 MiB could not transfer. Whole-file receive caps raised to the
256 MiB connection/allocation ceiling (chunk caps stay 64 MiB) and the
client ignores SIGPIPE so a server-side close surfaces as a clean error.
Unit tests added: config NULL-vs-empty round trip, incremental quick-check
skip/NEXT paths, delta oversize failure, partial-dir install, save-result
skip reporting.
The --inplace branch opened the destination with O_WRONLY|O_CREAT (no
O_TRUNC) and only restored metadata when the sender supplied it. Two
flaws resulted:
1. An existing destination file kept its original mode when no metadata
was sent, so setuid/setgid/sticky bits survived an overwrite (a root
sync could leave a root-owned setuid binary controlled by a client).
2. A shorter payload left stale trailing bytes from the previous version
because the file was never truncated to the new length.
In the inplace branch of file_to_disk_secure_impl:
- Always trim the file to the new payload length (ftruncate after the
write) so stale trailing bytes can never survive; sparse targets keep
their pre-size ftruncate.
- Always normalize the mode after a successful overwrite: apply the
metadata-derived safe mode when metadata is present (as before), else
fchmod to a safe default 0644, so setuid/setgid/sticky are cleared in
both cases.
- The --update newer-destination check still runs before any truncation
or chmod, preserving the skip semantics.
Adds unit tests in test_file.c: (a) setuid/sticky bits on an existing
destination are cleared after an inplace write with and without metadata,
(b) a shorter inplace payload leaves no trailing stale bytes.
The per-connection memory budget (MAX_CONNECTION_MEMORY, 256 MiB) only
charged wire buffers via receive_data_limited. Decompression buffers and
per-file chunk copies were not accounted for, and the multithreaded
receiver could enqueue up to 100 files (each up to 64 MiB uncompressed)
ahead of a slow disk writer, retaining ~6.4 GiB per connection. A client
sending highly compressible chunks with little bandwidth could OOM the
host while the reserve never tripped.
Bound the receive pipeline by aggregate payload bytes instead of item
count alone:
- Export MAX_CONNECTION_MEMORY from protocol.h.
- PipelineContextReceiver tracks queued_bytes (payload bytes received but
not yet released by the disk writer, i.e. queued or in the writer's
hand) under the existing mutex.
- receiver enqueue now blocks while the queue is full by count OR when
adding the file would push queued_bytes over the configured byte limit,
applying backpressure to the sender instead of failing the transfer.
- The disk writer releases the byte budget after each file is freed and
signals the not-full condition.
- The server sets the byte ceiling to
MAX_CONNECTION_MEMORY - 2*MAX_CHUNK_SIZE so that the queued payloads
plus the transient wire/decompression buffers of the one in-flight
chunk stay within the per-connection budget.
The single-threaded receive path is already bounded: it writes files to
disk before reading the next chunk, so its transient is at most one
chunk's wire + decompressed + copied payload (~3 * MAX_CHUNK_SIZE, below
the budget). Wire buffers remain charged exactly once by
receive_data_limited; this change does not double charge them.
Adds a deterministic unit test in test_multiprocessing.c proving that an
enqueue which would exceed the byte budget blocks until the writer
releases bytes.
The Phase 1 merge conflict resolutions introduced formatting that failed
the CI lint job (clang-format 18.1.3). Reformatted with the exact CI
version; no functional changes.
metadata_mtime_matches() required exact nanosecond equality at the default
modify_window=0, but destination write-time nsecs never match source
creation-time nsecs unless -M preserves metadata. Compare whole seconds
(rsync's default quick-check) at window 0; keep the subsecond-refined
tolerance for explicit window values.