Stage every successfully written file under a private 0700 .fastsync-stage
directory inside the receive root and atomically publish all staged files
only after the whole protocol stream (manifest/delete handling included)
has completed, immediately before the success/outcome frame. On any
abort/error before publication nothing is installed and staging is removed;
a publish failure aborts the transfer with best-effort cleanup of the
remainder (already-published files are not rolled back). Crash leftovers
are wiped when the next delayed transfer starts.
Wire: new delay_updates config flag (selection-options block), protocol
version bumped to 2.6.0, client/server validation rejects --inplace.
CLI/usage/validation updated. Works in single-threaded and -m modes
(exactly one write_thread stages files; the staged-file registry is
mutex-protected; publication runs once after both threads join).
--existing/--ignore-existing/--update decide against the final destination
at stage time; --backup is deferred to publication. remove_source_files
outcomes are only sent after publication so skipped/unpublished sources are
never deleted. Default (no flag) behavior is unchanged.
Tests: config wire round-trip, CLI parse, --inplace rejection, new
test_delay_updates unit suite (27 suites total), and integration
TestDelayUpdates covering single/-m parity, incremental reruns, remove
source files, receiver-skip ordering, and a deterministic publish-failure
abort path.
Receiver writes temps into a confined scratch dir and atomically renames
into place; EXDEV aborts; inplace/partial bypass; thread-safe temp names.
Uses existing wire field; no protocol bump. Reviewed (c-review APPROVE
WITH NITS, all fixed); PR #262.
Address c-review nits on the itemize/output feature:
- RSYNC_COMPAT.md: state that %b is the source length (always == %l) because
no wire-byte counter exists; keep Summary equal to the matrix (recounted:
54 implemented / 84 not-implemented, 147 rows total - four rows flipped).
- change_list.h/.c: document bytes_sent == size; note itemize/out-format lines
never interleave with each other but may interleave with legacy log
messages sharing the stream; mark the %M stat() path best-effort.
- change_render_format scan in format_uses_mtime now mirrors the tokenizer
(skips '%%' and unknown '%X' pairs) so a literal '%%M' no longer triggers
the stat() fallback.
- Integration tests: --list-only under -m; a changed file on a second
--incremental run emits exactly one '>f' line while unchanged files print
nothing; --log-file + --log-file-format under -m.
Add a rootless unit test that reaches the actual st_dev skip branch in both
the sequential and parallel (-m) scanners: a symlink nested under the scan
root points at a directory on /dev/shm (a different device than the build
fs) and, under --copy-links semantics, -x must drop that subtree while a
plain scan includes it. Skips only when no cross-device target exists.
Integration OneFileSystem test now cleans both dest dirs up front and
reports a busy test mountpoint instead of ignoring the umount result.
RSYNC_COMPAT.md notes that cross-filesystem mount-point subdirectories are
dropped entirely (rsync parity).
Implement rsync-style itemized output backed by one shared change-event
engine (src/client/change_list.c):
- -i/--itemize-changes prints ">f+++++++++ <path>" for files actually sent
(single-threaded and -m); unchanged files print nothing.
- --list-only prints an ls-style listing of files that would be transferred
without contacting the server or writing anything.
- --out-format=FORMAT prints a printf-style template per changed file
(tokens %%f %%n %%l %%b %%M %%%%; unknown escapes preserved).
- --log-file-format=FMT logs each transferred file when --log-file is set.
Events are emitted from the per-file sender path shared by both transfer
modes, so the single sender thread is the only reporter (no races).
Capture the transfer root's device (st_dev) at scanner creation and skip
descending into any subdirectory on a different device (a mount point).
Implemented sender/client-side only: sequential BFS and parallel (-m) root
scan apply the same scanner_same_filesystem decision; no wire/protocol change
and default behavior is unchanged. Unit tests cover the pure decision, same
device scanning in both modes, and CLI parsing; integration tests prove -x
leaves a single-filesystem tree byte-identical and, when root can mount a
tmpfs, skips a genuine cross-device subtree.
Review nits from independent review of the four fix branches:
- receive_delta_file STATUS_NEXT oversize branch now sets *failed=true
- receive_incremental_check oversize branch returns NULL (receiver sends the
single STATUS_ERROR) instead of double-sending
- add multithreaded -m --ignore-existing --remove-source-files integration
coverage so the writer-thread outcome path is exercised
- Correct misleading doc comments on set_string_option /
set_positive_int_option / set_nonneg_int_option (they return 0/-1,
not true/false).
- find_table_option_with_equals() now matches every OPTION_TABLE value
option (OPT_STRING/OPT_POS_INT/OPT_NONNEG_INT/OPT_ULL), so forms such
as --max-size=2G, --min-size=1K, --suffix=.bak, --timeout=30,
--max-depth=5, --backup-dir=X parse instead of dying as 'Unknown option'.
--max-size/--min-size now accept rsync-style binary suffixes (0 remains a
valid 'no limit' byte count). Existing special handling for
--compress-choice, --compress-level, --modify-window=, --chmod=,
--skip-compress=, --compress-threads= is preserved.
- Options that require a separate value (-p, --exclude, --include,
--delta-block, --delta-max, --server-port, --bwlimit, --chunk-size,
--log-file, --exclude-from, --include-from, -T, --skip-compress,
--compress-threads) now emit an explicit 'missing argument' diagnostic
instead of falling through to the generic 'Unknown option' branch when
given as the final argv entry.
- Add unit tests covering the = forms (--max-size=2G, --min-size=1K,
--suffix=.bak, --timeout=30, --max-depth=5, --backup-dir=X) and a clean
'missing argument' (not 'Unknown option') diagnostic for trailing
--exclude/--server-port/--skip-compress/-T.
- remove-source-files keeps sources skipped by --existing/--ignore-existing/
--update (rsync reference behavior)
- --backup keeps <file>~, --suffix .bak, and --backup-dir backups
- --partial --partial-dir installs completed files in the destination
- a 100 MB file transfers end to end (>64 MiB whole-file cap regression)
#251 --remove-source-files deletes sources that were skipped receiver-side
(--existing/--ignore-existing/--update). The receiver now reports a
per-file outcome for every processed data file when the sender requests
removal; the client only unlinks sources the receiver actually wrote.
add remove_source_files to the wire config and bump the protocol to 2.5.0.
#252 --backup/--suffix/--backup-dir broken by NULL-vs-empty wire loss. Receivers
canonicalize the empty wire string back to NULL for backup_dir, temp_dir,
partial_dir and suffix, and --suffix is received unconditionally.
#253 --partial --partial-dir never installed completed files. file_save_to_disk
now renames a fully written partial-dir file into the real destination.
#255 STATUS_CHECK read the entire old file before the size/mtime quick check.
Old contents are only read when a checksum compare or delta needs them.
#256 receive_delta_file failure paths did not set *failed, so the caller sent
STATUS_NEXT and waited for a body that never came. Every NULL return now
marks the transfer failed.
#257 files >64 MiB could not transfer. Whole-file receive caps raised to the
256 MiB connection/allocation ceiling (chunk caps stay 64 MiB) and the
client ignores SIGPIPE so a server-side close surfaces as a clean error.
Unit tests added: config NULL-vs-empty round trip, incremental quick-check
skip/NEXT paths, delta oversize failure, partial-dir install, save-result
skip reporting.
The --inplace branch opened the destination with O_WRONLY|O_CREAT (no
O_TRUNC) and only restored metadata when the sender supplied it. Two
flaws resulted:
1. An existing destination file kept its original mode when no metadata
was sent, so setuid/setgid/sticky bits survived an overwrite (a root
sync could leave a root-owned setuid binary controlled by a client).
2. A shorter payload left stale trailing bytes from the previous version
because the file was never truncated to the new length.
In the inplace branch of file_to_disk_secure_impl:
- Always trim the file to the new payload length (ftruncate after the
write) so stale trailing bytes can never survive; sparse targets keep
their pre-size ftruncate.
- Always normalize the mode after a successful overwrite: apply the
metadata-derived safe mode when metadata is present (as before), else
fchmod to a safe default 0644, so setuid/setgid/sticky are cleared in
both cases.
- The --update newer-destination check still runs before any truncation
or chmod, preserving the skip semantics.
Adds unit tests in test_file.c: (a) setuid/sticky bits on an existing
destination are cleared after an inplace write with and without metadata,
(b) a shorter inplace payload leaves no trailing stale bytes.
The per-connection memory budget (MAX_CONNECTION_MEMORY, 256 MiB) only
charged wire buffers via receive_data_limited. Decompression buffers and
per-file chunk copies were not accounted for, and the multithreaded
receiver could enqueue up to 100 files (each up to 64 MiB uncompressed)
ahead of a slow disk writer, retaining ~6.4 GiB per connection. A client
sending highly compressible chunks with little bandwidth could OOM the
host while the reserve never tripped.
Bound the receive pipeline by aggregate payload bytes instead of item
count alone:
- Export MAX_CONNECTION_MEMORY from protocol.h.
- PipelineContextReceiver tracks queued_bytes (payload bytes received but
not yet released by the disk writer, i.e. queued or in the writer's
hand) under the existing mutex.
- receiver enqueue now blocks while the queue is full by count OR when
adding the file would push queued_bytes over the configured byte limit,
applying backpressure to the sender instead of failing the transfer.
- The disk writer releases the byte budget after each file is freed and
signals the not-full condition.
- The server sets the byte ceiling to
MAX_CONNECTION_MEMORY - 2*MAX_CHUNK_SIZE so that the queued payloads
plus the transient wire/decompression buffers of the one in-flight
chunk stay within the per-connection budget.
The single-threaded receive path is already bounded: it writes files to
disk before reading the next chunk, so its transient is at most one
chunk's wire + decompressed + copied payload (~3 * MAX_CHUNK_SIZE, below
the budget). Wire buffers remain charged exactly once by
receive_data_limited; this change does not double charge them.
Adds a deterministic unit test in test_multiprocessing.c proving that an
enqueue which would exceed the byte budget blocks until the writer
releases bytes.
The Phase 1 merge conflict resolutions introduced formatting that failed
the CI lint job (clang-format 18.1.3). Reformatted with the exact CI
version; no functional changes.
metadata_mtime_matches() required exact nanosecond equality at the default
modify_window=0, but destination write-time nsecs never match source
creation-time nsecs unless -M preserves metadata. Compare whole seconds
(rsync's default quick-check) at window 0; keep the subsecond-refined
tolerance for explicit window values.