Blockers addressed together (shared scanner/delete-plan plumbing):
* #10: an empty in-scope source directory produced no plan keep entry, so the
receiver deleted the destination directory itself. The scanner now records
every traversed directory into a delete-plan sink, the plan sender keeps them,
and any directory whose plan the data stream never triggered is emitted after
the data so its extras are still removed. Differential tests cover
--delete-during and --delete-delay.
* #8: an invalid per-directory filter file was silently ignored when an earlier
merge file in the same directory existed; key the failure off the error text
(both sequential and parallel scanners) and fail the scan.
* #9: -R + --files-from receiver-protect rules recorded the source-relative
path; record the bare relative wire path in both scanners so the protected
destination mirror survives --delete.
* #5: the STATUS_STATS would-delete parser now validates each retained path and
enforces the shared MAX_MANIFEST_BYTES budget, and the --out-format dry-run
delete line is escaped like the itemize line.
* #11: drop the unused DELETE_PLAN_MAX_NAMES macro, log the delete-limit
warning once per session, roll back dir-merge names from a per-directory file
that fails to parse, and guard every filter error snprintf against err==NULL.
#10 leaves the empty directory itself kept and its extras removed, matching
rsync's final state on both per-directory timings.
Stream one delete plan per source directory from sender to receiver instead of
a single whole-tree keep-set manifest:
- --delete-during applies each directory's extras as its plan arrives, before
that directory's data (rsync's generator-order deletion).
- --delete-delay snapshots each directory's extras while the plan arrives and
commits the removals only after a fully-successful transfer, so files created
after the scan survive (matching rsync's delete-delay, not delete-after).
- Type conflicts (a destination file blocking a source directory, or vice
versa) are cleared immediately in both modes, so the nested write succeeds.
The plan carries the destination-relative directory, its kept child directory
names and its kept child file names; the first frame also carries the global
protected prefixes, size-skipped prefixes and --delete-missing-args paths.
--delete-before keeps the existing whole-tree early manifest; plain --delete and
--delete-after keep the end-of-transfer manifest commit.
Preserves the existing safety surface: protected/size-skipped prefixes and the
--delay-updates/basis skips are honored at any depth, deletion is scoped to the
synchronized directories (--files-from), MAX_SERVER_DELETE_COUNT and
--max-delete (partial + exit 25) are shared across plans, symlinks are never
followed, and paths are confined to the receive root.
- Scope the --delete extras walk to directories synchronized by the
transfer: add a synchronized-directory section to the delete manifest
(protocol 2.23.0) so --files-from subsets no longer delete untransmitted
paths outside listed directory subtrees (data-loss fix).
- Separate --max-size/--min-size prune protection from --delete-excluded so
size-pruned source mirrors survive (rsync parity).
- Unlink extraneous destination symlinks instead of skipping them.
- Make --max-delete partial (delete up to N, skip the rest) and exit 25;
accept negative values as unlimited.
- Draw --delete-missing-args deletions from the shared --max-delete budget.
- Honor --force during --delay-updates publication.
Add unit and integration regression tests; update the pinned config wire
golden and version strings for the 2.23.0 manifest/status additions.
Review fixes for Phase 7 Wave D.
#1 (HIGH): STATUS_DIR_TIMES entries no longer create directories. A new
receiver-only File.dir_time_only flag marks dir-time entries; file_save_to_disk_full
short-circuits them as FILE_SAVE_SKIPPED before any device/dir branch, so the sink
still accumulates metadata into the deferred DirTimeList but creates nothing. Empty
source dirs stay untransferred (-a), -m/--prune-empty-dirs semantics are preserved,
and a pre-existing regular file/symlink at an empty-dir mirror path no longer aborts
the transfer. dir_time_list_apply fstatat()s the leaf (AT_SYMLINK_NOFOLLOW) and skips
absent/non-directory paths QUIETLY; only a real existing directory is stamped.
Also initialize File.dir_time_only in file_create() (uninitialised garbage otherwise).
#2 (MED): send_dir_times() chunks entries into repeated STATUS_DIR_TIMES frames of at
most MAX_MANIFEST_ENTRIES, matching the receiver's per-frame bound; the tautological
> INT_MAX check is gone.
#3 (LOW): dir_time_list_add() assigns each grown array right after its realloc (no
dangling) and advances capacity only after both succeed.
#4 (LOW): RSYNC_COMPAT.md -- STATUS_MKDIR carries metadata, dir times are transmitted
via STATUS_DIR_TIMES and applied at the end, empty dirs are still never created; -m
rationale, -O row and Wave D notes updated. Summary counts untouched.
#5 (LOW): integration tests for the three #1 scenarios (empty-dir non-creation under
-a and -a -m, collision non-abort), scanner test now covers empty-dir capture, and
test_file_restore_symlink_metadata asserts the positive apply path when supported.
PROTOCOL_VERSION stays 2.17.0; config-frame layout unchanged.
Wave D of Phase 7. Make -O/--omit-dir-times and -J/--omit-link-times real by
preserving directory and symlink times, and mark --secluded-args as an explicit
Impossible/Divergence no-op.
Wire: PROTOCOL_VERSION 2.16.0 -> 2.17.0. Adds a terminal STATUS_DIR_TIMES frame
(int count + (wire path, metadata) pairs) sent after all file data and the
optional delete manifest. STATUS_MKDIR also carries metadata for --dirs entries.
Config-frame layout is unchanged.
Sender: the recursive scanner captures every traversed source directory (both
DirectoryScanner and the parallel scanner root + workers, appends mutex-guarded)
into a shared list; the single-threaded and -m paths transmit it last.
Receiver: a DirTimeList accumulates received directory metadata and applies it
with fd-relative no-follow utimensat only at the very end -- after all children,
after the commit-style --delete, and after --delay-updates publication -- in the
single-threaded success frame and in server.c after the -m threads join. -O skips
the application. Symlink metadata is applied at link creation with
utimensat/fchownat/fchmodat AT_SYMLINK_NOFOLLOW; -J suppresses only link times.
identity_apply_ownership_link shares the identity resolver with the fd path.
Docs: -O/-J rows -> Implemented; --secluded-args -> Impossible/Divergence;
--protocol accepted/rejected values and Phase-6/7 notes updated.
Tests: unit (scanner dir capture, DirTimeList apply, symlink metadata, protocol
version values) and integration (dir mtime round-trip + -O, symlink mtime
round-trip + -J, independent suppression), parameterized over single/multithread.
New flags parse onto the config; --delete-missing-args implies
--ignore-missing-args (order-independent) and does NOT imply --delete (rsync:
independent of other delete processing). The files-from preflight now classifies
listed-but-missing entries instead of hard-failing: under the flags each is
skipped (logged + counted, never silent) and the run succeeds for the rest,
including the all-missing case; an empty list stays a hard error. Under
--delete-missing-args the missing entries' destination mirrors (bare relative
path with -R, full source mirror otherwise) ride the manifest's third section in
both the single-threaded and -m senders; --dirs listed-but-missing entries are
skipped in the scanner.
The STATUS_MANIFEST frame now carries two count-delimited sections: the kept
paths and a protected-prefix list (excluded-on-source paths the walker must not
delete unless --delete-excluded opted out). The receiver's DeleteManifest is
passed through the commit/early paths unchanged. manifest_delete_extras
honors a client --max-delete (all-or-nothing) and produces a distinct error for
it versus the 100000-entry server bound. --force clears a non-empty directory
that blocks an incoming regular file (confined, symlink-safe) via a new
file_remove_tree_secure helper.
Deletion timing is now real and selected by the four rsync flags plus the
plain --delete default. Wire protocol bumps to 2.8.0: two new config
booleans (delete_during, delete_delay) are serialized and validated, joining
the existing delete_before/delete_after.
- Early modes (--delete-before, --delete-during/--del): the sender pre-scans
the whole tree (paths only), transmits the keep-set manifest BEFORE any
file data, and the receiver removes extras and acks STATUS_OK; the sender
only streams data after the deletion committed. Deletion is thus performed
even if a later transfer phase fails (rsync delete-before/during are
destructive by definition). FastSync streams in a single scan so it cannot
interleave per-directory like rsync delete-during; --delete-during selects
the same engine mode as --delete-before (documented divergence).
- Late/commit modes (plain --delete, --delete-after, --delete-delay): the
manifest closes the data stream and deletion is committed only after
STATUS_FINISHED proves the whole transfer succeeded, preserving FastSync's
commit-style safety. --delete-delay converges with --delete-after because
FastSync never snapshots the destination during data flow (documented).
- The STATUS_MANIFEST frame is now self-delimiting and position-independent.
Single-threaded receivers delete before the success frame; the -m receiver
hands the keep-set to server.c, which commits the deletion only after the
disk writer thread has drained (fixes a delete-vs-in-flight-temp race).
- Every timing flag implies --delete; at most one timing flag is allowed.
- Each timing flag implies --delete, matching rsync; conflicts are rejected.
#251 --remove-source-files deletes sources that were skipped receiver-side
(--existing/--ignore-existing/--update). The receiver now reports a
per-file outcome for every processed data file when the sender requests
removal; the client only unlinks sources the receiver actually wrote.
add remove_source_files to the wire config and bump the protocol to 2.5.0.
#252 --backup/--suffix/--backup-dir broken by NULL-vs-empty wire loss. Receivers
canonicalize the empty wire string back to NULL for backup_dir, temp_dir,
partial_dir and suffix, and --suffix is received unconditionally.
#253 --partial --partial-dir never installed completed files. file_save_to_disk
now renames a fully written partial-dir file into the real destination.
#255 STATUS_CHECK read the entire old file before the size/mtime quick check.
Old contents are only read when a checksum compare or delta needs them.
#256 receive_delta_file failure paths did not set *failed, so the caller sent
STATUS_NEXT and waited for a body that never came. Every NULL return now
marks the transfer failed.
#257 files >64 MiB could not transfer. Whole-file receive caps raised to the
256 MiB connection/allocation ceiling (chunk caps stay 64 MiB) and the
client ignores SIGPIPE so a server-side close surfaces as a clean error.
Unit tests added: config NULL-vs-empty round trip, incremental quick-check
skip/NEXT paths, delta oversize failure, partial-dir install, save-result
skip reporting.
The per-connection memory budget (MAX_CONNECTION_MEMORY, 256 MiB) only
charged wire buffers via receive_data_limited. Decompression buffers and
per-file chunk copies were not accounted for, and the multithreaded
receiver could enqueue up to 100 files (each up to 64 MiB uncompressed)
ahead of a slow disk writer, retaining ~6.4 GiB per connection. A client
sending highly compressible chunks with little bandwidth could OOM the
host while the reserve never tripped.
Bound the receive pipeline by aggregate payload bytes instead of item
count alone:
- Export MAX_CONNECTION_MEMORY from protocol.h.
- PipelineContextReceiver tracks queued_bytes (payload bytes received but
not yet released by the disk writer, i.e. queued or in the writer's
hand) under the existing mutex.
- receiver enqueue now blocks while the queue is full by count OR when
adding the file would push queued_bytes over the configured byte limit,
applying backpressure to the sender instead of failing the transfer.
- The disk writer releases the byte budget after each file is freed and
signals the not-full condition.
- The server sets the byte ceiling to
MAX_CONNECTION_MEMORY - 2*MAX_CHUNK_SIZE so that the queued payloads
plus the transient wire/decompression buffers of the one in-flight
chunk stay within the per-connection budget.
The single-threaded receive path is already bounded: it writes files to
disk before reading the next chunk, so its transient is at most one
chunk's wire + decompressed + copied payload (~3 * MAX_CHUNK_SIZE, below
the budget). Wire buffers remain charged exactly once by
receive_data_limited; this change does not double charge them.
Adds a deterministic unit test in test_multiprocessing.c proving that an
enqueue which would exceed the byte budget blocks until the writer
releases bytes.
- Reformat all C/H files to match .clang-format (LLVM style)
- Fix 26 cppcheck const-correctness warnings (constParameterPointer,
constVariablePointer, constVariable)
- Update function declarations in headers to match const parameters
Feature changes across 15 files:
- -a/--archive: enables -c -m -M (no -s, which slows transfers)
- -n/--dry-run: scan + print without connecting or transferring
- -p <port>: custom SSH port (passed as -p to ssh via execvp)
- --exclude <pattern>: glob-based filename filtering in scanner
- --delete: sender collects file manifest; receiver deletes unlisted files
Protocol: added STATUS_MANIFEST, use_delete field in Config wire format.
New utilities: glob_match(), delete_extras() with recursive directory walk.
Scanner: accepts exclude patterns; skips matching entries.
SSH transport: switched from execlp to execvp for dynamic port arg.
Server + multiprocessing: handle STATUS_MANIFEST in both single and
multithreaded receive paths.
Builds clean, all 7 tests pass.
- Merge io.h/c back into protocol.h/c
- Change send_data to take Data* argument
- Remove redundant file_load_data from send_chunk
- Move compression into file_send_single_calls
- Move file_receive from multiprocessing.c to file.c