164 Commits
Author SHA1 Message Date
TapTap 7585f46eb4 Release v2.29.0 (#312)
CI / lint (push) Successful in 1m50s
CI / parity-fast (push) Skipped
CI / parity-full (push) Successful in 21s
CI / sanitizers (address) (push) Successful in 55s
CI / sanitizers (undefined) (push) Successful in 43s
CI / build-and-test (push) Successful in 1m25s
CI / fuzz-build (push) Successful in 51s
CI / coverage (push) Successful in 46s
CI / valgrind (push) Successful in 2m41s
2026-09-23 02:05:14 +02:00
TapTap 00197102bf Release v2.29.0
CI / lint (push) Successful in 1m48s
CI / parity-fast (push) Skipped
CI / lint (pull_request) Successful in 1m48s
CI / parity-full (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / parity-full (push) Successful in 19s
CI / sanitizers (address) (push) Successful in 50s
CI / sanitizers (undefined) (push) Successful in 44s
CI / build-and-test (push) Successful in 1m19s
CI / coverage (push) Successful in 45s
CI / fuzz-build (push) Successful in 54s
CI / parity-fast (pull_request) Successful in 20s
CI / build-and-test (pull_request) Successful in 56s
CI / valgrind (push) Successful in 2m43s
- Transport I/O vtable over TCP/TLS; TLS multithreaded sendfile fixed
- Symlink-xattr wire block (protocol 2.29.0; config frame unchanged)
- No-wire parity burn-down: --inc-recursive, in-root --temp-dir,
  --delete-before phase-0, rsync-interoperable --fake-super, --devices
  per-entry failure, rsyncd.conf key subset (read-only default)
- Protocol version: 2.29.0
- Tested: unit, integration, ASan, UBSan, valgrind, differential parity
2026-09-23 01:59:54 +02:00
TapTap 13b257d1dd Merge PR #311: transport I/O vtable + symlink-xattr wire block (2.29.0)
CI / lint (push) Successful in 1m48s
CI / parity-fast (push) Skipped
CI / parity-full (push) Successful in 20s
CI / sanitizers (address) (push) Successful in 54s
CI / sanitizers (undefined) (push) Successful in 45s
CI / build-and-test (push) Successful in 1m20s
CI / fuzz-build (push) Successful in 48s
CI / coverage (push) Successful in 46s
CI / valgrind (push) Successful in 2m35s
2026-09-23 01:50:51 +02:00
TapTap f3d7672694 docs: bump release version to 2.29.0 and reconcile protocol docs
CI / lint (pull_request) Successful in 1m48s
CI / parity-full (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / parity-fast (pull_request) Successful in 20s
CI / build-and-test (pull_request) Successful in 57s
2026-09-23 01:45:30 +02:00
TapTap 18b821d32c Merge branch 'feat/transport-g2' into feat/transport-xattr 2026-09-23 01:36:25 +02:00
TapTap 6ad065925f Merge branch 'feat/transport-g1' into feat/transport-xattr 2026-09-23 01:36:25 +02:00
TapTap 46dcefe218 fix(protocol): classify TLS EOF before EINTR retry; harden current-ssl resolver; real TLS regression test 2026-09-23 01:36:08 +02:00
TapTap b88acdbd3c fix(xattr): whitelist path-based symlink apply; strengthen symlink xattr tests 2026-09-23 01:01:18 +02:00
TapTap c29bce54fc Merge branch 'feat/transport-c2' into feat/transport-xattr 2026-09-23 00:35:05 +02:00
TapTap d98e971fcc Merge branch 'feat/transport-c1' into feat/transport-xattr 2026-09-23 00:35:05 +02:00
TapTap 492ce0ce89 feat(xattr): carry and apply symlink xattrs (protocol 2.29.0) 2026-09-23 00:34:39 +02:00
TapTap ec6692ac41 protocol: add transport I/O vtable over TCP/TLS primitives
Introduce ProtocolIoOps (send/recv/has_pending), selected once by
protocol_session_init() and protocol_session_set_ssl(), and dispatch the
send, receive and status-read loops through session->ops instead of
branching on session->ssl at runtime.

Each op performs one transfer attempt and classifies the result
(PROTOCOL_IO_RETRY/CLOSED/ERROR), preserving the WANT_READ/WANT_WRITE
wait_events switching, the SSL_ERROR_SYSCALL/EINTR retry, the
SSL_pending poll gating and the deadline handling. The raw read()/write()
fallback lives in the plaintext ops.

Add unit tests: a socketpair session with a counting ops wrapper proving
the loops dispatch through the vtable, and a worker-thread test that
protocol_current_ssl() resolves the bound session's SSL when io_ssl is NULL.
2026-09-23 00:18:36 +02:00
TapTap 1f8d60e30d protocol: resolve TLS transport from the bound session, not thread-local io_ssl
file_send.c chose between sendfile() and the TLS-aware buffered path by
calling io_get_ssl(), which reads the thread-local io_ssl. A worker thread
that bound a TLS ProtocolSession via protocol_session_bind() never ran the
handshake in that thread, so io_ssl is NULL there and a TLS + --threads
transfer took the raw sendfile() path on an encrypted socket.

Add protocol_current_ssl(), which prefers the bound session's SSL and falls
back to io_ssl on the fd-shim path, and use it in file_send.c. Un-xfail
test_tls_with_multithreading.
2026-09-23 00:18:29 +02:00
TapTap 6db9827b87 Merge PR #310: restore dumpable flag for ASan/LSan
CI / lint (push) Successful in 1m46s
CI / parity-fast (push) Skipped
CI / parity-full (push) Successful in 21s
CI / sanitizers (address) (push) Successful in 54s
CI / sanitizers (undefined) (push) Successful in 43s
CI / build-and-test (push) Successful in 1m22s
CI / fuzz-build (push) Successful in 48s
CI / coverage (push) Successful in 44s
CI / valgrind (push) Successful in 2m17s
2026-09-22 23:50:54 +02:00
TapTap a07cc00bb0 test: restore dumpable flag after setuid drop so LSan can run under ASan
CI / lint (pull_request) Successful in 1m48s
CI / parity-full (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / parity-fast (pull_request) Successful in 18s
CI / build-and-test (pull_request) Successful in 55s
2026-09-22 23:45:41 +02:00
TapTap 3ae7685655 Merge PR #309: parity burn-down
CI / lint (push) Successful in 1m46s
CI / parity-fast (push) Skipped
CI / parity-full (push) Successful in 22s
CI / sanitizers (address) (push) Failing after 54s
CI / sanitizers (undefined) (push) Successful in 43s
CI / build-and-test (push) Successful in 1m22s
CI / fuzz-build (push) Successful in 49s
CI / coverage (push) Successful in 44s
CI / valgrind (push) Successful in 2m19s
2026-09-22 23:33:21 +02:00
TapTap 4f945a8e39 docs: reconcile RSYNC_COMPAT for parity-next review follow-ups
CI / lint (pull_request) Successful in 1m46s
CI / parity-full (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / parity-fast (pull_request) Successful in 17s
CI / build-and-test (pull_request) Successful in 54s
2026-09-22 23:27:54 +02:00
TapTap 15f38f5b76 Merge branch 'feat/parity-f4' into feat/parity-next 2026-09-22 23:20:21 +02:00
TapTap 787967d3ce Merge branch 'feat/parity-f2' into feat/parity-next 2026-09-22 23:20:21 +02:00
TapTap 06e7aef5c9 Merge branch 'feat/parity-f1' into feat/parity-next 2026-09-22 23:20:21 +02:00
TapTap ee265d78ba fix(delete-before): replay pre-scan list in --threads path 2026-09-22 23:19:55 +02:00
TapTap 47b1b9b915 fix(xattr): fake-super device round-trip, --devices continue-on-error, harden stat parse 2026-09-22 23:05:40 +02:00
TapTap bb6c788cf9 fix(daemon): rsync read-only module default; warn on unenforced security keys 2026-09-22 22:53:23 +02:00
TapTap 0b40c47d6a Merge branch 'feat/parity-b4' into feat/parity-next
# Conflicts:
#	RSYNC_COMPAT.md
2026-09-22 22:08:16 +02:00
TapTap 6f81005094 Merge branch 'feat/parity-b3' into feat/parity-next
# Conflicts:
#	RSYNC_COMPAT.md
2026-09-22 22:07:45 +02:00
TapTap 193d358c64 Merge branch 'feat/parity-b2' into feat/parity-next 2026-09-22 22:06:25 +02:00
TapTap 8792e3265a Merge branch 'feat/parity-b1' into feat/parity-next 2026-09-22 22:06:25 +02:00
TapTap 3ec0ffb644 fix(delete-before): reuse pre-scan file list in single-threaded data pass 2026-09-22 22:05:50 +02:00
TapTap 80e8d7f450 feat(xattr): rsync-interoperable --fake-super stat; fix --devices error parity 2026-09-22 22:03:41 +02:00
TapTap 86741725fd feat(daemon): accept rsync rsyncd.conf key subset and --dparam mapping 2026-09-22 22:02:09 +02:00
TapTap f96f1764af feat(cli): accept --inc-recursive no-op and in-root absolute --temp-dir 2026-09-22 21:58:54 +02:00
TapTap d780a4625e Merge PR #308: close valid residuals of issues #286-#297
CI / lint (push) Successful in 1m45s
CI / parity-fast (push) Skipped
CI / parity-full (push) Successful in 21s
CI / sanitizers (address) (push) Successful in 52s
CI / sanitizers (undefined) (push) Successful in 42s
CI / build-and-test (push) Successful in 1m16s
CI / fuzz-build (push) Successful in 47s
CI / coverage (push) Successful in 44s
CI / valgrind (push) Successful in 2m18s
2026-09-22 17:30:36 +02:00
TapTap 5d1303ffdf Merge branch 'fix/issues-f4' into fix/issues-triage
CI / lint (pull_request) Successful in 1m45s
CI / parity-full (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / parity-fast (pull_request) Successful in 17s
CI / build-and-test (pull_request) Successful in 54s
2026-09-22 17:18:59 +02:00
TapTap 06b53c5b3d Merge branch 'fix/issues-f3' into fix/issues-triage 2026-09-22 17:18:59 +02:00
TapTap bf09e8893e Merge branch 'fix/issues-f2' into fix/issues-triage 2026-09-22 17:18:59 +02:00
TapTap 034c27926f Merge branch 'fix/issues-f1' into fix/issues-triage 2026-09-22 17:18:59 +02:00
TapTap aa15099d22 fix(output): restore --info=flist header; dir metadata on plan path; root line for -d 2026-09-22 17:17:04 +02:00
TapTap ff4db23831 fix(identity): free TO name on glob success path; harden test oracle 2026-09-22 17:08:03 +02:00
TapTap 79e45441c0 fix(receiver): defer --dirs directory mode to avoid EACCES on children
file_save_directory_to_disk() applied the exact source mode (fchmod)
inline for explicit --dirs/STATUS_MKDIR entries.  A restrictive source
mode (e.g. 0555) then made the directory read-only before its children
were written, so a non-root receiver failed each child with EACCES.  The
recursive -a path never hit this because it defers directory metadata.

Remove the inline fchmod and let the existing deferred
dir_metadata_list_apply() stamp the exact mode at end of transfer, as the
recursive path does.  Keep the inline ownership and xattrs (a direct
file_save_to_disk_full() caller has no deferred pass) and document the
resulting intentional ordering.  Capture errno before output_escape() in
the inline timestamp diagnostic so strerror() reports the real error, and
add the missing trailing newline to tests/test_xattr.c.
2026-09-22 17:06:54 +02:00
TapTap 6537227467 test(transport): de-race fallback/fd tests and gate for valgrind 2026-09-22 17:05:54 +02:00
TapTap 309c9aed98 refactor: const-correct dir/root locals in progress itemize 2026-09-22 16:46:23 +02:00
TapTap 84594197ee Merge branch 'fix/issues-b1' into fix/issues-triage 2026-09-22 16:37:56 +02:00
TapTap 7bf25048f6 docs: correct parity claims for issues #286-#297 2026-09-22 16:37:21 +02:00
TapTap b6b5eee20e Merge branch 'fix/issues-a3' into fix/issues-triage 2026-09-22 16:20:55 +02:00
TapTap 8dcd87609d Merge branch 'fix/issues-a2' into fix/issues-triage 2026-09-22 16:20:55 +02:00
TapTap 19fe63bd59 Merge branch 'fix/issues-a1' into fix/issues-triage 2026-09-22 16:20:55 +02:00
TapTap 1f5f8dc5a7 fix(output): emit directory/root lines for -i and --out-format (#292) 2026-09-22 16:20:03 +02:00
TapTap 3787695ba6 fix(xattr): apply --dirs directory xattrs fd-relative (#286) 2026-09-22 16:05:18 +02:00
TapTap b7cb213c8c fix: FROM name globs for identity maps; transport fallback tests (#294, #219) 2026-09-22 16:03:18 +02:00
TapTap 8cc3dd993b Merge PR #307: structural refactor cycle (no behavior change)
CI / lint (push) Successful in 1m42s
CI / parity-fast (push) Skipped
CI / parity-full (push) Successful in 21s
CI / sanitizers (address) (push) Successful in 53s
CI / sanitizers (undefined) (push) Successful in 43s
CI / build-and-test (push) Successful in 1m16s
CI / fuzz-build (push) Successful in 48s
CI / coverage (push) Successful in 44s
CI / valgrind (push) Successful in 2m18s
2026-09-22 14:58:26 +02:00
TapTap 1167e7970b refactor(delete): rename basis helper to delete_basis_relative
CI / lint (pull_request) Successful in 1m43s
CI / parity-full (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / parity-fast (pull_request) Successful in 19s
CI / build-and-test (pull_request) Successful in 57s
2026-09-22 14:52:36 +02:00
TapTap e98729f00e refactor: const-correct delete-manifest API; apply clang-format 2026-09-22 14:24:39 +02:00
TapTap c0020364b2 Merge branch 'refactor/minor' into refactor/structural 2026-09-22 14:06:55 +02:00
TapTap d7ac940a6b Merge branch 'refactor/cfg' into refactor/structural 2026-09-22 14:06:55 +02:00
TapTap be20e836de fix: correct throttle legacy resolution; add EXDEV temp-dir coverage 2026-09-22 14:06:26 +02:00
TapTap b549887138 refactor(config): group CLI-parse state; drop old_args field 2026-09-22 14:03:10 +02:00
TapTap 1ffd4744c6 Merge branch 'refactor/delete' into refactor/structural 2026-09-22 13:53:20 +02:00
TapTap 494cef2a0b Merge branch 'refactor/pending' into refactor/structural 2026-09-22 13:53:20 +02:00
TapTap ed7527cc2c Merge branch 'refactor/handler' into refactor/structural 2026-09-22 13:53:20 +02:00
TapTap 934defa965 refactor(delete): consolidate delete engine into delete.c 2026-09-22 13:52:57 +02:00
TapTap ade9be8600 refactor(receiver): per-status dispatch and shared pending teardown 2026-09-22 13:49:28 +02:00
TapTap 707bb659e8 refactor(server): decompose handler into phases 2026-09-22 13:46:17 +02:00
TapTap 33980bc4c8 Merge branch 'refactor/filerecv' into refactor/structural 2026-09-22 13:39:34 +02:00
TapTap 6a9372f4f5 Merge branch 'refactor/client' into refactor/structural 2026-09-22 13:39:34 +02:00
TapTap 5e1d6b6e10 Merge branch 'refactor/scanner' into refactor/structural 2026-09-22 13:39:34 +02:00
TapTap d2d1b63f44 refactor(receive): split file_save/incremental_check/delete_commit out
Pure structural split of src/shared/file_receive.c into focused translation
units behind the unchanged file_receive.h facade:

- file_save.c   : save-to-disk, special nodes, --delay-updates staging
- incremental_check.c : xattr/delta/basis/fuzzy receive + check state machine
- delete_commit.c : manifest receive + delete budget walkers
- file_receive.c : wire receive dispatch + deferred dir metadata

The shared receive_file_xattrs helper and MAX_FILE_DATA_SIZE are declared in
incremental_check.h.  file_save_to_disk_full_ex is decomposed into static
helpers (validation, special dispatch, dir/symlink creation, path resolution,
pre-write policies, data install) routed through one cleanup epilogue.

No behavior change.
2026-09-22 13:38:54 +02:00
TapTap a6659472fe refactor(scanner): split into filter/sequential/parallel TUs
Move the filter rule-tree/context helpers and entry inspection into
scanner_filter.c, the parallel scanner into scanner_parallel.c, and keep
the sequential scanner in scanner.c.  Shared internal declarations live in
the new scanner_internal.h; scanner.h stays the public façade.

Decompose directory_scanner_next into static helpers (skipped-entry,
selection-protection, mount/-x, directory-finish and per-entry handlers)
with no semantic change.
2026-09-22 13:38:05 +02:00
TapTap 4638030288 refactor(client): split reporting/scan/manifest out of client_send
Move the stats/progress reporting, scanner-preparation/scan helpers and
manifest/list/dry-run senders out of the ~3.9k-line client_send.c into
client_report.c, client_scan.c and client_manifest.c, sharing declarations
through the new internal client_send_internal.h.  client_send.c keeps the
transfer orchestration and is now ~2.1k lines.

Decompose the monolithic send_files into static phase helpers
(send_files_prepare/_prepare_delete/_run/_finalize/_cleanup) driven by a
single SendFilesState; ownership, ordering and exit codes are unchanged.

No behavior change.
2026-09-22 13:34:55 +02:00
TapTap 07dfec629f Merge PR #306: codebase audit cycle (security, correctness, refactors, docs)
CI / lint (push) Successful in 2m31s
CI / parity-fast (push) Skipped
CI / parity-full (push) Successful in 23s
CI / sanitizers (undefined) (push) Successful in 47s
CI / sanitizers (address) (push) Successful in 1m13s
CI / build-and-test (push) Successful in 1m19s
CI / fuzz-build (push) Successful in 49s
CI / coverage (push) Successful in 44s
CI / valgrind (push) Successful in 2m18s
2026-09-21 22:23:33 +02:00
TapTap c062a0762e Merge branch 'fix/audit-docs3' into fix/audit-cycle
CI / lint (pull_request) Successful in 2m46s
CI / parity-full (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / parity-fast (pull_request) Successful in 31s
CI / build-and-test (pull_request) Successful in 54s
2026-09-21 22:11:53 +02:00
TapTap 283f9f0823 docs: reflect filter modifiers, inplace+partial-dir, credentials hardening
Update the docs for the audit follow-up fixes:
- --filter merge modifiers e/n/w/- are now accepted-and-consumed on
  merge/dir-merge rules (rejected on non-merge, x rejected everywhere);
  their semantics stay unimplemented, so the --filter row moves to Caveat
  and the tally becomes 119/11/27 = 157.
- --inplace + --partial-dir is rejected with rsync's message.
- secret_file_open() O_NOFOLLOW (symlinked credential paths fail closed;
  fd-backed paths exempt) and ~3 s bound-wait on FIFO reads.
- AGENTS setpriv wording corrected to the collected instance count.
- CHANGELOG [Unreleased] audit section extended with the follow-ups.
2026-09-21 22:11:19 +02:00
TapTap 5754b9a952 Merge branch 'fix/audit-misc2' into fix/audit-cycle 2026-09-21 22:02:05 +02:00
TapTap b8a0efef7b Merge branch 'fix/audit-filter2' into fix/audit-cycle 2026-09-21 22:02:05 +02:00
TapTap 2e77c09447 Merge branch 'fix/audit-creds2' into fix/audit-cycle 2026-09-21 22:02:05 +02:00
TapTap 06c4026b74 fix(filter): accept e/n/w/- merge modifiers on merge/dir-merge rules
The earlier modifier-rejection change rejected e/n/w on all rules, but rsync
3.4.1 accepts them (plus the '-' merge-only modifier) on merge and dir-merge
rules.  Restrict the rejection to non-merge rules and consume the merge-file
modifiers (e/n/w/-) so they no longer leak into the merge filename.

- is_merge_rule()/is_merge_modifier_char() gate the merge-only modifiers.
- scan vs consume sets: e/n/w still count as modifier-run chars on every rule
  (pure tokens like -new/-press stay rejected), but are only consumed on merge
  rules, preserving mixed-token parsing such as H,!secret -> ecret.
- '-' is accepted/consumed only on merge/dir-merge (e.g. dir-merge,- .rules).
- x remains rejected everywhere with its dedicated message.
- e/n/w/- semantics remain unimplemented and are documented as accepted-but-
  ignored in filter.h.

Tests: split the merge forms out of the rejection test into a new acceptance
test asserting the merge file is read and dir_merge_names keeps the modifier-
free basename; non-merge pure-modifier forms still rejected.
2026-09-21 22:01:44 +02:00
TapTap 9f47b13712 fix(credentials): bound-wait on FIFO reads so slow process substitution works
secret_file_open() opened secret files with O_NONBLOCK and only cleared it
for S_ISREG, so on a FIFO/process-substitution source (--password-file
<(...), --early-input <(...)) fgets() failed immediately with EAGAIN when
the writer had not yet produced data, breaking slow producers.

Keep O_NONBLOCK at open() (a writer-less FIFO must not block the open) and
route all three readers through a new secret_read_line() helper.  It
accumulates a line across reads and, on EAGAIN/EWOULDBLOCK (or a partial
line) with no newline and no EOF, clearerr()s and polls for readability
against one overall CLOCK_MONOTONIC deadline of
CREDENTIAL_FIFO_READ_TIMEOUT_MS (3000 ms); on timeout or a real read error
it fails with a clear message.  EOF finishes normally.  Regular files are
left blocking and read exactly as before.

Handles a line split across several write()s and keeps the owner/mode
fstat gate, O_NOFOLLOW and the /dev/fd/N exception unchanged.
2026-09-21 22:01:18 +02:00
TapTap b1eddf0133 fix: config leak, compression log, inplace+partial-dir rejection, umask/root test fixes 2026-09-21 21:59:58 +02:00
TapTap 338c27db73 fix: resolve cppcheck shadow/always-true findings 2026-09-21 21:28:36 +02:00
TapTap 8d46a26c04 Merge branch 'fix/audit-docs2' into fix/audit-cycle 2026-09-21 21:18:48 +02:00
TapTap 707ba272df Merge branch 'fix/audit-docs1' into fix/audit-cycle 2026-09-21 21:18:48 +02:00
TapTap bc71a3c0a5 docs: correct README build deps, delete defaults, scanner, flags; AGENTS deps/CI 2026-09-21 21:18:26 +02:00
TapTap cee9b7647c docs: fix parity tally, compat rows, changelog, handoff for audit cycle 2026-09-21 21:14:37 +02:00
TapTap 3adb6dddb5 chore: drop tracked scratch data and extend .gitignore 2026-09-21 21:11:04 +02:00
TapTap d77849774e Merge branch 'fix/audit-refb' into fix/audit-cycle 2026-09-21 21:09:20 +02:00
TapTap b5c6f8b60f Merge branch 'fix/audit-refa' into fix/audit-cycle 2026-09-21 21:09:20 +02:00
TapTap 134dcd74cc refactor: drop dead filter_rules_apply, unify set_error, dedup path_is_within 2026-09-21 21:09:05 +02:00
TapTap 42f2f845ac refactor: drop Config**, dead old-args plumbing, dedup constants, -Wformat-signedness 2026-09-21 19:50:33 +02:00
TapTap 0669335ca5 Merge branch 'fix/audit-logfmt' into fix/audit-cycle 2026-09-21 19:43:17 +02:00
TapTap 391f76cd56 fix(log): add printf format attributes and fix format mismatches 2026-09-21 19:42:44 +02:00
TapTap 2021afe613 Merge branch 'fix/audit-fdpaths' into fix/audit-cycle 2026-09-21 19:30:21 +02:00
TapTap ba914e8ab3 fix(credentials): allow fd-backed store paths without O_NOFOLLOW 2026-09-21 19:30:01 +02:00
TapTap df8fe1ae4e Merge branch 'fix/audit-signal' into fix/audit-cycle 2026-09-21 19:27:03 +02:00
TapTap 2c490d58b7 Merge branch 'fix/audit-creds' into fix/audit-cycle 2026-09-21 19:27:03 +02:00
TapTap d55dabff2e Merge branch 'fix/audit-help' into fix/audit-cycle 2026-09-21 19:27:03 +02:00
TapTap b478a59a81 fix(cli): correct help text, per-codec compression default, and stale test 2026-09-21 19:26:42 +02:00
TapTap 865941f850 fix(credentials): open secret files with O_NOFOLLOW|O_NONBLOCK; drop dup includes
secret_file_open() previously opened --password-file/--early-input with
plain O_RDONLY, so a symlinked path was followed before the owner/mode
fstat gate ran, and an empty/planted FIFO could block fgets forever.
Open with O_NOFOLLOW|O_NONBLOCK|O_CLOEXEC (mirroring the dummy-key
sidecar): ELOOP now fails closed, and a writer-less FIFO yields EOF/EAGAIN
instead of hanging. Clear O_NONBLOCK again for regular files, where it is
a no-op, so their stdio read path is unchanged.

file.c: drop the duplicate <fcntl.h>/<unistd.h> includes (kept the first
occurrences).

Tests: a symlinked password file is rejected, and a writer-less named
FIFO fails cleanly without hanging.
2026-09-21 19:25:58 +02:00
TapTap 221cefa7cc fix(signal): use sigaction and async-signal-safe handlers
Client: replace the non-async-signal-safe signal(3) call inside
client_signal_handler() with a precomputed SIG_DFL sigaction(2), which is
on the POSIX async-signal-safe list.  The handler stays installed while a
transfer is armed so a repeated Ctrl-C still leads to a graceful abort
rather than a hard kill mid-cleanup.

Server: cleanup() now only calls _exit(2) (async-signal-safe).  The former
server_delete()/daemon_conf_free()/credentials_free() teardown called
free()/close()/SSL_CTX_free() from signal context, which can deadlock or
corrupt the heap if the signal lands inside malloc/free.  Handlers are
installed with sigaction(2) instead of signal(3).  The normal shutdown
path in main() still performs the full teardown; the signal path relies on
process exit to reclaim the parent daemon's socket, anonymous shared
mapping and heap (no named/persistent parent resource is left behind).
2026-09-21 19:25:24 +02:00
TapTap bb9b59024a Merge branch 'fix/audit-tests' into fix/audit-cycle 2026-09-21 19:19:36 +02:00
TapTap 5baf120243 Merge branch 'fix/audit-misc' into fix/audit-cycle 2026-09-21 19:19:36 +02:00
TapTap 84b7e1fd2f Merge branch 'fix/audit-filterx' into fix/audit-cycle 2026-09-21 19:19:36 +02:00
TapTap 800a978e15 Merge branch 'fix/audit-config' into fix/audit-cycle 2026-09-21 19:19:36 +02:00
TapTap 0f40e747f0 fix(filter): reject the x modifier in the --filter list parser
The standalone filter_rule_parse() already rejected the rsync xattr-name
'x' modifier, but the list parser used by --filter/-f silently dropped the
flag for merge/dir-merge rules (and relied on a second parse for plain
rules).  Reject it explicitly in filter_list_parse_append_depth() with the
same diagnostic, so '-x', 'merge,x' and 'dir-merge,x' all fail cleanly.

Also reject the unimplemented rsync merge modifiers 'e', 'n' and 'w'
instead of folding them into the pattern, which previously produced
misleading errors such as "could not read merge file 'n file'".  Only a
token made up solely of modifier characters is treated as a modifier run,
so glued patterns ('-newfile', '-e2e') and mixed tokens ("H,!secret")
keep their historical parsing.

Adds tests/test_filter.c with focused rejection and supported-syntax
cases.
2026-09-21 19:19:06 +02:00
TapTap b07306d5bc fix(config): check ssh-dest allocation and enforce MAX_FILTER_RULES client-side 2026-09-21 18:53:07 +02:00
TapTap f2c89b6e7c fix: mutex leak, errno-after-free, log_perror misuse, status validation
- multiprocessing: destroy mutex_progress on the dir_entries_mutex
  init-failure path (init >= 7); drop bogus log_perror
- delete_plan: capture errno before free() in apply_deferred_path
- queue/array_list: log_message instead of log_perror for non-errno
  conditions
- protocol: reject unknown wire Status values via status_is_valid() in
  receive_status, receive_status_timed and the keepalive reader; declare
  protocol_receive_status_timed in protocol.h
- protocol: %llu for unsigned long long debug counters
- tests: out-of-range status rejection test
2026-09-21 18:52:46 +02:00
TapTap 379f127859 test: add client timeouts and de-flake default-port test 2026-09-21 18:50:55 +02:00
TapTap 3799200f71 Merge branch 'fix/audit-partial' into fix/audit-cycle 2026-09-21 18:45:48 +02:00
TapTap 91c4a967e6 Merge branch 'fix/audit-receiver' into fix/audit-cycle 2026-09-21 18:45:48 +02:00
TapTap 88c8968b4c Merge branch 'fix/audit-transport' into fix/audit-cycle 2026-09-21 18:45:48 +02:00
TapTap 082b31886a Merge branch 'fix/audit-compression' into fix/audit-cycle 2026-09-21 18:45:48 +02:00
TapTap 423a62e691 fix(receiver): confine --temp-dir scratch dir and gate setuid bits
Three receiver security fixes from the audit:

1. --temp-dir symlink escape (High): file_open_temp_dir() opened the
   client-controlled scratch dir with a bare open(), so a symlink planted
   under the receive root let a peer redirect receiver scratch files
   outside the authorized root.  The opened dir is now judged by the REAL
   path of its fd (via /proc/self/fd), and any target outside the
   authorized receive root is refused with a logged error (EACCES).  An
   in-root symlink (the EXDEV cross-filesystem fallback case) still works,
   and the no-root local batch path is unchanged.

2. setuid/setgid/sticky under SUPER_MODE_OFF (High): the special bits were
   applied under --perms (and via --chmod) even when the connection forbade
   super-user activities.  FileAttrPolicy gains super_permitted, set by
   file_attr_policy_from_config() from privilege_super_mode_permitted();
   metadata_mode_for_policy(), the symlink path, the special-node creation
   path, and the deferred directory-mode apply now strip the special bits
   when it is false.  Exact rsync semantics are preserved when permitted.

3. daemon umask (Low): daemonize() forced umask(0), so implied parent
   directories created without -p were world-writable 0777.  Set the
   conventional daemon umask 022 instead (rsync never forces 0); -p/-a mode
   preservation is unaffected because it restores modes via fchmod.

Tests: new unit tests for file_open_temp_dir confinement and the
masked/unmasked special-bit policy (incl. the --chmod path), a daemon
world-writable-dir regression test, an integration escape test, and a
root-only integration test asserting special bits are masked without
--allow-super.  The old cross-filesystem test encoded the vulnerable
behavior (symlink target outside the root) and is replaced by the escape
test; the EXDEV fallback code is retained for in-root links.
2026-09-21 18:45:29 +02:00
TapTap bc18ae205b fix(cli): --partial-dir implies --partial (rsync parity)
rsync 3.4.1 resolves --partial-dir after option parsing and sets keep_partial,
so --partial-dir=DIR alone retains an interrupted transfer's partial file.
FastSync only used the partial dir when --partial was also given, silently
discarding it otherwise.

Set Config->partial in cli_finalize_config whenever partial_dir is set.
Following rsync, an explicit --no-partial does NOT win (verified on rsync
3.4.1 in either option order); --inplace is guarded because it writes the
destination in place with no partial staging.

Tests: CLI unit coverage for the implication/precedence/inplace guard, and a
deterministic integration case that blocks the final install (non-empty
directory at the destination) and asserts the staged partial survives under
--partial-dir alone.
2026-09-21 18:37:39 +02:00
TapTap 10a61c6101 fix(compression): raise decompression ceiling to the protocol whole-file limit
MAX_DECOMPRESSED_SIZE was 100 MiB while the receiver advertises and the
sender compresses whole files up to MAX_RECEIVE_WHOLE_FILE_SIZE (256 MiB),
so -z on a 100-256 MiB regular file failed with 'Declared decompressed
size exceeds 104857600 bytes'.  Define the internal bomb-guard ceiling in
terms of the protocol constant so the two bounds cannot drift, and add
unit coverage for a 130 MiB payload (accepted) and an over-ceiling
declared size (still rejected).
2026-09-21 18:31:25 +02:00
TapTap bc1e1191af fix(io): pace sendfile with --bwlimit, retry poll EINTR, clamp SSL_read 2026-09-21 18:30:14 +02:00
TapTap 0fbb9de915 Merge PR #305: rsync-parity cycle 2.29 (120/10/27, no wire change)
CI / lint (push) Successful in 1m57s
CI / parity-fast (push) Skipped
CI / parity-full (push) Successful in 24s
CI / sanitizers (address) (push) Successful in 51s
CI / sanitizers (undefined) (push) Successful in 44s
CI / build-and-test (push) Successful in 1m16s
CI / fuzz-build (push) Successful in 46s
CI / coverage (push) Successful in 42s
CI / valgrind (push) Successful in 2m12s
2026-09-20 15:10:17 +02:00
TapTap 9b05972375 test: assert partial --max-delete survivor order matches rsync
CI / lint (pull_request) Successful in 1m58s
CI / parity-full (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / parity-fast (pull_request) Successful in 18s
CI / build-and-test (pull_request) Successful in 54s
2026-09-20 15:02:51 +02:00
TapTap d119f35066 docs: record parity cycle 2.29 (120/10/27) and deferred residuals 2026-09-20 15:02:44 +02:00
TapTap 00829fd265 parity: --info=mount/stats, --stats dir breakdown, --debug categories
--info=mount now prints rsync's mount-point skip line (matching rsync
3.4.1, which emits it for repeated -xx and drops the mount-point dir);
--info=stats enables the same block as --stats; -x is repeatable.
--stats counts traversed directories for the Number of files breakdown
even when no directory metadata is captured (-r without -t/-p).
--debug enables real output for flist/del/hash/deltasum/recv/filter/send
at their natural FastSync events (synthetic categories stay inert).

--stats and --debug rows keep their documented residual status.
2026-09-20 14:55:47 +02:00
TapTap ff261bc38a test: extend rsync order parity to dry-run and delete-delay 2026-09-20 14:55:47 +02:00
TapTap b235721f8b delete: rsync-exact abort boundary and -d per-directory plans
Transmit the complete --delete-during/--delete-delay per-directory plan
set before the first data frame, so a mid-transfer abort has already
applied every planned removal like rsync's generator; completed runs are
unchanged.  Route -d/--dirs through the same per-directory plans: the
generator records only directories whose direct children it enumerated,
so extras directly inside a listed directory are removed while an
untraversed subdirectory's mirror is shielded (rsync's -d DIR/ --delete).
Also shields a -x mount point's untraversed destination content.
2026-09-20 14:22:22 +02:00
TapTap 79a28cdb96 scanner: emit entries in rsync's sorted depth-first flist order
Buffer and sort each directory's inspected entries (non-directories
ascending, then directories ascending) and walk them depth-first via a
LIFO directory stack, so the sequential scanner's stream matches rsync
3.4.1's flist order.  This makes the --info=name transfer order and the
--delete-during/--delete-delay deletion sequence byte-identical to rsync
(differential tests in test_parity_order.py); --threads stays unordered
(no rsync analogue) and is documented as such.

Adds LIFO queue_push/queue_pop over the existing ring buffer.
2026-09-20 13:49:04 +02:00
TapTap 402cae80ad parity: rsync-exact relative basis-dir resolution and fuzzy eligibility
Resolve a relative --compare-dest/--copy-dest/--link-dest DIR against the
destination directory and append the file's transfer-relative name, as
rsync 3.4.1 does, instead of appending FastSync's source-mirrored wire
path (the historical spelling stays as a fallback for existing layouts).

Stop inheriting the ordinary delta engine's 16 KiB minimum and 10x size
ratio in the -y/--fuzzy candidate search: rsync's find_fuzzy has no
delta-size gate, so an oversized or sub-16-KiB sibling is now reused.
The ordinary delta path's bounds are unchanged.
2026-09-20 13:39:04 +02:00
TapTap 9691dba6f0 delete: reproduce rsync traversal order for extras removal
Collect each directory's entries up front and process extraneous
subdirectories first (descending name, depth-first), then extraneous
files (descending name), then descend into kept subdirectories in
ascending order.  Emit a trailing slash for deleted directories in
observers/dry-run output.  This matches rsync's delete order for
--delete-before/--delete-after/--delete-delay and for dry-run listings.
2026-09-20 13:15:04 +02:00
TapTap 558782d339 test: eliminate fork/write race in incremental-check server tests
CI / lint (push) Successful in 2m1s
CI / parity-fast (push) Skipped
CI / parity-full (push) Successful in 21s
CI / sanitizers (address) (push) Successful in 53s
CI / sanitizers (undefined) (push) Successful in 44s
CI / build-and-test (push) Successful in 1m11s
CI / fuzz-build (push) Successful in 47s
CI / coverage (push) Successful in 43s
CI / valgrind (push) Successful in 2m14s
The parent sends the file data body after the receiver's STATUS_NEXT, but
the forked child exited as soon as receive_incremental_check returned.  The
parent's send_data could then race the child's exit into a spurious EPIPE
(seen in the coverage job as test_server.c:1222), or the reverse: the parent
could be descheduled past the child's exit.

Keep the child alive until the parent closes its write end (drain to EOF),
and close the parent's write end before waitpid so the child can observe EOF.
Applied to the three tests sharing the pattern: size-mismatch, FIFO
destination, and FIFO basis.  Child exit status remains the authoritative
assertion.
2026-09-20 12:09:21 +02:00
TapTap 5597e74f6a docs(handoff): v2.28.0 released to main (PR #304, tag v2.28.0)
CI / lint (push) Successful in 2m0s
CI / parity-fast (push) Skipped
CI / parity-full (push) Successful in 19s
CI / sanitizers (address) (push) Successful in 51s
CI / sanitizers (undefined) (push) Successful in 43s
CI / build-and-test (push) Successful in 1m7s
CI / coverage (push) Failing after 31s
CI / fuzz-build (push) Successful in 47s
CI / valgrind (push) Successful in 2m13s
2026-09-20 01:23:49 +02:00
TapTap b4d54504f9 Release v2.28.0 (#304)
CI / lint (push) Successful in 2m1s
CI / parity-fast (push) Skipped
CI / parity-full (push) Successful in 19s
CI / sanitizers (address) (push) Successful in 50s
CI / sanitizers (undefined) (push) Successful in 43s
CI / build-and-test (push) Successful in 1m8s
CI / fuzz-build (push) Successful in 45s
CI / coverage (push) Successful in 43s
CI / valgrind (push) Successful in 2m13s
2026-09-20 01:17:25 +02:00
TapTap ee6523afac Release v2.28.0
CI / lint (push) Successful in 2m0s
CI / parity-fast (push) Skipped
CI / lint (pull_request) Successful in 2m0s
CI / parity-full (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / parity-full (push) Successful in 19s
CI / sanitizers (address) (push) Successful in 50s
CI / sanitizers (undefined) (push) Successful in 44s
CI / build-and-test (push) Successful in 1m8s
CI / coverage (push) Successful in 42s
CI / fuzz-build (push) Successful in 47s
CI / parity-fast (pull_request) Successful in 19s
CI / build-and-test (pull_request) Successful in 52s
CI / valgrind (push) Successful in 2m13s
- rsync-parity cycle: differential parity gate + tracks 1-6
- protocol 2.28.0 (batched wire changes)
- --delete defaults to delete-during; new --delete-commit, --verify-basis
- parity matrix 116/14/27 of 157
- tested: unit, integration, ASan, strict differential parity
2026-09-20 01:11:33 +02:00
TapTap 4163caa1d3 docs(handoff): record the rsync-parity 1-6 cycle (PR #303, protocol 2.28.0)
CI / lint (push) Successful in 1m57s
CI / parity-fast (push) Skipped
CI / parity-full (push) Successful in 20s
CI / sanitizers (address) (push) Successful in 49s
CI / sanitizers (undefined) (push) Successful in 42s
CI / build-and-test (push) Failing after 1m8s
CI / fuzz-build (push) Successful in 46s
CI / coverage (push) Successful in 42s
CI / valgrind (push) Successful in 2m13s
2026-09-19 17:21:26 +02:00
TapTap 10159dc120 Merge pull request 'feat(parity): rsync parity tracks 1-6 (protocol 2.28.0)' (#303) from feat/parity-2.28 into dev
CI / lint (push) Successful in 2m0s
CI / parity-fast (push) Skipped
CI / parity-full (push) Successful in 21s
CI / sanitizers (address) (push) Successful in 51s
CI / build-and-test (push) Successful in 1m7s
CI / sanitizers (undefined) (push) Successful in 43s
CI / fuzz-build (push) Successful in 47s
CI / coverage (push) Successful in 43s
CI / valgrind (push) Successful in 2m13s
2026-09-19 17:14:37 +02:00
TapTap cd7b96d0bb chore: remove manual-test scratch trees from branch
CI / lint (pull_request) Successful in 1m59s
CI / parity-full (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / parity-fast (pull_request) Successful in 17s
CI / build-and-test (pull_request) Successful in 49s
2026-09-19 17:10:12 +02:00
TapTap f6f49d536e test(delete): cover apply=false config-only carrier frame
Adds a unit test proving the config-only STATUS_DELETE_PLAN frame applies
--delete-missing-args exact deletions while walking no directory (the
--files-from-with-no-synced-dir fix).
2026-09-19 17:07:56 +02:00
TapTap 38d304103c fix(parity): review-wave fixes (basis stats over-report, filter-rule bounds)
- receiver stats: exclude basis-dir materializations (--link-dest/--copy-dest)
  from created/literal tallies; rsync reports 0 for a basis hit, so a fresh
  --link-dest --stats run now matches (differential test_link_dest_stats_matches_rsync)
- filter wire block: reject a pattern above the glob evaluation bound and lower
  MAX_FILTER_RULES to 1024, so a crafted rule list cannot amplify delete-walk
  glob work or install a rule that silently never protects
- negative tests for over-cap count and over-long pattern
- README: correct --delete default, --stats/--progress description, add
  --delete-commit; RSYNC_COMPAT stale version labels/overclaim fixed
2026-09-19 17:04:12 +02:00
TapTap 4b09213b88 chore: remove accidental scratch tree from branch 2026-09-19 16:30:21 +02:00
TapTap 36fd0774e8 feat(delete): default --delete to rsync delete-during; add --delete-commit
- plain --delete with no timing flag now selects delete-during (progressive
  deletion, matching rsync and avoiding the full old+new tree peak)
- new long-only FastSync --delete-commit restores the old atomic behavior
  (delete only after the whole transfer succeeds); timing-identical to
  --delete-after, implemented via the same wire bool
- timing flags are mutually exclusive; --delete-commit conflicts with other
  timings; --delete-before/--delete-during rows reworded per Phase-0 probes
- CHANGELOG migration note; tally unchanged 116/14/27
2026-09-19 16:29:48 +02:00
TapTap 711b7e50b3 docs(parity): --fuzzy name heuristic matches rsync; reclassify to caveat
Probe shows the candidate choice is observable only as --stats bandwidth
counters (tree/exit always identical). FastSync already uses rsync's
fuzzy_distance/find_filename_suffix name heuristic; the residual is the
narrower delta eligibility window (>=16 KiB, <=10x) vs rsync's wider one.
Row -> caveat; tally 116/14/27.
2026-09-19 16:06:24 +02:00
TapTap 67076bf218 feat(basis): rsync quick-check default + FastSync-only --verify-basis
- default basis match is rsync's metadata quick-check (size + mtime; size-only
  drops mtime; -I disables), no mandatory content digest
- new long-only --verify-basis (wire bool, protocol stays 2.28.0) restores the
  strict whole-file content equality
- --copy-dest re-applies source attributes; basis-hit 256 MiB cap removed by
  streaming the copy/hash; basis miss keeps the normal payload bound
- compare/copy/link-dest rows -> caveat; tally 116/13/28
2026-09-19 15:55:51 +02:00
TapTap 8ec8cb7203 test(parity): dest-only excluded entry protected under default --delete
Adds the exclude_protect_dest_only differential and flips --delete-excluded
to parity now that receiver-side rules protect a destination-only excluded
entry like rsync; tally 116/10/31.
2026-09-19 13:53:50 +02:00
TapTap 82959395fb feat(filter): receiver-side protect/risk engine for dest-only entries
- new bounded config-wire block (BLOCK_PROTECT_RULES) serializes the sender's
  compiled filter rules to the receiver (bounded count + 256 KiB patterns;
  strict action/sides validation)
- receiver evaluates protect/risk in the whole-tree extras walk and the
  per-directory delete plans, so a dest-only entry matching 'P' is kept like
  rsync; dry-run would-delete enumeration also honours it
- --filter flips to parity (115/11/31); per-dir merge receiver re-derivation
  remains the documented residual
2026-09-19 13:52:48 +02:00
TapTap c9f94ea46e docs(parity): --checksum-choice observable-equivalent to rsync; reclassify
Probe shows the block-checksum choice is not observable in the parity surface:
%c, Matched/Literal data and the destination tree are invariant across
xxh64/xxh128/xxh3/md5/md4/sha1 and the transfer,pre-transfer form; only %C
changes, and it is byte-identical to rsync. FastSync's fixed xxHash32 block
strong sum is collision-safe within the payload cap. Row -> parity (114/11/32).
2026-09-19 13:23:02 +02:00
TapTap eff9852038 feat(codecs): per-codec level defaults, RSYNC_*_LIST auto, zlibx reclassify
- rsync 3.4.1 per-codec defaults (zstd 3, zlib/zlibx 6, lz4 level ignored)
  and per-codec clamping; explicit --zl still wins
- auto resolves via whitespace-separated RSYNC_COMPRESS_LIST /
  RSYNC_CHECKSUM_LIST (first supported wins; all-unknown exits 4)
- zlibx reclassified: FastSync's zlib stream already excludes matched data,
  so its tree/stdout/exit match zlib
- --compress/-z and --compress-choice flip to parity (113/12/32)
2026-09-19 13:14:48 +02:00
TapTap c80098623f feat(progress): opt-in paths-only pre-count for rsync to-chk parity
- when --progress/--info=progress is requested, a metadata-only pre-scan
  builds the full file-list total and directory names so the to-chk
  denominator counts every regular/dir/link/special entry like rsync
- per-directory/symlink/special name lines emitted; sequential and --threads
- reuses the --delete-during/delay pre-scan when present; non-progress runs
  take no extra pass
- --progress stays a caveat (emission order still differs); single-file output
  remains byte-identical
2026-09-19 12:49:56 +02:00
TapTap 6119e1e75c feat(stats): receiver-observed created/literal counters (protocol 2.28.0)
- PROTOCOL_VERSION 2.27.0 -> 2.28.0; STATUS_STATS gains literal_bytes and
  the created reg/dir/link/special counters (golden wire updated)
- receiver reports which destination entries it newly created, including
  implicitly-created parent directories below the logical transfer root, so
  Number of created files carries rsync's per-type breakdown
- Literal data is now exact for a delta transfer (receiver counts the literal
  fragments it stored)
- differential-tested vs rsync 3.4.1 for fresh-create, update and delta
2026-09-19 10:45:23 +02:00
TapTap 25062352f6 fix(parity): dry-run delete protections, delete-delay actual-removal budget, --info name2
- -n/--delete sends the same protected/size-skipped/scope as a real run, so
  the read-only would-delete walk no longer over-reports (row -> caveat)
- --delete-delay charges --max-delete on actual removals and recursively
  re-scans a refilled deferred directory at commit; independent deferred cap
- --info=name emits the leading ./ root line and name2 'is uptodate' lines
- differential tests promoted from residual pins to rsync parity assertions
2026-09-19 09:25:07 +02:00
TapTap 00d628d4ba Merge pull request 'test(parity): empty allowlist, assert max_delete invariants' (#302) from fix/parity-allowlist into dev
CI / lint (push) Successful in 1m40s
CI / parity-fast (push) Skipped
CI / parity-full (push) Successful in 20s
CI / sanitizers (address) (push) Successful in 51s
CI / sanitizers (undefined) (push) Successful in 41s
CI / build-and-test (push) Successful in 1m8s
CI / fuzz-build (push) Successful in 47s
CI / coverage (push) Successful in 42s
CI / valgrind (push) Successful in 2m13s
2026-09-18 22:31:46 +02:00
TapTap 3545d88905 test(parity): assert max_delete invariants, empty the parity allowlist
CI / lint (pull_request) Successful in 1m40s
CI / parity-full (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / parity-fast (pull_request) Successful in 18s
CI / build-and-test (pull_request) Successful in 52s
The max_delete case allowlisted the tree aspect because the surviving extras after a partial --max-delete abort are deletion-order dependent. Under FASTSYNC_PARITY_STRICT a run where the orders coincide was reported as a stale entry (CI failure), while removing the entry made the order-dependent tree mismatch fail. Add a per-case 'compare_tree' flag: max_delete now asserts rc=25 plus the survivor count via extra_check instead of exact tree identity, so the allowlist can be empty. Burn-down reached zero.
2026-09-18 22:28:38 +02:00
TapTap 596a039454 Merge pull request 'fix(parity): ⚠️ residual burn-down + review fixes (protocol 2.27.0)' (#301) from feat/parity-fixes into dev
CI / lint (push) Successful in 1m42s
CI / parity-fast (push) Skipped
CI / parity-full (push) Failing after 17s
CI / build-and-test (push) Successful in 1m4s
CI / sanitizers (address) (push) Successful in 49s
CI / sanitizers (undefined) (push) Successful in 43s
CI / fuzz-build (push) Successful in 40s
CI / coverage (push) Successful in 40s
CI / valgrind (push) Successful in 2m12s
2026-09-18 22:19:25 +02:00
TapTap ae037cc27b Merge pull request 'test(parity): differential rsync 3.4.1 parity gate + CI' (#300) from feat/parity-gate into dev
CI / lint (push) Successful in 1m40s
CI / parity-fast (push) Skipped
CI / parity-full (push) Failing after 19s
CI / build-and-test (push) Successful in 58s
CI / sanitizers (address) (push) Successful in 51s
CI / sanitizers (undefined) (push) Successful in 45s
CI / fuzz-build (push) Successful in 43s
CI / coverage (push) Successful in 41s
CI / valgrind (push) Successful in 2m15s
2026-09-18 22:18:53 +02:00
TapTap 3d0672a721 docs(agents): note FASTSYNC_UNDER_VALGRIND for manual valgrind runs
CI / lint (pull_request) Successful in 1m40s
CI / parity-full (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / parity-fast (pull_request) Successful in 16s
CI / build-and-test (pull_request) Successful in 51s
2026-09-18 22:15:38 +02:00
TapTap b82aab72c5 Merge branch 'fix/parity-review-c' into feat/parity-fixes 2026-09-18 22:11:07 +02:00
TapTap 8e7764007d Merge branch 'fix/parity-review-b' into feat/parity-fixes
# Conflicts:
#	src/shared/delete_plan.h
#	tests/integration/test_delete_timing_parity.py
2026-09-18 22:11:02 +02:00
TapTap 1042d15db7 docs(parity): correct delete-delay budget, fuzzy, and test-review gaps
- --delete-delay: state that the reported count advances on actual removal
  while --max-delete is charged at plan/snapshot time (defer_add/planned).
  A new differential shows rsync instead charges on actual removals and
  recursively removes a queued directory, so the row moves to Caveat
  (matrix 111/13/33) and the residual is pinned by tests.
- Add a deterministic unit test (plan-time budget charge), a FastSync
  integration test (byte-barrier refill + --max-delete), and an rsync
  differential for a refilled deferred directory.
- Soften the --fuzzy summary: the tree is byte-exact by design, so it is
  pinned by the threshold suite, not a byte-level differential.
- README: describe what -m parity actually selects; fix the allowlist
  example to the real max_delete entry.
- xdist-safe delete-timing fixture names (timing and --threads mode).
- Renumber the duplicate HANDOFF item 9 to 10; add the missing final
  newline to test_checksum.c.
- Expose ignore_errors_allows_delete and unit-test the deletion gate
  without a privileged source directory; update the stale deleted-count
  doc comment.

No production behavior changes.
2026-09-18 22:10:14 +02:00
TapTap cbe37a77dd refactor(parity): address code-quality review findings
- cli: remove UB in --bwlimit scaling (range-check the double product before
  casting, drop atoi for the +/-1 form) and add huge/boundary unit tests
- test: widen the CI throttle wall-clock band to [1.5, 4.5]s with a 2s
  cross-tolerance so a loaded runner cannot flake it
- log: drop the unused LOG_INFO_BACKUP bit; --info=backup is accepted-but-
  silent like the other rsync-only categories
- client_send: remove the duplicate delete_display_path forward declaration
- utils: add non-allocating utils_strip_transfer_root and use it from
  scanner_note_nonreg and delete_display_path (was duplicated logic)
- scanner: lstat() instead of stat() when re-reading an empty dir's metadata
- file: drop the no-op else-if and the redundant ELOOP arm in
  file_ensure_directory_secure (symlinks are refused anyway)
- format/stats: document literal_data as whole-file accurate (delta upper
  bound) instead of claiming literal bytes sent
- docs: refresh stale protocol 2.26.0 labels to 2.27.0
2026-09-18 22:00:32 +02:00
TapTap a5d45ef266 fix(parity): init delete_suppressed; gate deleted-path retention
- Initialize PipelineContextSender.delete_suppressed=false: an uninitialized
  true silently skipped the --delete keep-set manifest under -m/--threads,
  so destination extras were never removed.
- Allocate/install the receiver deleted-path observer only when
  report_deletes is set (--info=del / -i / --out-format under --delete), cap
  the retained list at MAX_MANIFEST_ENTRIES, and free already-created lists
  on the receiver-pipeline create failure path.
- Validate report_deletes/report_stats/report_dest_info on receive.
- Correct stale comments (config.h report_deletes, delete_plan.h deleted
  count, multiprocessing.h stats locking, utils.h observer placement).
- Tests: sender delete_suppressed init, report_deletes gating (unit), and
  -m/--delete default delete-after keep-set removal (integration).
2026-09-18 21:56:44 +02:00
TapTap a960391b34 fix(cli): guard parse_bwlimit_value against NULL (cppcheck) 2026-09-18 21:26:28 +02:00
TapTap 134f8b027a Merge branch 'fix/parity-fs' into feat/parity-fixes
# Conflicts:
#	RSYNC_COMPAT.md
2026-09-18 21:19:31 +02:00
TapTap eb7e3fd2e0 test(parity): option-wave rsync differentials + docs (109/21/26)
- tests/integration/test_option_parity.py: bwlimit parse matrix + throttle
  rate, --info flist/name/nonreg/del(dry+real+itemize)/remove, real-setpriv
  --ignore-errors, rsync-daemon -M forwarding evidence, filter protect/risk
  destination-only divergence pin.
- RSYNC_COMPAT.md: tally 109/21/26; --bwlimit and --ignore-errors -> Parity,
  --filter and -M -> Divergent with differential rationale, --info residual
  narrowed to the categories with no client-observable event.
- HANDOFF/README: protocol 2.27.0, option-wave summary.
2026-09-18 21:16:12 +02:00
TapTap 6a129b54d4 feat(parity): real --info=del deletion lines + rsync throttle pacing
- Wire: config frame gains report_deletes (protocol 2.26.0 -> 2.27.0); the
  receiver lists actually-removed paths in the STATUS_STATS path list, so the
  sender prints rsync's `deleting PATH` / `*deleting   PATH` lines for a real
  --delete run (and -i/out-format).  Observers threaded through the manifest,
  missing-args and per-directory delete engines; golden wire len/hash updated.
- bwlimit: throttle now paces like rsync 3.4.1 -- ~100ms burst capacity and the
  sleep is no longer credited as refill, so 4 MiB at 1024/2048 KiB/s matches
  rsync within ~4% (was ~2x too fast).
2026-09-18 21:06:20 +02:00
TapTap c7b2c7eb2b docs(parity): dir-time entry is record-only; empty dirs carried by STATUS_MKDIR 2026-09-18 20:50:22 +02:00
TapTap e77dfbec70 fix(fs): differential-test and reclassify basis dirs, delay-updates, fuzzy, dry-run
Each row's exact residual reproduced against rsync 3.4.1:
- basis dirs (--compare/copy/link-dest): FastSync xxHash-verifies a basis hit
  while rsync --size-only installs the wrong same-size basis content.
- --delay-updates: the fixed .fastsync-stage name wipes an unrelated
  destination entry of that name even without --delete; rsync leaves it.
- --fuzzy: deterministic name/size heuristic (10x window), not rsync's matcher.
- --dry-run: would-delete report over-reports the updated file and an
  excluded-but-protected extra, and ordering differs.
Docs: RSYNC_COMPAT tally 109/16/31; HANDOFF item 9. clang-format + cppcheck +
ASan + full suite clean.
2026-09-18 20:49:51 +02:00
TapTap f3ac4df4d0 fix(parity): rsync bwlimit units, --info categories, --ignore-errors deletion semantics
- --bwlimit: faithful port of rsync 3.4.1 parse_size_arg (default KiB/s,
  binary K/M/G/T/P, decimal KB/MB, KiB/MiB, decimals, 0 = unlimited, 512-byte
  floor, (size+512)/1024 quantization).  Unit tests + docs.
- --info: wire del/remove/name/flist/nonreg/progress to real FastSync events in
  rsync's line format (deleting PATH, sender removed NAME, name lines,
  'sending incremental file list', skipping non-regular file "NAME"); name no
  longer aliases copy; --info=progress drives the progress path and report_stats.
- --ignore-errors: match rsync's default -- a source I/O error skips deletion
  unless --ignore-errors, while the readable tree still transfers and the run
  exits 23.  Covers all delete timings and both send paths.
2026-09-18 20:44:24 +02:00
TapTap 91197fd7cf fix(fs): rsync push iconv direction; gate empty-dir emission; reclassify --temp-dir
- --iconv now matches rsync's push direction: the destination charset is the
  client spec's REMOTE half, so a default receiver writes wire names verbatim;
  a server's own --iconv LOCAL overrides it (daemon charset analog). Updated
  unit + integration tests and added a default-server differential gate case.
- Empty-directory emission is gated behind a new ScannerOptions.emit_empty_dirs
  set only by the real sender, so low-level scanner helpers keep the historical
  file-only list.
- --temp-dir reclassified to Divergent: relative dirs match rsync exactly
  (resolved under the destination), but an absolute path is deliberately
  rejected by the confined receiver; differential test added.
- Docs/tally: 109 Parity / 22 Caveat / 25 Divergent.
2026-09-18 20:39:20 +02:00
TapTap 13708352ec fix(fs): recreate empty dirs, replace blocking file; -R --no-implied-dirs implied parents
- Recursive scans emit a directory entry for every traversed directory that
  produced no transferred child, so empty source dirs (and dirs emptied by
  filtering) are recreated like rsync; -m prunes them, --files-from/--list-only
  never emit implicit dirs.
- A directory entry now replaces a destination regular file (rsync removes the
  non-directory) instead of aborting; confined to the secure parent fd.
- -R --no-implied-dirs --files-from: stop refusing a listed file whose parent
  is not listed; create the implied parent with default attributes (no source
  metadata is captured for it), matching rsync 3.4.1.
- Differential gate: drop min_size/empty_dirs_recursive/dirs_plain allowlist
  entries (dirs_plain now hands FastSync the same trailing-slash source as
  rsync); add differential test for the files-from implied parent.
- Docs: -d row -> Parity, tally 108/24/24.

No wire change (PROTOCOL_VERSION stays 2.26.0).
2026-09-18 20:25:42 +02:00
TapTap d162d93570 Merge branch 'feat/parity-gate' into fix/parity-fs 2026-09-18 20:20:25 +02:00
TapTap 9fa1696eff fix(parity): actual-removal delete-delay counts, rsync-accurate stats/progress/%C
- --delete-delay: count/track only entries actually removed; a directory
  refilled before commit (ENOTEMPTY) no longer inflates Number of deleted
  files or the --max-delete budget (unit + integration + rsync differential).
- --stats: per-type Number of files breakdown; only stored regular files
  count as transferred; transferred/literal byte totals and Total file size
  (symlink target lengths) now match rsync for whole-file transfers.
- --progress: print the leading ./ root line and include the root entry in
  the to-chk denominator (single-file output byte-identical to rsync).
- --out-format %C: use the selected transfer checksum and render every
  algorithm exactly like rsync; checksum_digest_file gains md4/sha1/none.
  Reclassify --out-format to Divergent (protocol-specific %b/delta-%c).
- Docs: RSYNC_COMPAT tally 107/25/24, HANDOFF update. No wire change.
2026-09-18 20:13:29 +02:00
TapTap b705fb807f test(parity): add differential rsync-parity CI gate
CI / lint (pull_request) Successful in 1m42s
CI / parity-full (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / parity-fast (pull_request) Successful in 19s
CI / build-and-test (pull_request) Successful in 48s
Add tests/integration/test_differential_parity.py plus a shared
parity_harness.py and a data-driven parity_caveats.py allowlist.  The gate
runs real rsync 3.4.1 and FastSync over the same corpora, compares the
destination trees (paths, hashes, symlink targets, modes, hard-link
grouping) and the normalized -i/--stats/--out-format output, and fails on
any difference not listed in the allowlist.  Stale allowlist entries warn
(or fail under FASTSYNC_PARITY_STRICT=1) so the residual list shrinks.

Register parity/parity_ci markers and wire a fast PR job (parity_ci) plus a
push-only full job (parity, strict) into .gitea/workflows/ci.yaml.
Document the gate and the allowlist workflow in tests/integration/README.md.
2026-09-18 19:56:25 +02:00
136 changed files with 24521 additions and 8670 deletions
+41
View File
@@ -49,6 +49,47 @@ jobs:
if: github.event_name == 'push'
run: python3 -m pytest tests/integration/ -n 4 --dist=load -m "not setpriv" --durations=25 --tb=short -q
# Differential rsync-parity gate: runs real rsync 3.4.1 and FastSync over the
# same corpora and compares destinations + normalized output. The fast subset
# guards the ✅ surface on every PR; the full set (with FASTSYNC_PARITY_STRICT
# so a fixed caveat must be removed from the allowlist) burns the documented
# ⚠️/❌ residuals down on push. See tests/integration/README.md.
parity-fast:
runs-on: ubuntu-latest
container: gitea.tap-tap.win/taptap/fastsync-ci:v11
needs: lint
if: github.event_name == 'pull_request'
steps:
- name: Checkout
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Configure
run: cmake -B build -S . -DSTRICT_WARNINGS=ON
- name: Build
run: cmake --build build -j$(nproc)
- name: Differential parity (fast subset)
run: python3 -m pytest tests/integration/test_differential_parity.py -n 4 --dist=load -m parity_ci -q
parity-full:
runs-on: ubuntu-latest
container: gitea.tap-tap.win/taptap/fastsync-ci:v11
needs: lint
if: github.event_name == 'push'
steps:
- name: Checkout
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Configure
run: cmake -B build -S . -DSTRICT_WARNINGS=ON
- name: Build
run: cmake --build build -j$(nproc)
- name: Differential parity (full set)
run: FASTSYNC_PARITY_STRICT=1 python3 -m pytest tests/integration/test_differential_parity.py -n 4 --dist=load -m parity -q
sanitizers:
runs-on: ubuntu-latest
container: gitea.tap-tap.win/taptap/fastsync-ci:v11
+9
View File
@@ -12,3 +12,12 @@ build_docker2/
# Test/run artifacts
root/
test_partial_install_tmp/
# Editor/tooling + test caches/artifacts
.pytest_cache/
*.gcda
*.gcno
*.gcov
di/
test_data-manual/
*.log
+1 -1
View File
@@ -16,7 +16,7 @@ Ask the user or determine from context:
- **Minor** (x.Y.0) — new features, backward compatible
- **Patch** (x.y.Z) — bug fixes, no protocol changes
Current version: `PROTOCOL_VERSION "2.26.0"` in `src/shared/config.h`
Current version: `PROTOCOL_VERSION "2.29.0"` in `src/shared/config.h`
### Step 2: Check Protocol Version
+22 -6
View File
@@ -4,9 +4,9 @@ FastSync is a high-performance file synchronization system written in C11. It su
## Dependency installation
**CI rule:** never add `apt-get install` / `pip install` steps to CI workflows — use the custom Docker image instead. The image is built from the repo-root `Dockerfile` and is the same image CI uses: `gitea.tap-tap.win/taptap/fastsync-ci:v11`. It contains the full toolchain: gcc/g++, CMake, libzstd-dev, libssl-dev, make, git, cppcheck, clang-format, python3 + pytest + pytest-xdist, openssh-client, Node.js, plus `rsync` 3.4.1 (with zstd/xxhash/lz4), `acl` and `attr` (setfacl/getfacl, setfattr/getfattr) for drop-in parity tests.
**CI rule:** never add `apt-get install` / `pip install` steps to CI workflows — use the custom Docker image instead. The image is built from the repo-root `Dockerfile` and is the same image CI uses: `gitea.tap-tap.win/taptap/fastsync-ci:v11`. It contains the full toolchain: gcc/g++, CMake, libzstd-dev, zlib1g-dev, liblz4-dev, libxxhash-dev, libssl-dev, make, git, cppcheck, clang-format, python3 + pytest + pytest-xdist, openssh-client, Node.js, plus `rsync` 3.4.1 (with zstd/xxhash/lz4), `acl` and `attr` (setfacl/getfacl, setfattr/getfattr) for drop-in parity tests. (CMake hard-requires zstd, zlib, and lz4; xxHash is fetched via `FetchContent`.)
**Host rule:** for local development, use `nix-shell` (see `README.md`) which provides zstd, OpenSSL, CMake, and gcc. The Docker image can also be used locally for CI parity.
**Host rule:** for local development, use `nix-shell` (see `README.md`) which provides zstd, zlib, lz4, OpenSSL, CMake, and gcc. The Docker image can also be used locally for CI parity.
```bash
# Use the prebuilt CI image directly (faster, guaranteed CI parity)
@@ -38,11 +38,12 @@ If a dependency is missing from the CI image, add it to the `Dockerfile` (and re
When configuring for CI parity, use:
```bash
cmake -B build -S . -DSTRICT_WARNINGS=ON # -Wextra -Wpedantic -Werror
cmake -B build -S . -DSANITIZER=address # AddressSanitizer (ASan)
cmake -B build -S . -DSANITIZER=thread # ThreadSanitizer (TSan)
cmake -B build -S . -DSANITIZER=address # AddressSanitizer (ASan); in the CI matrix
cmake -B build -S . -DSANITIZER=undefined # UndefinedBehaviorSanitizer (UBSan); in the CI matrix
cmake -B build -S . -DSANITIZER=thread # ThreadSanitizer (TSan); local-only, NOT in CI
```
The CI workflow (`.gitea/workflows/ci.yaml`) runs lint (clang-format, cppcheck), then a **fast PR gate** — build + unit + a representative subset of integration tests marked `@pytest.mark.ci`, parallelized with pytest-xdist (`-n 4 --dist=load`). The full coverage jobs (full integration suite as `-m "not setpriv"`, sanitizer, fuzz, coverage, valgrind) run **only on push to `dev`/`main`**; pull requests skip them to keep PR CI under ~3 minutes. The two `setpriv` privilege tests are excluded from CI via a marker because their result depends on the runner/container uid and host mount permissions.
The CI workflow (`.gitea/workflows/ci.yaml`) runs lint (clang-format, cppcheck), then a **fast PR gate** — build + unit + a representative subset of integration tests marked `@pytest.mark.ci`, parallelized with pytest-xdist (`-n 4 --dist=load`). The full coverage jobs (full integration suite as `-m "not setpriv"`, the `address`+`undefined` sanitizer matrix, fuzz, coverage, valgrind) run **only on push to `dev`/`main`**; pull requests skip them to keep PR CI under ~3 minutes. TSan is not part of the CI matrix and is a local-only configuration. The `setpriv`-marked privilege tests (four decorated functions, collecting to eight instances because two are parametrized) are excluded from CI via a marker because their result depends on the runner/container uid and host mount permissions.
## Build
@@ -56,6 +57,21 @@ cmake -B build -S . && cmake --build build -j$(nproc)
./build/tests # unit tests
python3 -m pytest tests/integration/ -n 4 --dist=load -m "not setpriv" # full integration suite (CI excludes env-dependent privilege tests)
python3 -m pytest tests/integration/ -n 4 --dist=load -m ci # PR-gate subset only
# Differential rsync-parity gate (real rsync 3.4.1 vs FastSync)
python3 -m pytest tests/integration/test_differential_parity.py -n 4 --dist=load -m parity_ci # fast PR subset
python3 -m pytest tests/integration/test_differential_parity.py -n 4 --dist=load -m parity # full set
```
See `tests/integration/README.md` for the differential parity gate and its
`parity_caveats.py` allowlist (the residual burn-down mechanism).
Unit tests under valgrind must set `FASTSYNC_UNDER_VALGRIND=1` (CI does): the
tests use it to skip fork-based tests, because valgrind 3.22 does not expose
`vgpreload` in the guest's `/proc/self/maps`.
```bash
FASTSYNC_UNDER_VALGRIND=1 valgrind --leak-check=full --show-leak-kinds=definite --error-exitcode=1 ./build/tests
```
## CI Workflow — Waiting for Results
@@ -90,7 +106,7 @@ Two main branches: `dev` (integration) and `main` (stable releases).
### Rules
- **All PRs target `dev`** — never target `main` directly
- **`dev` is the default branch** in Gitea repo settings
- **`dev` is intended to be the default branch** in Gitea repo settings — verify in the repo settings, since this clone's `origin/HEAD` still points at `main`
- **`main` is protected** — only merged from `dev` via PR with 2 approvals + full CI pass
- **Feature/bug branches** branch from `dev`, PR back to `dev`
- **`dev` → `main` merges** happen on-demand or weekly, requiring full CI + review
+284
View File
@@ -4,8 +4,292 @@ All notable changes to FastSync are documented here. Versions match
`PROTOCOL_VERSION` (printed by `fastsync --version`); the client and server must
run the same version because the handshake is strict.
## [Unreleased]
## [2.29.0] - 2026-09-23
The rsync-parity cycle 2.29 (no wire change; `PROTOCOL_VERSION` stays 2.28.0).
`RSYNC_COMPAT.md` moves from **116 ✅ / 14 ⚠️ / 27 ❌** to
**120 ✅ / 10 ⚠️ / 27 ❌** of 157 rows.
An audit cycle follows on the same wire version (`PROTOCOL_VERSION` stays
2.28.0): a security-and-correctness pass over the parity-2.29 baseline, plus a
set of audit follow-ups (filter merge modifiers, the `--inplace`/`--partial-dir`
conflict, credential-file hardening, and small leak/log/test fixes). It fixes
a `--temp-dir` symlink escape, gates client-controlled special permission bits,
corrects `--partial-dir`/`--bwlimit`/`-z` behavior, handles unsupported filter
modifiers, and tightens client and wire validation. The only parity
reclassification is `--filter=RULE` moving ✅ → ⚠️, because its merge-only
`e`/`n`/`w`/`-` modifiers are now accepted and consumed but their semantics
remain unimplemented (accepted-but-ignored); the matrix is therefore **119 ✅ /
11 ⚠️ / 27 ❌** of 157 rows. The affected rows' notes and the summary tally in
`RSYNC_COMPAT.md` were updated. A following triage-fix cycle (see **Triage
fixes** below) moves `-F` and `-i` to ⚠️, for a final **117 ✅ / 13 ⚠️ / 27 ❌**
of 157 rows.
A no-wire parity burn-down cycle follows on 2.28.0: it accepts
`--inc-recursive`/`--no-inc-recursive` as inert no-ops, accepts an absolute
`--temp-dir` that canonicalizes inside the receive root, closes the
`--delete-before` phase-0 divergence (both the single-threaded and `--threads`
data passes replay the pre-scan list), makes `--fake-super` interoperable with
rsync's `user.rsync.%stat` key/grammar (regular files and char/block devices
faked as regular files), turns a failed device `mknod` into a continuing
per-entry failure, and accepts a practical subset of rsync's `rsyncd.conf`
grammar (modules are read-only by default, and accepted-but-unenforced
access-control keys emit a startup warning). The matrix moves to **119 ✅ /
14 ⚠️ / 24 ❌** of 157 rows.
A structural cycle then lands a transport I/O vtable over TCP/TLS (fixing the
TLS-multithreaded sendfile path and making the per-thread SSL resolution
explicit) and bumps the wire to **2.29.0**: the `STATUS_SYMLINK` frame grows an
optional symlink-xattr block (captured no-follow with `llistxattr`/`lgetxattr`,
applied no-follow with `lsetxattr`). Because the handshake is strict, 2.28.0 and
2.29.0 peers are incompatible. Note: Linux refuses to associate xattrs with a
symlink at all, so the symlink-xattr block is a no-op on Linux and is carried
for correctness on platforms/filesystems that do support it; the config-frame
layout is unchanged (golden length still 886).
### Changed
- **rsync-exact traversal order.** The sequential scanner now walks each
directory's entries in rsync 3.4.1's flist order (non-directories ascending,
then directories ascending, depth-first), so `--info=name`, the
`--delete-during`/`--delete-delay`/`-n` would-delete order and the partial
`--max-delete` survivor set match rsync byte-for-byte. `--threads` has no
rsync analogue and stays unordered.
- **Delete timing.** The complete `--delete-during`/`--delete-delay`
per-directory plan set is transmitted before the first data frame, so a
mid-transfer abort has already removed every planned extra like rsync's
generator; `-d/--dirs` uses per-directory plans (shielded untraversed
subdirectories) instead of the end-of-transfer commit. `-n`, `--delete`,
`--del`/`--delete-during` and `--delete-delay` are now ✅ Parity.
- **Basis directories.** A relative `--compare-dest`/`--copy-dest`/`--link-dest`
DIR resolves against the destination directory with the transfer-relative
name appended, exactly like rsync 3.4.1.
- **`-y`/`--fuzzy`.** The candidate search no longer inherits the ordinary delta
engine's 16 KiB minimum or 10× size-ratio bound, so an oversized or
sub-16-KiB sibling is reused exactly as rsync reuses it.
- `--info=mount` prints rsync's mount-point skip line (repeated `-xx` drops the
mount-point directory); `--info=stats` enables the `--stats` block; `-x` is
repeatable. `--stats` counts traversed directories for the `Number of files`
breakdown under a plain `-r` scan. `--debug` emits real output for
`flist`/`del`/`hash`/`deltasum`/`recv`/`filter`/`send`.
### Known residuals
- `--progress` and `--info` still need a receiver→sender event channel for the
root `./` line, ancestor-directory suppression, receiver-side `skip`/`backup`
wording, and symlink/empty-directory quick-checks.
- `--delete-before`'s phase-0 late-file divergence remains (rsync's pre-scan
fixes the file list before the data pass).
- A single file larger than 256 MiB cannot be streamed in the default path
(a general whole-file limit, not basis-specific).
- `--stats` byte totals and `--msgs2stderr` stay documented divergences.
### Security
- **`--temp-dir` symlink escape fixed.** The receiver's scratch directory was
opened with a bare `open()`, so a symlink planted under the receive root could
redirect receiver scratch files outside the authorized root. The opened
directory is now judged by the real path of its fd (`/proc/self/fd` via
`realpath`) and an escaping target is refused (`EACCES`, logged); an in-root
link to another filesystem (the `EXDEV` fallback case) still works.
- **Client-controlled special bits masked when super-user activities are not
permitted.** Setuid/setgid/sticky bits (`--perms`, `--chmod`, the symlink and
special-node paths, and deferred directory modes) are now stripped when the
connection forbids super activities (`--no-super`, a non-opted daemon module,
a privileged listener without `--allow-super`); exact rsync semantics are
preserved wherever super activities are permitted.
- **Daemon umask no longer forced to `0`.** `daemonize()` now sets the
conventional `022`, so implied parent directories created without `-p` are no
longer world-writable `0777`.
- **Daemon modules are read-only by default.** A `--daemon` module is now
served read-only unless it sets `read only = no` (or rsync's `write only =
yes`), matching rsync: a real `rsyncd.conf` that omits `read only` is no
longer silently writable. A global `read only` still sets the default for
later modules, and an explicit module value wins. This is a behavior change
for existing FastSync-native configs that relied on the old writable default;
add `read only = no` to keep them writable. An rsync `write only = yes` is
mapped to writability (FastSync is push-only, so a module can never be read
from the network).
- **Accepted-but-unenforced rsync security keys now warn at startup.** The
rsync keys FastSync recognizes but does not implement — `secrets file`,
`refuse options`, `exclude`/`include`/`filter`, `max size`/`min size`,
`pre-xfer exec`/`post-xfer exec`, `incoming chmod`/`outgoing chmod`,
`name converter`, `use chroot`, `uid`/`gid`, and the rest of the
access-control set — load for migration compatibility but now emit a
`WARN` naming the key (and module) so an operator does not believe the
restriction is enforced. `auth users`/`secrets file` stay fail-closed: a
module declaring `auth users` still requires a FastSync credential store.
- **Credentials and signal handling hardened.** Secret files are opened with
`O_NOFOLLOW|O_NONBLOCK` (while allowing fd-backed store paths and bound-waiting
a FIFO read for ~3 s so a slow process substitution works but a connected-but-
silent FIFO cannot hang), and signal handlers use `sigaction` with
async-signal-safe bodies.
### Fixed
- **`-z` on 100–256 MiB files.** The decompressor's internal ceiling was 100 MiB
while the receiver advertises and the sender compresses whole files up to
`MAX_RECEIVE_WHOLE_FILE_SIZE` (256 MiB), so `-z` on a 100–256 MiB regular file
failed with `Declared decompressed size exceeds 104857600 bytes`. The ceiling
is now defined in terms of the protocol whole-file bound (still an
allocation-clamped bomb guard).
- **`--bwlimit` now paces `--sendfile`.** The plaintext-TCP `--sendfile` fast
path bypassed the protocol's token bucket, so the limit was ignored there. It
now throttles through the same per-session leaky bucket as the TLS path.
- **`--partial-dir` implies `--partial`.** Matching rsync 3.4.1 (which sets
`keep_partial` after option parsing), `--partial-dir=DIR` alone retains an
interrupted transfer's partial and wins over an explicit `--no-partial`;
`--inplace` still bypasses the partial machinery, and combining `--inplace`
with `--partial-dir` is now rejected up front with rsync's message
(`--inplace cannot be used with --partial-dir`).
- **Filter modifiers handled.** The `x` xattr-name modifier is rejected with a
clear error everywhere. The merge-only `e`/`n`/`w` and `-` modifiers are now
accepted and consumed on `merge`/`dir-merge` rules (so they no longer leak
into the merge filename) while still being rejected on non-merge rules,
matching rsync; their semantics remain unimplemented (accepted-but-ignored).
Glued patterns (`-newfile`, `-e2e`) and mixed tokens (`H,!secret`) keep their
historical parsing.
- **Credential-file reads hardened.** Secret files (`--password-file`/
`--early-input`/`--hash-credentials` input) are opened with `O_NOFOLLOW`, so a
symlinked credential path now fails closed (`ELOOP`) instead of being followed
before the owner/mode gate; literal fd-backed paths (`/dev/fd/<digits>`,
`/proc/self/fd/<digits>`) are exempt so process substitution still works. A
FIFO/process-substitution read now waits under a bounded ~3 s deadline for its
writer, so a slow producer works while a connected-but-silent FIFO fails
instead of hanging.
- **Miscellaneous correctness fixes:** `--filter` rule count is checked
client-side against `MAX_FILTER_RULES` before any network I/O (the receiver
still re-checks the expanded count); unknown wire `Status` values are rejected
as protocol errors; a mutex leak on an init-failure path, an `errno` read
after `free()` in deferred delete application, `log_perror` misuse for
non-`errno` conditions, and a `NULL` `server_host`/`ssh_destination`
allocation path were fixed (the `config_create` failure now releases through
`config_delete`); the decompression-limit log now prints the effective bound
rather than the compile-time ceiling; the daemon umask and root test fixtures
were hardened; `SSL_read` length is clamped and `sendfile` `poll()` retries on
`EINTR`.
### Refactored / Docs
- Dropped dead `filter_rules_apply` and dead `--old-args` plumbing, unified
`set_error`, deduplicated `path_is_within` and shared constants, and added
printf format attributes (fixing format mismatches). `RSYNC_COMPAT.md`,
`CHANGELOG.md` and `HANDOFF.md` were updated for the audit cycle; the
`RSYNC_COMPAT.md` summary tally was corrected to match the rows.
### Triage fixes
- **`--dirs` directory xattrs applied inline.** A `-d/--dirs` transfer now
applies captured directory `-X`/`-A` xattrs fd-relative on the directory entry
instead of dropping them, so directory xattrs survive the non-recursive path
(`src/shared/file_save.c`, `tests/test_xattr.c`).
- **Directory/root itemize and `--out-format` lines.** `-i`/`--itemize-changes`
and `--out-format` now emit the transfer-root `./` line and per-directory
`cd...`/`.d..t...` lines, rendered by the shared itemize code. This matches
rsync's fresh-transfer output; because the root line is unconditional and an
incremental re-run may itemize directories/symlinks that rsync's quick-check
leaves silent, `-i` is now a ⚠️ Caveat row.
- **FROM name globs for identity maps.** `--usermap`/`--groupmap` `FROM` tokens
now accept `*`/`?`/`[...]` globs, expanded sender-side against the passwd/group
database and collapsed into bounded numeric ranges (`MAX_IDENTITY_MAP`),
matching rsync.
- **Transport fallback unit tests.** Added unit coverage for the TCP/TLS
transport fallback paths (`tests/test_transport_tcp.c`,
`tests/test_transport_tls.c`).
- **Docs corrections.** `RSYNC_COMPAT.md`/`README.md` corrected stale parity
claims for issues #286–#297: the `-F` and `-i` reclassifications, the
`--munge-links` direction, the accepted checksum/compression name sets,
`--bwlimit` parsing, `--stop-at` grammar, `--trust-sender`, symlink xattrs, and
the native/non-interoperable batch and credential notes. The summary tally is
now **117 ✅ / 13 ⚠️ / 27 ❌** of 157 rows.
## [2.28.0] - 2026-09-20
The rsync-parity cycle. `PROTOCOL_VERSION` moves `2.26.0 → 2.27.0 → 2.28.0`;
client and server must run the same version (the handshake is strict). See
`RSYNC_COMPAT.md` for the per-option matrix, now **116 ✅ / 14 ⚠️ / 27 ❌** of
157 rows.
### Added
- **Differential rsync 3.4.1 parity gate** (`tests/integration/
test_differential_parity.py`, `parity_harness.py`, `parity_caveats.py`): runs
real `rsync` and FastSync over generated corpora and diffs the destination
tree, normalized stdout and exit code. A fast subset runs on pull requests and
the full strict set on push; the residual allowlist is empty.
- FastSync-only long option **`--verify-basis`**: require a
`--compare-dest`/`--copy-dest`/`--link-dest` hit to match the source by
whole-file digest instead of trusting the size+mtime quick-check.
- FastSync-only long option **`--delete-commit`** (implies `--delete`): the old
atomic late whole-tree commit.
- `--bwlimit` now parses rsync's units exactly and paces like rsync's leaky
bucket; `--ignore-errors` reproduces rsync's skip-unreadable-subdir and
IO-error-suppressed deletion (exit 23).
- `--info=name/flist/del/remove/nonreg/progress` emit rsync's line format,
including real-run `deleting`/`*deleting` lines carried by a new
`report_deletes` wire bool.
- Receiver-observed `--stats` counters: `Number of created files` now carries
rsync's `(reg/dir/link/special)` breakdown and `Literal data` is exact for a
delta transfer (extended `STATUS_STATS`).
- `--progress` uses an opt-in paths-only pre-count so the `to-chk` denominator
counts every entry like rsync, and emits per-directory/symlink/special names.
- Receiver-side `protect`/`risk` filter engine (new bounded filter-rule wire
block): `--filter='P ...'` now shields a destination-only entry like rsync.
- `auto` for `--compress-choice`/`--checksum-choice` honors
`RSYNC_COMPRESS_LIST`/`RSYNC_CHECKSUM_LIST`, and per-codec compression-level
defaults match rsync.
- Empty source directories are recreated recursively; `-R --no-implied-dirs
--files-from` places listed files under missing implied parents; `--iconv`
matches rsync's push direction; `--delete-delay` reports actual removals and
recursively removes a refilled deferred directory.
### Changed
- **`--delete` now defaults to delete-during (rsync `--del`) timing.** With no
explicit timing flag, a plain `--delete` removes each directory's extras as
that directory is processed instead of committing one whole-tree deletion only
after the entire transfer succeeds. This matches rsync, frees destination
space progressively, and avoids the whole-old+new-tree peak that could
`ENOSPC` a tight destination. The client maps the default onto the existing
`delete_during` wire boolean, so `PROTOCOL_VERSION` stays `2.28.0`.
- Basis directories (`--compare-dest`/`--copy-dest`/`--link-dest`) now default
to rsync's metadata quick-check (equal size and mtime; `--size-only` drops the
mtime leg) instead of FastSync's historical always-verify content hash.
`--copy-dest` re-applies the source attributes, and basis materialization is
streamed so the 256 MiB whole-file cap no longer applies to a basis hit.
- The per-directory `STATUS_DELETE_PLAN` frame gained a one-int `apply` flag:
the one-shot per-run config block (protected prefixes, size-pruned mirrors,
`--delete-missing-args` exact paths) is now always transmitted first on a
config-only carrier (`apply=false`), fixing a latent bug where a
`--delete-missing-args` run whose `--files-from` list synchronized no directory
never sent its exact deletions.
### Notes
- `--delete`/`--delete-during` remain caveats for the mid-transfer abort
boundary (rsync's generator removes all planned extras ahead of its throttled
sender; FastSync removes only reached directories — final trees agree).
`--delete-before`, `--progress`, `--stats`, `--fuzzy` and the basis rows keep
their documented residuals in `RSYNC_COMPAT.md`; `--filter` and
`--delete-excluded` are now parity, including protection of a destination-only
excluded entry under default `--delete`.
### Migration
- Scripts that relied on plain `--delete` deleting nothing until the transfer
fully succeeded must pass **`--delete-commit`** (or `--delete-after`) to keep
that behavior. Plain `--delete` now removes reached directories' extras during
the transfer, exactly like rsync's default; on a completed run the final tree
is unchanged.
- Deployments that relied on FastSync's stricter basis verification should pass
**`--verify-basis`**; the default now trusts the size+mtime quick-check like
rsync.
## [2.26.0] - 2026-09-17
### Added
- **Parity-completion wave.** Closed the remaining rsync-parity gaps against
+14 -3
View File
@@ -1,6 +1,6 @@
cmake_minimum_required(VERSION 3.22)
project(FastFileTransfer VERSION 2.26.0)
project(FastFileTransfer VERSION 2.29.0)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
set(CMAKE_C_STANDARD 11)
@@ -26,9 +26,9 @@ elseif(NOT SANITIZER STREQUAL "none")
endif()
# --- Strict warnings option ---
option(STRICT_WARNINGS "Enable strict warnings (Wextra, Wpedantic, Werror)" OFF)
option(STRICT_WARNINGS "Enable strict warnings (Wextra, Wpedantic, Wformat-signedness, Werror)" OFF)
if(STRICT_WARNINGS)
add_compile_options(-Wextra -Wpedantic -Werror)
add_compile_options(-Wextra -Wpedantic -Wformat-signedness -Werror)
endif()
# --- Coverage option ---
@@ -99,17 +99,21 @@ set(SHARED_SRCS
src/shared/daemon_limits.c
src/shared/data.c
src/shared/delay_updates.c
src/shared/delete.c
src/shared/delete_commit.c
src/shared/delete_plan.c
src/shared/delta.c
src/shared/file.c
src/shared/file_list.c
src/shared/file_receive.c
src/shared/file_save.c
src/shared/file_send.c
src/shared/file_store.c
src/shared/filter.c
src/shared/format.c
src/shared/hardlink.c
src/shared/identity.c
src/shared/incremental_check.c
src/shared/log.c
src/shared/metadata.c
src/shared/motd.c
@@ -136,9 +140,14 @@ set(SERVER_MAIN_SRCS src/server/server.c)
# Client implementation (no main): everything except the CLI entry point.
set(CLIENT_CORE_SRCS
src/client/change_list.c
src/client/client_manifest.c
src/client/client_report.c
src/client/client_scan.c
src/client/client_send.c
src/client/client_validation.c
src/client/scanner.c
src/client/scanner_filter.c
src/client/scanner_parallel.c
src/client/usage.c
)
set(CLIENT_MAIN_SRCS src/client/client_cli.c)
@@ -221,10 +230,12 @@ set(TEST_SRCS
tests/test_daemon_limits.c
tests/test_data.c
tests/test_delay_updates.c
tests/test_delete_plan.c
tests/test_delta.c
tests/test_file.c
tests/test_file_list.c
tests/test_file_sendfile.c
tests/test_filter.c
tests/test_format.c
tests/test_fuzz_smoke.c
tests/test_glob.c
+215 -22
View File
@@ -1,16 +1,31 @@
# FastSync — Session Handoff (2026-09-17)
# FastSync — Session Handoff (2026-09-21)
## Current status
- **Release `v2.21.0`** tagged (`919a729`, "Release v2.21.0"); full CI green
(run 552: lint, build-and-test, ASan, UBSan, fuzz-build, coverage, valgrind).
`dev` has the release commit plus later doc-only merges (a README refresh and
this handoff).
- **Release PR #284 (`dev` -> `main`)** open, CI green (run 553).
`main` is protected: it needs review/approval to merge.
https://gitea.tap-tap.win/TapTap/FastSync/pulls/284
- **`PROTOCOL_VERSION` = `"2.26.0"`** (`src/shared/config.h`); CMake
`project(FastFileTransfer VERSION 2.26.0)`.
- Working tree clean; no wave worktrees remain.
- **Release `v2.28.0`** is tagged and merged to `main`: tag `v2.28.0` points at
`ee6523a`, and the PR #304 merge commit `b4d54504` is on `main`.
- **`dev` is at `0fbb9de`** — the merge of parity cycle 2.29 (PR #305). The old
`558782d` (incremental-check flake fix) is an ancestor.
- **`PROTOCOL_VERSION` = `"2.28.0"`** (`src/shared/config.h`); CMake
`project(FastFileTransfer VERSION 2.28.0)`.
- **Parity cycle 2.29 is merged to `dev`** (PR #305), no wire change. It closed
the scanner-order, delete-timing, relative-basis and fuzzy-eligibility
residuals and improved the `--info`/`--stats`/`--debug` partials. Parity
matrix: **120 ✅ / 10 ⚠️ / 27 ❌ = 157**. Remaining ⚠️ rows: `--info`,
`--debug`, `--msgs2stderr`, `--stats`, `--progress`, `--delete-before`, the
three basis-dir options, and `-y`/`--fuzzy`.
- **Audit cycle complete on branch `fix/audit-cycle`** (branched from `dev` @
`0fbb9de`), integration PR to `dev` pending. No wire change
(`PROTOCOL_VERSION` stays 2.28.0). It lands the receiver/client security and
correctness fixes — `--temp-dir` symlink-escape confinement, special-bit
masking under a super-off policy, daemon `umask(022)`, the `-z` decompression
ceiling raised to the 256 MiB whole-file bound, `--bwlimit` pacing the
plaintext `--sendfile` path, `--partial-dir` implying `--partial`, rejection
of unsupported filter modifiers (`x`/`e`/`n`/`w`), client-side
`MAX_FILTER_RULES` enforcement, unknown wire `Status` rejection, and the
accompanying refactors/docs. The parity matrix is unchanged at
**120 ✅ / 10 ⚠️ / 27 ❌ = 157**; this docs pass (worktree `fix/audit-docs2`)
corrects the `RSYNC_COMPAT.md` summary tally to match the rows.
## What landed this session
1. **Wave 8 (refactors):** Config X-macro wire table; single-owner `authorized_root`;
@@ -39,19 +54,197 @@
docs state push-only / remote-source unsupported.
5. **Preserve-attribute split (protocol 2.22.0)** landed on `feat/preserve-attr-split`: per-attribute `-p/-t/-o/-g` + `--no-*` negations, `-a` = `-rlptgoD`, and the 2.21.0 → 2.22.0 wire bump.
6. **Rsync-parity wave (protocol 2.23.0)** on `feat/rsync-parity`: rsync short options/clustering/attached values (`-r`/`-b`/`-L`/`-B`, `-av`, `-aAX`, `-B1000`, `-essh`, `-MOPT`), `-c` checksum quick-check, `--checksum-choice`/`--compress-choice` validation and seed randomization, rsync timeout/max-alloc defaults, temp-dir confinement + `EXDEV` fallback, ownership/mapping parity (numeric-ids modifier, map ranges/`*`/empty-FROM, `--chown`+map conflicts, fake-super resolved-owner record), verbatim symlink storage with rsync `--safe-links`/`--munge-links`, socket recreation under `--specials`, `--chmod` 3.4.1 semantics, and delete scoping + `--max-delete` partial/exit-25. Wire: appended delete-manifest synchronized-directory section and `STATUS_DELETE_LIMIT`.
7. **Parity-completion wave (protocol 2.24.0 → 2.26.0)** on `feat/parity-completion`: per-directory delete plans (`STATUS_DELETE_PLAN`) for `--delete-during`/`--delete-delay`; receiver `STATUS_STATS` counters feeding `--stats`/`--progress` and `--out-format %b/%c/%C`, plus `-n --delete` lines; `lz4`/`zlib`/`zlibx` compression and `md4`/`sha1`/`none` checksums with `auto` negotiation (default `xxh128`/`zstd`); general `-R`/`--no-implied-dirs`/`-d`; the full filter grammar (`merge`/`dir-merge`/`hide`/`show`/`protect`/`risk`/`clear` + modifiers) and corrected `-F`/`-FF`; receiver-side `--chown`/map TO-name resolution; absolute basis dirs + `--link-dest` relink; receiver-side `--ignore-existing` short-circuit; `--preallocate` over `--sparse` via `fallocate(2)`; `--iconv=.`/`-`/`--no-iconv`; lone `-h` help; aliases `--ignore-non-existing`/`--protect-args`/`--msgs2stderr`; and the full `--info`/`--debug` vocabulary. `RSYNC_COMPAT.md` reclassifies the matrix to 106 ✅ / 27 ⚠️ / 23 ❌.
7. **Parity-completion wave (protocol 2.24.0 → 2.26.0)** on `feat/parity-completion`: per-directory delete plans (`STATUS_DELETE_PLAN`) for `--delete-during`/`--delete-delay`; receiver `STATUS_STATS` counters feeding `--stats`/`--progress` and `--out-format %b/%c/%C`, plus `-n --delete` lines; `lz4`/`zlib`/`zlibx` compression and `md4`/`sha1`/`none` checksums with `auto` negotiation (default `xxh128`/`zstd`); general `-R`/`--no-implied-dirs`/`-d`; the full filter grammar (`merge`/`dir-merge`/`hide`/`show`/`protect`/`risk`/`clear` + modifiers) and corrected `-F`/`-FF`; receiver-side `--chown`/map TO-name resolution; absolute basis dirs + `--link-dest` relink; receiver-side `--ignore-existing` short-circuit; `--preallocate` over `--sparse` via `fallocate(2)`; `--iconv=.`/`-`/`--no-iconv`; lone `-h` help; aliases `--ignore-non-existing`/`--protect-args`/`--msgs2stderr`; and the full `--info`/`--debug` vocabulary. `RSYNC_COMPAT.md` reclassifies the matrix to 106 ✅ / 27 ⚠️ / 23 ❌; the later rsync-parity-stats pass (`fix/parity-stats`) moves it to 107 ✅ / 25 ⚠️ / 24 ❌ (see item 8).
8. **rsync-parity-stats pass** on `fix/parity-stats` (no wire change, `PROTOCOL_VERSION` stays `2.26.0`): `--delete-delay` now reports only entries it actually removes, while the `--max-delete` budget is charged at plan/snapshot time (`planned`, via `defer_add`) to bound the deferred list (a refilled deferred directory that survives `ENOTEMPTY` is not reported but still consumes budget); `--stats` gained the `(reg/dir/link/special)` `Number of files` breakdown and now counts only regular files actually stored for `Number of regular files transferred`/transferred size/literal data (up-to-date re-runs report 0); `Total file size` includes symlink target lengths; `--progress` prints the leading `./` root line and counts it in `to-chk` so a single-file transfer matches rsync; and `%C` uses the selected transfer checksum with `checksum_digest_file` supporting md4/sha1/none, byte-identical to rsync for every algorithm. `--out-format` reclassified ❌ (`%b`/delta-`%c` are protocol-specific). Differential + regression tests added; full suite + ASan + clang-format + cppcheck clean.
9. **Option-parity wave (protocol 2.26.0 → 2.27.0, on `fix/parity-options`):**
`--bwlimit` now ports rsync 3.4.1's units/quantization and paces like its
leaky bucket; `--ignore-errors` reproduces rsync's default (an I/O error
skips deletion unless the flag is set; the readable tree still transfers and
the run exits 23) across every delete timing; the `--info` categories with a
FastSync event (`name`/`flist`/`del`/`remove`/`nonreg`/`progress`) emit
rsync's line format, with real-run `deleting`/`*deleting` lines carried over
the new trailing config bool `report_deletes` (golden wire updated by
`tests/test_config.c`). Two residuals were reclassified **divergent**: `-M`
over daemon/TCP (no argv channel in FastSync's binary config handshake;
rsync-daemon differential pins the rsync behavior) and receiver-side
`protect`/`risk` re-derivation for destination-only entries (would need a
receiver filter engine; differential pins the divergence — **reversed by
track 4a below**, which adds that engine). The options pass
stands at **110 ✅ / 21 ⚠️ / 26 ❌**. New `tests/integration/test_option_parity.py`
holds the rsync differentials (bwlimit parse+rate, info lines, real-setpriv
`--ignore-errors`, rsync-daemon `-M`, filter-protect pin).
10. **rsync-parity-fs pass** on `fix/parity-fs` (no wire change of its own; integrated
on top of the 2.27.0 options wave): recursive transfers now recreate empty source directories (and
`-m/--prune-empty-dirs` still suppresses them), a directory entry replaces a
blocking destination regular file, and `-R --no-implied-dirs --files-from`
places a listed file under a missing implied parent with default attributes
instead of refusing (real rsync 3.4.1 parity, differential-tested). `--iconv`
now reproduces rsync's push direction (destination charset = the spec's REMOTE
half; a server `--iconv` overrides), and `-T/--temp-dir` relative semantics are
confirmed identical while the absolute-path confinement is a deliberate
divergence. The basis-dir options, `--delay-updates` and `--dry-run` were
reclassified to ❌ after a differential test reproduced each exact residual
(basis content verification, fixed staging-name collision, and dry-run
would-delete over-report). `--fuzzy` was also reclassified to ❌ (deterministic
heuristic with a 10× size window, not rsync's matcher), but its residual is the
candidate-selection heuristic itself: the final tree is byte-exact by design, so
it is pinned by the `TestFuzzy` threshold suite rather than a byte-level rsync
differential. (Track 5b later found the name heuristic is rsync's own and moved
the row ❌ → ⚠️, leaving only the narrower delta size window; see entry 15.) The parity-review pass then moved `--delete-delay` to ⚠️ (the
plan-time `--max-delete` charge and non-recursive deferred removal differ from
rsync when a snapshotted entry fails removal). Differential-gate allowlist
entries `min_size`/`empty_dirs_recursive`/`dirs_plain` were removed. The
integrated stats+options+fs branch stands at **111 ✅ / 13 ⚠️ / 33 ❌ = 157**;
full suite + ASan + clang-format + cppcheck clean.
11. **No-wire parity track 1** on `feat/parity-2.28` (no protocol change):
`-n --delete` now sends the same filter-excluded + size-pruned protected
prefixes and synchronized-directory scope as a real run (dry-run would-delete
matches rsync for source-derived protections; the destination-only exclude
residual was later closed by track 4a, readdir ordering remains);
`--delete-delay` now charges
`--max-delete` on actual removals and re-scans a queued directory at commit
to remove content created after the plan, with an independent deferred-list
cap (only partial-delete ordering remains); and `--info=name2` emits `NAME is
uptodate` plus the leading `./` root name line for `--info=name` (only the
root-line trigger condition and receiver-side `skip` wording remain). Matrix
now **111 ✅ / 14 ⚠️ / 32 ❌ = 157**; differential + unit tests added in
`test_features.py`, `test_option_parity.py`, the unit test
`tests/test_delete_plan.c`,
`test_delete_delay_budget_parity.py`, `test_delete_timing_parity.py`.
12. **No-wire parity track 2b** on `feat/parity-2.28` (no protocol change):
`--progress`/`-P`/`--info=progress` (when not `--quiet`) now run an opt-in
paths-only metadata pre-count (no file reads/hashing) that supplies rsync's
full file-list total for the `to-chk` denominator and the directory names,
and emits per-directory/symlink/special name lines, in both the sequential
and `--threads` paths. `--delete-during`/`--delete-delay` reuse their
keep-set pre-scan instead of a second walk; non-progress runs are
unaffected. Differential tests (`progress`/`progress_threads` over a new
`multidir` corpus) match rsync's name set and `to-chk` denominator on a
fresh transfer, and the single-file byte-identical test still passes;
emission order (rsync's sorted depth-first vs FastSync's readdir/BFS stream)
plus re-run over-naming (unconditional `./`, ancestor dirs named with a
transferred child, and no quick-check for symlinks/empty dirs) remain the
caveats, so the row stays ⚠️ and the matrix is unchanged at
**111 ✅ / 14 ⚠️ / 32 ❌ = 157**.
13. **Wire parity track 4a** on `feat/parity-2.28` (`PROTOCOL_VERSION` stays
`2.28.0`): the receiver now has a delete-time filter engine. The sender
compiles its root-level selection rules exactly as the scanner does
(`filter_base_build`) and streams them as one bounded, self-describing
config-frame block (action, sides, anchored, dir-only, negate, owner,
pattern; bounded rule count and pattern bytes, unknown action/sides is a
protocol error). The receiver reconstructs `protect_rules` and applies them
first-match-wins to each extraneous destination path in every delete timing
(the whole-tree commit walker, the `--delete-during`/`--delete-delay`
per-directory plans, and the `-n` would-delete enumeration), so a
`P *.log` rule protects a destination-only `extra.log` like rsync (with
`risk` cancelling); the sender-derived protected-prefix behavior is
preserved when no rules are sent and `--delete-excluded` semantics are
unchanged. Per-directory merge (`:`/`.`) receiver re-derivation remains the
residual. `TestFilterProtect` (real + dry-run) plus differential cases
`filter_protect`, `filter_protect_during`, `filter_protect_delay` added and
the `--filter=RULE` row moves ❌ → ✅: matrix now
**115 ✅ / 11 ⚠️ / 31 ❌ = 157**; unit tests, the three named integration
files, clang-format and cppcheck clean.
14. **Wire parity track 5a** on `feat/parity-2.28` (`PROTOCOL_VERSION` stays
`2.28.0` by project decision): the three basis-dir options now default to
rsync's metadata quick-check (equal size + equal mtime, or size alone under
`--size-only`; `-I` disables matching) instead of FastSync's historical
xxHash64 content equality, so a same-size/different-content basis is trusted
exactly as rsync trusts it. A new FastSync-only, long-only `--verify-basis`
flag restores the strict whole-file content equality; its bool is appended to
the basis block of the config frame (golden wire frame 882 → 886 bytes).
`--verify-basis` streams the confined basis descriptor to hash it, and a
basis hit is no longer capped at the 256 MiB whole-file payload bound:
`--copy-dest` streams the basis through a bounded buffer and `--link-dest`'s
copy fallback streams from the basis, so an over-limit hit materializes (a
basis MISS still falls back to the normal transfer and keeps its own bound).
A `--copy-dest` hit re-applies the SOURCE attributes (the sender transmits
the source metadata with the basis check frame), matching rsync's
"copy then fix attributes"; a `--link-dest` success keeps the shared inode's
attributes (writing through it would mutate the basis). Differential cases
`copy_dest` and `verify_basis` added; `test_basis_dir_size_only_content_residual`
converted to a passing parity assertion; `TestBasisDestDirs` updated for the
new default + `--verify-basis`; unit tests cover the quick-check/verify
decision and the same-size/different-content handshake. The
`--compare-dest`/`--copy-dest`/`--link-dest` rows move ❌ → ⚠️ (relative-DIR
resolution base and over-limit MISS refusal): matrix now
**116 ✅ / 13 ⚠️ / 28 ❌ = 157**.
15. **No-wire parity track 5b** on `feat/parity-2.28` (`PROTOCOL_VERSION` stays
`2.28.0` by project decision): `-y`/`--fuzzy` reclassified ❌ → ⚠️. A probe
against real rsync 3.4.1 (pinned `-B8192`, repeated-content 64 KiB corpus)
showed the name heuristic is already rsync's (`util1.c fuzzy_distance` /
`find_filename_suffix` + the exact size+mtime pass) and the output is always
byte-exact; the only residual is candidate ELIGIBILITY, because FastSync's
`delta_should_attempt` gate caps the size ratio at 10× and requires both
files ≥ 16 KiB while rsync will reuse a basis from 0.25× to 10000× and below
16 KiB. The choice is observable only as `--stats` bandwidth counters. Added
differential case `fuzzy_basis` (same-suffix sibling, one name edit,
identical content, block size pinned) asserting tree **and** normalized
`--stats` parity where the choices coincide, plus `TestFuzzy` pinning the
window boundary on both sides (>10× and <16 KiB siblings declined by
FastSync while rsync uses them, both trees byte-identical). Matrix now
**116 ✅ / 14 ⚠️ / 27 ❌ = 157**.
16. **Lockstep delete-default track 6** on `feat/parity-2.28` (`PROTOCOL_VERSION`
stays `2.28.0`): plain `--delete` now defaults to rsync's delete-during
(`--del`) timing, normalized on the client onto the existing `delete_during`
wire bool. The old late whole-tree commit is opt-in via `--delete-after` or
the FastSync-only long `--delete-commit` (identical `delete_after` timing).
`-d/--dirs` still falls back to the end commit, `--delay-updates` still
deletes before publication, and `--files-from`/`-R` scope is unchanged. The
`STATUS_DELETE_PLAN` frame gained a one-int `apply` flag so the per-run
config block (including `--delete-missing-args` exact paths) is always
transmitted, on a config-only carrier when the scope allows no directory
plan — fixing a latent bug with a file-only `--files-from` list. Differential
cases `delete`/`delete_commit`/`filter_protect_after` plus the extended
`test_delete_timing_parity.py` (plain `--delete` mid-abort removes reached
extras, `--delete-commit` defers) pass; full `-m "not setpriv"` suite,
clang-format and cppcheck clean. Matrix unchanged at
**116 ✅ / 14 ⚠️ / 27 ❌ = 157** (the `--delete`/`--delete-during` rows stay
⚠️ for the abort boundary; `--delete-after` stays ✅).
17. **Audit cycle** on `fix/audit-cycle` (from `dev` @ `0fbb9de`;
`PROTOCOL_VERSION` stays `2.28.0`): a security/correctness pass over the
parity-2.29 baseline. It raises the decompression ceiling to the 256 MiB
protocol whole-file bound (`-z` on 100–256 MiB files now works), paces the
plaintext-TCP `--sendfile` path with `--bwlimit`, confines the `--temp-dir`
scratch dir by the fd's real path (symlink escape refused), masks
client-controlled setuid/setgid/sticky bits when super activities are not
permitted, sets the daemon umask to `022`, makes `--partial-dir` imply
`--partial`, rejects the unsupported filter modifiers (`x`/`e`/`n`/`w`),
enforces `MAX_FILTER_RULES` client-side, rejects unknown wire `Status`
values, and hardens credentials/signal handling (with the accompanying
refactors and docs). No row changes classification, so the matrix stays
**120 ✅ / 10 ⚠️ / 27 ❌ = 157**. This docs pass is on `fix/audit-docs2`.
## Next steps
1. **Merge PR #284** (`dev` -> `main`) once reviewed (protected branch).
2. **Deferred security items** (documented, not implemented):
- Pre-auth config/daemon-auth handshake has no aggregate wall-clock deadline
(per-message timeout only) — slowloris holds connection slots.
- Per-source registry fails open when the shared table is full (per-module/global
caps and host ACLs still apply); consider fail-closed or larger/evicting table.
- SCRAM-like daemon auth has no TLS channel binding (and is not RFC 5802).
- `cleanup()` signal handler calls non-async-signal-safe teardown; daemon `umask(0)`.
- Wire protocol assumes homogeneous word size/endianness (lengths are native
`size_t`) — document or move to fixed-width framing.
1. **Open and merge the audit-cycle PR** (`fix/audit-cycle`, including this
`fix/audit-docs2` docs pass) into `dev` once reviewed. `dev` is the default
branch; all PRs target `dev`, never `main` directly.
2. **Remaining deferred items:**
- **Large structural refactors:** delete-engine consolidation
(`delete_extras_fd`/`manifest_delete_extras`/the delete-plan path),
god-function splits, and translation-unit splits.
- **`--progress`/`--info` receiver→sender event channel:** the root `./`
line, ancestor-directory suppression, receiver-side `skip`/`backup` echo,
and symlink/empty-dir quick-check feedback.
- **`--delete-before` phase-0 keep-set** (rsync fixes the file list before
the data pass; FastSync keeps its pre-scan snapshot race).
- **>256 MiB single-file streaming** (B4, the general whole-file limit).
- **Wire native-size framing:** lengths are native `size_t` and the protocol
assumes homogeneous word size/endianness — document or move to fixed-width
framing.
- **SCRAM-like daemon auth channel binding:** no TLS channel binding today
(and it is not RFC 5802).
- Still-open security nits: the pre-auth config/daemon-auth handshake has no
aggregate wall-clock deadline (per-message timeout only — slowloris holds
connection slots); the per-source registry fails open when the shared table
is full (per-module/global caps and host ACLs still apply).
3. **Out of scope / intentional:** pull (remote source) mode is **not** planned —
FastSync is push-only; see `RSYNC_COMPAT.md#direction`.
+113 -68
View File
@@ -75,8 +75,9 @@ matrix is classified as parity, caveat, or divergent in
owner, group, devices, and special files — and does not imply compression or
multithreading (see [Client](#client)). Ownership application is still
privilege-gated: a receiver that cannot `chown` logs a warning and skips it.
Under `-p` the source mode is copied exactly, including setuid/setgid/sticky
and group/other-write bits (strict rsync parity; see
Under `-p` the source mode is copied exactly, including group/other-write
bits; setuid/setgid/sticky bits are copied only when super-user activities are
permitted, and are masked under `SUPER_MODE_OFF`/`--no-super` (see
[`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)).
- Symlink transfer stores targets **verbatim** (`-l`/`--links`), including
absolute and `..`-bearing targets, matching rsync. The receiver does not
@@ -113,13 +114,18 @@ matrix is classified as parity, caveat, or divergent in
(`-B1000`, `-essh`, `-MOPT`, `--opt=value`) are accepted, matching rsync.
- `-r`, `-b`, `-L`, and `-B` are parsed with the rsync short names.
- `--stats` prints the counters FastSync can observe plus the receiver-only
counters (`Matched data`, deleted files) reported over the wire; rsync's
per-type `Number of files` breakdown is not reproduced. `--progress` prints
rsync-style per-file blocks (without rsync's leading `./` line).
counters reported over the wire (`Matched data`, deleted files, and the
created/literal counters); `Number of files` and `Number of created files`
carry rsync's per-type breakdown. `--progress` prints rsync-style per-file
blocks including the leading `./` line, and (when progress is requested) a
paths-only pre-count supplies rsync's `to-chk` denominator.
- Codecs match rsync 3.4.1: `zstd`/`lz4`/`zlib`/`zlibx` compression and
`xxh128`/`xxh3`/`xxh64`/`md5`/`md4`/`sha1`/`none` checksums, negotiated with
`auto`; `zlibx` behaves as `zlib`, and the transfer checksum is not separately
selectable.
`xxh128`/`xxh3`/`xxh64`/`md5`/`md4`/`sha1`/`none` checksums. `auto` honors
`RSYNC_COMPRESS_LIST`/`RSYNC_CHECKSUM_LIST` and otherwise follows rsync's
compiled-in order. An omitted `--compress-level` uses the codec's rsync
default (zstd 3, zlib/zlibx 6, lz4 ignored); `zlib`/`zlibx` share the
literal-only zlib path (rsync's zlibx semantics), and the transfer checksum is
not separately selectable.
The detailed flag matrix is maintained in
[`RSYNC_COMPAT.md`](RSYNC_COMPAT.md). It reports each row as **parity**,
@@ -167,7 +173,7 @@ This produces `./build/client` and `./build/server`. `compile_commands.json` is
| `--preserve` | Preserve mode and mtime (`-p` + `-t`; add `-o`/`-g` for owner/group or `-U`/`--atimes` for atime; `-N`/`--crtimes` captures birth time but cannot apply it) |
| `-U, --atimes` | Preserve access times. Captured with the metadata payload; does not enable ownership. |
| `-N, --crtimes` | Capture birth time; cannot be applied (documented divergence) |
| `-p, --perms` | Preserve permission bits. Strict rsync parity: the source mode is copied exactly, including setuid/setgid/sticky and group/other-write bits |
| `-p, --perms` | Preserve permission bits. The source mode is copied exactly, including group/other-write bits; setuid/setgid/sticky are copied only when super-user activities are permitted (`SUPER_MODE_OFF`/`--no-super` masks them) |
| `-t, --times` | Preserve modification times |
| `-o, --owner` | Preserve the source owner (privilege-gated; mapped by name on the receiver with a numeric fallback) |
| `-g, --group` | Preserve the source group (privilege-gated; mapped by name on the receiver with a numeric fallback) |
@@ -181,7 +187,7 @@ This produces `./build/client` and `./build/server`. `compile_commands.json` is
| `--groupmap=MAP` | Map group names when applying ownership |
| `--numeric-ids` | Apply source numeric uid/gid directly instead of mapping by name |
| `--copy-as=USER[:GROUP]` | Force every written entry to USER[:GROUP] (requires a privileged receiver) |
| `--fake-super` | Record the resolved owner plus mode/time in a reserved `user.fastsync.stat` xattr and replay mode/time; never performs a real chown |
| `--fake-super` | Record the resolved owner plus full mode/rdev in rsync's reserved `user.rsync.%stat` xattr (rsync 3.4.1 grammar) and replay the permission bits; never performs a real chown |
| `--super` | Permit the receiver to attempt confined super-user activities (device nodes) |
| `-D` | Preserve device and special files (implies `--devices --specials`) |
| `--devices` | Recreate device nodes on the destination (privileged; skipped without `CAP_MKNOD`) |
@@ -199,34 +205,37 @@ This produces `./build/client` and `./build/server`. `compile_commands.json` is
| `-u, --update` | Skip files newer than the source on the receiver |
| `--incremental` | Skip files unchanged since last transfer (size + mtime). Auto-enables `--preserve`. Incompatible with `--chunk-serialization`. |
| `--existing` | Skip files not already present at the destination; update existing files normally. |
| `--compare-dest <dir>` | Extra comparison basis: unchanged files are not transferred (requires/implies `--incremental`) |
| `--copy-dest <dir>` | Like `--compare-dest`, but copies the unchanged file from DIR into the destination |
| `--link-dest <dir>` | Like `--copy-dest`, but hard-links the unchanged file from DIR (repeatable; earlier DIRs win) |
| `--delete` | Delete files on receiver not present in source (default timing: delete-after, i.e. only after the whole transfer succeeded). Scoped to the synchronized directories, so `--files-from` subsets are safe |
| `--ignore-existing` | Skip files that already exist on the receiver; like rsync it does not apply to directories or symlinks. |
| `--compare-dest <dir>` | Extra comparison basis: unchanged files are not transferred (requires/implies `--incremental`; a basis MISS above the 256 MiB whole-file payload bound is refused — see [`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)) |
| `--copy-dest <dir>` | Like `--compare-dest`, but copies the unchanged file from DIR into the destination (same basis-size caveat; see [`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)) |
| `--link-dest <dir>` | Like `--copy-dest`, but hard-links the unchanged file from DIR (repeatable; earlier DIRs win; same basis-size caveat; see [`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)) |
| `--verify-basis` | FastSync-only: require a basis hit (`--compare-dest`/`--copy-dest`/`--link-dest`) to match the source by whole-file digest instead of trusting the size+mtime quick-check (default matches rsync) |
| `--delete` | Delete files on receiver not present in source (default timing: delete-during, matching rsync, so destination space is freed progressively). Scoped to the synchronized directories, so `--files-from` subsets are safe |
| `--delete-before` | Delete extras before the transfer starts (implies `--delete`) |
| `--delete-during`, `--del` | Delete extras once the keep-set is known, before data is applied (implies `--delete`) |
| `--delete-delay` | Delete extras only after a successful transfer (implies `--delete`) |
| `--delete-after` | Explicit delete-after timing (implies `--delete`) |
| `--delete-commit` | FastSync-only: keep the pre-2.28 atomic timing — delete only after the whole transfer succeeded (identical timing to `--delete-after`) |
| `--delete-excluded` | Also delete filter-excluded destination mirrors (size-pruned mirrors stay protected) |
| `--max-delete <n>` | Delete at most n destination entries; the rest are skipped and the run exits 25 (partial), matching rsync |
| `--delay-updates` | Put updated files into place only at the end of the transfer (`--force` is honored at publication) |
| `-T, --temp-dir <dir>` | Scratch directory for temp files before the atomic install; confined to the receive root (relative only), with an `EXDEV` non-atomic copy fallback |
| `--delay-updates` | Put updated files into place only at the end of the transfer (`--force` is honored at publication; the fixed `.fastsync-stage` staging name diverges from rsync — see [`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)) |
| `-T, --temp-dir <dir>` | Scratch directory for temp files before the atomic install; confined to the receive root (a relative path resolves below it; an absolute path is accepted only when it canonicalizes inside it), with an `EXDEV` non-atomic copy fallback |
| `-n, --dry-run` | Report what would be transferred without mutating the destination. Since protocol 2.21.0 a server-routed target contacts the receiver and reports would-transfer based on receiver state; a plain local destination keeps the client-side scan. Never mutates or deletes. |
| `-v, --verbose` | Enable debug logging |
| `-q, --quiet` | Suppress non-error output |
| `--progress` | Show rsync-style per-file progress blocks from the receiver's wire counters (FastSync does not print rsync's leading `./` line) |
| `--progress` | Show rsync-style per-file progress blocks from the receiver's wire counters; the root `./` line is printed whenever progress is active (rsync prints it only when the transfer root is created) |
| `-P` | Enables partial-transfer mode + progress output; interrupted writes retain the already-written temp for resumption |
| `--stats` | Print transfer statistics at end (bytes, files, timing), including the receiver-only counters reported over the wire; rsync's per-type `Number of files` breakdown is not reproduced |
| `--stats` | Print transfer statistics at end (bytes, files, timing), including the receiver-only counters reported over the wire; `Number of files` and `Number of created files` carry rsync's per-type breakdown (deleted files are reported as a single total) |
| `-i, --itemize-changes` | Print an rsync-style per-file change line |
| `--out-format=FORMAT` | Output format for changed files (`%f %n %l %b %M %%`) |
| `--out-format=FORMAT` | Output format for changed files (`%f %n %l %b %c %C %i %M %%`) |
| `--list-only` | List source files instead of transferring |
| `--fsync` | Fsync every written file before publication |
| `-h, --human-readable` | Format transfer byte/rate counts with rsync's decimal (base-1000) units |
| `--max-depth <n>` | Maximum directory depth to recurse (0 = unlimited, default: 0) |
| `--log-file <path>` | Write log messages to file instead of stderr |
| `--write-batch=FILE` | Run the normal live transfer and also emit a self-contained batch file of the source tree |
| `--only-write-batch=FILE` | Emit the batch file only (no destination, no server) |
| `--read-batch=FILE` | Apply a batch file to the destination (no source, no server) |
| `--write-batch=FILE` | Run the normal live transfer and also emit a self-contained batch file of the source tree (FastSync-native format, not rsync-interoperable) |
| `--only-write-batch=FILE` | Emit the batch file only (no destination, no server); FastSync-native format, not rsync-interoperable |
| `--read-batch=FILE` | Apply a batch file to the destination (no source, no server); FastSync-native format, not rsync-interoperable |
| `--source-dir <path>` | Source directory (overrides `FASTSYNC_SOURCE_DIR`) |
| `--dest-dir <path>` | Server destination directory (overrides `FASTSYNC_DEST_DIR`) |
| `--save-to-disk` | Write received files to disk |
@@ -239,12 +248,12 @@ This produces `./build/client` and `./build/server`. `compile_commands.json` is
| `-4, --ipv4` | Force IPv4 for destination resolution |
| `-6, --ipv6` | Force IPv6 for destination resolution |
| `--sockopts=OPTS` | Comma-separated OPT=VAL socket options applied before connect (`TCP_NODELAY`, `SO_KEEPALIVE`, `SO_RCVBUF`, `SO_SNDBUF`, `SO_REUSEADDR`) |
| `--bwlimit <KB/s>` | Bandwidth limit in kilobytes per second |
| `--bwlimit <RATE>` | Bandwidth limit, using rsync's exact `parse_size_arg` grammar: a bare value is KiB/s; `K`/`M`/`G`/`T`/`P` are binary suffixes; `KB`/`MB` are decimal and `KiB`/`MiB` binary; decimals are accepted and quantized to whole KiB; `0` (or empty) means no limit. Also paces `--sendfile` transfers |
| `--chunk-size <n>` | Chunk size in bytes (default: 10485760) |
| `--timeout <sec>` | I/O timeout in seconds, applied to both the socket (`SO_RCVTIMEO`/`SO_SNDTIMEO`) and the per-message protocol poll deadline. Default `0` = disabled (matching rsync); `0` disables it. `--no-timeout` is the negation. The value is not sent on the wire; the server side keeps its own safe floor. |
| `--contimeout <sec>` | Connection timeout in seconds (default: 60, matching rsync); `0` disables it (`--no-contimeout` is the negation) |
| `--stop-after=MINS` | Stop the transfer after MINS minutes (a positive integer); whatever was already transferred is kept |
| `--stop-at=TIME` | Stop at an absolute time (`HH:MM`, `HH:MM:SS`, or `now+N[smhd]`); an early stop skips the late `--delete` keep-set |
| `--stop-at=TIME` | Stop at an absolute time. Accepts rsync's `parse_time` forms (`Y-M-DTh:m`, `Y/M/DTh:m`, `Y-M-D`, `M-D`, `D`, `h:m`, `:m`, `T h:m`; omitted fields resolve to the next matching point in the local timezone), plus `now+N[smhd]` and FastSync's `HH:MM`/`HH:MM:SS` clock-time spelling. An early stop skips the late `--delete` keep-set |
| `-b, --backup` | Backup existing destination files before overwriting |
| `--backup-dir <dir>` | Target directory for backups (requires `--backup`) |
| `--tls` | Enable TLS encryption |
@@ -307,17 +316,22 @@ transfer is never aborted.
`timeout`, `contimeout`, `quiet`, `stats`, `max_depth`, and `log_file` are
client-only.
5. **Queue** — thread-safe bounded queue with condition variables.
6. **DirectoryScanner** — recursive BFS traversal with exclude and include
pattern support, max-depth enforcement.
6. **DirectoryScanner** — recursive traversal that buffers and sorts each
directory (non-directories ascending, then directories ascending) and walks
depth-first in rsync flist order, with exclude and include pattern support and
max-depth enforcement.
### Key Algorithms
1. **File scanning** — BFS directory traversal; entries matched against exclude
and include patterns, with max-depth enforced.
1. **File scanning** — sorted depth-first traversal in rsync flist order (each
directory's non-directories ascending, then its directories ascending);
entries matched against exclude and include patterns, with max-depth
enforced. The `--threads` parallel scanner remains unordered.
2. **Chunking** — files accumulated until the `chunk_size` threshold (default
10 MiB) is reached, then flushed.
3. **Compression** — streaming zstd via `ZSTD_compressStream2()` /
`ZSTD_decompressStream()`.
`ZSTD_decompressStream()`, with lz4 and zlib/zlibx codecs also supported
(selectable with `--compress-choice`).
4. **Network protocol** — status-code-driven exchange with metadata packing,
keep-alive, and abort support.
5. **Incremental check** — the client sends `STATUS_CHECK` + path + size +
@@ -372,6 +386,8 @@ Received files are written to a temporary path (suffixed with `.tmp`) and then a
- C11 compiler
- CMake >= 3.22
- zstd library
- zlib library
- lz4 library
- OpenSSL (development headers and libraries)
- pthreads
- SSH client (for SSH transport mode only)
@@ -380,12 +396,12 @@ Received files are written to a temporary path (suffixed with `.tmp`) and then a
**Ubuntu/Debian:**
```bash
sudo apt install cmake build-essential libzstd-dev libssl-dev openssh-client
sudo apt install cmake build-essential libzstd-dev zlib1g-dev liblz4-dev libssl-dev openssh-client
```
**Nix:**
```bash
nix-shell # provides zstd, openssl, cmake, gcc
nix-shell # provides zstd, zlib, lz4, openssl, cmake, gcc
```
## Building
@@ -505,8 +521,8 @@ features without changing the meaning of ordinary compatibility options.
|---|---|
| `-j`, `--threads[=N]` | Enable the multithreaded scanner/loader/sender pipeline. `N` (1–256) sets the parallel scanner worker count; bare `-j`/`--threads` uses the default. |
| `-z [level]`, `--compress [level]` | Enable streaming compression (default `zstd`), levels 1-22. |
| `--compress-level <n>` | Set the compression level. |
| `--zc <alg>` | Alias for `--compress-choice`. FastSync supports `zstd` (default), `lz4`, `zlib`, `zlibx`, `none`, and `auto`; `zlibx` behaves as `zlib`. |
| `--compress-level <n>` | Set the compression level (1-22). Omitted, each codec uses its rsync default: zstd 3, zlib/zlibx 6, lz4 ignored. |
| `--zc <alg>` | Alias for `--compress-choice`. FastSync supports `zstd` (default), `lz4`, `zlib`, `zlibx`, `none`, and `auto`; `zlib`/`zlibx` share the same literal-only zlib path. |
| `--zl <n>` | Alias for `--compress-level`. |
| `--skip-compress <list>` | Skip compression for `/`- or `,`-separated suffixes; defaults to rsync 3.4.1's built-in list. Incompatible with `--chunk-serialization`. |
| `--compress-threads <n>` | Use `n` zstd compression workers. Requires compression and a zstd build with threaded support; the setting affects sender CPU work only. |
@@ -519,9 +535,9 @@ features without changing the meaning of ordinary compatibility options.
| `--server-host <host>` | Select the TCP server host. |
| `--server-port <port>` | Select the TCP server port (`--port <port>` and `--port=<port>` are rsync-friendly aliases). |
| `--tls` | Enable TLS for TCP transport. |
| `--bwlimit <KB/s>` | Apply token-bucket bandwidth limiting. |
| `--progress` | Show rsync-style per-file progress blocks from the receiver's wire counters (FastSync omits rsync's leading `./` line). |
| `--stats` | Print transfer statistics, including the receiver-only counters reported over the wire; rsync's per-type `Number of files` breakdown is not reproduced. |
| `--bwlimit <RATE>` | Apply token-bucket bandwidth limiting with rsync's exact `parse_size_arg` grammar (bare = KiB/s, `K`/`M`/`G`/`T`/`P` binary, `KB`/`MB` decimal, `KiB`/`MiB` binary, decimals quantized to whole KiB, `0`/empty = no limit; also paces `--sendfile` transfers). |
| `--progress` | Show rsync-style per-file progress blocks from the receiver's wire counters; the root `./` line is printed whenever progress is active (rsync prints it only when the transfer root is created). |
| `--stats` | Print transfer statistics, including the receiver-only counters reported over the wire; `Number of files`/`Number of created files` carry rsync's per-type breakdown (deleted files are a single total). |
| `--timeout <seconds>` | Set the socket **and** per-message protocol I/O timeout. Default `0` = disabled (matching rsync); `0` disables it. |
| `--contimeout <seconds>` | Connection timeout (default 60, matching rsync); `0` disables it. |
@@ -555,22 +571,27 @@ remote SSH argv is already built injection-safe.
| `--size-only` | Skip incremental files matching in size, ignoring mtime. |
| `-I, --ignore-times` | Transfer files even when size and mtime match. |
| `-u, --update` | Skip files newer than the source on the receiver. |
| `--ignore-existing` | Skip files that already exist on the receiver; like rsync it does not apply to directories or symlinks. |
| `-@, --modify-window <sec>` | Modification-time tolerance (seconds) for the incremental/basis quick-check; `0` requires an exact mtime match. |
| `-W, --whole-file` | Transfer changed files without delta processing (`--no-whole-file` clears it). |
| `-B <n>, --block-size <n>` | Delta block size in bytes (alias `--delta-block`). |
| `-d, --dirs` | Transfer the named directory entries without recursing into their contents (aliases `--old-dirs`/`--old-d`). |
| `-R, --relative` | Use rsync's relative path semantics (including the `/./` cut); with `--files-from`, preserve each listed entry's relative path below the destination root. |
| `--files-from <file>` | Read the source file list from FILE (paths relative to the source root). |
| `--delay-updates` | Put updated files into place only at the end of the transfer. |
| `--compare-dest <dir>` | Extra comparison basis: unchanged files are not transferred (requires/implies `--incremental`). |
| `--copy-dest <dir>` | Like `--compare-dest`, but copies the unchanged file from DIR into the destination. |
| `--link-dest <dir>` | Like `--copy-dest`, but hard-links the unchanged file from DIR (repeatable; earlier DIRs win). |
| `-0, --from0` | Treat entries in `--files-from` files as NUL-delimited instead of newline-delimited. |
| `--delay-updates` | Put updated files into place only at the end of the transfer (the fixed `.fastsync-stage` staging name diverges from rsync; see [`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)). |
| `--compare-dest <dir>` | Extra comparison basis: unchanged files are not transferred (requires/implies `--incremental`; a basis MISS above the 256 MiB whole-file payload bound is refused — see [`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)). |
| `--copy-dest <dir>` | Like `--compare-dest`, but copies the unchanged file from DIR into the destination (same basis-size caveat; see [`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)). |
| `--link-dest <dir>` | Like `--copy-dest`, but hard-links the unchanged file from DIR (repeatable; earlier DIRs win; same basis-size caveat; see [`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)). |
| `--verify-basis` | FastSync-only: require a basis hit to match the source by whole-file digest instead of trusting the size+mtime quick-check (default matches rsync). |
| `--preallocate` | Allocate destination file space up front (fail-fast on a full disk). |
| `--append` | Resume a shorter destination by appending only its tail (prefix not verified; requires `--incremental`). |
| `--append-verify` | Like `--append`, but verifies the retained prefix checksum first (falls back to a full transfer on mismatch). |
| `--delete` | Request removal of destination entries absent from the source. The server must allow deletion. Default timing is delete-after: extras are removed only after the whole transfer succeeded. Scoped to the synchronized directories, so `--files-from` subsets are safe. |
| `--delete` | Request removal of destination entries absent from the source. The server must allow deletion. Default timing is delete-during (matching rsync's `--del`): extras are removed per directory as the transfer proceeds, so destination space is freed progressively. Scoped to the synchronized directories, so `--files-from` subsets are safe. |
| `--delete-before` | Delete extras before the transfer starts (implies `--delete`). |
| `--delete-during`, `--del` | Delete extras once the keep-set manifest is known, before data is applied (implies `--delete`; early mode, same engine behaviour as `--delete-before`). |
| `--delete-delay` | Delete extras only after a successful transfer (implies `--delete`; commit mode, same behaviour as `--delete-after`). |
| `--delete-during`, `--del` | Delete each directory's extras as that directory is processed (implies `--delete`). Since protocol 2.24.0 the sender streams a per-directory `STATUS_DELETE_PLAN` frame as it reaches each source directory; this is also the default timing of a plain `--delete`. |
| `--delete-delay` | Record extras per directory during the scan but remove them only after a successful transfer (implies `--delete`). Uses the same per-directory `STATUS_DELETE_PLAN` frames as `--delete-during`, applied late. |
| `--delete-commit` | FastSync-only: atomic delete-after timing (only after the whole transfer succeeded). |
| `--delete-after` | Explicit delete-after timing: delete only after the transfer succeeded (implies `--delete`). |
| `--delete-excluded` | Also delete filter-excluded destination mirrors (size-pruned mirrors stay protected). |
| `--max-delete <n>` | Delete at most n destination entries; the rest are skipped and the run exits 25 (partial), matching rsync. |
@@ -585,18 +606,18 @@ remote SSH argv is already built injection-safe.
| `--max-alloc <SIZE>` | Maximum single allocation (binary units; default 1G; `0` = no local limit). |
| `--max-depth <n>` | Limit recursive scanning depth; zero means unlimited. |
| `-b, --backup` | Back up overwritten files. |
| `-T, --temp-dir <dir>` | Scratch directory for temp files before the atomic install (confined to the receive root; `EXDEV` falls back to a non-atomic copy). |
| `-T, --temp-dir <dir>` | Scratch directory for temp files before the atomic install (confined to the receive root: relative resolves below it, absolute must canonicalize inside it; `EXDEV` falls back to a non-atomic copy). |
| `--backup-dir <dir>` | Store backups under a separate directory (requires `--backup`). |
| `--suffix <suffix>` | Set the backup filename suffix (default: `~`). |
| `--partial` | Select partial-transfer handling. On failed/interrupted writes the already-written temp file is retained (best-effort) for resumption. With `--partial --partial-dir <dir>`, completed files are written under the partial directory and installed atomically. |
| `--partial-dir <dir>` | Set a relative partial-transfer directory below the server destination root. Use with `--partial`. |
| `--inplace` | Write directly to the destination instead of using a temporary file. |
| `--partial-dir <dir>` | Set a relative partial-transfer directory below the server destination root. Implies `--partial`. Rejected together with `--inplace` (`--inplace cannot be used with --partial-dir`, matching rsync), because the inplace path bypasses partial/temp staging. |
| `--inplace` | Write directly to the destination instead of using a temporary file. Cannot be combined with `--partial-dir`. |
| `--fsync` | Fsync every written file before publication. |
| `--write-batch=FILE` | Run the normal live transfer and also emit a self-contained batch file of the source tree. |
| `--only-write-batch=FILE` | Emit the batch file only (no destination, no server). |
| `--read-batch=FILE` | Apply a batch file to the destination (no source, no server). |
| `--write-batch=FILE` | Run the normal live transfer and also emit a self-contained batch file of the source tree (FastSync-native format, not rsync-interoperable). |
| `--only-write-batch=FILE` | Emit the batch file only (no destination, no server); FastSync-native format, not rsync-interoperable. |
| `--read-batch=FILE` | Apply a batch file to the destination (no source, no server); FastSync-native format, not rsync-interoperable. |
| `--stop-after=MINS` | Stop the transfer after MINS minutes; whatever was already transferred is kept. |
| `--stop-at=TIME` | Stop at an absolute time (`HH:MM`, `HH:MM:SS`, or `now+N[smhd]`). An early stop skips the late `--delete` keep-set. |
| `--stop-at=TIME` | Stop at an absolute time. Accepts rsync's `parse_time` forms (`Y-M-DTh:m`, `Y/M/DTh:m`, `Y-M-D`, `M-D`, `D`, `h:m`, `:m`, `T h:m`; omitted fields resolve to the next matching point in the local timezone), plus `now+N[smhd]` and FastSync's `HH:MM`/`HH:MM:SS` clock-time spelling. An early stop skips the late `--delete` keep-set. |
### Metadata and links
@@ -605,8 +626,11 @@ remote SSH argv is already built injection-safe.
| `--preserve` | Preserve mode and mtime (long form only; equivalent to `-p` + `-t`). Add `-o`/`-g` for owner/group, `-U`/`--atimes` for atime, or an identity flag (`--chown`/`--usermap`/`--groupmap`/`--numeric-ids`/`--copy-as`) for mapped ownership. |
| `-U`, `--atimes` | Preserve access times. Captured with the metadata payload; does not enable ownership. |
| `-N`, `--crtimes` | Capture birth time and transmit it; it cannot be applied because no portable filesystem call can set a birth time (documented divergence). |
| `-p`, `--perms` | Preserve permission bits. One of the four per-attribute preserve flags (with `-t`/`-o`/`-g`); under `-p` the source mode is copied exactly (setuid/setgid/sticky and group/other-write included), matching rsync. |
| `-p`, `--perms` | Preserve permission bits. One of the four per-attribute preserve flags (with `-t`/`-o`/`-g`); under `-p` the source mode is copied exactly (group/other-write included; setuid/setgid/sticky included only when super-user activities are permitted, masked under `SUPER_MODE_OFF`/`--no-super`), matching rsync otherwise. |
| `-t`, `--times` | Preserve modification times. Independent of the other attributes; `-O`/`--omit-dir-times` suppresses directories only. |
| `-O`, `--omit-dir-times` | Do not apply modification times to directories. |
| `-J`, `--omit-link-times` | Do not apply times to symlinks. |
| `--open-noatime` | Open source files with `O_NOATIME` so reading for a transfer does not update their access time (client-only). |
| `-o`, `--owner` | Preserve the source owner (uid). Mapped by name on the receiver with a raw-numeric fallback (only numeric ids cross the wire); application is privilege-gated. |
| `-g`, `--group` | Preserve the source group (gid). Same name-mapping/numeric-fallback and privilege gating as `-o`. |
| `--no-perms`, `--no-times`, `--no-owner`, `--no-group` | Negate each per-attribute flag (also `--no-p`/`--no-t`/`--no-o`/`--no-g`); `--no-preserve` clears all four. |
@@ -619,7 +643,7 @@ remote SSH argv is already built injection-safe.
| `--groupmap=MAP` | Map group names when applying ownership (same syntax as `--usermap`). |
| `--numeric-ids` | Mapping modifier: apply the source numeric uid/gid directly instead of mapping by name (combine with `-o`/`-g`, `-a`, or a map). |
| `--copy-as=USER[:GROUP]` | Force every written entry to USER[:GROUP]; requires a privileged receiver. |
| `--fake-super` | Record the resolved owner plus mode/time in a reserved `user.fastsync.stat` xattr and replay mode/time; never performs a real chown. |
| `--fake-super` | Record the resolved owner plus full mode/rdev in rsync's reserved `user.rsync.%stat` xattr (rsync 3.4.1 grammar) and replay the permission bits; never performs a real chown. |
| `--super` | Permit the receiver to attempt confined super-user activities (device nodes). |
| `--no-super` | Forbid those super-user activities even when the receiver is root. |
| `-l`, `--links` | Copy symlinks as symlinks; the target is stored verbatim (absolute and `..`-bearing targets included), matching rsync. |
@@ -633,6 +657,7 @@ remote SSH argv is already built injection-safe.
| `-D` | Preserve device and special files (implies `--devices --specials`). |
| `--devices` | Recreate device nodes on the destination (privileged; skipped without `CAP_MKNOD`). |
| `--specials` | Recreate special files: FIFOs and unix sockets. |
| `--copy-devices` | Copy a source device's content as an ordinary regular file on the destination (rsync's non-privileged safe mode) instead of recreating the device node. |
| `-S`, `--sparse` | Sparse-file handling: receiver preserves holes (zero runs are written as holes; no wire change). |
### Output and logging
@@ -641,12 +666,17 @@ remote SSH argv is already built injection-safe.
|---|---|
| `-v`, `--verbose` | Enable debug logging. |
| `-q`, `--quiet` | Suppress non-error output. |
| `--progress` | Show rsync-style per-file progress blocks (not rsync's leading `./` line). |
| `--stats` | Print transfer statistics, including the receiver-only counters reported over the wire. |
| `-i`, `--itemize-changes` | Print an rsync-style per-file change line. |
| `--out-format=FORMAT` | Output format for changed files (`%f %n %l %b %M %%`). |
| `--progress` | Show rsync-style per-file progress blocks; the root `./` line is printed whenever progress is active (rsync prints it only when the transfer root is created). |
| `--stats` | Print transfer statistics, including the receiver-only counters reported over the wire; `Number of files`/`Number of created files` carry rsync's per-type breakdown (deleted files are a single total). |
| `-i, --itemize-changes` | Print an rsync-style per-file change line. |
| `--out-format=FORMAT` | Output format for changed files (`%f %n %l %b %c %C %i %M %%`). |
| `--list-only` | List source files instead of transferring. |
| `--outbuf=MODE` | stdout/stderr buffering: `N` (none/unbuffered), `L` (line-buffered), or `B` (block-buffered, default). |
| `--log-file <path>` | Write log output to a file. |
| `--log-file-format=FORMAT` | Per-file log-line format (requires `--log-file`). |
| `--stderr=MODE` | Route logging to stderr: `errors` or `all`. |
| `--msgs2stderr` | Route all messages to stderr (deprecated spelling of `--stderr=all`). |
| `--no-msgs2stderr` | Select errors-only stderr (deprecated spelling; the default). |
| `-V`, `--version` | Print the FastSync protocol version. |
| `--help` | Print command usage. |
@@ -659,7 +689,7 @@ remote SSH argv is already built injection-safe.
| `--fastsync-server-path <path>` | Remote FastSync server path for SSH mode (client-only; never crosses the wire). |
| `--rsync-path <path>` | Alias for `--fastsync-server-path`. |
| `-M`, `--remote-option=OPT` | Append OPT to the remote server invocation over SSH (repeatable; rejected for daemon/TCP destinations). |
| `--trust-sender` | Receiver-local: trust the remote sender's file list and skip path re-validation (does not affect symlink targets). |
| `--trust-sender` | Receiver-local: trust the remote sender's file list and skip path re-validation (does not affect symlink targets). **On the client this flag alone is inert** — it is never sent on the wire; the server must be started with its own `--trust-sender`, or the client must forward it with `-M--trust-sender` (SSH only). |
| `--timeout <sec>` | Socket + per-message I/O timeout; default `0` = disabled. |
| `--contimeout <sec>` | Connection timeout; default 60; `0` disables. |
| `--source-dir <path>` | Set the source directory explicitly. |
@@ -671,6 +701,11 @@ remote SSH argv is already built injection-safe.
| `-4`, `--ipv4` | Force IPv4 for destination resolution. |
| `-6`, `--ipv6` | Force IPv6 for destination resolution. |
| `--sockopts=OPTS` | Comma-separated OPT=VAL socket options applied before connect. |
| `--blocking-io` | SSH transport only: leave the socket without read/write timeouts so it blocks naturally (no effect on TCP). |
| `--protocol=NUM` | Force the wire protocol version; must equal the current `PROTOCOL_VERSION` (FastSync cannot speak older/virtual wire formats). |
| `--old-args` | Accepted for rsync CLI compatibility; no effect (the remote server path is always safely quoted). |
| `--iconv=LOCAL[,REMOTE]` | Convert file-name charsets at the wire boundary (`LOCAL` is our names' charset, `REMOTE` the peer's, defaulting to `LOCAL`). |
| `--no-iconv` | Disable `--iconv` charset conversion (same as `--iconv=-`). |
| `--tls` | Enable TLS. Requires `--cert`, `--key`, and `--ca`. |
| `--cert <path>` | TLS certificate file. |
| `--key <path>` | TLS private key file. |
@@ -697,14 +732,14 @@ remote SSH argv is already built injection-safe.
| `-6`, `--ipv6` | Bind an IPv6 socket. |
| `--allow-delete` | Permit client delete manifests. Deletion is refused by default. This also gates `--force` (which can recursively replace/remove a destination directory tree). |
| `--allow-super` | Standalone TCP listener only: keep super-user activities enabled for a **root** receiver. Without it a root standalone server forces `SUPER_MODE_OFF`, so client `--devices`/`--write-devices`/`--super` and client-chosen ownership requests are skipped/refused. Rejected with `--stdio` (the SSH remote argv is client-composed; use a forced command if the default must hold). No effect when not root. Daemon modules opt in per module with `client owner = yes`. |
| `--trust-sender` | Trust the remote sender's file list: skip the receiver's up-front path-traversal re-validation (fewer checks, faster, potentially unsafe; off by default). It does not affect symlink targets, which are stored verbatim either way. |
| `--trust-sender` | Trust the remote sender's file list: skip the receiver's up-front path-traversal re-validation (fewer checks, faster, potentially unsafe; off by default). It does not affect symlink targets, which are stored verbatim either way. A client `--trust-sender` is never sent over the wire — the server must set this flag itself, or the client must forward it via `-M--trust-sender`. |
| `--no-super` | Operator veto: never attempt super-user activities (ownership, device nodes) even as root, and refuse any client `--copy-as`/`--super` request. |
| `--allow-unauthenticated` | Permit plaintext/anonymous network clients; an auth-required module still accepts only opted-in loopback plaintext. |
| `--iconv=LOCAL[,REMOTE]` | Declare this server's LOCAL charset for file-name conversion. |
| `--password-file=FILE` | Credential store for modules that declare `auth users`. Requires `--daemon`. |
| `--early-input=FILE` | Second credential store layered over `--password-file`. Requires `--daemon`. |
| `--hash-credentials <file>` | Read `<file>`'s `user:password` lines and print PBKDF2 credential-store lines to stdout, then exit. Cannot be combined with `--daemon` or `--stdio`. |
| `--iterations N` | PBKDF2 iteration count for `--hash-credentials` (default 600000, range 100000–10000000). Requires `--hash-credentials`. |
| `--password-file=FILE` | Credential store for modules that declare `auth users`. Requires `--daemon`. FastSync-native SCRAM/PBKDF2 format, not rsync-interoperable. |
| `--early-input=FILE` | Second credential store layered over `--password-file`. Requires `--daemon`. FastSync-native format, not rsync-interoperable. |
| `--hash-credentials <file>` | Read `<file>`'s `user:password` lines and print PBKDF2 credential-store lines to stdout, then exit. Cannot be combined with `--daemon` or `--stdio`. FastSync-native, not rsync-interoperable. |
| `--iterations N` | PBKDF2 iteration count for `--hash-credentials` (default 600000, range 100000–10000000). Requires `--hash-credentials`. FastSync-native, not rsync-interoperable. |
| `-v`, `--verbose` | Enable debug logging. |
| `--help` | Print server usage. |
@@ -733,9 +768,17 @@ and `address`, the global section accepts:
- `hosts allow` / `hosts deny` — comma- and/or whitespace-separated host access
patterns.
A `[module]` requires `path`, and may also set `read only`, `client owner`,
`auth users`, `max connections` (0 = unlimited; enforced per module across all
connection children), and its own `hosts allow`/`hosts deny`.
A `[module]` requires `path`, and may also set `read only`, `write only`,
`client owner`, `auth users`, `max connections` (0 = unlimited; enforced per
module across all connection children), and its own `hosts allow`/`hosts deny`.
Like rsync, a module is **read-only by default**: a bare `[module]` with only a
`path` refuses a write transfer. Opt a module into writability explicitly with
`read only = no` or `write only = yes`; a global `read only` value in the
section before the first `[module]` sets the default for later modules, and a
module's own `read only`/`write only = yes` always wins over it. An
rsync-style `write only = yes` is mapped to writability because FastSync is
push-only (a module can never be read from the network).
The per-host cap and the shared auth lockout identify a source by its numeric
peer IP. **Loopback peers (127.0.0.0/8, IPv6 `::1`) are exempt**: every local
@@ -789,7 +832,7 @@ before the module list, before authentication, and the connecting peer address
## Protocol and Security
FastSync protocol version `2.26.0` is shared by the client and server. The
FastSync protocol version `2.29.0` is shared by the client and server. The
current protocol is sender-driven and includes configuration negotiation,
including the maximum allocation limit, incremental checks, checksums,
manifests, keep-alives, abort handling, per-file remove-source results, and
@@ -862,8 +905,10 @@ The project will reach the drop-in replacement goal in stages:
completion wave's scope; the tests live in `tests/integration/` and skip
cleanly when rsync is unavailable.
3. `-a` implements full rsync `-rlptgoD`; under `-p` the source mode is copied
exactly (no masking). Ownership application stays privilege-gated, as in
rsync.
exactly, including group/other-write bits, with setuid/setgid/sticky copied
only when super-user activities are permitted (masked under
`SUPER_MODE_OFF`/`--no-super`). Ownership application stays privilege-gated,
as in rsync.
4. Symlink (verbatim storage), sparse-file, metadata, delete-policy (including
`--max-delete` partial + exit 25, per-directory `--delete-during`/
`--delete-delay`), codecs, and resumable-write semantics are implemented;
+272 -150
View File
File diff suppressed because one or more lines are too long
+3
View File
@@ -8,3 +8,6 @@ markers =
daemon_detach: real double-fork backgrounding path (--daemon without
--no-detach); slower/fragile, so it runs in the full suite but not the
fast PR gate
parity: differential rsync-parity case (full set; runs on push to
dev/main)
parity_ci: fast differential rsync-parity subset (runs on the PR gate)
+133 -27
View File
@@ -1,5 +1,6 @@
#include "change_list.h"
#include "checksum.h"
#include "log.h"
#include "utils.h"
#include <fcntl.h>
#include <limits.h>
@@ -70,7 +71,16 @@ static bool strbuf_append(StrBuf* buf, const char* text) {
bool change_list_enabled(const Config* config) {
return config != NULL && (config->itemize_changes || config->out_format != NULL ||
(config->log_file != NULL && config->log_file_format != NULL));
(config->log_file != NULL && config->log_file_format != NULL) ||
(config->info_level & LOG_INFO_NAME) != 0);
}
/* Emitted once, lazily, ahead of the first --info=name entry: rsync prints the
* transfer-root `./` name line when the root directory is (re)created. */
static bool name_root_printed = false;
void change_reset_name_root(void) {
name_root_printed = false;
}
/* ---- Itemize code ---- */
@@ -131,6 +141,10 @@ static void itemize_code(const Config* config, const ChangeEvent* event, char co
update = 'h';
else if (created)
update = (event->is_directory || event->is_symlink || event->is_special) ? 'c' : '>';
else if (event->is_directory)
/* rsync: an existing directory that only has attribute changes carries no
transfer, so the update column is `.` rather than `>`. */
update = '.';
else
update = '>';
code[0] = update;
@@ -158,12 +172,15 @@ static void itemize_code(const Config* config, const ChangeEvent* event, char co
code[11] = '\0';
}
/* rsync %n: the transfer-relative name, with a trailing slash for directories. */
/* rsync %n: the transfer-relative name, with a trailing slash for directories.
* The transfer root is `.` (so `%n` renders `./`), matching rsync's root entry. */
static bool append_name(StrBuf* buf, const ChangeEvent* event) {
if (!strbuf_append(buf, event->name != NULL ? event->name : ""))
const char* name = event->name != NULL ? event->name : "";
if (event->is_directory && name[0] == '\0')
return strbuf_append(buf, "./");
if (!strbuf_append(buf, name))
return false;
if (event->is_directory && (event->name == NULL || event->name[0] == '\0' ||
event->name[strlen(event->name) - 1] != '/'))
if (event->is_directory && name[strlen(name) - 1] != '/')
return strbuf_append_char(buf, '/');
return true;
}
@@ -192,27 +209,55 @@ char* change_render_itemize(const Config* config, const ChangeEvent* event) {
return line.data;
}
/* ---- --out-format / --log-file-format ---- */
/* rsync 3.4.1's `%C` uses the negotiated transfer checksum; with the default
* "auto" choice on both ends that is xxh128. FastSync's internal XXH64 default
* is not an rsync algorithm, so map it to xxh128 for parity. */
static ChecksumAlgo out_format_checksum_algo(const Config* config) {
switch ((ChecksumAlgo)config->checksum_algo) {
case CHECKSUM_ALGO_MD5:
return CHECKSUM_ALGO_MD5;
case CHECKSUM_ALGO_XXH3:
return CHECKSUM_ALGO_XXH3;
case CHECKSUM_ALGO_XXH128:
return CHECKSUM_ALGO_XXH128;
case CHECKSUM_ALGO_XXH64:
default:
return CHECKSUM_ALGO_XXH128;
/* rsync's `--info=name` line for an updated entry: the transfer-relative name
* (trailing slash for directories) plus the ` -> target` / ` => target` link
* suffix. `--info=name` does not alter an itemize/out-format run. */
static char* change_render_name(const ChangeEvent* event) {
StrBuf line = {0};
bool ok = append_name(&line, event) && append_link_suffix(&line, event);
if (!ok) {
strbuf_free(&line);
return NULL;
}
if (line.data == NULL) {
line.data = str_dup("");
if (!line.data)
return NULL;
}
return line.data;
}
/* Render a digest as rsync's sum_as_hex: for xxh128 the HIGH 64-bit half is
* printed before the low half; every other algorithm prints its bytes in order. */
/* rsync's `--info=name2` line for an unchanged entry: `NAME is uptodate`. */
static char* change_render_name_uptodate(const ChangeEvent* event) {
char* name = change_render_name(event);
if (name == NULL)
return NULL;
size_t length = strlen(name);
char* line = malloc(length + sizeof(" is uptodate"));
if (line == NULL) {
free(name);
return NULL;
}
memcpy(line, name, length);
memcpy(line + length, " is uptodate", sizeof(" is uptodate"));
free(name);
return line;
}
/* ---- --out-format / --log-file-format ---- */
/* rsync 3.4.1's `%C` uses the negotiated TRANSFER checksum (the first name of a
* two-name "transfer,pre-transfer" --checksum-choice), not the pre-transfer
* whole-file digest FastSync compares against on the wire. The default "auto"
* resolves to xxh128, so an explicit selection and the default both render the
* selected algorithm's digest. */
static ChecksumAlgo out_format_checksum_algo(const Config* config) {
return (ChecksumAlgo)config->cli.checksum_transfer_algo;
}
/* Render a digest as rsync's sum_as_hex: xxh128 prints the HIGH 64-bit half
* before the low half, and xxh64/xxh3 print their 64-bit value big-endian; every
* other algorithm prints its bytes in order. */
static void digest_to_hex(ChecksumAlgo algo, const uint8_t* digest, size_t len, char* out) {
if (algo == CHECKSUM_ALGO_XXH128 && len == 16) {
uint64_t low = 0;
@@ -222,6 +267,12 @@ static void digest_to_hex(ChecksumAlgo algo, const uint8_t* digest, size_t len,
snprintf(out, len * 2 + 1, "%016llx%016llx", (unsigned long long)high, (unsigned long long)low);
return;
}
if ((algo == CHECKSUM_ALGO_XXH64 || algo == CHECKSUM_ALGO_XXH3) && len == 8) {
uint64_t value = 0;
memcpy(&value, digest, sizeof(value));
snprintf(out, len * 2 + 1, "%016llx", (unsigned long long)value);
return;
}
static const char hex[] = "0123456789abcdef";
for (size_t i = 0; i < len; i++) {
out[i * 2] = hex[(digest[i] >> 4) & 0xf];
@@ -260,6 +311,9 @@ static void fill_event_checksum(const Config* config, const File* file, ChangeEv
if (file->path == NULL)
return;
ChecksumAlgo algo = out_format_checksum_algo(config);
/* rsync renders `--checksum-choice=none` as a blank 2-character column. */
if (algo == CHECKSUM_ALGO_NONE)
return;
uint8_t digest[CHECKSUM_MAX_DIGEST_LEN];
size_t len = 0;
/* rsync's %C is the transfer checksum, which is always seeded with 0 (it is
@@ -329,9 +383,10 @@ char* change_render_format(const char* format, const Config* config, const Chang
if (event->checksum_known) {
ok = strbuf_append(&line, event->checksum);
} else {
/* rsync pads a non-regular / untransferred entry with spaces. */
/* rsync pads a non-regular / untransferred / `none` entry with spaces;
`none` renders as a blank 2-character column. */
ChecksumAlgo algo = out_format_checksum_algo(config);
int width = checksum_digest_len(algo) * 2;
int width = algo == CHECKSUM_ALGO_NONE ? 2 : checksum_digest_len(algo) * 2;
for (int i = 0; i < width && ok; i++)
ok = strbuf_append_char(&line, ' ');
}
@@ -438,10 +493,24 @@ static void print_escaped_line(FILE* stream, const char* line, bool eight_bit_ou
void change_emit(const Config* config, const ChangeEvent* event) {
if (event == NULL || !change_list_enabled(config))
return;
if (event->decision == CHANGE_UP_TO_DATE)
return;
bool to_stdout = config->itemize_changes || config->out_format != NULL;
bool to_log = config->log_file != NULL && config->log_file_format != NULL;
bool progress_active = config->show_progress || (config->info_level & LOG_INFO_PROGRESS);
if (event->decision == CHANGE_UP_TO_DATE) {
/* --info=name2 prints `NAME is uptodate` for entries the receiver already
had. An itemize/out-format run reports them through its own format (or
not at all), the progress stream has no frame for them, and neither the
itemize nor the log-file stream previously reported an up-to-date entry,
so nothing else here changes. */
if (!to_stdout && (config->info_level & LOG_INFO_NAME_UPTODATE) != 0 && !progress_active) {
char* line = change_render_name_uptodate(event);
if (line != NULL) {
print_escaped_line(stdout, line, config->eight_bit_output);
free(line);
}
}
return;
}
if (to_stdout) {
char* line = config->out_format != NULL
? change_render_format(config->out_format, config, event)
@@ -450,6 +519,20 @@ void change_emit(const Config* config, const ChangeEvent* event) {
print_escaped_line(stdout, line, config->eight_bit_output);
free(line);
}
} else if ((config->info_level & LOG_INFO_NAME) != 0 && !progress_active) {
/* --info=name without -i/--out-format: print the updated entry's name. The
--progress path owns the name line when progress output is active (it
emits the same names before the progress frames), so do not duplicate.
The transfer-root `./` line precedes the first such name. */
if (!name_root_printed) {
name_root_printed = true;
fputs("./\n", stdout);
}
char* line = change_render_name(event);
if (line != NULL) {
print_escaped_line(stdout, line, config->eight_bit_output);
free(line);
}
}
if (to_log) {
char* line = change_render_format(config->log_file_format, config, event);
@@ -610,6 +693,29 @@ void change_emit_file_sent(const Config* config, const File* file) {
change_emit_file_sent_bytes(config, file, payload, 0);
}
void change_emit_file_uptodate(const Config* config, const File* file) {
if (file == NULL || !change_list_enabled(config))
return;
ChangeEvent event;
memset(&event, 0, sizeof(event));
event.decision = CHANGE_UP_TO_DATE;
event.is_directory = false;
event.is_symlink = file->is_symlink;
event.is_special = file->is_special;
event.is_hardlink = file->link_group != 0 && !file->link_first;
event.symlink_target = file->symlink_target;
event.hardlink_target = file->hardlink_target;
event.size = file->data != NULL ? file->data->size : 0;
event.dest = file->dest_state;
char* name = NULL;
char* path = NULL;
fill_event_from_file(config, file, &event, &name, &path);
if (name != NULL && path != NULL)
change_emit(config, &event);
free(name);
free(path);
}
void change_emit_dir_sent(const Config* config, const File* file) {
if (file == NULL || !change_list_enabled(config))
return;
+9
View File
@@ -102,4 +102,13 @@ void change_emit_file_sent(const Config* config, const File* file);
/* Build and emit a CHANGE_SENT event for an explicit directory entry (-d). */
void change_emit_dir_sent(const Config* config, const File* file);
/* Build and emit a CHANGE_UP_TO_DATE event for a file the receiver already had.
* With --info=name2 it renders rsync's "NAME is uptodate" line (no output
* otherwise). */
void change_emit_file_uptodate(const Config* config, const File* file);
/* Reset the lazy transfer-root `./` line emitted ahead of the first
* --info=name entry. Call once at the start of a transfer. */
void change_reset_name_root(void);
#endif
+344 -59
View File
@@ -23,6 +23,7 @@
#include <langinfo.h>
#include <limits.h>
#include <locale.h>
#include <math.h>
#include <time.h>
#include <signal.h>
#include <stdbool.h>
@@ -53,13 +54,26 @@ bool client_abort_pending(void) {
}
#ifndef FASTSYNC_TEST_BUILD
/* SIG_DFL disposition used by the handler's "not armed" fallback. It is built
* once at load time so the handler can restore the default action with
* sigaction(2) -- which is async-signal-safe -- instead of signal(3), which is
* not. The zero-initialized sa_mask is the empty set. */
static const struct sigaction client_default_action = {
.sa_handler = SIG_DFL,
.sa_flags = 0,
};
/* Signal handler: perform NO work beyond storing the flag. Logging, protocol
* I/O and the STATUS_ABORT frame are all done later on the normal send path,
* which is not async-signal-safe. When no transfer is armed, fall back to the
* default action so local-only modes remain interruptible. */
* which is not async-signal-safe. When no transfer is armed, restore the
* default disposition (async-signal-safe sigaction) and re-raise so local-only
* modes remain interruptible. The handler deliberately stays installed while a
* transfer is armed -- rather than using SA_RESETHAND -- so a second Ctrl-C
* during the graceful abort keeps setting the flag instead of hard-killing the
* process mid-cleanup. */
static void client_signal_handler(int signo) {
if (!client_abort_armed) {
signal(signo, SIG_DFL);
sigaction(signo, &client_default_action, NULL);
raise(signo);
return;
}
@@ -156,20 +170,26 @@ static int set_positive_int_option(int* dest, const char* value, const char* opt
* name is a hard error with rsync's exit code 4, never a silent no-op. */
static int set_compression_choice(Config* config, const char* value) {
if (!value) {
config->cli_exit_code = 4;
config->cli.cli_exit_code = 4;
return -1;
}
int algo;
if (strcasecmp(value, "auto") == 0)
algo = (int)compression_negotiate_default();
else
if (strcasecmp(value, "auto") == 0) {
algo = compression_choice_resolve();
if (algo < 0) {
log_message(LOG_LEVEL_ERROR, "RSYNC_COMPRESS_LIST names no supported compression algorithm");
config->cli.cli_exit_code = 4;
return -1;
}
} else {
algo = compression_algo_from_name(value);
}
if (algo < 0) {
log_message(LOG_LEVEL_ERROR,
"--compress-choice '%s' is not a supported algorithm; FastSync supports zstd, "
"lz4, zlib, zlibx, none or auto",
value);
config->cli_exit_code = 4;
config->cli.cli_exit_code = 4;
return -1;
}
const char* canonical = compression_algo_name((CompressionAlgo)algo);
@@ -206,7 +226,7 @@ static int resolve_checksum_name(const char* name, size_t len, int* out) {
* resolves to FastSync's negotiated default (xxh128). */
static int set_checksum_choice(Config* config, const char* value) {
if (!value) {
config->cli_exit_code = 4;
config->cli.cli_exit_code = 4;
return -1;
}
const char* comma = strchr(value, ',');
@@ -224,19 +244,28 @@ static int set_checksum_choice(Config* config, const char* value) {
"--checksum-choice '%s' is invalid; FastSync supports xxh64 (or xxhash), xxh128, "
"xxh3, md5, md4, sha1, none or auto, optionally as 'transfer,pre-transfer'",
value);
config->cli_exit_code = 4;
config->cli.cli_exit_code = 4;
return -1;
}
ChecksumAlgo negotiated = checksum_negotiate_default();
int negotiated = -1;
if (rc1 == 1 || rc2 == 1) {
negotiated = checksum_choice_resolve();
if (negotiated < 0) {
log_message(LOG_LEVEL_ERROR, "RSYNC_CHECKSUM_LIST names no supported checksum algorithm");
config->cli.cli_exit_code = 4;
return -1;
}
}
if (rc1 == 1)
transfer = (int)negotiated;
transfer = negotiated;
if (!name2)
pre = transfer;
else if (rc2 == 1)
pre = (int)negotiated;
pre = negotiated;
config->checksum_algo = pre;
config->checksum_transfer_algo = transfer;
config->cli.checksum_transfer_algo = transfer;
config->cli.checksum_choice_set = true;
/* rsync: "none" for the transfer checksum forces --whole-file. */
if (transfer == (int)CHECKSUM_ALGO_NONE)
config->whole_file = true;
@@ -488,9 +517,8 @@ static bool split_flag_level(const char* token, char* name, size_t name_size, in
* of rsync's `symsafe`, `hlink`, and `own`. */
static bool is_accepted_debug_category(const char* name) {
static const char* const categories[] = {
"acl", "backup", "bind", "chdir", "cmd", "connect", "del", "deltasum",
"dup", "exit", "filter", "flist", "fuzzy", "genr", "hash", "hl",
"hlink", "iconv", "nstr", "own", "owner", "recv", "send", "time",
"acl", "backup", "bind", "chdir", "cmd", "connect", "dup", "exit", "fuzzy",
"genr", "hl", "hlink", "iconv", "nstr", "own", "owner", "time",
};
for (size_t i = 0; i < sizeof(categories) / sizeof(categories[0]); i++) {
if (strcmp(name, categories[i]) == 0)
@@ -501,7 +529,9 @@ static bool is_accepted_debug_category(const char* name) {
static bool is_accepted_info_category(const char* name) {
static const char* const categories[] = {
"backup", "del", "flist", "mount", "nonreg", "progress", "remove", "syms", "symsafe",
"backup",
"syms",
"symsafe",
};
for (size_t i = 0; i < sizeof(categories) / sizeof(categories[0]); i++) {
if (strcmp(name, categories[i]) == 0)
@@ -552,6 +582,18 @@ static int parse_debug_flags(const char* value, Config* config) {
flag = LOG_DEBUG_PACK;
} else if (strcmp(name, "util") == 0) {
flag = LOG_DEBUG_UTIL;
} else if (strcmp(name, "flist") == 0) {
flag = LOG_DEBUG_FLIST;
} else if (strcmp(name, "del") == 0) {
flag = LOG_DEBUG_DEL;
} else if (strcmp(name, "hash") == 0 || strcmp(name, "deltasum") == 0) {
flag = LOG_DEBUG_HASH;
} else if (strcmp(name, "recv") == 0) {
flag = LOG_DEBUG_RECV;
} else if (strcmp(name, "filter") == 0) {
flag = LOG_DEBUG_FILTER;
} else if (strcmp(name, "send") == 0) {
flag = LOG_DEBUG_SEND;
} else if (is_accepted_debug_category(name)) {
continue;
} else {
@@ -608,14 +650,41 @@ static int parse_info_flags(const char* value, Config* config) {
free(flags);
return 1;
}
if (strcmp(name, "copy") == 0 || strcmp(name, "name") == 0)
if (strcmp(name, "copy") == 0)
flag = LOG_INFO_COPY;
else if (strcmp(name, "misc") == 0)
else if (strcmp(name, "name") == 0) {
/* name level 2 adds rsync's "is uptodate" lines. */
if (level == 0)
parsed &= ~(uint32_t)(LOG_INFO_NAME | LOG_INFO_NAME_UPTODATE);
else {
parsed |= LOG_INFO_NAME;
if (level >= 2)
parsed |= LOG_INFO_NAME_UPTODATE;
else
parsed &= ~(uint32_t)LOG_INFO_NAME_UPTODATE;
}
continue;
} else if (strcmp(name, "misc") == 0)
flag = LOG_INFO_MISC;
else if (strcmp(name, "skip") == 0)
flag = LOG_INFO_SKIP;
else if (strcmp(name, "stats") == 0)
else if (strcmp(name, "stats") == 0) {
flag = LOG_INFO_STATS;
/* `--info=stats` requests the same transfer-statistics block as
`--stats`; `--info=stats0` turns it back off. */
config->stats = level > 0;
} else if (strcmp(name, "del") == 0)
flag = LOG_INFO_DEL;
else if (strcmp(name, "remove") == 0)
flag = LOG_INFO_REMOVE;
else if (strcmp(name, "flist") == 0)
flag = LOG_INFO_FLIST;
else if (strcmp(name, "nonreg") == 0)
flag = LOG_INFO_NONREG;
else if (strcmp(name, "mount") == 0)
flag = LOG_INFO_MOUNT;
else if (strcmp(name, "progress") == 0)
flag = LOG_INFO_PROGRESS;
else if (is_accepted_info_category(name))
continue;
else {
@@ -893,8 +962,18 @@ static const OptionEntry OPTION_TABLE[] = {
/* rsync -r/--recursive: FastSync is always recursive, so this is a
* faithful no-op (accepted silently, never consumes an argument). */
{"--recursive", "-r", OPT_NOOP, 0},
/* rsync's incremental-recursion scan-mode switch. FastSync always performs
* a single full recursive scan, so both spellings are accepted as no-ops:
* the destination is identical whichever mode the caller requests.
* --no-inc-recursive is handled before the generic --no-* negation branch
* (see cli_handle_pre_negation) but is registered here for discoverability. */
{"--inc-recursive", NULL, OPT_NOOP, 0},
{"--no-inc-recursive", NULL, OPT_NOOP, 0},
{"--update", "-u", OPT_FLAG, offsetof(Config, update)},
{"--old-args", NULL, OPT_FLAG, offsetof(Config, old_args)},
/* rsync's --old-args: accepted for CLI compatibility as a documented no-op
* (the remote server path is always safely quoted; see usage.c). It is
* recognized but stores no Config field. */
{"--old-args", NULL, OPT_NOOP, 0},
{"--rsh", "-e", OPT_STRING, offsetof(Config, rsh_command)},
{"--blocking-io", NULL, OPT_FLAG, offsetof(Config, blocking_io)},
{"--links", "-l", OPT_FLAG, offsetof(Config, follow_symlinks)},
@@ -956,6 +1035,12 @@ static const OptionEntry OPTION_TABLE[] = {
{"--delete-during", "--del", OPT_FLAG, offsetof(Config, delete_during)},
{"--delete-delay", NULL, OPT_FLAG, offsetof(Config, delete_delay)},
{"--delete-after", NULL, OPT_FLAG, offsetof(Config, delete_after)},
/* FastSync-only long spelling of the late whole-tree commit, which selects
the same timing as rsync's --delete-after in FastSync (the whole-tree
keep-set manifest is committed only after the entire transfer succeeded).
Plain --delete now defaults to delete-during, so this restores the old
FastSync behavior; it maps onto the same delete_after wire field. */
{"--delete-commit", NULL, OPT_FLAG, offsetof(Config, delete_after)},
{"--delete-excluded", NULL, OPT_FLAG, offsetof(Config, delete_excluded)},
{"--max-delete", NULL, OPT_SIGNED_INT, offsetof(Config, max_delete)},
{"--ignore-errors", NULL, OPT_FLAG, offsetof(Config, ignore_errors)},
@@ -1013,6 +1098,11 @@ static const OptionEntry OPTION_TABLE[] = {
* --remote-option is parsed. --trust-sender is a local receiver policy and
* never travels to the remote peer. */
{"--trust-sender", NULL, OPT_FLAG, offsetof(Config, trust_sender)},
/* FastSync-only (not an rsync option): require a basis-hit's content to
* match the source by whole-file digest instead of trusting rsync's
* size+mtime quick-check. Long-only; crosses the wire so the receiver
* performs the extra read/hash. */
{"--verify-basis", NULL, OPT_FLAG, offsetof(Config, verify_basis)},
};
/* Only boolean options with no required argument are safe to negate. */
@@ -1055,6 +1145,7 @@ static const NegatableOption NEGATABLE_OPTIONS[] = {
{"xattrs", "X", offsetof(Config, preserve_xattrs)},
{"acls", "A", offsetof(Config, preserve_acls)},
{"fake-super", NULL, offsetof(Config, fake_super)},
{"verify-basis", NULL, offsetof(Config, verify_basis)},
};
static bool opt_is(const char* arg, const char* name, const char* alias) {
@@ -1115,12 +1206,12 @@ static int apply_negation(Config* config, const char* arg) {
config->preserve_times = false;
config->preserve_owner = false;
config->preserve_group = false;
config->metadata_explicitly_disabled = true;
config->cli.metadata_explicitly_disabled = true;
/* --no-preserve is an explicit opt-out of the whole bundle: record it so
* the --incremental/--delta auto-preserve in cli_finalize_config does not
* silently re-enable perms/times. */
config->preserve_perms_explicit_off = true;
config->preserve_times_explicit_off = true;
config->cli.preserve_perms_explicit_off = true;
config->cli.preserve_times_explicit_off = true;
return 0;
}
*(bool*)((char*)config + entry->offset) = false;
@@ -1128,9 +1219,9 @@ static int apply_negation(Config* config, const char* arg) {
* auto-preserve the OTHER attribute without undoing this one. A later
* -p/-t sets the attribute directly; this flag only gates the implication. */
if (entry->offset == offsetof(Config, preserve_perms))
config->preserve_perms_explicit_off = true;
config->cli.preserve_perms_explicit_off = true;
else if (entry->offset == offsetof(Config, preserve_times))
config->preserve_times_explicit_off = true;
config->cli.preserve_times_explicit_off = true;
return 0;
}
@@ -1160,7 +1251,13 @@ static int apply_table_option(Config* config, const OptionEntry* entry, const ch
void* field = (char*)config + entry->offset;
switch (entry->kind) {
case OPT_FLAG:
*(bool*)field = true;
/* -x/--one-file-system is repeatable in rsync: `-xx` increments the level so
the scanner drops mount-point directories instead of recreating them
empty. Everything else is a plain boolean. */
if (entry->offset == offsetof(Config, one_file_system))
(*(int*)field)++;
else
*(bool*)field = true;
return 0;
case OPT_NOOP:
return 0;
@@ -1297,6 +1394,12 @@ static bool cli_handle_pre_negation(CliParseCtx* ctx) {
ctx->no_delta = true;
else if (strcmp(arg, "--no-incremental") == 0)
ctx->no_incremental = true;
/* Real rsync option names that merely start with "--no-" and are inert
* no-ops (e.g. --no-inc-recursive) are registered as OPT_NOOP entries;
* accept them before the generic negation table would reject the name. */
const OptionEntry* noop = find_table_option(arg);
if (noop && noop->kind == OPT_NOOP)
return true;
if (apply_negation(config, arg) != 0) {
ctx->exit_code = -1;
return true;
@@ -1353,7 +1456,7 @@ static bool cli_handle_range_time_options(CliParseCtx* ctx) {
ctx->exit_code = -1;
return true;
}
config->stop_at_set = true;
config->cli.stop_at_set = true;
return true;
}
if (strcmp(arg, "--stop-at") == 0) {
@@ -1368,7 +1471,7 @@ static bool cli_handle_range_time_options(CliParseCtx* ctx) {
ctx->exit_code = -1;
return true;
}
config->stop_at_set = true;
config->cli.stop_at_set = true;
return true;
}
const char* threads_prefix = "--compress-threads=";
@@ -1438,6 +1541,8 @@ static bool cli_handle_table_option(CliParseCtx* ctx) {
ctx->exit_code = -1;
return true;
}
if (entry->offset == offsetof(Config, compression_level))
config->cli.compression_level_set = true;
if (entry->offset == offsetof(Config, chmod_spec)) {
mode_t ignored;
if (!chmod_apply(0, config->chmod_spec, &ignored)) {
@@ -1450,7 +1555,7 @@ static bool cli_handle_table_option(CliParseCtx* ctx) {
defaults to 127.0.0.1, so a value check cannot distinguish it). Used
by --dry-run to route an explicit remote target to the server. */
if (entry->offset == offsetof(Config, server_host))
config->server_host_set = true;
config->cli.server_host_set = true;
}
} else if (apply_table_option(config, entry, NULL) != 0) {
ctx->exit_code = -1;
@@ -1464,7 +1569,9 @@ static bool cli_handle_table_option(CliParseCtx* ctx) {
if (entry->offset == offsetof(Config, per_dir_filter) && config->per_dir_filter_count < INT_MAX)
config->per_dir_filter_count++;
/* A delete-timing flag selects when --delete removes extras, so it
implies --delete exactly like the rsync options do. */
implies --delete exactly like the rsync options do. --delete-commit (the
FastSync-only late-commit spelling) is mapped onto delete_after and so is
covered here too. */
if (entry->offset == offsetof(Config, delete_before) ||
entry->offset == offsetof(Config, delete_during) ||
entry->offset == offsetof(Config, delete_delay) ||
@@ -1722,6 +1829,7 @@ static bool cli_handle_transfer_flags(CliParseCtx* ctx) {
return true;
}
config->compression_level = (int)level;
config->cli.compression_level_set = true;
log_info_message(LOG_INFO_MISC, "Set Compression level to %ld", level);
ctx->i++;
}
@@ -1785,7 +1893,7 @@ static int set_server_port_option(Config* config, const char* value, const char*
return -1;
}
config->server_port = port;
config->server_port_set = true;
config->cli.server_port_set = true;
return 0;
}
@@ -1819,22 +1927,130 @@ static int set_log_file_option(Config* config, const char* log_path) {
return 0;
}
/* Apply a --bwlimit value (kilobytes per second). Returns 0 on success, -1 on
* error. */
/* Faithful port of rsync 3.4.1's `parse_size_arg(bwlimit_arg, 'K', "bwlimit",
* 512, -1, True)`: a default KiB suffix, binary (1024) multipliers unless a
* `b`/`B` decimal suffix or explicit `iB` is given, an optional decimal
* fraction, the P/T/G/M/K suffixes, and the special rules that a value of 0
* means "no limit" while any other value below 512 bytes is rejected. The
* parsed byte count is then quantized to whole KiB exactly like rsync's
* `bwlimit = (size + 512) / 1024`. Returns 0 on success, -1 on a parse error. */
static int parse_bwlimit_value(const char* value, unsigned long long* bytes_per_sec_out) {
if (!value || !bytes_per_sec_out)
return -1;
const char* arg = value;
int reps;
long long mult;
while (*arg >= '0' && *arg <= '9')
arg++;
if (*arg != '\0' && (*arg == '.' || *arg == localeconv()->decimal_point[0]))
for (arg++; *arg >= '0' && *arg <= '9'; arg++) {
}
char suffix = *arg && *arg != '+' && *arg != '-' ? *arg++ : 'K';
switch (suffix) {
case 'b':
case 'B':
reps = 0;
break;
case 'k':
case 'K':
reps = 1;
break;
case 'm':
case 'M':
reps = 2;
break;
case 'g':
case 'G':
reps = 3;
break;
case 't':
case 'T':
reps = 4;
break;
case 'p':
case 'P':
reps = 5;
break;
default:
log_message(LOG_LEVEL_ERROR, "--bwlimit=%s is invalid", value);
return -1;
}
if (*arg == 'b' || *arg == 'B') {
mult = 1000;
arg++;
} else if (*arg == '\0' || *arg == '+' || *arg == '-') {
mult = 1024;
} else if ((arg[0] == 'i' || arg[0] == 'I') && (arg[1] == 'b' || arg[1] == 'B')) {
mult = 1024;
arg += 2;
} else {
log_message(LOG_LEVEL_ERROR, "--bwlimit=%s is invalid", value);
return -1;
}
long long base = 1;
for (int i = 0; i < reps; i++) {
if (base > LLONG_MAX / mult) {
log_message(LOG_LEVEL_ERROR, "--bwlimit=%s is too large", value);
return -1;
}
base *= mult;
}
/* rsync multiplies the numeric prefix (atof) by mult^reps in a signed
* ssize_t, which is undefined on overflow. Scale in double and range-check
* before converting, so a huge value is rejected as "too large" (where
* rsync's overflow happens to land on a negative result) without invoking
* signed-overflow UB. */
double scaled = (double)base * strtod(value, NULL);
/* (double)LLONG_MAX rounds up to 2^63, which is itself out of range for the
* cast, so reject at >= that bound; LLONG_MIN == -2^63 is exactly
* representable and thus castable, so the lower bound stays strict. */
if (!isfinite(scaled) || scaled >= (double)LLONG_MAX || scaled < (double)LLONG_MIN) {
log_message(LOG_LEVEL_ERROR, "--bwlimit=%s is too large", value);
return -1;
}
long long size = (long long)scaled;
if ((*arg == '+' || *arg == '-') && arg[1] == '1' && arg != value) {
/* The only form accepted here is "+1"/"-1" (a longer number leaves a
trailing byte and is rejected below), so apply the delta directly and
guard the one overflow direction. */
if (*arg == '+') {
if (size == LLONG_MAX) {
log_message(LOG_LEVEL_ERROR, "--bwlimit=%s is too large", value);
return -1;
}
size += 1;
} else {
size -= 1;
}
arg += 2;
}
if (*arg != '\0' || size < 0) {
log_message(LOG_LEVEL_ERROR, "--bwlimit=%s is %s", value, size < 0 ? "too large" : "invalid");
return -1;
}
if (size != 0 && size < 512) {
log_message(LOG_LEVEL_ERROR, "--bwlimit=%s is too small (min: 512 or 0 for unlimited)", value);
return -1;
}
long long kib = size == 0 ? 0 : (size + 512) / 1024;
if (kib > (long long)(ULLONG_MAX / 1024)) {
log_message(LOG_LEVEL_ERROR, "--bwlimit=%s is too large", value);
return -1;
}
*bytes_per_sec_out = (unsigned long long)kib * 1024;
return 0;
}
/* Apply a --bwlimit value using rsync 3.4.1's units/semantics. Returns 0 on
* success, -1 on error. */
static int set_bwlimit_option(const char* value) {
unsigned long long kbps;
if (parse_ull_arg(value, &kbps, "--bwlimit") != 0)
unsigned long long bytes_per_sec;
if (parse_bwlimit_value(value, &bytes_per_sec) != 0)
return -1;
if (kbps == 0) {
log_message(LOG_LEVEL_ERROR, "--bwlimit must be a positive integer");
return -1;
}
if (kbps > ULLONG_MAX / 1024) {
log_message(LOG_LEVEL_ERROR, "--bwlimit value too large");
return -1;
}
io_set_bwlimit(kbps * 1024);
log_info_message(LOG_INFO_MISC, "Set bandwidth limit to %llu KB/s", kbps);
io_set_bwlimit(bytes_per_sec);
log_info_message(LOG_INFO_MISC, "Set bandwidth limit to %llu KB/s", bytes_per_sec / 1024);
return 0;
}
@@ -2408,7 +2624,32 @@ static bool cli_handle_outbuf_option(CliParseCtx* ctx) {
* load --files-from once every argument has been seen. Returns 0 on success,
* -1 on error. */
static int cli_finalize_config(Config* config, bool verbose, bool no_delta, bool no_incremental) {
set_log_level(config->quiet ? LOG_LEVEL_ERROR : (verbose ? LOG_LEVEL_DEBUG : LOG_LEVEL_WARNING));
/* An explicit --debug=FLAGS enables the debug log level by itself (rsync
behaviour); -v enables every other INFO-level message. */
bool debug_enabled = verbose || config->debug_level != 0;
set_log_level(config->quiet ? LOG_LEVEL_ERROR
: (debug_enabled ? LOG_LEVEL_DEBUG : LOG_LEVEL_WARNING));
/* rsync's plain --delete defaults to delete-during (--del): each directory's
extras are removed as that directory is processed, so space is freed
progressively and a tight destination never has to hold the whole old+new
tree at once. The late whole-tree commit FastSync historically used is
still selected explicitly by --delete-after or by the FastSync-only long
spelling --delete-commit (an exact alias for --delete-after). Resolve the
default on the client, before validation and before the config crosses the
wire, so exactly one timing flag is ever set; an explicit timing (including
--delete-commit) always wins. */
if (config->use_delete && !config->delete_before && !config->delete_during &&
!config->delete_delay && !config->delete_after)
config->delete_during = true;
/* rsync parity: --partial-dir=DIR chooses where an interrupted transfer's
partial file is kept, so it implies --partial. rsync applies the
implication after option parsing, so it wins over an explicit --no-partial
regardless of the order the two options appear in (verified on rsync
3.4.1). --inplace is the exception: the destination file is written in
place with no partial/temp staging, so the partial machinery is bypassed
and the implication is skipped to leave --inplace behavior untouched. */
if (config->partial_dir && !config->inplace)
config->partial = true;
if (config->compress_choice) {
int algo = compression_algo_from_name(config->compress_choice);
if (algo >= 0) {
@@ -2416,14 +2657,49 @@ static int cli_finalize_config(Config* config, bool verbose, bool no_delta, bool
config->use_compression = (algo != (int)COMPRESSION_ALGO_NONE);
}
}
if (config->use_compression && config->compression_algo == (int)COMPRESSION_ALGO_NONE)
config->compression_algo = (int)compression_negotiate_default();
/* A bare -z (no --compress-choice) resolves like rsync's "auto": the
* RSYNC_COMPRESS_LIST preference list first, then the compiled-in order. A
* list that names no supported codec is rsync's failed negotiation (exit 4). */
if (config->use_compression && !config->compress_choice) {
int resolved = compression_choice_resolve();
if (resolved < 0) {
log_message(LOG_LEVEL_ERROR, "RSYNC_COMPRESS_LIST names no supported compression algorithm");
config->cli.cli_exit_code = 4;
return -1;
}
config->compression_algo = resolved;
if (resolved == (int)COMPRESSION_ALGO_NONE)
config->use_compression = false;
}
/* Apply rsync's per-codec compression level: an explicit --compress-level is
* clamped to the codec's range, otherwise the codec's own default is used. */
if (config->use_compression) {
CompressionAlgo algo = (CompressionAlgo)config->compression_algo;
config->compression_level = config->cli.compression_level_set
? compression_clamp_level(algo, config->compression_level)
: compression_default_level(algo);
log_debug_message(LOG_DEBUG_UTIL, "Client compression: %s (level %d)",
compression_algo_name(algo), config->compression_level);
}
/* The negotiated checksum is always resolved (rsync negotiates one for the
* delta strong sum even without --checksum): RSYNC_CHECKSUM_LIST first, then
* the compiled-in order. An explicit --checksum-choice already set it. */
if (!config->cli.checksum_choice_set) {
int resolved = checksum_choice_resolve();
if (resolved < 0) {
log_message(LOG_LEVEL_ERROR, "RSYNC_CHECKSUM_LIST names no supported checksum algorithm");
config->cli.cli_exit_code = 4;
return -1;
}
config->checksum_algo = resolved;
config->cli.checksum_transfer_algo = resolved;
}
/* rsync parity: "none" as the pre-transfer checksum cannot be combined with
* --checksum (exit 4). The check runs here because --checksum may appear on
* either side of --checksum-choice. */
if (config->checksum && config->checksum_algo == (int)CHECKSUM_ALGO_NONE) {
log_message(LOG_LEVEL_ERROR, "Invalid checksum-choice for --checksum: none");
config->cli_exit_code = 4;
config->cli.cli_exit_code = 4;
return -1;
}
@@ -2502,11 +2778,11 @@ static int cli_finalize_config(Config* config, bool verbose, bool no_delta, bool
* explicitly negated them (--no-perms/--no-times/--no-preserve). This runs
* BEFORE the derived use_metadata bit so the transport frame is still sent
* for the incremental/delta handshake even when both attributes were negated
* via --no-preserve (metadata_explicitly_disabled handles that opt-out). */
if (preserve_implied && !config->metadata_explicitly_disabled) {
if (!config->preserve_perms_explicit_off)
* via --no-preserve (cli.metadata_explicitly_disabled handles that opt-out). */
if (preserve_implied && !config->cli.metadata_explicitly_disabled) {
if (!config->cli.preserve_perms_explicit_off)
config->preserve_perms = true;
if (!config->preserve_times_explicit_off)
if (!config->cli.preserve_times_explicit_off)
config->preserve_times = true;
}
@@ -2549,8 +2825,17 @@ static int cli_finalize_config(Config* config, bool verbose, bool no_delta, bool
}
}
}
config->report_stats = config->stats || config->show_progress || format_needs_wire ||
(config->dry_run && config->use_delete);
/* --info=del on a real --delete run asks the receiver to report the paths it
actually removed; the report rides the STATUS_STATS path list, so the wire
stats frame must be negotiated too. --debug=del needs the same paths, so
it opts into the existing report (no new wire field). */
config->report_deletes =
config->use_delete && !config->dry_run &&
((config->info_level & LOG_INFO_DEL) != 0 || config->itemize_changes ||
config->out_format != NULL || (config->debug_level & LOG_DEBUG_DEL) != 0);
config->report_stats = config->stats || config->show_progress ||
(config->info_level & LOG_INFO_PROGRESS) || format_needs_wire ||
config->report_deletes || (config->dry_run && config->use_delete);
return 0;
}
@@ -2884,7 +3169,7 @@ int main(int argc, char* argv[]) {
int parse_ret = parse_args(config, argc, argv, positional_args, &positional_count);
if (parse_ret != 0) {
if (parse_ret < 0)
exit_code = config->cli_exit_code ? config->cli_exit_code : 1;
exit_code = config->cli.cli_exit_code ? config->cli.cli_exit_code : 1;
goto cleanup;
}
@@ -3027,7 +3312,7 @@ int main(int argc, char* argv[]) {
exit_code = 1;
}
} else if (config->use_multithreading) {
exit_code = send_files_multithreaded(&config);
exit_code = send_files_multithreaded(config);
} else {
exit_code = send_files(config);
}
+682
View File
@@ -0,0 +1,682 @@
#include "client_send_internal.h"
#include "array_list.h"
#include "change_list.h"
#include "charset.h"
#include "config.h"
#include "data.h"
#include "delta.h"
#include "file.h"
#include "format.h"
#include "log.h"
#include "protocol.h"
#include "scanner.h"
#include "transport_tls.h"
#include "utils.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/stat.h>
#include <time.h>
/* True when --dry-run should contact a receiver rather than running the
* client-side local manifest. Any target a real run would reach over the wire
* selects the server-contacting path: a remote (SSH host:path), a daemon
* (host::module/path), an explicit --server-host, --server-port/--port, TLS, or
* a source-bind --address. A plain local destination (none of these) keeps the
* original client-side behavior, which never dials the default 127.0.0.1:8080. */
bool dry_run_targets_server(const Config* config) {
if (!config)
return false;
if (config->transport == TRANSPORT_SSH)
return true;
if (config->module && config->module[0] != '\0')
return true;
if (config->cli.server_host_set || config->cli.server_port_set)
return true;
if (config->use_tls)
return true;
if (config->address != NULL)
return true;
return false;
}
bool add_chunk_to_manifest(ArrayList* manifest, const Chunk* chunk) {
if (!manifest)
return true;
for (int i = 0; i < chunk->element_count; i++) {
const char* path = file_wire_path(chunk->items[i]);
if (*path == '/')
path++;
char* entry = str_dup(path);
if (!entry) {
log_message(LOG_LEVEL_ERROR, "Failed to allocate manifest entry");
return false;
}
if (!array_list_add(manifest, entry)) {
free(entry);
return false;
}
}
return true;
}
/* Print dry-run manifest showing files that would be transferred. Returns 0 on success. */
int send_dry_run_manifest(const Config* config) {
int skipped = 0;
ArrayList* missing_dest = NULL;
if (config->delete_missing_args) {
missing_dest = array_list_create(free);
if (!missing_dest)
return -1;
}
if (!files_from_list_check(config, missing_dest, &skipped)) {
if (missing_dest)
array_list_delete(missing_dest);
return -1;
}
PreparedScanner prepared;
if (!prepare_scanner(config, 0, &prepared)) {
if (missing_dest)
array_list_delete(missing_dest);
return -1;
}
DirectoryScanner* scanner =
directory_scanner_create_with_options(config->send_directory, &prepared.options);
if (!scanner) {
prepared_scanner_destroy(&prepared);
if (missing_dest)
array_list_delete(missing_dest);
return -1;
}
Chunk* chunk;
int file_count = 0;
unsigned long long total_bytes = 0;
char size_buffer[32];
if (!config->quiet)
printf("Dry run: files to be transferred\n");
while ((chunk = directory_scanner_next(scanner)) != NULL) {
for (int i = 0; i < chunk->element_count; i++) {
if (!config->quiet) {
char* escaped_path =
output_escape(file_wire_path(chunk->items[i]), config->eight_bit_output);
if (!escaped_path) {
chunk_destroy(chunk);
directory_scanner_destroy(scanner);
prepared_scanner_destroy(&prepared);
if (missing_dest)
array_list_delete(missing_dest);
return -1;
}
if (config->human_readable)
printf(
" %s (%s)\n", escaped_path,
display_bytes(chunk->items[i]->data->size, true, size_buffer, sizeof(size_buffer)));
else
printf(" %s (%zu bytes)\n", escaped_path, chunk->items[i]->data->size);
free(escaped_path);
}
total_bytes += chunk->items[i]->data->size;
file_count++;
}
chunk_destroy(chunk);
}
directory_scanner_destroy(scanner);
prepared_scanner_destroy(&prepared);
/* --delete-missing-args: the missing entries' destination mirrors render as
would-be deletions (rsync's dry-run also lists its *deleting lines). */
if (missing_dest && !config->quiet) {
for (int i = 0; i < missing_dest->size; i++) {
char* escaped = output_escape((char*)missing_dest->items[i], config->eight_bit_output);
printf(" %s (missing; would be deleted)\n", escaped ? escaped : "<allocation failed>");
free(escaped);
}
}
if (missing_dest)
array_list_delete(missing_dest);
if (!config->quiet) {
if (config->human_readable)
printf("Total: %d files, %s\n", file_count,
display_bytes(total_bytes, true, size_buffer, sizeof(size_buffer)));
else
printf("Total: %d files, %.1f MB\n", file_count, (double)total_bytes / (double)BYTES_PER_MIB);
}
return 0;
}
typedef struct {
char* name; /* transfer-relative name ("" == the source root) */
mode_t mode;
unsigned long long size;
time_t mtime;
long mtime_nsec;
bool is_dir;
bool is_symlink;
char* link_target;
} ListEntry;
static void list_entries_destroy(ListEntry* entries, size_t count) {
if (entries == NULL)
return;
for (size_t i = 0; i < count; i++) {
free(entries[i].name);
free(entries[i].link_target);
}
free(entries);
}
static int compare_list_entries(const void* left, const void* right) {
const ListEntry* a = (const ListEntry*)left;
const ListEntry* b = (const ListEntry*)right;
return strcmp(a->name, b->name);
}
/* Relative path of an entry below `root` ("" for the root itself). Mirrors
* change_list's relative_name for list-only rendering. */
static char* list_relative_name(const char* root, const char* full) {
if (root == NULL || full == NULL)
return str_dup(full != NULL ? full : "");
size_t root_len = strlen(root);
while (root_len > 1 && root[root_len - 1] == '/')
root_len--;
if (strncmp(root, full, root_len) == 0) {
if (full[root_len] == '\0')
return str_dup("");
if (full[root_len] == '/')
return str_dup(full + root_len + 1);
}
return str_dup(full);
}
/* --list-only: print an ls-style listing of the entries that WOULD be
* transferred and exit without contacting the server or writing anything.
* Names are transfer-relative (rsync prints `a.txt`, `sub/b.txt`, `.`) and
* directory entries are included. Returns 0 on success, 1 on error. */
int send_list_only(const Config* config) {
int skipped = 0;
if (!files_from_list_check(config, NULL, &skipped))
return 1;
PreparedScanner prepared;
if (!prepare_scanner(config, 0, &prepared))
return 1;
prepared.options.use_metadata = true; /* capture mode + mtime for the listing */
prepared.options.list_dirs = true;
DirectoryScanner* scanner =
directory_scanner_create_with_options(config->send_directory, &prepared.options);
if (!scanner) {
prepared_scanner_destroy(&prepared);
return 1;
}
ListEntry* entries = NULL;
size_t count = 0;
size_t capacity = 0;
bool oom = false;
/* rsync lists the source root itself (as "."). Only when the source is a
* directory and no --files-from subset is in effect. */
if (config->files_from_set == NULL && config->send_directory != NULL) {
struct stat st;
if (stat(config->send_directory, &st) == 0 && S_ISDIR(st.st_mode)) {
capacity = 64;
entries = calloc(capacity, sizeof(ListEntry));
if (entries == NULL) {
oom = true;
} else if ((entries[0].name = str_dup("")) == NULL) {
/* A NULL name would be dereferenced by qsort/render: fail the listing. */
oom = true;
} else {
entries[0].mode = st.st_mode;
entries[0].mtime = st.st_mtime;
entries[0].mtime_nsec = st.st_mtim.tv_nsec;
entries[0].size = (unsigned long long)st.st_size;
entries[0].is_dir = true;
count = 1;
}
}
}
Chunk* chunk;
while (!oom && (chunk = directory_scanner_next(scanner)) != NULL) {
for (int i = 0; i < chunk->element_count; i++) {
File* f = chunk->items[i];
if (f == NULL)
continue;
if (count == capacity) {
size_t new_capacity = capacity > 0 ? capacity * 2 : 64;
if (new_capacity <= capacity) {
oom = true;
break;
}
ListEntry* grown = realloc(entries, new_capacity * sizeof(ListEntry));
if (!grown) {
oom = true;
break;
}
entries = grown;
memset(entries + capacity, 0, (new_capacity - capacity) * sizeof(ListEntry));
capacity = new_capacity;
}
char* name = list_relative_name(config->send_directory, file_wire_path(f));
if (!name) {
oom = true;
break;
}
mode_t mode = 0;
time_t mtime = 0;
long mtime_nsec = 0;
if (f->metadata != NULL) {
mode = f->metadata->mode;
mtime = f->metadata->mtime_sec;
mtime_nsec = f->metadata->mtime_nsec;
} else {
struct stat st;
if (lstat(f->path, &st) == 0) {
mode = st.st_mode;
mtime = st.st_mtime;
mtime_nsec = st.st_mtim.tv_nsec;
}
}
entries[count].name = name;
entries[count].mode = mode;
entries[count].mtime = mtime;
entries[count].mtime_nsec = mtime_nsec;
if (f->is_symlink)
entries[count].size = f->symlink_target != NULL ? strlen(f->symlink_target) : 0;
else if (f->is_dir) {
struct stat dir_st;
entries[count].size = stat(f->path, &dir_st) == 0 ? (unsigned long long)dir_st.st_size : 0;
} else
entries[count].size = f->data != NULL ? f->data->size : 0;
entries[count].is_dir = f->is_dir;
entries[count].is_symlink = f->is_symlink;
entries[count].link_target =
f->is_symlink && f->symlink_target ? str_dup(f->symlink_target) : NULL;
count++;
}
chunk_destroy(chunk);
}
bool failed = oom || directory_scanner_failed(scanner) || directory_scanner_had_io_error(scanner);
directory_scanner_destroy(scanner);
prepared_scanner_destroy(&prepared);
if (failed) {
list_entries_destroy(entries, count);
if (oom)
log_message(LOG_LEVEL_ERROR, "memory allocation failed while listing");
return 1;
}
if (count > 1)
qsort(entries, count, sizeof(ListEntry), compare_list_entries);
for (size_t i = 0; i < count; i++) {
ChangeEvent event;
memset(&event, 0, sizeof(event));
event.name = entries[i].name;
event.path = entries[i].name;
event.mode = entries[i].mode;
event.size = entries[i].size;
event.mtime_sec = entries[i].mtime;
event.mtime_nsec = entries[i].mtime_nsec;
event.is_directory = entries[i].is_dir;
event.is_symlink = entries[i].is_symlink;
event.symlink_target = entries[i].link_target;
char* line = change_render_list_line(config, &event);
if (line != NULL) {
char* escaped = output_escape(line, config->eight_bit_output);
printf("%s\n", escaped != NULL ? escaped : line);
free(escaped);
free(line);
}
}
list_entries_destroy(entries, count);
return 0;
}
/* Send the delete manifest to the server. Returns 0 on success, -1 on
failure. It carries FOUR sections: the keep-set paths, the protected
excluded prefixes, the --delete-missing-args exact-delete paths, and the
destination-relative directories the sender synchronized this run.
When --delete-excluded is given `protected` is empty: excluded destination
mirrors are then ordinary extras and are removed. When
--delete-missing-args is active `missing_args` holds the destination mirrors
of missing --files-from entries: each is an explicit receiver-side deletion
request, independent of the extras walk. `synced_dirs` confines the extras
walk to entries directly inside a synchronized directory. A NULL
keep-set / protected / missing / dirs list transmits an empty section. All
four sections are unbounded on the sender; the receiver enforces
MAX_MANIFEST_ENTRIES per section and a single MAX_MANIFEST_BYTES budget
shared across the sections, rejecting (with STATUS_ERROR) an over-budget
frame. A heavily filtered source whose exclusion list is large therefore
fails the run cleanly on the receiver rather than being truncated. */
int send_delete_manifest(int fd, ArrayList* manifest, ArrayList* protected_prefixes,
ArrayList* size_skipped, ArrayList* missing_args, ArrayList* synced_dirs) {
if (!send_status(fd, STATUS_MANIFEST))
return -1;
int keep_count = manifest ? manifest->size : 0;
if (!send_int(fd, keep_count))
return -1;
for (int i = 0; i < keep_count; i++) {
if (!send_wire_str(fd, (char*)manifest->items[i]))
return -1;
}
/* The receiver has ONE protected-prefix section; filter-excluded prefixes
(dropped under --delete-excluded) and size-pruned prefixes (always
protected) are concatenated into it. */
int protected_count =
(protected_prefixes ? protected_prefixes->size : 0) + (size_skipped ? size_skipped->size : 0);
if (!send_int(fd, protected_count))
return -1;
if (protected_prefixes) {
for (int i = 0; i < protected_prefixes->size; i++) {
if (!send_wire_str(fd, (char*)protected_prefixes->items[i]))
return -1;
}
}
if (size_skipped) {
for (int i = 0; i < size_skipped->size; i++) {
if (!send_wire_str(fd, (char*)size_skipped->items[i]))
return -1;
}
}
int missing_count = missing_args ? missing_args->size : 0;
if (!send_int(fd, missing_count))
return -1;
for (int i = 0; i < missing_count; i++) {
if (!send_wire_str(fd, (char*)missing_args->items[i]))
return -1;
}
int dirs_count = synced_dirs ? synced_dirs->size : 0;
if (!send_int(fd, dirs_count))
return -1;
for (int i = 0; i < dirs_count; i++) {
if (!send_wire_str(fd, (char*)synced_dirs->items[i]))
return -1;
}
return 0;
}
/* Transmit the keep-set manifest and wait for the receiver's verdict. Used by
--delete-before/--delete-during, where the extras are removed on the receiver
BEFORE the first byte of file data is sent: the receiver acknowledges with
STATUS_OK once the bounded delete committed, or STATUS_ERROR if it could not
(in which case the sender aborts without streaming any data). The ACK may
take much longer than an ordinary per-message round trip because the receiver
performs the whole bounded deletion walk (up to MAX_SERVER_DELETE_COUNT
unlinks) before replying, so the wait uses a generous explicit deadline
instead of the default 60 s receive window. */
#define DELETE_ACK_TIMEOUT_SEC 3600
/* While waiting for the (potentially slow) receiver-side deletion, send a
* STATUS_KEEPALIVE at most this often so the connection is demonstrably alive
* and neither side's per-message timeout trips. */
#define DELETE_ACK_KEEPALIVE_SEC 10
bool send_delete_manifest_early(Client* client, ArrayList* manifest, ArrayList* protected_prefixes,
ArrayList* size_skipped, ArrayList* missing_args,
ArrayList* synced_dirs) {
if (!client || !manifest)
return false;
if (send_delete_manifest(client->file_descriptor, manifest, protected_prefixes, size_skipped,
missing_args, synced_dirs) != 0)
return false;
Status ack;
/* The wait is long (up to an hour) and runs inline on this thread: a helper
* thread would race the non-thread-safe protocol send path, so keepalives are
* emitted from this wait loop itself. A Ctrl-C/SIGTERM abort flag also ends
* the wait; the caller then best-effort sends STATUS_ABORT. */
if (!receive_status_keepalive(client->file_descriptor, &ack, DELETE_ACK_TIMEOUT_SEC,
DELETE_ACK_KEEPALIVE_SEC, client_abort_pending)) {
/* A Ctrl-C/SIGTERM abort ends the wait above; tell the receiver before the
caller tears the connection down (best-effort). */
if (client_abort_pending()) {
log_info_message(LOG_INFO_MISC,
"Abort requested while awaiting delete ack; sending STATUS_ABORT");
send_status(client->file_descriptor, STATUS_ABORT);
}
return false;
}
if (ack != STATUS_OK) {
log_server_rejection("Server failed to delete files before the transfer");
return false;
}
return true;
}
/* Server-contacting --dry-run. Connects to the configured remote/daemon and
* runs the normal per-file incremental decision WITHOUT transmitting any file
* data: the receiver (which also sees dry_run=true on the wire) answers
* STATUS_OK for an up-to-date file and STATUS_DRY_RUN_TRANSFER for a file it
* would otherwise write, mutating nothing on either side. The would-transfer
* set and the same trailer as the local dry-run are printed. A
* --compare-dest exact basis hit with no destination copy is reported as a
* skip by the receiver.
*
* Only regular files take the receiver-consulted check; directory / symlink /
* special / hard-link-sibling entries have no per-file content check, so they
* are reported conservatively as would-transfer and their frames are never
* sent (which is what keeps the receiver mutation-free). --delete* is
* deliberately NOT transmitted in dry-run, so no deletion can occur; the
* would-delete manifest report is a documented follow-up.
*
* Returns 0 on success, 1 on error. */
int send_dry_run_remote(Config* config) {
int from_skipped = 0;
ArrayList* missing_args = NULL;
if (config->delete_missing_args) {
missing_args = array_list_create(free);
if (!missing_args)
return 1;
}
if (!files_from_list_check(config, missing_args, &from_skipped)) {
if (missing_args)
array_list_delete(missing_args);
return 1;
}
if (missing_args)
array_list_delete(missing_args);
/* A live session may follow, so arm graceful abort handling. */
client_set_abort_armed(true);
Client* client = connect_transfer_client(config);
if (!client) {
if (config->transport == TRANSPORT_TCP)
log_message(LOG_LEVEL_ERROR, "could not connect to server%s",
config->use_tls ? " via TLS" : "");
client_set_abort_armed(false);
return 1;
}
ProtocolSession session;
protocol_session_init(&session, client->file_descriptor, client->file_descriptor);
protocol_session_set_io_timeout(&session, config->timeout);
protocol_session_set_ssl(&session, (SSL*)client->ssl);
protocol_session_bind(&session);
int ret = 1;
time_t dry_start = time(NULL);
ReceiverStats dry_stats;
memset(&dry_stats, 0, sizeof(dry_stats));
PreparedScanner prepared;
memset(&prepared, 0, sizeof(prepared));
DirectoryScanner* scanner = NULL;
ArrayList* dry_manifest = NULL;
ArrayList* dry_dirs = NULL;
ArrayList* dry_excluded = NULL;
ArrayList* dry_size_skipped = NULL;
if (!config_send(client->file_descriptor, config))
goto dry_fail;
receive_daemon_motd(client, config);
if (!prepare_scanner(config, 0, &prepared))
goto dry_fail;
/* -n --delete: build the same keep-set manifest, protected prefixes, and
synchronized-directory scope a real run would send, so the receiver's
read-only extras walk enumerates exactly the deletions a real run makes. */
if (config->use_delete) {
dry_manifest = array_list_create(free);
dry_dirs = array_list_create(free);
dry_size_skipped = array_list_create(free);
if (!dry_manifest || !dry_dirs || !dry_size_skipped)
goto dry_fail;
if (!config->delete_excluded) {
dry_excluded = array_list_create(free);
if (!dry_excluded)
goto dry_fail;
prepared.options.excluded_paths = dry_excluded;
}
prepared.options.size_skipped_paths = dry_size_skipped;
/* A --files-from subset confines the extras walk to the directories the
scan synchronized; a full recursive transfer marks the root itself. */
if (config->files_from_set == NULL) {
char* root_marker = delete_scope_root_marker(config);
if (!root_marker || !array_list_add(dry_dirs, root_marker)) {
free(root_marker);
goto dry_fail;
}
} else {
prepared.options.synced_dirs = dry_dirs;
}
}
scanner = directory_scanner_create_with_options(config->send_directory, &prepared.options);
if (!scanner)
goto dry_fail;
int file_count = 0;
unsigned long long total_bytes = 0;
char size_buffer[32];
if (!config->quiet)
printf("Dry run: files to be transferred\n");
Chunk* chunk;
while ((chunk = directory_scanner_next(scanner)) != NULL) {
if (dry_manifest && !add_chunk_to_manifest(dry_manifest, chunk)) {
chunk_destroy(chunk);
goto dry_fail;
}
for (int i = 0; i < chunk->element_count; i++) {
File* f = chunk->items[i];
if (!f)
continue;
unsigned long long fsize = f->data ? f->data->size : 0;
bool would;
if (f->is_dir || f->is_symlink || f->is_special ||
(f->link_group != 0 && !f->link_first && f->hardlink_target != NULL)) {
/* No receiver-side content check exists for these frame types; a real
run would (re)create them, so report would-transfer and send no
frame (the receiver must stay mutation-free). */
would = true;
} else if (fsize > MAX_RECEIVE_WHOLE_FILE_SIZE && !config->use_incremental &&
!config_has_basis(config)) {
/* A non-incremental run streams a >whole-file-limit source without the
STATUS_CHECK handshake, so no read-only receiver decision is possible
(and none is needed: a real run would transfer it). */
would = true;
} else {
DeltaSignature* sig = NULL;
unsigned long long resume_offset = 0;
int rc = incremental_check(client, f, config, &sig, &resume_offset);
delta_signature_destroy(sig);
if (rc < 0) {
chunk_destroy(chunk);
goto dry_fail;
}
if (rc == 1)
continue; /* up to date; nothing to report */
if (rc != 4) {
log_message(LOG_LEVEL_ERROR, "Unexpected receiver reply during dry-run");
chunk_destroy(chunk);
goto dry_fail;
}
would = true;
}
if (would) {
if (!config->quiet) {
char* escaped_path = output_escape(file_wire_path(f), config->eight_bit_output);
if (!escaped_path) {
chunk_destroy(chunk);
goto dry_fail;
}
if (config->human_readable)
printf(" %s (%s)\n", escaped_path,
display_bytes(fsize, true, size_buffer, sizeof(size_buffer)));
else
printf(" %s (%llu bytes)\n", escaped_path, fsize);
free(escaped_path);
}
total_bytes += fsize;
file_count++;
}
}
chunk_destroy(chunk);
}
bool io_error = directory_scanner_had_io_error(scanner);
if (directory_scanner_failed(scanner))
goto dry_fail;
if (io_error)
log_message(LOG_LEVEL_WARNING, "source scan hit an unreadable directory");
/* Send the keep-set manifest (no data frames) so the receiver can enumerate
the destination extras; an early-timing delete ACKs before it will accept
the terminal FINISHED. */
bool early_delete = config->use_delete && config_delete_timing_early(config);
if (dry_manifest) {
if (send_delete_manifest(client->file_descriptor, dry_manifest, dry_excluded, dry_size_skipped,
NULL, dry_dirs) != 0)
goto dry_fail;
if (early_delete) {
Status ack;
if (!receive_status_keepalive(client->file_descriptor, &ack, DELETE_ACK_TIMEOUT_SEC,
DELETE_ACK_KEEPALIVE_SEC, client_abort_pending) ||
ack != STATUS_OK)
goto dry_fail;
}
}
/* Terminate the stream so the receiver emits its success frame; no data frame
is ever sent in dry-run. */
if (!send_status(client->file_descriptor, STATUS_FINISHED))
goto dry_fail;
Status status;
if (!receive_status(client->file_descriptor, &status))
goto dry_fail;
if (status == STATUS_STATS) {
ArrayList* would_delete = array_list_create(free);
if (!would_delete)
goto dry_fail;
if (!receive_stats_record(client->file_descriptor, &dry_stats, would_delete)) {
array_list_delete(would_delete);
goto dry_fail;
}
print_delete_reports(config, would_delete);
array_list_delete(would_delete);
if (!receive_status(client->file_descriptor, &status))
goto dry_fail;
}
if (status != STATUS_OK)
goto dry_fail;
if (!config->quiet) {
if (config->human_readable)
printf("Total: %d files, %s\n", file_count,
display_bytes(total_bytes, true, size_buffer, sizeof(size_buffer)));
else
printf("Total: %d files, %.1f MB\n", file_count, (double)total_bytes / (double)BYTES_PER_MIB);
}
{
TransferStats dry_transfer;
memset(&dry_transfer, 0, sizeof(dry_transfer));
dry_transfer.flist_reg = (unsigned long long)file_count;
dry_transfer.total_file_size = total_bytes;
dry_transfer.transferred_regular = (unsigned long long)file_count;
dry_transfer.transferred_file_size = total_bytes;
dry_transfer.literal_data = total_bytes;
report_transfer_stats(config, &dry_transfer, dry_start, &dry_stats);
}
ret = io_error ? 1 : 0;
dry_fail:
if (dry_manifest)
array_list_delete(dry_manifest);
if (dry_dirs)
array_list_delete(dry_dirs);
if (dry_excluded)
array_list_delete(dry_excluded);
if (dry_size_skipped)
array_list_delete(dry_size_skipped);
if (scanner)
directory_scanner_destroy(scanner);
prepared_scanner_destroy(&prepared);
disconnect_transfer_client(client);
protocol_session_unbind();
client_set_abort_armed(false);
return ret;
}
File diff suppressed because it is too large Load Diff
+470
View File
@@ -0,0 +1,470 @@
#include "client_send_internal.h"
#include "array_list.h"
#include "charset.h"
#include "config.h"
#include "delete_plan.h"
#include "file.h"
#include "file_list.h"
#include "filter.h"
#include "hardlink.h"
#include "log.h"
#include "scanner.h"
#include "utils.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/stat.h>
/* Build the scanner options for one scan. Returns false and logs on failure. */
bool prepare_scanner(const Config* config, int num_threads, PreparedScanner* out) {
if (!out)
return false;
out->base_filters = NULL;
out->hardlinks = NULL;
out->relative_prefix = NULL;
memset(&out->options, 0, sizeof(out->options));
int rule_count = config->filters ? config->filters->size : 0;
const char** texts = NULL;
if (rule_count > 0) {
texts = malloc((size_t)rule_count * sizeof(char*));
if (!texts) {
log_message(LOG_LEVEL_ERROR, "memory allocation failed for filter rules");
return false;
}
for (int i = 0; i < rule_count; i++)
texts[i] = (const char*)config->filters->items[i];
}
if (rule_count > 0 || config->cvs_exclude) {
char err[160];
out->base_filters = filter_base_build(texts, rule_count, config->cvs_exclude,
config->delete_excluded, err, sizeof(err));
free(texts);
if (!out->base_filters) {
log_message(LOG_LEVEL_ERROR, "invalid filter rule: %s", err);
return false;
}
} else {
free(texts);
}
ScannerOptions* options = &out->options;
options->use_metadata = config->use_metadata;
options->preserve_atimes = config->preserve_atimes;
options->preserve_crtimes = config->preserve_crtimes;
options->preserve_xattrs = config->preserve_xattrs;
options->preserve_acls = config->preserve_acls;
options->chunk_size = config->chunk_size;
/* --exclude/--include are compiled, in command-line order, into the SAME
* ordered filter rule list as --filter/-f (see config_add_selection_rule), so
* the legacy per-kind arrays are deliberately NOT passed to the scanner:
* doing so would re-apply them with the old "excludes first, then includes as
* a mandatory whitelist" precedence and defeat rsync's first-match-wins
* ordering. The arrays remain populated purely for the Config API surface. */
options->exclude_patterns = NULL;
options->exclude_count = 0;
options->include_patterns = NULL;
options->include_count = 0;
options->max_size = config->max_size;
options->min_size = config->min_size;
options->max_depth = config->max_depth;
options->num_threads = num_threads;
options->follow_symlinks = config->follow_symlinks;
options->copy_links = config->copy_links;
options->safe_links = config->safe_links;
options->copy_unsafe_links = config->copy_unsafe_links;
options->copy_dirlinks = config->copy_dirlinks;
options->munge_links = config->munge_links;
options->checksum = config->checksum;
options->one_file_system = config->one_file_system;
options->preserve_devices = config->preserve_devices;
options->preserve_specials = config->preserve_specials;
options->copy_devices = config->copy_devices;
options->file_list = (const FileListSet*)config->files_from_set;
options->base_filters = out->base_filters;
options->per_dir_filters = config->per_dir_filter;
options->delete_excluded = config->delete_excluded;
options->exclude_per_dir_filter_files = config->per_dir_filter_count >= 2;
options->dirs = config->dirs;
options->relative = config->relative;
/* A real recursive transfer recreates empty source directories (rsync
parity); low-level scanner users leave this off. */
options->emit_empty_dirs = true;
/* --no-implied-dirs only has meaning with -R (rsync): without it the option
is a documented no-op, so the scanner must not suppress directory
metadata. */
options->no_implied_dirs = config->no_implied_dirs && config->relative;
/* -R/--relative outside --files-from reconstructs every destination path from
* the source spec (rsync's '/./' cut point). With --files-from the listed
* entry already supplies the bare relative path, so no prefix is built. */
if (config->relative && config->files_from_set == NULL && config->send_directory) {
out->relative_prefix = scanner_relative_prefix(config->send_directory);
if (!out->relative_prefix) {
log_message(LOG_LEVEL_ERROR, "memory allocation failed building --relative path prefix");
filter_rule_list_free(out->base_filters);
out->base_filters = NULL;
return false;
}
options->relative_prefix = out->relative_prefix;
}
options->prune_empty_dirs = config->prune_empty_dirs;
options->ignore_io_errors = config->ignore_errors;
options->ignore_missing_args = config->ignore_missing_args || config->delete_missing_args;
options->note_nonreg = (config->info_level & LOG_INFO_NONREG) != 0 && !config->quiet;
options->note_mount = (config->info_level & LOG_INFO_MOUNT) != 0 && !config->quiet;
options->send_directory = config->send_directory;
options->eight_bit_output = config->eight_bit_output;
options->excluded_paths = NULL;
options->excluded_mutex = NULL;
options->size_skipped_paths = NULL;
options->synced_dirs = NULL;
options->hardlinks = NULL;
/* Set by the real send paths; NULL for the metadata-only scans (progress
pre-count, batch) that must not perturb the sender's --stats counter. */
options->dir_count = NULL;
/* P7 Wave D: capture source directory metadata when a directory attribute is
requested (-p for modes, -t for times unless -O omits them). Whether they
are APPLIED is decided receiver-side. */
options->capture_dir_times = dir_metadata_should_capture(config);
options->dir_entries = NULL;
options->dir_entries_mutex = NULL;
if (config->preserve_hard_links) {
out->hardlinks = hardlink_table_create();
if (!out->hardlinks) {
filter_rule_list_free(out->base_filters);
out->base_filters = NULL;
return false;
}
options->hardlinks = out->hardlinks;
}
return true;
}
void prepared_scanner_destroy(PreparedScanner* prepared) {
if (!prepared)
return;
filter_rule_list_free(prepared->base_filters);
prepared->base_filters = NULL;
hardlink_table_destroy(prepared->hardlinks);
prepared->hardlinks = NULL;
free(prepared->relative_prefix);
prepared->relative_prefix = NULL;
}
/* -R/--relative implied directories: rsync transmits the metadata of the
* parent directories implied by the source path (every prefix component above
* the source root) so the receiver applies their attributes to the created
* parents. FastSync's scan only covers the source root and below, so append
* one metadata-only directory entry per implied ancestor. --no-implied-dirs
* suppresses this exactly like rsync. A missing ancestor is never fatal. */
bool append_implied_dir_times(const Config* config, ArrayList* dir_entries) {
if (!dir_entries || !config->relative || config->files_from_set != NULL ||
config->no_implied_dirs || !config->send_directory)
return true;
char* prefix = scanner_relative_prefix(config->send_directory);
if (!prefix)
return true;
int ncomp = 0;
for (const char* s = prefix; *s;) {
while (*s == '/')
s++;
if (!*s)
break;
while (*s && *s != '/')
s++;
ncomp++;
}
if (ncomp <= 1) {
free(prefix);
return true;
}
char* fs = str_dup(config->send_directory);
if (!fs) {
free(prefix);
return true;
}
size_t flen = strlen(fs);
while (flen > 1 && fs[flen - 1] == '/')
fs[--flen] = '\0';
bool ok = true;
/* Walk the source path upwards one component at a time (fs is truncated in
place, so each step targets the next implied ancestor). */
for (int depth = ncomp - 2; depth >= 0 && ok; depth--) {
char* slash = strrchr(fs, '/');
if (!slash || slash == fs)
break;
*slash = '\0';
char* p = prefix;
int c = 0;
while (c <= depth) {
while (*p == '/')
p++;
while (*p && *p != '/')
p++;
c++;
}
char saved = *p;
*p = '\0';
struct stat st;
if (stat(fs, &st) == 0 && S_ISDIR(st.st_mode)) {
File* file = file_create(fs);
if (!file) {
ok = false;
} else {
file->is_dir = true;
file->metadata =
file_metadata_create(fs, &st, config->preserve_atimes, config->preserve_crtimes);
file->send_path = str_dup(prefix);
if (!file->metadata || !file->send_path || !array_list_add(dir_entries, file)) {
file_destroy(file);
ok = false;
}
}
}
*p = saved;
}
free(fs);
free(prefix);
return ok;
}
/* The delete-walk root scope for a full (non---files-from) transfer: rsync
* confines --delete to the directories it actually transferred. A plain
* recursive run mirrors the source under the receive root, so "." (the whole
* tree) is correct; an -R run transfers only the reconstructed prefix subtree,
* so the walk is scoped to that prefix instead. Returns a malloc'd wire path
* (or "."), or NULL on allocation failure. */
char* delete_scope_root_marker(const Config* config) {
if (config->relative && config->files_from_set == NULL && config->send_directory) {
char* prefix = scanner_relative_prefix(config->send_directory);
if (!prefix)
return NULL;
if (prefix[0] != '\0')
return prefix;
free(prefix);
}
return str_dup(".");
}
/* The -R destination prefix that confines a per-directory delete walk, or NULL
* when the whole receive root is in scope. The marker was installed into
* `synced_dirs` by delete_scope_root_marker(); for a plain recursive transfer
* it is "." (whole root) and for --files-from the list is not a single prefix. */
const char* delete_plan_walk_root(const Config* config, const ArrayList* synced_dirs) {
if (!config || config->files_from_set != NULL || !config->relative || !config->send_directory)
return NULL;
if (!synced_dirs || synced_dirs->size != 1)
return NULL;
const char* marker = (const char*)synced_dirs->items[0];
if (marker[0] == '\0' || strcmp(marker, ".") == 0)
return NULL;
return marker;
}
/* The destination-relative mirror path for a missing --files-from entry: where
a PRESENT entry with the same name would have been written. With -R that is
the entry's bare relative path (the bare wire path the receiver uses);
otherwise it is the full source mirror below the destination root
(`send_directory` joined to the entry, leading '/' stripped), exactly the
path the manifest records for a present sibling. Returns an owned string, or
NULL on allocation failure. */
static char* files_from_missing_dest_path(const Config* config, const char* entry) {
if (config->relative)
return str_dup(entry);
char* joined = path_cat(config->send_directory, entry);
if (!joined)
return NULL;
const char* rel = *joined == '/' ? joined + 1 : joined;
char* dup = str_dup(rel);
free(joined);
return dup;
}
/* --files-from semantics: every listed entry must resolve under the source
* root, otherwise rsync reports a hard error instead of silently transferring
* nothing. An entry of "." (the whole tree) and listed-but-empty directories
* are valid. An empty list is valid too: rsync transfers nothing and exits 0.
* With --ignore-missing-args
* (implied by --delete-missing-args) a listed-but-missing entry is instead
* skipped: nothing is transferred for it, it never enters the keep-set and the
* run succeeds for the rest (an all-missing non-empty list succeeds
* transferring nothing, matching rsync). With --delete-missing-args
* `missing_dest` (when non-NULL) collects the entry's destination-relative
* mirror for the receiver's exact-deletion request. Runs before any
* transfer so the failure/skip is surfaced uniformly in the single-threaded,
* -m, dry-run and --list-only paths. */
bool files_from_list_check(const Config* config, ArrayList* missing_dest, int* skipped_out) {
*skipped_out = 0;
const FileListSet* set = (const FileListSet*)config->files_from_set;
if (!set)
return true;
if (!config->send_directory) {
log_message(LOG_LEVEL_ERROR, "--files-from requires a source directory");
return false;
}
if (set->count == 0) {
/* rsync treats an empty --files-from list as "nothing to transfer" and
exits 0 (the source directory is still a valid source arg), so this is
not an error. Nothing passes the (empty) allow-set, so no file is sent
and no keep-set entry is produced. */
return true;
}
bool ignore = config->ignore_missing_args || config->delete_missing_args;
for (int i = 0; i < set->count; i++) {
const char* entry = set->entries[i];
if (entry[0] == '\0')
continue; /* "." == list the whole tree */
char* full = path_cat(config->send_directory, entry);
if (!full) {
log_message(LOG_LEVEL_ERROR, "memory allocation failed while validating --files-from");
return false;
}
struct stat st;
if (lstat(full, &st) != 0) {
free(full);
if (ignore) {
(*skipped_out)++;
char* escaped_entry = output_escape(entry, log_get_8_bit_output());
log_info_message(LOG_INFO_MISC, "skipping missing --files-from entry '%s'",
escaped_entry ? escaped_entry : "<allocation failed>");
free(escaped_entry);
if (config->delete_missing_args && missing_dest) {
char* mirror = files_from_missing_dest_path(config, entry);
if (!mirror || !array_list_add(missing_dest, mirror)) {
free(mirror);
log_message(LOG_LEVEL_ERROR, "memory allocation failed while validating --files-from");
return false;
}
}
continue;
}
char* escaped_entry = output_escape(entry, log_get_8_bit_output());
char* escaped_src = output_escape(config->send_directory, log_get_8_bit_output());
log_message(LOG_LEVEL_ERROR, "--files-from entry '%s' not found in source '%s'",
escaped_entry ? escaped_entry : "<allocation failed>",
escaped_src ? escaped_src : "<allocation failed>");
free(escaped_entry);
free(escaped_src);
return false;
}
free(full);
}
if (*skipped_out > 0) {
if (config->delete_missing_args) {
/* --list-only never deletes and a --dry-run only shows intent, so the
summary must not claim a real deletion happened in those modes. */
if (config->list_only)
log_message(LOG_LEVEL_WARNING,
"--delete-missing-args: %d missing --files-from entr%s skipped (--list-only "
"never deletes)",
*skipped_out, *skipped_out == 1 ? "y" : "ies");
else if (config->dry_run)
log_message(LOG_LEVEL_WARNING,
"--delete-missing-args: %d missing --files-from entr%s would be deleted from "
"the destination (dry run)",
*skipped_out, *skipped_out == 1 ? "y" : "ies");
else
log_message(
LOG_LEVEL_WARNING,
"--delete-missing-args: %d missing --files-from entr%s will be deleted from the "
"destination",
*skipped_out, *skipped_out == 1 ? "y" : "ies");
} else if (config->ignore_missing_args)
log_message(LOG_LEVEL_WARNING,
"--ignore-missing-args: ignored %d missing --files-from entr%s", *skipped_out,
*skipped_out == 1 ? "y" : "ies");
}
return true;
}
/* Walk the whole source tree once collecting only destination-relative wire
paths, loading and sending nothing. --delete-before/--delete-during need the
complete keep-set manifest before the first data byte, so it is built by a
dedicated pre-scan pass and transmitted early; the data pass then re-scans
with a fresh scanner. --delete-before additionally replays this very scan as
its data pass (rsync's single file list), so `chunks_out` (optional) retains
the scanned Chunk objects for the caller to send instead of destroying them;
the caller owns the list and must give it a chunk_destroy destructor. A
source I/O error is fatal unless the options carry --ignore-errors, in which
case the scan continues past the unreadable directory and *io_error_out
reports it (the caller still performs the deletion but reports the run as
errored). */
bool scan_paths_only(const Config* config, const ScannerOptions* options, ArrayList* manifest,
DeletePlanSender* plans, bool* io_error_out,
unsigned long long* non_dir_count_out, ArrayList* chunks_out,
bool emit_nonreg) {
if (io_error_out)
*io_error_out = false;
if (non_dir_count_out)
*non_dir_count_out = 0;
ScannerOptions local = *options;
/* The pre-scan is normally a paths-only pass with no client output: it must
not emit --info=nonreg lines because the data pass re-scans and emits them
once. When the caller replays this scan as the data pass (--delete-before)
there is no later scan, so it opts in and the lines are emitted here. */
local.note_nonreg = emit_nonreg && options->note_nonreg;
DirectoryScanner* scanner = directory_scanner_create_with_options(config->send_directory, &local);
if (!scanner)
return false;
bool ok = true;
Chunk* chunk;
while ((chunk = directory_scanner_next(scanner)) != NULL) {
if (non_dir_count_out) {
for (int i = 0; i < chunk->element_count; i++) {
const File* f = chunk->items[i];
if (f && !f->is_dir)
(*non_dir_count_out)++;
}
}
if (manifest && !add_chunk_to_manifest(manifest, chunk)) {
ok = false;
chunk_destroy(chunk);
break;
}
if (plans) {
for (int i = 0; i < chunk->element_count; i++) {
File* f = chunk->items[i];
if (!f)
continue;
const char* path = file_wire_path(f);
if (!delete_plan_sender_add(plans, path, f->is_dir)) {
ok = false;
break;
}
}
if (!ok) {
chunk_destroy(chunk);
break;
}
}
if (chunks_out) {
/* Retain the chunk for the caller's data pass; ownership moves with it. */
if (!array_list_add(chunks_out, chunk)) {
ok = false;
chunk_destroy(chunk);
break;
}
} else {
chunk_destroy(chunk);
}
}
if (ok) {
/* Keep every traversed source directory, including empty ones, so a plan
no longer removes the destination directory itself. Their own plans are
emitted after the data stream (no file frame triggers them). */
if (plans && options->plan_dirs) {
for (int i = 0; i < options->plan_dirs->size; i++) {
if (!delete_plan_sender_add(plans, (const char*)options->plan_dirs->items[i], true)) {
ok = false;
break;
}
}
}
}
if (ok && directory_scanner_failed(scanner))
ok = false;
if (io_error_out)
*io_error_out = directory_scanner_had_io_error(scanner);
directory_scanner_destroy(scanner);
return ok;
}
+559 -1770
View File
File diff suppressed because it is too large Load Diff
+7 -1
View File
@@ -22,7 +22,13 @@ void client_set_abort_armed(bool armed);
* never free it, and the caller retains ownership (freeing it with
* config_delete() once the call returns). */
int send_files(Config* config);
int send_files_multithreaded(Config** config);
int send_files_multithreaded(Config* config);
/* rsync's --ignore-errors deletion gate: with no I/O error during the scan the
* deletion phase always proceeds; with one it is suppressed unless
* `--ignore-errors` was given. Exposed so the decision can be unit-tested
* without a privileged (mode-000) source directory. See client_send.c. */
bool ignore_errors_allows_delete(const Config* config, bool had_io_error);
/* Phase 6 residual-batch (client-only). See client_send.c. */
int write_batch_from_source(const Config* config, const char* batch_path);
int apply_batch_to_dest(const Config* config, const char* batch_path, const char* dest_root);
+95
View File
@@ -0,0 +1,95 @@
#ifndef CLIENT_SEND_INTERNAL_H
#define CLIENT_SEND_INTERNAL_H
/* Declarations shared between the client_send.c transfer orchestration and the
* reporting (client_report.c), scanner-preparation (client_scan.c) and
* manifest/list/dry-run (client_manifest.c) translation units that were split
* out of it. Nothing here is part of the public client_send.h facade. */
#include "array_list.h"
#include "client_send.h"
#include "config.h"
#include "delete_plan.h"
#include "delta.h"
#include "format.h"
#include "log.h"
#include "scanner.h"
#include <stdatomic.h>
#include <stdbool.h>
#include <stddef.h>
#include <time.h>
/* One mebibyte in bytes; the unit used by the --stats/--progress lines.
Always cast to double when dividing so the output stays fractional. */
#define BYTES_PER_MIB (1024ULL * 1024ULL)
/* Compiled scanner inputs that are shared read-only across scanner instances
* and, in -m mode, across worker threads. `base_filters` owns the compiled
* command-line + -C rules; the FileListSet allow-set lives in the Config.
* `hardlinks` owns the --hard-links/-H link-group detection table (NULL when
* off) and is shared (mutex-guarded) across every scanner/worker of one scan. */
typedef struct {
ScannerOptions options;
FilterRuleList* base_filters; /* owned; may be NULL */
HardLinkTable* hardlinks; /* owned; may be NULL */
char* relative_prefix; /* owned -R prefix; may be NULL */
} PreparedScanner;
/* client_scan.c */
bool prepare_scanner(const Config* config, int num_threads, PreparedScanner* out);
void prepared_scanner_destroy(PreparedScanner* prepared);
bool append_implied_dir_times(const Config* config, ArrayList* dir_entries);
char* delete_scope_root_marker(const Config* config);
const char* delete_plan_walk_root(const Config* config, const ArrayList* synced_dirs);
bool files_from_list_check(const Config* config, ArrayList* missing_dest, int* skipped_out);
bool scan_paths_only(const Config* config, const ScannerOptions* options, ArrayList* manifest,
DeletePlanSender* plans, bool* io_error_out,
unsigned long long* non_dir_count_out, ArrayList* chunks_out,
bool emit_nonreg);
/* client_report.c */
void log_server_rejection(const char* context);
const char* display_bytes(unsigned long long bytes, bool human_readable, char* buffer,
size_t buffer_size);
unsigned long long dir_count_for_stats(const Config* config, const ArrayList* dir_entries,
atomic_ullong* counter);
void report_transfer_stats(const Config* config, const TransferStats* stats, time_t start,
const ReceiverStats* recv);
void transfer_stats_note_entry(TransferStats* stats, const File* file);
void transfer_stats_note_transferred(TransferStats* stats, const File* file);
bool info_flag_enabled(const Config* config, LogInfoFlag flag);
void print_delete_reports(const Config* config, const ArrayList* paths);
const char* delete_display_path(const Config* config, const char* path);
bool progress_requested(const Config* config);
void client_progress_cleanup(void);
void client_progress_begin(const Config* config);
void client_progress_file(const Config* config, const File* file);
void client_progress_name(const Config* config, const File* file);
/* Emit a transferred entry's ancestor directories (as -i/--out-format change
* lines or --progress name lines) before the entry's own line. */
void client_change_emit_ancestors(const Config* config, const File* file);
void client_progress_uptodate(const Config* config, const File* file);
void client_progress_prepare(const Config* config, const ArrayList* plan_dirs,
unsigned long long plan_non_dir_count);
bool receive_stats_record(int fd, ReceiverStats* stats, ArrayList* would_delete);
/* client_send.c */
void receive_daemon_motd(Client* client, const Config* config);
Client* connect_transfer_client(const Config* config);
void disconnect_transfer_client(Client* client);
int incremental_check(Client* client, File* file, const Config* config, DeltaSignature** out_sig,
unsigned long long* resume_offset);
/* client_manifest.c */
bool dry_run_targets_server(const Config* config);
bool add_chunk_to_manifest(ArrayList* manifest, const Chunk* chunk);
int send_dry_run_manifest(const Config* config);
int send_list_only(const Config* config);
int send_dry_run_remote(Config* config);
int send_delete_manifest(int fd, ArrayList* manifest, ArrayList* protected_prefixes,
ArrayList* size_skipped, ArrayList* missing_args, ArrayList* synced_dirs);
bool send_delete_manifest_early(Client* client, ArrayList* manifest, ArrayList* protected_prefixes,
ArrayList* size_skipped, ArrayList* missing_args,
ArrayList* synced_dirs);
#endif
+18 -1
View File
@@ -53,7 +53,7 @@ bool validate_config(const Config* config) {
return false;
}
if (config->compression_threads > 0 && !config->use_compression) {
log_message(LOG_LEVEL_ERROR, "--compress-threads requires compression (-c or -z)");
log_message(LOG_LEVEL_ERROR, "--compress-threads requires compression (-z/--compress)");
return false;
}
if (config->transport == TRANSPORT_SSH && config->use_sendfile) {
@@ -75,6 +75,13 @@ bool validate_config(const Config* config) {
log_message(LOG_LEVEL_ERROR, "-4/--ipv4 and -6/--ipv6 are mutually exclusive");
return false;
}
/* rsync 3.4.1 rejects --inplace together with --partial-dir (exit 1): the
inplace write path bypasses partial staging, so a partial-dir name would be
silently ignored. Match rsync's message and refuse before any I/O. */
if (config->inplace && config->partial_dir) {
log_message(LOG_LEVEL_ERROR, "--inplace cannot be used with --partial-dir");
return false;
}
if (config->log_file_format && !config->log_file) {
log_message(LOG_LEVEL_ERROR, "--log-file-format requires --log-file");
return false;
@@ -102,6 +109,16 @@ bool validate_config(const Config* config) {
log_message(LOG_LEVEL_ERROR, "%s", invariants_error);
return false;
}
/* The receiver rejects a protect-rule block with more than MAX_FILTER_RULES
entries as an opaque protocol error; reject an over-limit --filter set here,
before any network I/O, with an actionable message. send_protect_entries()
re-checks the final built count because cvs-exclude / merge rules can
expand it beyond config->filters->size. */
if (config->filters && config->filters->size > MAX_FILTER_RULES) {
log_message(LOG_LEVEL_ERROR, "too many filter rules: %d (maximum %d)", config->filters->size,
MAX_FILTER_RULES);
return false;
}
/* --protocol: FastSync has exactly one wire format, so the forced version
must equal the current PROTOCOL_VERSION exactly. Rejected here, before any
network I/O, rather than letting the server hit its own mismatch check. */
+516 -1549
View File
File diff suppressed because it is too large Load Diff
+58 -5
View File
@@ -10,6 +10,7 @@
#include "stop_condition.h"
#include <dirent.h>
#include <stdbool.h>
#include <stddef.h>
#include <stdatomic.h>
#include <sys/types.h>
#include <threads.h>
@@ -50,7 +51,7 @@ typedef struct {
bool copy_dirlinks;
bool munge_links;
bool checksum;
bool one_file_system;
int one_file_system;
/* Phase 4 special/devices: whether device nodes (--devices) and special files
* (--specials) are preserved via recreation, and whether --copy-devices
* copies a device's content as an ordinary regular file. */
@@ -121,9 +122,29 @@ typedef struct {
* it) and to emit its plan after the data stream, when no file frame would
* otherwise trigger it. Guarded by `excluded_mutex`. */
ArrayList* plan_dirs;
/* --ignore-errors: an unreadable directory during the scan is recorded as an
* I/O error and skipped instead of aborting the scan. Client-only. */
/* --ignore-errors: an unreadable subdirectory no longer aborts the scan (it
* is always skipped so the rest of the tree transfers); this flag is kept so
* the client can distinguish the option state when deciding deletion policy.
* Client-only. */
bool ignore_io_errors;
/* --info=nonreg: print rsync's `skipping non-regular file "NAME"` line for a
* non-regular entry that is not being preserved. Client-only. */
bool note_nonreg;
/* --info=mount: print rsync's `[sender] skipping mount-point dir NAME` when
* -xx drops a mount-point directory. Client-only. */
bool note_mount;
/* --stats directory accounting for a `-r` run (no -t/-p): a shared counter of
* traversed directories that are NOT otherwise represented by an inline
* directory entry (rsync still counts every directory in `Number of files`).
* Incremented when a directory is opened and decremented when an empty
* directory is emitted inline (so it is counted exactly once). Atomic
* because the parallel scanner's workers share it; NULL disables the
* accounting. Client-only. */
atomic_ullong* dir_count;
/* Source root and 8-bit-output policy used to render a `--info=nonreg` name
* relative to the transfer root. Borrowed read-only. */
const char* send_directory;
bool eight_bit_output;
/* --ignore-missing-args (implied by --delete-missing-args): an explicitly
* --files-from-listed entry that does not exist under the source is skipped
* instead of failing (the --dirs generator is the only scanner path that
@@ -151,6 +172,17 @@ typedef struct {
bool capture_dir_times;
ArrayList* dir_entries;
mtx_t* dir_entries_mutex;
/* Recreate empty source directories on a recursive transfer: emit a
* payload-less directory entry for every traversed directory that produced
* no transferred/descended child. Off by default so low-level scanner users
* (unit helpers, --list-only) see only the historical file list; the real
* sender sets it in prepare_scanner. */
bool emit_empty_dirs;
/* --no-implied-dirs with -R + --files-from: a directory that is only an
* implied parent of a listed entry (not itself listed, nor below a listed
* directory) must not carry source metadata; it is created with default
* attributes at the destination, matching rsync. */
bool no_implied_dirs;
} ScannerOptions;
/* Internal per-scanner filter state. FilterNode chains represent the ordered
@@ -168,6 +200,22 @@ typedef struct {
int current_depth;
dev_t root_dev;
bool failed;
/* rsync-order traversal: each opened directory's entries are inspected once
and buffered (an internal SortedEntry[] owned here) sorted as rsync's flist
orders them -- non-directories ascending, then directories ascending. The
entries are walked in order and child directories are collected in
`pending_dirs` (an ArrayList of DirEntry*, owned here) and pushed onto the
LIFO `directories` stack in reverse at directory exhaustion, so the emitted
stream is depth-first like rsync. `sorted_*` are reset per directory. */
void* sorted_entries;
size_t sorted_count;
size_t sorted_index;
void* pending_dirs;
/* Recursive scan: whether the open directory yielded any transferred or
descended entry. When it did not, closing it emits a directory entry so
the empty source directory is recreated at the destination (rsync
parity). */
bool current_dir_produced;
/* Phase 2 (files-from / filter layer). */
char* root_path; /* transfer root (fs path) for rel computation */
char* current_rel; /* rel path of the open directory ("" == root) */
@@ -186,6 +234,10 @@ typedef struct {
--ignore-errors the scan continues past it and the caller decides what to
do; `failed` is reserved for fatal errors that always abort the scan. */
bool io_error;
/* The transfer ROOT could not be opened. It is always fatal, even under
--ignore-errors, but the client still maps it to rsync's partial-transfer
exit (23) rather than a generic failure. */
bool root_io_error;
} DirectoryScanner;
typedef struct {
@@ -205,7 +257,8 @@ typedef struct {
int completed;
Chunk* initial_chunk;
ProtocolSession* allocation_session;
FilterNode* root_filter_node; /* root .rsync-filter context (owned by ps) */
FilterNode* root_filter_node; /* root .rsync-filter context (owned by ps) */
const ScannerOptions* options; /* borrowed scan options (--info=nonreg output) */
} ParallelScanner;
DirectoryScanner* directory_scanner_create(const char* root_directory, bool use_metadata,
@@ -224,7 +277,7 @@ void directory_scanner_destroy(DirectoryScanner* scanner);
/* --one-file-system (-x) decision: a directory entry may be descended into
* only when the option is disabled or the entry lives on the same device as
* the transfer root. Exposed so tests can exercise the rule directly. */
bool scanner_same_filesystem(bool one_file_system, dev_t root_device, dev_t entry_device);
bool scanner_same_filesystem(int one_file_system, dev_t root_device, dev_t entry_device);
/* Relative path of an on-disk path below `root` ("" == the root itself, NULL
* when `fs_path` is not under `root`). Handles trailing slashes and a root of
+675
View File
@@ -0,0 +1,675 @@
#include "log.h"
#include "scanner.h"
#include "scanner_internal.h"
#include "array_list.h"
#include "chunk.h"
#include "file.h"
#include "queue.h"
#include "utils.h"
#include <dirent.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/stat.h>
#include <sys/sysmacros.h>
#include <threads.h>
#include <unistd.h>
#include <limits.h>
#include "xattr.h"
/* A chain node: `own` holds the .rsync-filter rules of one directory, `parent`
* the context that directory inherited (nearest ancestor with a filter file).
* The chain for a directory's contents runs from that directory's own node up
* to the root; the command-line base rules are evaluated after the whole
* chain. */
struct FilterNode {
FilterNode* parent;
FilterRuleList* own;
};
void filter_node_destroy(void* item) {
if (item) {
FilterNode* node = (FilterNode*)item;
if (node->own)
filter_rule_list_free(node->own);
free(node);
}
}
FilterNode* filter_node_alloc(FilterNode* parent, FilterRuleList* own) {
FilterNode* node = malloc(sizeof(FilterNode));
if (!node)
return NULL;
node->parent = parent;
node->own = own;
return node;
}
/* Evaluate a rule chain for one entry. rsync precedence, highest first: the
* innermost (current) directory's .rsync-filter rules, then each ancestor's,
* then the root's, and finally the command-line base rules (--filter/-C). The
* sender-side verdict decides whether the entry is hidden from the transfer;
* the receiver-side verdict decides whether its destination mirror is protected
* from --delete. Each side takes the FIRST matching rule independently. */
typedef struct {
bool hide; /* sender-side exclude matched */
bool protect; /* receiver-side exclude matched */
} FilterOutcome;
static void chain_rules_outcome(const FilterRuleList* base, const FilterNode* node, const char* rel,
const char* leaf, bool is_dir, FilterOutcome* out) {
memset(out, 0, sizeof(*out));
bool sender_decided = false;
bool receiver_decided = false;
const FilterNode* n = node;
while (!sender_decided || !receiver_decided) {
const FilterRuleList* list = n ? n->own : base;
if (list) {
if (!sender_decided) {
FilterAction action = filter_rules_apply_side(list, rel, leaf, is_dir, FILTER_SIDE_SENDER);
if (action != FILTER_ACTION_NONE) {
out->hide = action == FILTER_ACTION_EXCLUDE;
sender_decided = true;
}
}
if (!receiver_decided) {
FilterAction action =
filter_rules_apply_side(list, rel, leaf, is_dir, FILTER_SIDE_RECEIVER);
if (action != FILTER_ACTION_NONE) {
out->protect = action == FILTER_ACTION_PROTECT;
receiver_decided = true;
}
}
}
if (!n)
break;
n = n->parent;
}
}
static bool entry_allowed(const FilterRuleList* base, const FilterNode* node, const char* rel,
const char* leaf, bool is_dir, bool exclude_filter_files,
bool* protect_out) {
/* -FF: per-directory .rsync-filter files are never transferred (single -F
transfers them, matching rsync). */
if (exclude_filter_files && !is_dir && strcmp(leaf, ".rsync-filter") == 0) {
if (protect_out)
*protect_out = false;
return false;
}
FilterOutcome outcome;
chain_rules_outcome(base, node, rel, leaf, is_dir, &outcome);
if (protect_out)
*protect_out = outcome.protect;
return !outcome.hide;
}
void dir_entry_destroy(void* item) {
if (item) {
DirEntry* de = (DirEntry*)item;
free(de->path);
free(de);
}
}
DirEntry* dir_entry_create(const char* path, int depth, FilterNode* context) {
DirEntry* de = malloc(sizeof(DirEntry));
if (!de)
return NULL;
de->path = str_dup(path);
if (!de->path) {
free(de);
return NULL;
}
de->depth = depth;
de->context = context;
return de;
}
/* Apply rsync's symlink-resolution precedence to one S_ISLNK entry:
* --copy-links dereferences every symlink;
* --copy-unsafe-links dereferences only targets unsafe_symlink() flags;
* -k/--copy-dirlinks dereferences only a symlink whose referent is a dir;
* --safe-links (receiver-side in rsync; modelled here) ignores an unsafe
* target that would otherwise be carried; with --munge-links
* every stored target becomes absolute, so --safe-links then
* ignores every symlink, exactly as rsync documents;
* -l/--links carries the link.
* `link_rel` is the symlink's transfer-relative path (incl. name) and is used
* only for the lexical unsafe test. `target` receives the raw link value. */
LinkAction scanner_link_action(const ScannerOptions* options, const char* path,
const char* link_rel, char* target, size_t target_size) {
if (!options->follow_symlinks && !options->copy_links && !options->safe_links &&
!options->copy_unsafe_links && !options->copy_dirlinks)
return LINK_ACTION_SKIP;
ssize_t length = readlink(path, target, target_size - 1);
if (length < 0)
return LINK_ACTION_SKIP;
target[length] = '\0';
bool unsafe = file_symlink_unsafe(target, link_rel);
if (options->copy_links || (options->copy_unsafe_links && unsafe))
return LINK_ACTION_DEREF;
if (options->copy_dirlinks) {
struct stat ref;
if (stat(path, &ref) == 0 && S_ISDIR(ref.st_mode))
return LINK_ACTION_DEREF;
}
if (options->safe_links && (unsafe || options->munge_links))
return LINK_ACTION_SKIP_PROTECTED;
if (!options->follow_symlinks || target[0] == '\0')
return LINK_ACTION_SKIP;
return LINK_ACTION_CARRY;
}
/* --one-file-system (-x) decision. Only directories can carry a different
* device than their parent (mount points), so this is checked when a child
* directory is about to be descended into. */
bool scanner_same_filesystem(int one_file_system, dev_t root_device, dev_t entry_device) {
return one_file_system <= 0 || entry_device == root_device;
}
/* Build a payload-less directory File carrying the captured metadata (when
* requested). Used by -x mount-point emission and --list-only directory
* entries. Returns NULL on allocation failure. */
File* scanner_build_dir_file(const char* path, const struct stat* stats,
const ScannerOptions* options) {
File* dir = file_create(path);
if (dir == NULL)
return NULL;
dir->is_dir = true;
if (options->use_metadata) {
dir->metadata =
file_metadata_create(dir->path, stats, options->preserve_atimes, options->preserve_crtimes);
if (!dir->metadata) {
file_destroy(dir);
return NULL;
}
}
return dir;
}
/* Relative path of an on-disk path below `root`. The transfer root may be
* given with a trailing slash; the returned rel path never has one and is ""
* for the root itself. A root of "/" is handled (its children start at "/").
* Exposed so tests can exercise the mapping directly. */
char* scanner_path_relative(const char* root, const char* fs_path) {
size_t root_len = strlen(root);
while (root_len > 1 && root[root_len - 1] == '/')
root_len--;
if (strncmp(root, fs_path, root_len) != 0)
return NULL;
if (root_len == 1 && root[0] == '/') {
if (fs_path[1] == '\0')
return str_dup("");
return str_dup(fs_path + 1);
}
if (fs_path[root_len] == '\0')
return str_dup("");
if (fs_path[root_len] != '/')
return NULL;
return str_dup(fs_path + root_len + 1);
}
/* -R/--relative destination-relative prefix reconstructed from a source spec:
* everything after the first '.' path component (rsync's '/./' cut point),
* with leading/trailing slashes removed; or the whole spec (normalized) when
* there is no cut. Returns "" for the receive root. Exposed for tests. */
char* scanner_relative_prefix(const char* spec) {
if (!spec || spec[0] == '\0')
return NULL;
const char* after = spec;
if (spec[0] == '.' && spec[1] == '/') {
after = spec + 2;
} else {
const char* cut = strstr(spec, "/./");
if (cut)
after = cut + 3;
}
size_t cap = strlen(spec) + 1;
char* out = malloc(cap);
if (!out)
return NULL;
size_t len = 0;
for (const char* s = after; *s;) {
while (*s == '/')
s++;
const char* comp = s;
while (*s && *s != '/')
s++;
size_t clen = (size_t)(s - comp);
if (clen == 0 || (clen == 1 && comp[0] == '.'))
continue;
if (len)
out[len++] = '/';
memcpy(out + len, comp, clen);
len += clen;
}
out[len] = '\0';
return out;
}
/* Relative path of a child entry below the current directory. */
char* child_rel_path(const char* parent_rel, const char* name) {
if (!parent_rel || parent_rel[0] == '\0')
return str_dup(name);
return path_cat(parent_rel, name);
}
/* Destination-relative wire path for an entry under an -R prefix. */
char* scanner_prefix_send_path(const char* prefix, const char* rel) {
if (prefix[0] == '\0')
return str_dup(rel);
if (rel[0] == '\0')
return str_dup(prefix);
return path_cat(prefix, rel);
}
/* Apply the --files-from allow-set and the filter layer to one entry. On
* return `*protect_out` is true when a receiver-side rule protects the entry's
* destination mirror from deletion. */
bool entry_passes_selection(const FileListSet* file_list, const FilterRuleList* base,
const FilterNode* node, const char* rel, const char* leaf, bool is_dir,
bool per_dir_filters, bool exclude_filter_files, bool* protect_out) {
if (protect_out)
*protect_out = false;
if (file_list && !file_list_affects(file_list, rel))
return false;
if (base || per_dir_filters)
return entry_allowed(base, node, rel, leaf, is_dir, exclude_filter_files, protect_out);
return true;
}
/* Best-effort capture of the file's whitelisted xattrs (-X/-A). A failure to
* read xattrs is non-fatal: the file is transferred without them. A symlink
* entry reads the LINK's own xattrs (never the referent's) with the no-follow
* variant; on Linux the VFS refuses xattrs on symlinks, so that yields NULL. */
void scanner_capture_xattrs(const DirectoryScanner* scanner, File* file) {
if (!scanner || !file || !(scanner->options.preserve_xattrs || scanner->options.preserve_acls))
return;
file->xattrs = file->is_symlink
? xattr_capture_path_nofollow(file->path, scanner->options.preserve_acls)
: xattr_capture_path(file->path, scanner->options.preserve_acls);
}
/* Apply --hard-links (-H) detection to one regular File. On a sibling (a
* later member of an already-seen source inode) the File keeps the group id
* and the first member's wire path but carries NO data payload (size 0); the
* first member is left untouched (data present, link_first). Allocation
* failure is fatal: the scanner is marked failed. */
void scanner_assign_hardlink(DirectoryScanner* scanner, HardLinkTable* table, File* file,
const struct stat* stats) {
if (!table || !file || !stats)
return;
int gid;
bool is_first;
char* first_path = NULL;
if (!hardlink_table_assign(table, file_wire_path(file), stats->st_dev, stats->st_ino, &gid,
&is_first, &first_path)) {
if (scanner)
scanner->failed = true;
return;
}
file->link_group = gid;
file->link_first = is_first;
if (!is_first) {
file->hardlink_target = first_path;
file->data->size = 0;
} else {
free(first_path);
}
}
/* Phase 4 special/devices decision for one non-regular entry, matching rsync:
- a char/block device is RECREATED as a node under -D/--devices, unless
--copy-devices asks for its content to be copied into a regular file;
- a FIFO/socket is RECREATED under --specials;
- when the matching flag is absent the entry is SKIPPED ("skipping
non-regular file"), exactly like rsync's default, instead of being
silently copied as a zero-length regular file;
- anything else (regular/directory) is left to the normal data path. */
ScannerSpecial scanner_prepare_special(bool preserve_devices, bool preserve_specials,
bool copy_devices, File* file, const struct stat* stats) {
if (!file || !stats)
return SCANNER_SPECIAL_REGULAR;
bool is_device = S_ISCHR(stats->st_mode) || S_ISBLK(stats->st_mode);
bool is_fifo = S_ISFIFO(stats->st_mode);
bool is_socket = S_ISSOCK(stats->st_mode);
if (!is_device && !is_fifo && !is_socket)
return SCANNER_SPECIAL_REGULAR;
if (is_device && copy_devices)
return SCANNER_SPECIAL_REGULAR; /* copy device content as a regular file */
bool preserve = is_device ? preserve_devices : preserve_specials;
if (!preserve)
return SCANNER_SPECIAL_SKIP;
file->is_special = true;
file->data->size = 0;
file->data->data = NULL;
if (is_device) {
file->rdev_major = (int32_t)major(stats->st_rdev);
file->rdev_minor = (int32_t)minor(stats->st_rdev);
}
return SCANNER_SPECIAL_RECREATE;
}
/* Append `rel` to the caller's exclusion sink, taking `mtx` when shared across
parallel worker threads. Returns false on allocation failure (list left
unchanged). */
bool excluded_sink_append(ArrayList* list, mtx_t* mtx, const char* rel) {
if (!list)
return true;
char* dup = str_dup(rel);
if (!dup)
return false;
if (mtx)
mtx_lock(mtx);
bool ok = array_list_add(list, dup);
if (mtx)
mtx_unlock(mtx);
if (!ok)
free(dup);
return ok;
}
/* Record one pruned filesystem path in a delete-protection sink. The stored
form is the entry's wire/destination-relative path (a single leading '/'
removed, exactly how manifest keep entries are stored), so the receiver's
walker prefixes match the destination layout. An allocation failure is a
fatal scan error. */
static void scanner_record_protected(DirectoryScanner* scanner, const char* fs_path,
ArrayList* sink) {
if (!sink || !fs_path)
return;
const char* rel = *fs_path == '/' ? fs_path + 1 : fs_path;
if (!excluded_sink_append(sink, scanner->options.excluded_mutex, rel))
scanner->failed = true;
}
/* rsync's `--info=nonreg` line for a non-regular entry that is not being
* preserved: `skipping non-regular file "NAME"`. The name is the path relative
* to the transfer root, so it matches rsync's displayed name. */
void scanner_note_nonreg(const ScannerOptions* options, const char* fs_path) {
if (!options || !options->note_nonreg || !fs_path)
return;
const char* rel = utils_strip_transfer_root(fs_path, options->send_directory);
char* escaped = output_escape(rel, options->eight_bit_output);
printf("skipping non-regular file \"%s\"\n", escaped ? escaped : rel);
free(escaped);
fflush(stdout);
}
/* rsync 3.4.1's `--info=mount` line, emitted when `-xx` drops a mount-point
* directory: `[sender] skipping mount-point dir NAME` (the client is the
* sender). Plain `-x` keeps the empty directory and prints nothing, matching
* rsync. */
void scanner_note_mount(const ScannerOptions* options, const char* fs_path) {
if (!options || !options->note_mount || !fs_path)
return;
const char* rel = utils_strip_transfer_root(fs_path, options->send_directory);
char* escaped = output_escape(rel, options->eight_bit_output);
printf("[sender] skipping mount-point dir %s\n", escaped ? escaped : rel);
free(escaped);
fflush(stdout);
}
/* --debug=filter: a selection/filter decision dropped an entry. */
void scanner_note_filter(const ScannerOptions* options, const char* name) {
if (!options || !log_debug_enabled(LOG_DEBUG_FILTER) || !name)
return;
log_debug_message(LOG_DEBUG_FILTER, "filter: excluded %s", name);
}
/* Account for a directory that will not be represented by an inline directory
* entry. Paired with scanner_dir_count_uncount for empty directories that are
* emitted inline, so every traversed directory is counted exactly once. */
void scanner_dir_count_count(const ScannerOptions* options) {
if (options && options->dir_count)
atomic_fetch_add(options->dir_count, 1);
}
void scanner_dir_count_uncount(const ScannerOptions* options) {
if (options && options->dir_count)
atomic_fetch_sub(options->dir_count, 1);
}
/* A user-selection exclusion (--filter/-C/per-dir or --exclude/--include). */
void scanner_record_excluded(DirectoryScanner* scanner, const char* fs_path) {
scanner_record_protected(scanner, fs_path, scanner->options.excluded_paths);
}
/* A --max-size/--min-size prune (always protected, even under --delete-excluded). */
void scanner_record_size_skipped(DirectoryScanner* scanner, const char* fs_path) {
scanner_record_protected(scanner, fs_path, scanner->options.size_skipped_paths);
}
/* Record a directory the scan synchronized. `fs_path` is its absolute path and
`rel` its path relative to the transfer root ("" for the root); the stored
form matches the wire layout (the bare relative path in -R+--files-from, else
the source path with a leading '/' removed, with "." for the receive root).
Returns false on allocation failure. */
bool scanner_record_synced_dir(const ScannerOptions* options, const char* fs_path, const char* rel,
bool relative_mode) {
if (!options->synced_dirs && !options->plan_dirs)
return true;
if (!file_list_dir_in_scope(options->file_list, rel))
return true;
char* prefixed = NULL;
const char* dest;
if (relative_mode) {
dest = rel;
} else if (options->relative_prefix) {
prefixed = scanner_prefix_send_path(options->relative_prefix, rel);
if (!prefixed)
return false;
dest = prefixed;
} else {
dest = fs_path;
}
if (dest[0] == '/')
dest++;
if (dest[0] == '\0')
dest = ".";
bool ok = true;
if (options->synced_dirs)
ok = excluded_sink_append(options->synced_dirs, options->excluded_mutex, dest);
/* The delete-plan keep set needs an entry for every traversed source
directory, including empty ones, so its destination mirror is kept rather
than deleted as an extra; the receive root (".") is implicit. */
if (ok && options->plan_dirs && strcmp(dest, ".") != 0)
ok = excluded_sink_append(options->plan_dirs, options->excluded_mutex, dest);
free(prefixed);
return ok;
}
/* Read every per-directory filter file that applies to `dir_path` (its
* .rsync-filter when -F is active, plus each registered "dir-merge NAME") into a
* fresh list. Returns NULL on allocation/parse failure (message in `err`);
* returns an empty list (and *any_exists=false) when no file exists. */
FilterRuleList* read_dir_filters(const ScannerOptions* options, const char* dir_path,
const char* rel, bool* any_exists, char* err, size_t err_size) {
if (err && err_size > 0)
err[0] = '\0';
const FilterRuleList* base = options->base_filters;
bool have_names = options->per_dir_filters || (base && base->dir_merge_count > 0);
if (any_exists)
*any_exists = false;
if (!have_names)
return NULL;
FilterRuleList* own = filter_rule_list_create();
if (!own) {
snprintf(err, err_size, "memory allocation failed");
return NULL;
}
FilterParseOptions opts = {.delete_excluded = options->delete_excluded, .cvs_exclude = false};
bool exists = false;
if (options->per_dir_filters) {
if (!filter_file_append(own, dir_path, ".rsync-filter", rel, &opts, &exists, err, err_size))
goto fail;
if (exists && any_exists)
*any_exists = true;
}
if (base) {
for (int i = 0; i < base->dir_merge_count; i++) {
if (!filter_file_append(own, dir_path, base->dir_merge_names[i], rel, &opts, &exists, err,
err_size))
goto fail;
if (exists && any_exists)
*any_exists = true;
}
}
return own;
fail:
filter_rule_list_free(own);
return NULL;
}
/* Merge the open directory's own per-directory filter files (the default
* .rsync-filter when -F is active, plus every "dir-merge NAME" registered on the
* base rule list) into the inherited context, returning the context used for
* this directory's entries. On a parse error the scanner is marked failed.
* Returns 0 on success, -1 on failure. */
int open_directory_filter_context(DirectoryScanner* scanner, const FilterNode* inherited) {
char err[256];
bool any_exists = false;
FilterRuleList* own = read_dir_filters(&scanner->options, scanner->current_path,
scanner->current_rel ? scanner->current_rel : "",
&any_exists, err, sizeof(err));
if (!own) {
/* read_dir_filters() leaves `err` set on a parse/allocation failure even
when an earlier merge file in the same directory existed (any_exists true);
key off the error text rather than any_exists so an invalid per-directory
filter file can never be silently ignored. */
if (err[0] == '\0') {
scanner->current_node = (FilterNode*)inherited;
return 0;
}
char* escaped_path = output_escape(scanner->current_path, log_get_8_bit_output());
log_message(LOG_LEVEL_ERROR, "invalid per-directory filter in %s: %s",
escaped_path ? escaped_path : "<allocation failed>", err);
free(escaped_path);
scanner->failed = true;
return -1;
}
if (any_exists && (own->count > 0 || own->dir_merge_count > 0)) {
FilterNode* node = filter_node_alloc((FilterNode*)inherited, own);
if (!node || !array_list_add(scanner->filter_nodes, node)) {
filter_node_destroy(node);
scanner->failed = true;
return -1;
}
scanner->current_node = node;
} else {
filter_rule_list_free(own);
scanner->current_node = (FilterNode*)inherited;
}
return 0;
}
/* Inspect symlinks, resolve the entry type, and apply file filters once for both scanners.
* `link_rel` is the entry's path relative to the transfer root (including its
* name), used for the lexical rsync unsafe-symlink test. */
int scanner_inspect_entry(const ScannerOptions* options, const char* containing_dir,
const char* link_rel, const char* name, ScannerEntry* entry) {
entry->excluded = false;
entry->size_excluded = false;
entry->referent_error = false;
entry->is_symlink = false;
entry->link_target = NULL;
entry->path = path_cat(containing_dir, name);
if (!entry->path)
return -1;
struct stat link_stats;
if (lstat(entry->path, &link_stats) != 0) {
free(entry->path);
return 0;
}
if (!S_ISLNK(link_stats.st_mode))
goto regular;
char link_target[4096];
switch (scanner_link_action(options, entry->path, link_rel, link_target, sizeof(link_target))) {
case LINK_ACTION_SKIP:
goto skip;
case LINK_ACTION_SKIP_PROTECTED:
/* --safe-links ignored the link, but rsync still counts it as present in
the transfer, so its destination mirror survives --delete. Record it as
an excluded path (the same delete-protection channel as a filter prune). */
entry->excluded = true;
goto skip;
case LINK_ACTION_DEREF:
if (stat(entry->path, &entry->stats) != 0) {
/* rsync reports "symlink has no referent" and continues with a partial
transfer (exit 23); record the error so the run exits 23 too. */
char* escaped = output_escape(entry->path, log_get_8_bit_output());
log_message(LOG_LEVEL_WARNING, "symlink has no referent: %s",
escaped ? escaped : "<allocation failed>");
free(escaped);
entry->referent_error = true;
goto skip;
}
entry->is_directory = S_ISDIR(entry->stats.st_mode);
if (entry->is_directory)
return 1;
goto apply_filters;
case LINK_ACTION_CARRY:
break;
}
/* Carry the link as a symlink. --munge-links is applied by the RECEIVER (it
prefixes every stored target with /rsyncd-munged/); when the SOURCE already
holds a munged value the sender strips it so the receiver re-munges a clean
target, round-tripping a munged tree exactly like rsync. */
entry->is_symlink = true;
entry->stats = link_stats;
entry->is_directory = false;
entry->link_target = str_dup(link_target);
if (!entry->link_target)
goto skip;
if (options->munge_links)
file_symlink_unmunge(entry->link_target);
goto apply_filters;
regular:
/* Not a symlink: the lstat() above already described this entry, and lstat
and stat are identical for every non-symlink, so reuse that result instead
of issuing a redundant stat() on the scanner hot path. stat() is still
used on the dereference paths above/below for actual symlinks (copy-links,
safe/copy-unsafe links, and -k symlinks-to-directories). */
entry->stats = link_stats;
entry->is_directory = S_ISDIR(link_stats.st_mode);
if (entry->is_directory)
return 1;
apply_filters:
for (int i = 0; i < options->exclude_count; i++)
if (glob_match(options->exclude_patterns[i], name)) {
entry->excluded = true;
goto skip;
}
if (options->include_count > 0) {
bool included = false;
for (int i = 0; i < options->include_count; i++)
if (glob_match(options->include_patterns[i], name))
included = true;
if (!included) {
entry->excluded = true;
goto skip;
}
}
if ((options->max_size > 0 && (unsigned long long)entry->stats.st_size > options->max_size) ||
(options->min_size > 0 && (unsigned long long)entry->stats.st_size < options->min_size)) {
entry->excluded = true;
entry->size_excluded = true;
goto skip;
}
return 1;
skip:
free(entry->path);
entry->path = NULL;
free(entry->link_target);
entry->link_target = NULL;
return 0;
}
+107
View File
@@ -0,0 +1,107 @@
#ifndef SCANNER_INTERNAL_H
#define SCANNER_INTERNAL_H
/* Internal declarations shared between the scanner translation units
* (scanner_filter.c, scanner.c, scanner_parallel.c). Nothing here is part of
* the public scanner façade (scanner.h); every symbol stays internal to the
* client module. */
#include "array_list.h"
#include "file.h"
#include "scanner.h"
#include <stdbool.h>
#include <stddef.h>
#include <sys/stat.h>
typedef struct {
char* path;
int depth;
FilterNode* context; /* inherited per-directory filter context */
} DirEntry;
/* How rsync's readlink_stat()/generator resolves one source symlink. */
typedef enum {
LINK_ACTION_SKIP, /* not transferred (no link option) */
LINK_ACTION_SKIP_PROTECTED, /* ignored as unsafe by --safe-links; rsync keeps
it in the transfer, so its destination mirror
must be protected from --delete */
LINK_ACTION_DEREF, /* follow the referent (--copy-links, an unsafe
target under --copy-unsafe-links, or -k dir) */
LINK_ACTION_CARRY, /* transmit the link itself (-l) */
} LinkAction;
typedef struct {
char* path;
struct stat stats;
bool is_directory;
/* True when the entry should be carried through as a SYMLINK (is_symlink)
rather than a dereferenced file/directory. When true, `link_target` holds
the owned target string to transmit (sender-munged under --munge-links);
ownership transfers to the File built from this entry. */
bool is_symlink;
char* link_target;
/* True when the entry was pruned by a user selection rule (--filter/-C/per-dir
rules or the --exclude/--include layer) rather than skipped for another
reason (unreadable, symlink policy, not applicable). */
bool excluded;
/* True when the entry was skipped specifically by --max-size/--min-size.
Size pruning protects the destination mirror even under --delete-excluded,
so it is recorded into a separate sink from `excluded`. */
bool size_excluded;
/* True when a symlink selected for dereferencing (-L/--copy-links or an
unsafe target under --copy-unsafe-links) had no usable referent (a broken
link or a stat() failure). rsync still reports this as a partial transfer
(exit 23) even though the entry is skipped, so the scanner records it as a
non-fatal I/O error. */
bool referent_error;
} ScannerEntry;
typedef enum {
SCANNER_SPECIAL_REGULAR, /* ordinary file: transfer content */
SCANNER_SPECIAL_RECREATE, /* is_special node to recreate on the receiver */
SCANNER_SPECIAL_SKIP, /* non-regular entry not requested: skip */
} ScannerSpecial;
/* scanner_filter.c */
void filter_node_destroy(void* item);
FilterNode* filter_node_alloc(FilterNode* parent, FilterRuleList* own);
void dir_entry_destroy(void* item);
DirEntry* dir_entry_create(const char* path, int depth, FilterNode* context);
LinkAction scanner_link_action(const ScannerOptions* options, const char* path,
const char* link_rel, char* target, size_t target_size);
File* scanner_build_dir_file(const char* path, const struct stat* stats,
const ScannerOptions* options);
char* child_rel_path(const char* parent_rel, const char* name);
char* scanner_prefix_send_path(const char* prefix, const char* rel);
bool entry_passes_selection(const FileListSet* file_list, const FilterRuleList* base,
const FilterNode* node, const char* rel, const char* leaf, bool is_dir,
bool per_dir_filters, bool exclude_filter_files, bool* protect_out);
void scanner_capture_xattrs(const DirectoryScanner* scanner, File* file);
void scanner_assign_hardlink(DirectoryScanner* scanner, HardLinkTable* table, File* file,
const struct stat* stats);
ScannerSpecial scanner_prepare_special(bool preserve_devices, bool preserve_specials,
bool copy_devices, File* file, const struct stat* stats);
bool excluded_sink_append(ArrayList* list, mtx_t* mtx, const char* rel);
void scanner_note_nonreg(const ScannerOptions* options, const char* fs_path);
void scanner_note_mount(const ScannerOptions* options, const char* fs_path);
void scanner_note_filter(const ScannerOptions* options, const char* name);
void scanner_dir_count_count(const ScannerOptions* options);
void scanner_dir_count_uncount(const ScannerOptions* options);
void scanner_record_excluded(DirectoryScanner* scanner, const char* fs_path);
void scanner_record_size_skipped(DirectoryScanner* scanner, const char* fs_path);
bool scanner_record_synced_dir(const ScannerOptions* options, const char* fs_path, const char* rel,
bool relative_mode);
FilterRuleList* read_dir_filters(const ScannerOptions* options, const char* dir_path,
const char* rel, bool* any_exists, char* err, size_t err_size);
int open_directory_filter_context(DirectoryScanner* scanner, const FilterNode* inherited);
int scanner_inspect_entry(const ScannerOptions* options, const char* containing_dir,
const char* link_rel, const char* name, ScannerEntry* entry);
/* scanner.c */
bool scanner_capture_dir_time(ArrayList* dir_entries, mtx_t* mutex, const char* root_path,
const char* fs_path, bool relative_mode, const char* relative_prefix,
bool preserve_atimes, bool preserve_crtimes, bool preserve_xattrs,
bool preserve_acls, bool no_implied_dirs,
const FileListSet* file_list);
#endif
+705
View File
@@ -0,0 +1,705 @@
#include "log.h"
#include "scanner.h"
#include "scanner_internal.h"
#include "array_list.h"
#include "chunk.h"
#include "file.h"
#include "queue.h"
#include "utils.h"
#include <dirent.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/stat.h>
#include <sys/sysmacros.h>
#include <threads.h>
#include <unistd.h>
#include <limits.h>
#include "xattr.h"
typedef struct {
ParallelScanner* ps;
char** dirs;
int dir_count;
char* root_dir; /* the transfer root, for relative-path computation */
ScannerOptions options;
ProtocolSession* allocation_session;
} ParallelWorkerArg;
static int parallel_worker_thread(void* arg) {
ParallelWorkerArg* wa = (ParallelWorkerArg*)arg;
ProtocolSession* allocation_session = wa->allocation_session;
if (allocation_session)
protocol_session_bind(allocation_session);
for (int i = 0; i < wa->dir_count; i++) {
DirectoryScanner* ds = directory_scanner_create_with_options(wa->dirs[i], &wa->options);
if (!ds) {
mtx_lock(&wa->ps->result_mutex);
wa->ps->failed = true;
atomic_store(&wa->ps->cancelled, true);
cnd_broadcast(&wa->ps->result_not_empty);
cnd_broadcast(&wa->ps->result_not_full);
mtx_unlock(&wa->ps->result_mutex);
for (int j = i; j < wa->dir_count; j++)
free(wa->dirs[j]);
break;
}
/* Root .rsync-filter rules (parsed by the parallel scanner) apply to the
* contents of every assigned subdirectory. Relative paths (used by the
* allow-set and per-directory rules) are computed against the transfer
* root, not the subdirectory the worker is seeded with. Exclusion
* recording shares one caller-owned list across the workers. */
free(ds->root_path);
ds->root_path = str_dup(wa->root_dir);
ds->seed_node = wa->ps->root_filter_node;
ds->options.excluded_mutex = &wa->ps->result_mutex;
Chunk* chunk;
while ((chunk = directory_scanner_next(ds)) != NULL) {
if (!queue_enqueue_multithreaded_cancel(wa->ps->result_queue, chunk, &wa->ps->result_mutex,
&wa->ps->result_not_empty, &wa->ps->result_not_full,
&wa->ps->cancelled)) {
chunk_destroy(chunk);
break;
}
}
if (directory_scanner_failed(ds)) {
mtx_lock(&wa->ps->result_mutex);
wa->ps->failed = true;
atomic_store(&wa->ps->cancelled, true);
cnd_broadcast(&wa->ps->result_not_empty);
cnd_broadcast(&wa->ps->result_not_full);
mtx_unlock(&wa->ps->result_mutex);
} else if (directory_scanner_had_io_error(ds)) {
/* --ignore-errors path: an unreadable directory was skipped, not fatal. */
mtx_lock(&wa->ps->result_mutex);
wa->ps->io_error = true;
mtx_unlock(&wa->ps->result_mutex);
}
directory_scanner_destroy(ds);
free(wa->dirs[i]);
}
ParallelScanner* ps = wa->ps;
free(wa->root_dir);
free(wa->dirs);
free(wa);
mtx_lock(&ps->result_mutex);
ps->completed++;
if (ps->completed >= ps->expected_threads) {
ps->done = true;
cnd_signal(&ps->result_not_empty);
}
mtx_unlock(&ps->result_mutex);
if (allocation_session)
protocol_session_unbind();
return thrd_success;
}
static void parallel_scanner_creation_failed(ParallelScanner* ps) {
mtx_lock(&ps->result_mutex);
ps->failed = true;
atomic_store(&ps->cancelled, true);
ps->expected_threads = ps->created_threads;
if (ps->completed >= ps->expected_threads)
ps->done = true;
cnd_broadcast(&ps->result_not_empty);
cnd_broadcast(&ps->result_not_full);
mtx_unlock(&ps->result_mutex);
}
/* Initialize result queue and synchronization primitives. Returns true on success. */
static bool parallel_scanner_init(ParallelScanner* ps) {
ps->result_queue = queue_create(100, chunk_destroy);
if (!ps->result_queue)
return false;
atomic_init(&ps->cancelled, false);
int init = 0;
bool ok = true;
if (mtx_init(&ps->result_mutex, mtx_plain) != thrd_success)
ok = false;
if (ok) {
init++;
if (cnd_init(&ps->result_not_empty) != thrd_success)
ok = false;
}
if (ok) {
// cppcheck-suppress unreadVariable
init++;
if (cnd_init(&ps->result_not_full) != thrd_success)
ok = false;
}
if (!ok) {
if (init >= 3)
cnd_destroy(&ps->result_not_full);
if (init >= 2)
cnd_destroy(&ps->result_not_empty);
if (init >= 1)
mtx_destroy(&ps->result_mutex);
queue_destroy(ps->result_queue);
ps->result_queue = NULL;
return false;
}
return true;
}
/* Split files into chunks of roughly chunk_size bytes. Returns the first chunk (also stored
* chunks beyond the first are enqueued on `queue`). Nulls out consumed entries in `files`.
* Sets *failed on allocation/enqueue errors. */
static Chunk* batch_files(ArrayList* files, unsigned long long chunk_size, Queue* queue,
bool* failed) {
Chunk* first = NULL;
if (files->size <= 0)
return NULL;
ArrayList* batch = array_list_create(NULL);
if (!batch) {
*failed = true;
return NULL;
}
unsigned long long batch_size = 0;
for (int i = 0; i < files->size; i++) {
File* f = (File*)files->items[i];
if (!array_list_add(batch, f)) {
*failed = true;
break;
}
batch_size += f->data->size;
if (batch_size >= chunk_size || i == files->size - 1) {
void** items = array_list_to_array(batch);
if (!items) {
*failed = true;
array_list_delete(batch);
batch = NULL;
break;
}
Chunk* c = chunk_create((File**)items, batch->size);
free(items);
if (!c) {
*failed = true;
array_list_delete(batch);
batch = NULL;
break;
}
int batch_start = i - batch->size + 1;
for (int j = batch_start; j <= i; j++)
files->items[j] = NULL;
batch->item_destroyer = NULL;
array_list_delete(batch);
batch = NULL;
if (!first) {
first = c;
} else {
if (!queue_enqueue(queue, c)) {
chunk_destroy(c);
*failed = true;
}
}
if (i < files->size - 1) {
batch = array_list_create(NULL);
if (!batch) {
*failed = true;
break;
}
batch_size = 0;
}
}
}
if (batch) {
batch->item_destroyer = NULL;
array_list_delete(batch);
}
return first;
}
/* Scan one root-directory entry into either the subdirs or files list. */
static void scan_root_entry(const ScannerOptions* options, const FilterNode* root_node,
const char* root_directory, const struct dirent* entry,
ArrayList* root_files, ArrayList* subdirs, dev_t root_dev,
ParallelScanner* ps) {
ScannerEntry inspected;
int inspection =
scanner_inspect_entry(options, root_directory, entry->d_name, entry->d_name, &inspected);
if (inspection < 0) {
ps->failed = true;
return;
}
if (inspection == 0) {
if (inspected.referent_error)
ps->io_error = true;
ArrayList* sink = NULL;
if (inspected.excluded)
sink = inspected.size_excluded ? options->size_skipped_paths : options->excluded_paths;
if (sink) {
/* A root-level prune protects the destination mirror of the entry's wire
path: under -R + --files-from that is the bare relative name, otherwise
it is the full source path with a leading '/' removed (matching the
send_path/file_wire_path the scanner hands the sender). */
if (options->relative && options->file_list != NULL) {
if (!excluded_sink_append(sink, options->excluded_mutex, entry->d_name))
ps->failed = true;
} else if (options->relative_prefix) {
char* wrel = scanner_prefix_send_path(options->relative_prefix, entry->d_name);
if (!wrel) {
ps->failed = true;
} else {
if (!excluded_sink_append(sink, options->excluded_mutex, wrel))
ps->failed = true;
free(wrel);
}
} else {
char* abs_path = path_cat(root_directory, entry->d_name);
if (!abs_path) {
ps->failed = true;
} else {
const char* rel = *abs_path == '/' ? abs_path + 1 : abs_path;
if (!excluded_sink_append(sink, options->excluded_mutex, rel))
ps->failed = true;
free(abs_path);
}
}
}
return;
}
char* cur_path = inspected.path;
struct stat st = inspected.stats;
bool is_dir = inspected.is_directory;
char* rel = str_dup(entry->d_name);
if (!rel) {
free(cur_path);
ps->failed = true;
return;
}
bool protect = false;
bool passes = entry_passes_selection(options->file_list, options->base_filters, root_node, rel,
entry->d_name, is_dir, options->per_dir_filters,
options->exclude_per_dir_filter_files, &protect);
/* -R + --files-from: root-level files keep their bare relative send path. */
bool use_rel = options->relative && options->file_list != NULL;
if (!passes || protect) {
/* --files-from subset pruning is not a filter exclusion; -R bare-wire-path
exclusions are never recorded (see ScannerOptions.excluded_paths). */
bool files_from_prune = options->file_list && !file_list_affects(options->file_list, rel);
if ((!files_from_prune && !use_rel) || protect) {
const char* rel_path;
char* prefixed = NULL;
if (use_rel) {
/* -R + --files-from: the destination/wire path is the bare relative
name, not the source path. */
rel_path = rel;
} else if (options->relative_prefix) {
prefixed = scanner_prefix_send_path(options->relative_prefix, entry->d_name);
if (!prefixed) {
free(rel);
free(cur_path);
ps->failed = true;
return;
}
rel_path = prefixed;
} else {
rel_path = *cur_path == '/' ? cur_path + 1 : cur_path;
}
if (options->excluded_paths &&
!excluded_sink_append(options->excluded_paths, options->excluded_mutex, rel_path))
ps->failed = true;
free(prefixed);
}
if (!passes) {
scanner_note_filter(options, entry->d_name);
free(rel);
free(cur_path);
return;
}
}
if (is_dir) {
if (!scanner_same_filesystem(options->one_file_system, root_dev, st.st_dev)) {
if (options->one_file_system > 1) {
/* -xx: drop the mount-point directory entirely (rsync) and print the
--info=mount line when enabled. */
scanner_note_mount(options, cur_path);
free(rel);
free(cur_path);
return;
}
/* -x/--one-file-system: emit the mount-point directory entry (empty) but
do not descend into it (see the sequential scanner for the same rule). */
File* mount = file_create(cur_path);
free(cur_path);
if (mount == NULL) {
free(rel);
ps->failed = true;
return;
}
mount->is_dir = true;
if (options->use_metadata) {
mount->metadata = file_metadata_create(mount->path, &st, options->preserve_atimes,
options->preserve_crtimes);
if (!mount->metadata) {
free(rel);
file_destroy(mount);
ps->failed = true;
return;
}
}
if (options->relative_prefix) {
mount->send_path = scanner_prefix_send_path(options->relative_prefix, rel);
if (!mount->send_path) {
free(rel);
file_destroy(mount);
ps->failed = true;
return;
}
}
free(rel);
if (!array_list_add(root_files, mount)) {
file_destroy(mount);
ps->failed = true;
}
return;
}
free(rel);
if (!array_list_add(subdirs, cur_path)) {
free(cur_path);
ps->failed = true;
}
return;
}
File* file = file_create(cur_path);
free(cur_path);
if (!file) {
free(rel);
free(inspected.link_target);
inspected.link_target = NULL;
ps->failed = true;
return;
}
if (inspected.is_symlink) {
file->is_symlink = true;
file->symlink_target = inspected.link_target;
inspected.link_target = NULL;
} else {
file->data->size = st.st_size;
}
if (use_rel) {
file->send_path = rel;
rel = NULL;
} else if (options->relative_prefix) {
file->send_path = scanner_prefix_send_path(options->relative_prefix, rel);
free(rel);
rel = NULL;
if (!file->send_path) {
file_destroy(file);
ps->failed = true;
return;
}
}
ScannerSpecial special = scanner_prepare_special(
options->preserve_devices, options->preserve_specials, options->copy_devices, file, &st);
if (special == SCANNER_SPECIAL_SKIP) {
scanner_note_nonreg(ps->options, file->path);
free(rel);
file_destroy(file);
return;
}
if (options->hardlinks && S_ISREG(st.st_mode)) {
int gid;
bool is_first;
char* first_path = NULL;
if (!hardlink_table_assign((HardLinkTable*)options->hardlinks, file_wire_path(file), st.st_dev,
st.st_ino, &gid, &is_first, &first_path)) {
ps->failed = true;
} else {
file->link_group = gid;
file->link_first = is_first;
if (!is_first) {
file->hardlink_target = first_path;
file->data->size = 0;
} else {
free(first_path);
}
}
}
if (options->use_metadata)
file->metadata =
file_metadata_create(file->path, &st, options->preserve_atimes, options->preserve_crtimes);
if (options->use_metadata && !file->metadata) {
free(rel);
file_destroy(file);
ps->failed = true;
return;
}
if ((options->preserve_xattrs || options->preserve_acls) &&
!(file->link_group != 0 && !file->link_first))
file->xattrs = file->is_symlink
? xattr_capture_path_nofollow(file->path, options->preserve_acls)
: xattr_capture_path(file->path, options->preserve_acls);
if (!array_list_add(root_files, file)) {
free(rel);
file_destroy(file);
ps->failed = true;
return;
}
free(rel);
}
/* Scan the root directory itself, collecting root files and subdirectories.
* Returns false if the root directory could not be opened. */
static bool scan_root_directory(ParallelScanner* ps, const char* root_directory,
const ScannerOptions* options, const FilterNode* root_node,
dev_t root_dev, ArrayList* root_files, ArrayList* subdirs) {
DIR* dir = opendir(root_directory);
if (!dir) {
log_perror("Could not open root directory for parallel scan");
return false;
}
/* The parallel scanner opens the transfer root directly (not through
open_next_directory), so record it as synchronized here. */
if (!scanner_record_synced_dir(options, root_directory, "",
options->relative && options->file_list != NULL)) {
closedir(dir);
ps->failed = true;
return false;
}
log_debug_message(LOG_DEBUG_FLIST, "flist: scanning %s", root_directory);
const struct dirent* entry;
while ((entry = readdir(dir)) != NULL) {
if (strcmp(entry->d_name, ".") == 0 || strcmp(entry->d_name, "..") == 0)
continue;
scan_root_entry(options, root_node, root_directory, entry, root_files, subdirs, root_dev, ps);
}
closedir(dir);
return true;
}
/* Spawn worker threads, one per group of subdirectories. */
static void spawn_parallel_workers(ParallelScanner* ps, ArrayList* subdirs,
const ScannerOptions* options, const char* root_directory,
unsigned long long cs) {
if (subdirs->size <= 0)
return;
int n = options->num_threads > 0 ? options->num_threads : 4;
if (n > subdirs->size)
n = subdirs->size;
ps->num_threads = n;
ps->expected_threads = n;
ps->threads = calloc(n, sizeof(thrd_t));
if (!ps->threads) {
ps->num_threads = 0;
ps->expected_threads = 0;
ps->failed = true;
return;
}
int dirs_per_thread = subdirs->size / n;
int remainder = subdirs->size % n;
int start = 0;
ps->num_threads = 0;
for (int t = 0; t < n; t++) {
int count = dirs_per_thread + (t < remainder ? 1 : 0);
if (count == 0)
break;
ParallelWorkerArg* wa = calloc(1, sizeof(ParallelWorkerArg));
if (!wa) {
parallel_scanner_creation_failed(ps);
break;
}
wa->ps = ps;
wa->dirs = calloc(count, sizeof(char*));
wa->root_dir = str_dup(root_directory);
if (!wa->dirs || !wa->root_dir) {
free(wa->root_dir);
free(wa->dirs);
free(wa);
parallel_scanner_creation_failed(ps);
break;
}
bool dup_ok = true;
for (int j = 0; j < count; j++) {
wa->dirs[j] = str_dup((char*)subdirs->items[start + j]);
if (!wa->dirs[j])
dup_ok = false;
}
if (!dup_ok) {
for (int j = 0; j < count; j++)
free(wa->dirs[j]);
free(wa->root_dir);
free(wa->dirs);
free(wa);
parallel_scanner_creation_failed(ps);
break;
}
wa->dir_count = count;
wa->options = *options;
wa->options.chunk_size = cs;
wa->allocation_session = ps->allocation_session;
start += count;
if (thrd_create(&ps->threads[t], parallel_worker_thread, wa) != thrd_success) {
for (int j = 0; j < count; j++)
free(wa->dirs[j]);
free(wa->root_dir);
free(wa->dirs);
free(wa);
parallel_scanner_creation_failed(ps);
break;
}
ps->num_threads++;
ps->created_threads++;
}
}
ParallelScanner* parallel_scanner_create_with_options(const char* root_directory,
const ScannerOptions* options,
ProtocolSession* allocation_session) {
if (!root_directory || !options)
return NULL;
ParallelScanner* ps = calloc(1, sizeof(ParallelScanner));
if (!ps)
return NULL;
if (!parallel_scanner_init(ps)) {
free(ps);
return NULL;
}
ps->allocation_session = allocation_session;
ps->options = options;
ArrayList* root_files = array_list_create(file_destroy);
ArrayList* subdirs = array_list_create(free);
if (!root_files || !subdirs) {
array_list_delete(root_files);
array_list_delete(subdirs);
parallel_scanner_destroy(ps);
return NULL;
}
dev_t root_dev = 0;
if (options->one_file_system) {
struct stat root_stats;
if (stat(root_directory, &root_stats) != 0) {
log_perror("Could not stat source directory");
array_list_delete(root_files);
array_list_delete(subdirs);
parallel_scanner_destroy(ps);
return NULL;
}
root_dev = root_stats.st_dev;
}
/* Build the root directory's per-directory filter context once; workers seed
* their scanners with it so per-dir rules behave identically to the sequential
* scanner. */
FilterNode* root_node = NULL;
{
char err[256];
bool any_exists = false;
FilterRuleList* own =
read_dir_filters(options, root_directory, "", &any_exists, err, sizeof(err));
if (!own) {
/* A parse/allocation failure must fail the scan even when an earlier
merge file in the same directory existed (see the sequential scanner). */
if (err[0] != '\0') {
log_message(LOG_LEVEL_ERROR, "invalid per-directory filter in %s: %s", root_directory, err);
array_list_delete(root_files);
array_list_delete(subdirs);
parallel_scanner_destroy(ps);
return NULL;
}
/* no files exist: leave root_node NULL */
} else if (any_exists && (own->count > 0 || own->dir_merge_count > 0)) {
root_node = filter_node_alloc(NULL, own);
if (!root_node) {
filter_rule_list_free(own);
array_list_delete(root_files);
array_list_delete(subdirs);
parallel_scanner_destroy(ps);
return NULL;
}
} else {
filter_rule_list_free(own);
}
}
ps->root_filter_node = root_node;
if (!scan_root_directory(ps, root_directory, options, root_node, root_dev, root_files, subdirs)) {
array_list_delete(root_files);
array_list_delete(subdirs);
parallel_scanner_destroy(ps);
return NULL;
}
/* The root itself is a traversed directory (rsync counts it in
`Number of files`); the worker DirectoryScanners account for every
subdirectory below it. */
scanner_dir_count_count(options);
/* P7 Wave D: the parallel scanner never runs a DirectoryScanner over the
transfer root itself (it hands the root's immediate subdirectories to
workers), so capture the root's directory time here. */
if (options->capture_dir_times &&
!scanner_capture_dir_time(
options->dir_entries, options->dir_entries_mutex, root_directory, root_directory,
options->relative && options->file_list != NULL, options->relative_prefix,
options->preserve_atimes, options->preserve_crtimes, options->preserve_xattrs,
options->preserve_acls, options->no_implied_dirs, options->file_list)) {
array_list_delete(root_files);
array_list_delete(subdirs);
parallel_scanner_destroy(ps);
return NULL;
}
unsigned long long cs = options->chunk_size > 0 ? options->chunk_size : DESIRED_CHUNK_SIZE;
ps->initial_chunk = batch_files(root_files, cs, ps->result_queue, &ps->failed);
array_list_delete(root_files);
spawn_parallel_workers(ps, subdirs, options, root_directory, cs);
array_list_delete(subdirs);
return ps;
}
Chunk* parallel_scanner_next(ParallelScanner* ps) {
if (ps->initial_chunk) {
Chunk* c = ps->initial_chunk;
ps->initial_chunk = NULL;
return c;
}
if (ps->num_threads == 0) {
mtx_lock(&ps->result_mutex);
if (!queue_is_empty(ps->result_queue)) {
Chunk* chunk = queue_dequeue(ps->result_queue);
mtx_unlock(&ps->result_mutex);
return chunk;
}
ps->done = true;
mtx_unlock(&ps->result_mutex);
return NULL;
}
Chunk* chunk = queue_dequeue_multithreaded(
ps->result_queue, &ps->result_mutex, &ps->result_not_empty, &ps->result_not_full, &ps->done);
return chunk;
}
bool parallel_scanner_failed(const ParallelScanner* ps) {
return ps == NULL || ps->failed;
}
bool parallel_scanner_had_io_error(const ParallelScanner* ps) {
return ps != NULL && ps->io_error;
}
void parallel_scanner_destroy(ParallelScanner* ps) {
if (!ps)
return;
mtx_lock(&ps->result_mutex);
ps->done = true;
atomic_store(&ps->cancelled, true);
cnd_broadcast(&ps->result_not_empty);
cnd_broadcast(&ps->result_not_full);
mtx_unlock(&ps->result_mutex);
for (int i = 0; i < ps->num_threads; i++)
thrd_join(ps->threads[i], NULL);
free(ps->threads);
if (ps->root_filter_node)
filter_node_destroy(ps->root_filter_node);
if (ps->initial_chunk)
chunk_destroy(ps->initial_chunk);
queue_destroy(ps->result_queue);
mtx_destroy(&ps->result_mutex);
cnd_destroy(&ps->result_not_empty);
cnd_destroy(&ps->result_not_full);
free(ps);
}
+49 -25
View File
@@ -19,11 +19,16 @@ void print_usage(void) {
printf("\n");
printf("Options:\n");
printf(" -c, --checksum Verify content by checksum instead of size+mtime\n");
printf(" -z, --compress [level] Enable compression (level 1-22, default 5)\n");
printf(" -z, --compress [level] Enable compression. The default level is\n");
printf(" per-codec: zstd 3 (range 1-22), zlib/zlibx 6, lz4\n");
printf(" ignores the level\n");
printf(" -a, --archive rsync archive mode (-rlptgoD): links, perms, times,\n");
printf(" owner, group, devices and specials; not\n");
printf(" compression/multithreading\n");
printf(" -r, --recursive Recurse into directories (FastSync is always recursive)\n");
printf(" --inc-recursive Accepted for rsync CLI compatibility; no effect (FastSync\n");
printf(" always performs a full scan, so the destination is identical)\n");
printf(" --no-inc-recursive Accepted for rsync CLI compatibility; no effect\n");
printf(" -n, --dry-run Show what would be transferred\n");
printf(" --remove-source-files Remove regular source files after successful transfer\n");
printf(" -p, --perms Preserve permission bits\n");
@@ -36,8 +41,8 @@ void print_usage(void) {
printf(" arguments, e.g. -e \"ssh -p 2222\"\n");
printf(" --rsync-path <path> Alias for --fastsync-server-path (path to the\n");
printf(" fastsync server binary on the remote side)\n");
printf(" --blocking-io Leave the SSH transport socket without read/write\n");
printf(" timeouts so it blocks naturally\n");
printf(" --blocking-io SSH transport only: leave the socket without read/write\n");
printf(" timeouts so it blocks naturally (no effect on TCP)\n");
printf(" --outbuf=MODE stdout/stderr buffering: N (none/unbuffered),\n");
printf(" L (line-buffered), or B (block-buffered, default)\n");
printf(" --progress Show transfer progress\n");
@@ -49,6 +54,7 @@ void print_usage(void) {
printf(" converted before transmission and back on receipt; a\n");
printf(" name that cannot be represented in the target charset\n");
printf(" fails that transfer cleanly (rsync-compatible)\n");
printf(" --no-iconv Disable --iconv charset conversion (same as --iconv=-)\n");
printf(" --protocol=NUM Force the wire protocol version (must equal the current\n");
printf(" PROTOCOL_VERSION; FastSync cannot speak older/virtual\n");
printf(" wire formats)\n");
@@ -62,8 +68,7 @@ void print_usage(void) {
printf(" NOTE: the FastSync batch format is NOT interoperable with rsync's batch\n");
printf(" files (different container format); do not mix the two tools.\n");
printf(" --delete Delete files on receiver not in source\n");
printf(" (default timing: delete only after the whole\n");
printf(" transfer has succeeded)\n");
printf(" (default timing: delete-during, like rsync --del)\n");
printf(" --delete-before Delete extras before the transfer starts\n");
printf(" (implies --delete)\n");
printf(" --delete-during Delete a directory's extras as that directory is\n");
@@ -72,7 +77,9 @@ void print_usage(void) {
printf(" --delete-delay Record the extras during the scan but remove them\n");
printf(" only after a successful transfer (implies --delete)\n");
printf(" --delete-after Delete only after the whole transfer succeeded\n");
printf(" (the default --delete timing; implies --delete)\n");
printf(" (implies --delete)\n");
printf(" --delete-commit FastSync-only: restore the late whole-tree commit\n");
printf(" (identical to --delete-after; implies --delete)\n");
printf(" --delete-excluded Also delete destination files that were excluded on\n");
printf(" the source (default protects them, matching rsync)\n");
printf(" --max-delete=NUM Delete at most NUM destination entries per run; if the\n");
@@ -90,10 +97,11 @@ void print_usage(void) {
printf(" entry's destination mirror receiver-side. Independent of\n");
printf(" --delete (it does not imply --delete; a non-empty directory\n");
printf(" mirror is removed only with --force or --delete)\n");
printf(" -m, --prune-empty-dirs Do not transfer empty directory entries (--dirs mode);\n");
printf(" recursive transfers never send empty dirs\n");
printf(" -m, --prune-empty-dirs Do not create empty directories (a recursive transfer\n");
printf(" otherwise recreates them, like rsync)\n");
printf(" Note: each timing flag implies --delete. Combining a timing flag with\n");
printf(" --no-delete (in either order) is rejected as a config error.\n");
printf(" --no-delete (in either order) is rejected as a config error, as is more\n");
printf(" than one timing flag.\n");
printf(" --ignore-existing Skip files that already exist on receiver\n");
printf(" --delay-updates Put updated files into place only at the end of transfer\n");
printf(" --dirs, -d, --old-dirs, --old-d Transfer the named directory entries without\n");
@@ -103,8 +111,9 @@ void print_usage(void) {
printf(" -R, --relative With --files-from, preserve each listed entry's relative path\n");
printf(" below the destination root instead of mirroring the full\n");
printf(" source path (no effect without --files-from)\n");
printf(" --no-implied-dirs With -R --files-from, refuse to place a listed file whose\n");
printf(" parent directory is not itself listed\n");
printf(" --no-implied-dirs With -R, do not apply the source metadata of a listed file's\n");
printf(" implied parent directories (they are still created with\n");
printf(" default attributes)\n");
printf(" --mkpath Create the destination root directory on the server when it\n");
printf(" does not exist yet\n");
printf(" --exclude <pattern>, --exclude=<pattern> Exclude files matching pattern\n");
@@ -137,6 +146,9 @@ void print_usage(void) {
printf(" into the destination instead of transferring its data\n");
printf(" --link-dest <dir> Like --copy-dest, but hard-links the unchanged file from DIR\n");
printf(" into the destination (repeatable; earlier DIRs win)\n");
printf(" --verify-basis FastSync-only: require a basis hit's content to match the\n");
printf(" source by whole-file digest instead of trusting rsync's\n");
printf(" size+mtime (or --size-only) quick-check\n");
printf(" --checksum-choice, --cc <alg> Whole-file checksum algorithm for --incremental/\n");
printf(" --checksum compares. Accepted: xxh128 (default), xxh3, xxh64\n");
printf(" (aka xxhash), md5, md4, sha1, or none. A two-name\n");
@@ -156,7 +168,7 @@ void print_usage(void) {
printf(" --no-delta, or --no-incremental)\n");
printf(" --no-fuzzy Disable --fuzzy\n");
printf(" -B <n>, --block-size <n>, --delta-block <n>\n");
printf(" Delta block size in bytes (default: %d)\n", DELTA_BLOCK_SIZE_DEFAULT);
printf(" Delta block size in bytes (default: %u)\n", DELTA_BLOCK_SIZE_DEFAULT);
printf(" --delta-max <n> Max file size for delta transfer (default: %llu)\n",
DELTA_MAX_FILE_SIZE);
printf(" -j, --threads[=N] Enable the multithreaded scanner/loader/sender\n");
@@ -186,16 +198,21 @@ void print_usage(void) {
printf(" -U, --atimes Preserve access times\n");
printf(" -N, --crtimes Capture birth time; cannot be applied (documented\n");
printf(" divergence)\n");
printf(" -O, --omit-dir-times Do not apply modification times to directories\n");
printf(" -J, --omit-link-times Do not apply times to symlinks\n");
printf(" --open-noatime Open source files with O_NOATIME so reading for a\n");
printf(" transfer does not update their access time\n");
printf(" -X, --xattrs Preserve user extended attributes (user.* only;\n");
printf(" privileged security.*/trusted.* namespaces are\n");
printf(" never captured or applied)\n");
printf(" -A, --acls Preserve POSIX ACLs (the system.posix_acl_* xattrs;\n");
printf(" setting an ACL the receiver is not permitted to\n");
printf(" set is warned and skipped, never fatal)\n");
printf(" --fake-super Store the source uid/gid/mode/mtime in a reserved\n");
printf(" user.fastsync.stat xattr on each written file and\n");
printf(" re-apply it (fd-relative) on a privileged run; the\n");
printf(" recording format diverges from rsync's user.rsync.%%stat%%\n");
printf(" --fake-super Store the source mode/rdev/uid/gid in rsync's\n");
printf(" reserved user.rsync.%%stat xattr on each written\n");
printf(" file (interoperable with rsync); it never performs a\n");
printf(" real chown, so an unprivileged receiver records the\n");
printf(" privileged stat for a later restore\n");
printf(" --super Permit the receiver to attempt super-user activities\n");
printf(" (char/block device-node creation, --write-devices)\n");
printf(" within the confined receive root. Never elevates\n");
@@ -243,7 +260,9 @@ void print_usage(void) {
printf(" reusable digest is sent (keep the file mode 0600)\n");
printf(" --no-motd Suppress display of the daemon's MOTD (the server\n");
printf(" still sends it; the client just does not show it)\n");
printf(" --bwlimit <KB/s> Bandwidth limit in kilobytes per second\n");
printf(" --bwlimit=RATE Limit socket I/O bandwidth (default unit KiB/s,\n");
printf(" rsync-style: 0 = no limit; K/M/G/T/P suffixes are\n");
printf(" binary, KB/MB decimal, KiB/MiB binary; decimals allowed)\n");
printf(" --tls Enable TLS encryption\n");
printf(" --cert <path> TLS certificate file (PEM)\n");
printf(" --key <path> TLS private key file (PEM)\n");
@@ -279,8 +298,12 @@ void print_usage(void) {
printf(" -x, --one-file-system Do not cross filesystem boundaries\n");
printf(" --log-file <path>, --log-file=<path> Write log messages to file\n");
printf(" --stderr=MODE Route logging to stderr: errors or all\n");
printf(" --msgs2stderr Route all messages to stderr (deprecated spelling of\n");
printf(" --stderr=all)\n");
printf(" --no-msgs2stderr Select errors-only stderr (deprecated spelling; the\n");
printf(" default)\n");
printf(" --partial Keep partial files on interrupted transfer\n");
printf(" --partial-dir <dir> Directory for partial files\n");
printf(" --partial-dir <dir> Directory for partial files (implies --partial)\n");
printf(" -T, --temp-dir <dir> Scratch dir for temp files before atomic install.\n");
printf(" Confined to the receive root: a relative dir resolves below\n");
printf(" it and an absolute/traversal dir is rejected. The dir must\n");
@@ -329,7 +352,8 @@ void print_usage(void) {
printf(" --append-verify Like --append, but verifies the retained prefix checksum\n");
printf(" before appending (falls back to a full transfer on mismatch)\n");
printf(" --fsync Fsync every written file before publication\n");
printf(" --compress-level <n> Compression level (default: 5)\n");
printf(" --compress-level <n> Compression level (per-codec default: zstd 3,\n");
printf(" zlib/zlibx 6, lz4 ignores it)\n");
printf(" --zl <n> Alias for --compress-level\n");
printf(" --skip-compress=LIST Skip compression for suffixes in LIST (separated by\n");
printf(" '/' as in rsync, or ','); a leading dot is optional. The\n");
@@ -341,19 +365,19 @@ void print_usage(void) {
}
void print_debug_usage(void) {
printf("Emitting debug flags: IO,PROTO,PACK,UTIL,ALL,NONE\n");
printf("Emitting debug flags: IO,PROTO,PACK,UTIL,FLIST,DEL,HASH,DELTASUM,\n");
printf("RECV,FILTER,SEND,ALL,NONE\n");
printf("Also accepted for rsync CLI parity (silent): ACL,BACKUP,BIND,CHDIR,\n");
printf("CONNECT,CMD,DEL,DELTASUM,DUP,EXIT,FILTER,FLIST,FUZZY,GENR,HASH,HLINK,\n");
printf("ICONV,NSTR,OWN,RECV,SEND,TIME.\n");
printf("CONNECT,CMD,DUP,EXIT,FUZZY,GENR,HLINK,ICONV,NSTR,OWN,TIME.\n");
printf("Flags may be comma-separated, for example: --debug=io,proto\n");
printf("An optional level suffix is accepted (e.g. --debug=io2); level 0\n");
printf("silences that item. Unknown names are rejected.\n");
}
void print_info_usage(void) {
printf("Emitting info flags: COPY,NAME,MISC,SKIP,STATS,ALL,NONE\n");
printf("Also accepted for rsync CLI parity (silent): BACKUP,DEL,FLIST,MOUNT,\n");
printf("NONREG,PROGRESS,REMOVE,SYMSAFE.\n");
printf("Emitting info flags: COPY,MISC,SKIP,STATS,DEL,REMOVE,NAME,FLIST,\n");
printf("NONREG,PROGRESS,MOUNT,ALL,NONE\n");
printf("Also accepted for rsync CLI parity (silent): BACKUP,SYMS,SYMSAFE.\n");
printf("Flags may be comma-separated, for example: --info=name,stats\n");
printf("An optional level suffix is accepted (e.g. --info=stats2); level 0\n");
printf("silences that item. Unknown names are rejected.\n");
+394 -195
View File
@@ -53,7 +53,13 @@ bool receiver_send_final_success(int fd, const Config* config, const ReceiverOut
return send_status(fd, final_status);
size_t count = outcomes ? outcomes->count : 0;
for (size_t i = 0; i < count; i++) {
Status per_file = outcomes->entries[i] == FILE_SAVE_WRITTEN ? STATUS_NEXT : STATUS_OK;
Status per_file;
if (outcomes->entries[i] == FILE_SAVE_WRITTEN)
per_file = STATUS_NEXT;
else if (outcomes->entries[i] == FILE_SAVE_FAILED)
per_file = STATUS_ERROR;
else
per_file = STATUS_OK;
if (!send_status(fd, per_file))
return false;
}
@@ -61,13 +67,18 @@ bool receiver_send_final_success(int fd, const Config* config, const ReceiverOut
}
bool receiver_send_stats_frame(int fd, const Config* config, const ReceiverStats* stats,
const struct ArrayList* would_delete) {
const struct ArrayList* would_delete,
const struct ArrayList* deleted_paths) {
if (!config->report_stats)
return true;
ReceiverStats local;
memset(&local, 0, sizeof(local));
const ReceiverStats* out = stats ? stats : &local;
size_t count = would_delete ? (size_t)would_delete->size : 0;
/* The path list carries the dry-run would-delete set for a -n run and the
actually-removed set for a real --info=del run. */
const struct ArrayList* paths =
config->dry_run ? would_delete : (config->report_deletes ? deleted_paths : NULL);
size_t count = paths ? (size_t)paths->size : 0;
if (count > (size_t)MAX_MANIFEST_ENTRIES)
count = MAX_MANIFEST_ENTRIES;
ReceiverStats record = *out;
@@ -76,7 +87,7 @@ bool receiver_send_stats_frame(int fd, const Config* config, const ReceiverStats
!send_int(fd, (int)count))
return false;
for (size_t i = 0; i < count; i++) {
const char* path = (const char*)would_delete->items[i];
const char* path = (const char*)paths->items[i];
if (!send_wire_str(fd, path ? path : ""))
return false;
}
@@ -90,6 +101,25 @@ static void receiver_tally_deleted(const ReceiverSink* sink, size_t deleted) {
sink->stats->deleted_files += deleted;
}
/* Observer for --info=del: record each truly-removed destination-relative path
in the ArrayList passed as the observer context, so the terminal STATUS_STATS
frame can list it. A failed append is best-effort (the deletion already
happened; output is cosmetic). Shared by the single-threaded receiver and
the -m pipeline's deferred commit. */
void receiver_record_deleted_path(void* context, const char* rel_path) {
ArrayList* paths = context;
if (!paths || !rel_path)
return;
/* Bound the retained list like the keep-set manifest: only MAX_MANIFEST_ENTRIES
paths are ever transmitted in the terminal STATUS_STATS frame, so recording
more only grows memory. A hostile/huge deletion set is therefore capped. */
if ((size_t)paths->size >= (size_t)MAX_MANIFEST_ENTRIES)
return;
char* copy = str_dup(rel_path);
if (copy && !array_list_add(paths, copy))
free(copy);
}
static bool receiver_process_chunk(Chunk* chunk, const ReceiverSink* sink) {
if (!chunk || !sink || !sink->store_file)
return false;
@@ -279,16 +309,276 @@ int receiver_process(Config* config, int file_descriptor, const ReceiverSink* si
return receiver_process_pending(config, file_descriptor, sink, NULL, NULL);
}
/* Per-connection state threaded through the status handlers below. The parked
keep-set / per-directory session live here so one teardown helper can release
them on every exit path. */
typedef struct {
Config* config;
int fd;
const ReceiverSink* sink;
DeleteManifest** pending_manifest;
DeletePlanSession** pending_plans;
/* Parked keep-set for the late/commit timing. Every exit path frees it
exactly once; the only exception is the successful FINISHED handoff, which
transfers ownership to *pending_manifest (used by the -m receiver). */
DeleteManifest* deferred_manifest;
/* Per-directory delete session for --delete-during/--delete-delay. During the
loop it applies plans inline (during) or snapshots their extras (delay); on
a successful FINISHED it is either committed here or handed to
*pending_plans so the -m caller commits after its disk writer drained. */
DeletePlanSession* plan_session;
bool early_delete;
bool per_dir_delete;
bool delete_limit_noted;
} ReceiverPendingState;
/* Outcome of one frame handler. NEXT reads the following status frame; FAIL
tears the connection down without a peer STATUS_ERROR; ERROR tears it down
and (when the sink owns error reporting) emits STATUS_ERROR. */
typedef enum {
RECEIVER_STEP_NEXT,
RECEIVER_STEP_FAIL,
RECEIVER_STEP_ERROR,
} ReceiverStep;
static ReceiverStep receiver_handle_keepalive(ReceiverPendingState* state) {
if (!send_status(state->fd, STATUS_KEEPALIVE))
return RECEIVER_STEP_FAIL;
return RECEIVER_STEP_NEXT;
}
static ReceiverStep receiver_handle_abort(ReceiverPendingState* state) {
(void)state;
log_message(LOG_LEVEL_INFO, "Received abort from client, cleaning up");
return RECEIVER_STEP_FAIL;
}
static ReceiverStep receiver_handle_check(ReceiverPendingState* state) {
bool skipped = false;
bool would_transfer = false;
File* file = receive_incremental_check_ex(state->fd, state->config, &skipped, &would_transfer);
if (state->config->dry_run) {
/* Server-contacting --dry-run: the reply has already been sent
(STATUS_OK = up to date, STATUS_DRY_RUN_TRANSFER = would transfer) and
nothing may be stored. Both flags false means a genuine protocol
error (STATUS_ERROR already sent or sent by receive_error below). */
if (!skipped && !would_transfer)
return RECEIVER_STEP_ERROR;
} else if (!skipped && (!file || !state->sink->store_file(file, state->sink->context))) {
return RECEIVER_STEP_ERROR;
}
return RECEIVER_STEP_NEXT;
}
static ReceiverStep receiver_handle_chunk(ReceiverPendingState* state) {
Chunk* chunk = receive_chunk_data(state->fd, state->config);
if (!chunk || !receiver_process_chunk(chunk, state->sink))
return RECEIVER_STEP_ERROR;
return RECEIVER_STEP_NEXT;
}
static ReceiverStep receiver_handle_check_batch(ReceiverPendingState* state) {
if (!receiver_process_batch(state->config, state->fd))
return RECEIVER_STEP_FAIL;
return RECEIVER_STEP_NEXT;
}
static ReceiverStep receiver_handle_mkdir(ReceiverPendingState* state) {
File* dir = file_receive_directory(state->fd, state->config);
if (!dir || !state->sink->store_file(dir, state->sink->context))
return RECEIVER_STEP_ERROR;
return RECEIVER_STEP_NEXT;
}
static ReceiverStep receiver_handle_dir_times(const ReceiverPendingState* state) {
if (!receiver_process_dir_times(state->fd, state->config, state->sink))
return RECEIVER_STEP_ERROR;
return RECEIVER_STEP_NEXT;
}
static ReceiverStep receiver_handle_hardlink(ReceiverPendingState* state) {
File* file = file_receive_hardlink(state->fd);
if (!file || !state->sink->store_file(file, state->sink->context))
return RECEIVER_STEP_ERROR;
return RECEIVER_STEP_NEXT;
}
static ReceiverStep receiver_handle_symlink(ReceiverPendingState* state) {
File* sym = file_receive_symlink(state->fd, state->config);
if (!sym || !state->sink->store_file(sym, state->sink->context))
return RECEIVER_STEP_ERROR;
return RECEIVER_STEP_NEXT;
}
static ReceiverStep receiver_handle_special(ReceiverPendingState* state) {
File* file = file_receive_special(state->fd);
if (!file || !state->sink->store_file(file, state->sink->context))
return RECEIVER_STEP_ERROR;
return RECEIVER_STEP_NEXT;
}
static ReceiverStep receiver_handle_manifest(ReceiverPendingState* state) {
Config* config = state->config;
int fd = state->fd;
const ReceiverSink* sink = state->sink;
DeleteManifest* manifest = receive_manifest_entries(fd);
if (!manifest)
return RECEIVER_STEP_FAIL; /* receive_manifest_entries already sent STATUS_ERROR */
if (config->dry_run) {
/* Server-contacting --dry-run mutates nothing, so a keep-set manifest
is consumed and discarded. The early-delete mode still needs its ACK
so a sender blocked on the delete handshake is not left hanging.
When would-delete reporting is armed, enumerate (read-only) the
destination extras so the terminal STATUS_STATS frame can list them. */
if (config->use_delete && sink->would_delete) {
size_t count = 0;
if (!manifest_would_delete_list(config, manifest, sink->would_delete, &count))
log_message(LOG_LEVEL_WARNING, "dry-run: could not enumerate would-delete paths");
}
delete_manifest_free(manifest);
if (state->early_delete && !send_status(fd, STATUS_OK))
return RECEIVER_STEP_FAIL;
return RECEIVER_STEP_NEXT;
}
if (state->early_delete) {
/* --delete-before: the whole-tree manifest is authoritative the moment
it arrives, before any file data. Delete now and acknowledge so the
sender only starts streaming once the deletion committed (or failed).
A later transfer failure does not restore these deletions. A
--max-delete-capped commit still succeeds and the transfer proceeds;
the terminal success frame reports the cap. */
size_t deleted = 0;
DeletePathObserver observer =
(config->report_deletes && sink->deleted_paths) ? receiver_record_deleted_path : NULL;
DeleteCommitResult deletion =
(config->use_delete || config->delete_missing_args)
? manifest_delete_all_observed(config, manifest, &deleted, observer,
(void*)sink->deleted_paths)
: DELETE_COMMIT_OK;
receiver_tally_deleted(sink, deleted);
delete_manifest_free(manifest);
if (deletion == DELETE_COMMIT_ERROR) {
send_status(fd, STATUS_ERROR);
return RECEIVER_STEP_FAIL;
}
if (deletion == DELETE_COMMIT_LIMIT_REACHED && sink->note_delete_limit)
sink->note_delete_limit(sink->context);
if (!send_status(fd, STATUS_OK))
return RECEIVER_STEP_FAIL;
} else if (config->use_delete || config->delete_missing_args) {
/* Plain --delete / --delete-after and the --delete-missing-args
exact-path deletions: hold the manifest and commit it only after
STATUS_FINISHED. The per-directory modes never send this frame. */
if (state->deferred_manifest) {
log_message(LOG_LEVEL_ERROR, "Received a second delete manifest");
delete_manifest_free(state->deferred_manifest);
state->deferred_manifest = NULL;
delete_manifest_free(manifest);
send_status(fd, STATUS_ERROR);
return RECEIVER_STEP_FAIL;
}
state->deferred_manifest = manifest;
} else {
delete_manifest_free(manifest);
}
return RECEIVER_STEP_NEXT;
}
static ReceiverStep receiver_handle_delete_plan(ReceiverPendingState* state) {
Config* config = state->config;
int fd = state->fd;
const ReceiverSink* sink = state->sink;
if (!state->per_dir_delete) {
log_message(LOG_LEVEL_ERROR, "Received a per-directory delete plan without a per-dir "
"delete timing");
send_status(fd, STATUS_ERROR);
return RECEIVER_STEP_FAIL;
}
if (!state->plan_session) {
state->plan_session = delete_plan_session_create(config);
if (state->plan_session && config->report_deletes && sink->deleted_paths)
delete_plan_session_set_delete_observer(state->plan_session, receiver_record_deleted_path,
(void*)sink->deleted_paths);
}
if (!state->plan_session || delete_plan_session_receive(state->plan_session, config, fd) != 0)
return RECEIVER_STEP_FAIL;
if (delete_plan_session_limit_reached(state->plan_session) && !state->delete_limit_noted &&
sink->note_delete_limit) {
sink->note_delete_limit(sink->context);
state->delete_limit_noted = true;
}
return RECEIVER_STEP_NEXT;
}
static ReceiverStep receiver_handle_file(ReceiverPendingState* state) {
File* file = file_receive(state->config, state->fd);
if (!file) {
log_message(LOG_LEVEL_ERROR, "Failed to receive file");
return RECEIVER_STEP_ERROR;
}
if (!state->sink->store_file(file, state->sink->context))
return RECEIVER_STEP_ERROR;
return RECEIVER_STEP_NEXT;
}
/* One dispatch per admitted frame type; STATUS_NEXT (and any other
data-bearing status) falls through to the regular file receiver. */
static ReceiverStep receiver_dispatch_status(ReceiverPendingState* state, Status status) {
switch (status) {
case STATUS_KEEPALIVE:
return receiver_handle_keepalive(state);
case STATUS_ABORT:
return receiver_handle_abort(state);
case STATUS_CHECK:
return receiver_handle_check(state);
case STATUS_CHUNK:
return receiver_handle_chunk(state);
case STATUS_CHECK_BATCH:
return receiver_handle_check_batch(state);
case STATUS_MKDIR:
return receiver_handle_mkdir(state);
case STATUS_DIR_TIMES:
return receiver_handle_dir_times(state);
case STATUS_HARDLINK:
return receiver_handle_hardlink(state);
case STATUS_SYMLINK:
return receiver_handle_symlink(state);
case STATUS_SPECIAL:
return receiver_handle_special(state);
case STATUS_MANIFEST:
return receiver_handle_manifest(state);
case STATUS_DELETE_PLAN:
return receiver_handle_delete_plan(state);
default:
return receiver_handle_file(state);
}
}
/* Release the parked keep-set / per-directory session exactly once on every
failure exit. Never commit a deletion for a failed stream. */
static void receiver_drop_pending(ReceiverPendingState* state) {
if (state->deferred_manifest) {
delete_manifest_free(state->deferred_manifest);
state->deferred_manifest = NULL;
}
if (state->plan_session) {
delete_plan_session_destroy(state->plan_session);
state->plan_session = NULL;
}
}
/* Runs the whole receive loop. The delete manifest may legitimately arrive
either FIRST (--delete-before / --delete-during: the sender transmits the
validated keep-set before any file data) or LAST (plain --delete /
--delete-after / --delete-delay: the manifest closes the data stream). In
validated keep-set before any file data) or LAST (--delete-after /
--delete-commit / --delete-delay: the manifest closes the data stream). In
the early modes the receiver deletes as soon as the manifest has been read
and acknowledges with STATUS_OK so the sender only starts streaming once the
deletion has committed (or failed); in the late modes the manifest is held
and the deletion is committed only after the terminal STATUS_FINISHED proves
the whole transfer succeeded. See receiver_process_pending() for how the -m
receiver defers that commit until its disk writer has drained. */
the whole transfer succeeded. A plain --delete defaults to the per-directory
delete-during plan mode (no manifest at all). See the per-frame handlers
above for how the -m receiver defers that commit until its disk writer has
drained. */
int receiver_process_pending(Config* config, int file_descriptor, const ReceiverSink* sink,
DeleteManifest** pending_manifest, DeletePlanSession** pending_plans) {
Status status;
@@ -303,158 +593,29 @@ int receiver_process_pending(Config* config, int file_descriptor, const Receiver
last_progress = session_start;
if (!receiver_note_status(&session_start, &last_progress, status, file_descriptor, sink))
return -1;
bool early_delete = config_delete_timing_early(config);
bool per_dir_delete = config_delete_timing_per_dir(config);
/* Parked keep-set for the late/commit timing. Every exit path below frees it
exactly once; the only exception is the successful FINISHED handoff, which
transfers ownership to *pending_manifest (used by the -m receiver). */
DeleteManifest* deferred_manifest = NULL;
/* Per-directory delete session for --delete-during/--delete-delay. During the
loop it applies plans inline (during) or snapshots their extras (delay); on
a successful FINISHED it is either committed here or handed to
*pending_plans so the -m caller commits after its disk writer drained. */
DeletePlanSession* plan_session = NULL;
bool delete_limit_noted = false;
ReceiverPendingState state = {
.config = config,
.fd = file_descriptor,
.sink = sink,
.pending_manifest = pending_manifest,
.pending_plans = pending_plans,
.deferred_manifest = NULL,
.plan_session = NULL,
.early_delete = config_delete_timing_early(config),
.per_dir_delete = config_delete_timing_per_dir(config),
.delete_limit_noted = false,
};
bool notify_peer = false;
while (status == STATUS_NEXT || status == STATUS_CHUNK || status == STATUS_CHECK ||
status == STATUS_KEEPALIVE || status == STATUS_ABORT || status == STATUS_CHECK_BATCH ||
status == STATUS_MKDIR || status == STATUS_MANIFEST || status == STATUS_HARDLINK ||
status == STATUS_SYMLINK || status == STATUS_SPECIAL || status == STATUS_DIR_TIMES ||
status == STATUS_DELETE_PLAN) {
if (status == STATUS_KEEPALIVE) {
if (!send_status(file_descriptor, STATUS_KEEPALIVE))
goto fail;
goto next_status;
}
if (status == STATUS_ABORT) {
log_message(LOG_LEVEL_INFO, "Received abort from client, cleaning up");
ReceiverStep step = receiver_dispatch_status(&state, status);
if (step == RECEIVER_STEP_FAIL)
goto fail;
}
if (status == STATUS_CHECK) {
bool skipped = false;
bool would_transfer = false;
File* file = receive_incremental_check_ex(file_descriptor, config, &skipped, &would_transfer);
if (config->dry_run) {
/* Server-contacting --dry-run: the reply has already been sent
(STATUS_OK = up to date, STATUS_DRY_RUN_TRANSFER = would transfer) and
nothing may be stored. Both flags false means a genuine protocol
error (STATUS_ERROR already sent or sent by receive_error below). */
if (!skipped && !would_transfer)
goto receive_error;
} else if (!skipped && (!file || !sink->store_file(file, sink->context))) {
goto receive_error;
}
} else if (status == STATUS_CHUNK) {
Chunk* chunk = receive_chunk_data(file_descriptor, config);
if (!chunk || !receiver_process_chunk(chunk, sink))
goto receive_error;
} else if (status == STATUS_CHECK_BATCH) {
if (!receiver_process_batch(config, file_descriptor))
goto fail;
goto next_status;
} else if (status == STATUS_MKDIR) {
File* dir = file_receive_directory(file_descriptor, config);
if (!dir || !sink->store_file(dir, sink->context))
goto receive_error;
} else if (status == STATUS_DIR_TIMES) {
if (!receiver_process_dir_times(file_descriptor, config, sink))
goto receive_error;
} else if (status == STATUS_HARDLINK) {
File* file = file_receive_hardlink(file_descriptor);
if (!file || !sink->store_file(file, sink->context))
goto receive_error;
} else if (status == STATUS_SYMLINK) {
File* sym = file_receive_symlink(file_descriptor, config);
if (!sym || !sink->store_file(sym, sink->context))
goto receive_error;
} else if (status == STATUS_SPECIAL) {
File* file = file_receive_special(file_descriptor);
if (!file || !sink->store_file(file, sink->context))
goto receive_error;
} else if (status == STATUS_MANIFEST) {
DeleteManifest* manifest = receive_manifest_entries(file_descriptor);
if (!manifest)
goto fail; /* receive_manifest_entries already sent STATUS_ERROR */
if (config->dry_run) {
/* Server-contacting --dry-run mutates nothing, so a keep-set manifest
is consumed and discarded. The early-delete mode still needs its ACK
so a sender blocked on the delete handshake is not left hanging.
When would-delete reporting is armed, enumerate (read-only) the
destination extras so the terminal STATUS_STATS frame can list them. */
if (config->use_delete && sink->would_delete) {
size_t count = 0;
if (!manifest_would_delete_list(config, manifest, sink->would_delete, &count))
log_message(LOG_LEVEL_WARNING, "dry-run: could not enumerate would-delete paths");
}
delete_manifest_free(manifest);
if (early_delete && !send_status(file_descriptor, STATUS_OK))
goto fail;
goto next_status;
}
if (early_delete) {
/* --delete-before: the whole-tree manifest is authoritative the moment
it arrives, before any file data. Delete now and acknowledge so the
sender only starts streaming once the deletion committed (or failed).
A later transfer failure does not restore these deletions. A
--max-delete-capped commit still succeeds and the transfer proceeds;
the terminal success frame reports the cap. */
size_t deleted = 0;
DeleteCommitResult deletion = (config->use_delete || config->delete_missing_args)
? manifest_delete_all_counted(config, manifest, &deleted)
: DELETE_COMMIT_OK;
receiver_tally_deleted(sink, deleted);
delete_manifest_free(manifest);
if (deletion == DELETE_COMMIT_ERROR) {
send_status(file_descriptor, STATUS_ERROR);
goto fail;
}
if (deletion == DELETE_COMMIT_LIMIT_REACHED && sink->note_delete_limit)
sink->note_delete_limit(sink->context);
if (!send_status(file_descriptor, STATUS_OK))
goto fail;
} else if (config->use_delete || config->delete_missing_args) {
/* Plain --delete / --delete-after and the --delete-missing-args
exact-path deletions: hold the manifest and commit it only after
STATUS_FINISHED. The per-directory modes never send this frame. */
if (deferred_manifest) {
log_message(LOG_LEVEL_ERROR, "Received a second delete manifest");
delete_manifest_free(deferred_manifest);
deferred_manifest = NULL;
delete_manifest_free(manifest);
send_status(file_descriptor, STATUS_ERROR);
goto fail;
}
deferred_manifest = manifest;
} else {
delete_manifest_free(manifest);
}
goto next_status;
} else if (status == STATUS_DELETE_PLAN) {
if (!per_dir_delete) {
log_message(LOG_LEVEL_ERROR, "Received a per-directory delete plan without a per-dir "
"delete timing");
send_status(file_descriptor, STATUS_ERROR);
goto fail;
}
if (!plan_session)
plan_session = delete_plan_session_create(config);
if (!plan_session || delete_plan_session_receive(plan_session, config, file_descriptor) != 0)
goto fail;
if (delete_plan_session_limit_reached(plan_session) && !delete_limit_noted &&
sink->note_delete_limit) {
sink->note_delete_limit(sink->context);
delete_limit_noted = true;
}
goto next_status;
} else {
File* file = file_receive(config, file_descriptor);
if (!file) {
log_message(LOG_LEVEL_ERROR, "Failed to receive file");
goto receive_error;
}
if (!sink->store_file(file, sink->context))
goto receive_error;
}
next_status:
if (step == RECEIVER_STEP_ERROR)
goto receive_error;
if (!receive_status(file_descriptor, &status))
goto receive_error;
if (!receiver_note_status(&session_start, &last_progress, status, file_descriptor, sink))
@@ -473,17 +634,19 @@ int receiver_process_pending(Config* config, int file_descriptor, const Receiver
disk writer may still be draining; the caller commits after the writer has
joined so no extra file is removed unless the transfer is known to have
succeeded. */
if (deferred_manifest) {
if (pending_manifest) {
*pending_manifest = deferred_manifest;
deferred_manifest = NULL;
if (state.deferred_manifest) {
if (state.pending_manifest) {
*state.pending_manifest = state.deferred_manifest;
state.deferred_manifest = NULL;
} else {
size_t deleted = 0;
DeleteCommitResult deletion =
manifest_delete_all_counted(config, deferred_manifest, &deleted);
DeletePathObserver observer =
(config->report_deletes && sink->deleted_paths) ? receiver_record_deleted_path : NULL;
DeleteCommitResult deletion = manifest_delete_all_observed(
config, state.deferred_manifest, &deleted, observer, (void*)sink->deleted_paths);
receiver_tally_deleted(sink, deleted);
delete_manifest_free(deferred_manifest);
deferred_manifest = NULL;
delete_manifest_free(state.deferred_manifest);
state.deferred_manifest = NULL;
if (deletion == DELETE_COMMIT_ERROR) {
send_status(file_descriptor, STATUS_ERROR);
goto fail;
@@ -497,25 +660,28 @@ int receiver_process_pending(Config* config, int file_descriptor, const Receiver
nothing yet and applies its decompressed snapshot here. The -m receiver
hands the session to its caller instead, which commits after the disk
writer drained. */
if (plan_session) {
if (pending_plans) {
*pending_plans = plan_session;
plan_session = NULL;
if (state.plan_session) {
if (config->report_deletes && sink->deleted_paths)
delete_plan_session_set_delete_observer(state.plan_session, receiver_record_deleted_path,
(void*)sink->deleted_paths);
if (state.pending_plans) {
*state.pending_plans = state.plan_session;
state.plan_session = NULL;
} else if (config->dry_run) {
/* Central dry-run no-op: never commit a deletion for a -n run. */
delete_plan_session_destroy(plan_session);
plan_session = NULL;
delete_plan_session_destroy(state.plan_session);
state.plan_session = NULL;
} else {
DeleteCommitResult deletion = delete_plan_session_commit(plan_session, config);
bool limit = delete_plan_session_limit_reached(plan_session);
receiver_tally_deleted(sink, delete_plan_session_deleted(plan_session));
delete_plan_session_destroy(plan_session);
plan_session = NULL;
DeleteCommitResult deletion = delete_plan_session_commit(state.plan_session, config);
bool limit = delete_plan_session_limit_reached(state.plan_session);
receiver_tally_deleted(sink, delete_plan_session_deleted(state.plan_session));
delete_plan_session_destroy(state.plan_session);
state.plan_session = NULL;
if (deletion == DELETE_COMMIT_ERROR) {
send_status(file_descriptor, STATUS_ERROR);
goto fail;
}
if (limit && !delete_limit_noted && sink->note_delete_limit)
if (limit && !state.delete_limit_noted && sink->note_delete_limit)
sink->note_delete_limit(sink->context);
}
}
@@ -529,26 +695,14 @@ int receiver_process_pending(Config* config, int file_descriptor, const Receiver
}
return 0;
receive_error:
notify_peer = true;
fail:
/* Failure exits that must not (or already did) report a STATUS_ERROR. The
parked keep-set/session is dropped: never commit a deletion for a failed
stream. */
if (deferred_manifest) {
delete_manifest_free(deferred_manifest);
deferred_manifest = NULL;
}
if (plan_session)
delete_plan_session_destroy(plan_session);
return -1;
receive_error:
if (deferred_manifest) {
delete_manifest_free(deferred_manifest);
deferred_manifest = NULL;
}
if (plan_session)
delete_plan_session_destroy(plan_session);
if (sink->send_error)
receiver_drop_pending(&state);
if (notify_peer && sink->send_error)
send_status(file_descriptor, STATUS_ERROR);
return -1;
}
@@ -569,11 +723,20 @@ typedef struct {
--delete would-delete path list collected while processing the manifest. */
ReceiverStats stats;
ArrayList* would_delete;
/* --info=del: actually-removed paths collected during the delete commit. */
ArrayList* deleted_paths;
/* Per-run count of entries that failed to materialize without aborting the
stream (currently ONLY a --devices mknod EPERM/EACCES). A nonzero count
makes the terminal frame carry a non-OK status so the client exits
non-zero, matching rsync's continue-and-exit-partial behavior. */
size_t failed_entries;
} ReceiverSaveContext;
static bool receiver_save_file(File* file, void* context_pointer) {
ReceiverSaveContext* context = context_pointer;
FileSaveResult result = FILE_SAVE_ERROR;
bool created = false;
unsigned created_dirs = 0;
if (context->config->dry_run) {
/* Defense in depth: a dry-run receiver mutates nothing even if a data
frame reaches the sink (the sender is not supposed to send one). */
@@ -583,12 +746,22 @@ static bool receiver_save_file(File* file, void* context_pointer) {
--remove-source-files sender keeps its source. */
result = FILE_SAVE_SKIPPED;
} else {
result = file_save_to_disk_full(context->config->receive_root_directory, file, context->config);
result = file_save_to_disk_full_ex(context->config->receive_root_directory, file,
context->config, &created, &created_dirs);
}
/* Wire-stats tally: bytes reconstructed from the basis file (delta matches)
count as matched data in the end-of-transfer report. */
if (result != FILE_SAVE_ERROR && file->matched_bytes > 0)
context->stats.matched_data += file->matched_bytes;
/* --devices parity: a device node the receiver could not mknod (EPERM/EACCES)
is counted per-run but does not abort the transfer. The terminal frame
turns a nonzero count into a non-OK status so the client exits non-zero. */
if (result == FILE_SAVE_FAILED)
context->failed_entries++;
/* Protocol 2.28.0: receiver-observed literal bytes and the created-entry
breakdown (regular/dir/link/special) for the `--stats` report. */
if (result == FILE_SAVE_WRITTEN)
receiver_stats_note_saved(&context->stats, file, created, created_dirs);
/* A directory's metadata is deferred, never applied inline: collect it now
and apply it at the end. -O/--omit-dir-times and --preserve_perms/-times
are honored by dir_metadata_list_apply's caller (see
@@ -618,10 +791,27 @@ static void receiver_note_delete_limit(void* context_pointer) {
context->delete_limit_reached = true;
}
/* Terminal status for a run. A capped --delete limit wins (rsync exit 25);
otherwise any per-entry failure (for example an unprivileged --devices
mknod) makes the terminal frame non-OK so the client exits non-zero. rsync
reports 23 here; mapping the client's exact exit code to 23 is a separate,
pre-existing concern. A clean run keeps STATUS_OK. */
static Status receiver_final_status(bool delete_limit_reached, size_t failed_entries) {
if (delete_limit_reached)
return STATUS_DELETE_LIMIT;
return failed_entries > 0 ? STATUS_ERROR : STATUS_OK;
}
static bool receiver_send_success_frame(int fd, void* context_pointer) {
ReceiverSaveContext* context = context_pointer;
Status final_status = context->delete_limit_reached ? STATUS_DELETE_LIMIT : STATUS_OK;
if (!receiver_send_stats_frame(fd, context->config, &context->stats, context->would_delete))
if (context->failed_entries > 0)
log_message(LOG_LEVEL_WARNING,
"%zu entr%s failed to materialize; continuing (partial transfer)",
context->failed_entries, context->failed_entries == 1 ? "y" : "ies");
Status final_status =
receiver_final_status(context->delete_limit_reached, context->failed_entries);
if (!receiver_send_stats_frame(fd, context->config, &context->stats, context->would_delete,
context->deleted_paths))
return false;
/* Server-contacting --dry-run: nothing was staged or written, so there is
nothing to publish and no directory times to stamp. */
@@ -651,8 +841,15 @@ int receiver_receive_files(Config* config, int file_descriptor) {
ReceiverSaveContext context = {.config = config, .outcomes = {0}};
dir_time_list_init(&context.dir_times);
context.would_delete = array_list_create(free);
if (!context.would_delete)
/* report_deletes (--info=del / -i / --out-format under --delete) is the only
reason to retain the actually-removed paths; a plain --delete must not
str_dup every removal. NULL is handled by every consumer. */
context.deleted_paths = config->report_deletes ? array_list_create(free) : NULL;
if (!context.would_delete || (config->report_deletes && !context.deleted_paths)) {
array_list_delete(context.would_delete);
array_list_delete(context.deleted_paths);
return -1;
}
ReceiverSink sink = {receiver_save_file,
&context,
true,
@@ -660,12 +857,14 @@ int receiver_receive_files(Config* config, int file_descriptor) {
receiver_send_success_frame,
receiver_note_delete_limit,
&context.stats,
context.would_delete};
context.would_delete,
context.deleted_paths};
int ret = receiver_process(config, file_descriptor, &sink);
if (ret != 0 && config->delay_updates && config->delay_context)
delay_updates_cleanup(config->delay_context);
receiver_outcomes_destroy(&context.outcomes);
dir_time_list_free(&context.dir_times);
array_list_delete(context.would_delete);
array_list_delete(context.deleted_paths);
return ret;
}
+11 -1
View File
@@ -46,11 +46,20 @@ typedef struct {
carries the -n/--dry-run --delete path list. */
ReceiverStats* stats;
struct ArrayList* would_delete;
/* When --info=del requested it, receiver-owned strings for every path the
deletion commit ACTUALLY removed, sent in the terminal STATUS_STATS frame's
path list so the sender can print rsync's `deleting PATH` lines. */
struct ArrayList* deleted_paths;
} ReceiverSink;
bool receiver_outcomes_append(ReceiverOutcomes* outcomes, unsigned char code);
void receiver_outcomes_destroy(ReceiverOutcomes* outcomes);
/* DeletePathObserver implementation for --info=del: `context` is an ArrayList*
that receives owned copies of every truly-removed destination-relative path.
Shared by the single-threaded receiver and the -m pipeline's deferred commit. */
void receiver_record_deleted_path(void* context, const char* rel_path);
/* Send the terminal success frame. `final_status` is usually STATUS_OK, or
STATUS_DELETE_LIMIT when a --max-delete commit was capped. */
bool receiver_send_final_success(int fd, const Config* config, const ReceiverOutcomes* outcomes,
@@ -60,7 +69,8 @@ bool receiver_send_final_success(int fd, const Config* config, const ReceiverOut
non-NULL, a count and that many wire strings) when the wire config requested
report_stats. A no-op otherwise. */
bool receiver_send_stats_frame(int fd, const Config* config, const ReceiverStats* stats,
const struct ArrayList* would_delete);
const struct ArrayList* would_delete,
const struct ArrayList* deleted_paths);
int receiver_process(Config* config, int file_descriptor, const ReceiverSink* sink);
/* receiver_process with an escape hatch for the commit-style (late) deletion:
+40 -2
View File
@@ -29,8 +29,10 @@ PipelineContextReceiver* pipeline_context_receiver_create(Config* config, Queue*
context->deferred_manifest = NULL;
context->deferred_plans = NULL;
context->delete_limit_reached = false;
context->failed_entries = 0;
memset(&context->stats, 0, sizeof(context->stats));
context->would_delete = NULL;
context->deleted_paths = NULL;
atomic_init(&context->cancelled, false);
int init = 0;
if (mtx_init(&context->mutex, mtx_plain) != thrd_success)
@@ -46,6 +48,15 @@ PipelineContextReceiver* pipeline_context_receiver_create(Config* config, Queue*
context->would_delete = array_list_create(free);
if (!context->would_delete)
goto fail;
/* The actually-removed path list is only needed to render rsync's
`deleting PATH` lines, which the client requests via report_deletes
(--info=del / -i / --out-format under --delete). A plain --delete run must
not allocate it or observe every removal. */
if (config->report_deletes) {
context->deleted_paths = array_list_create(free);
if (!context->deleted_paths)
goto fail;
}
return context;
fail:
@@ -56,6 +67,12 @@ fail:
cnd_destroy(&context->condition_not_full);
if (init >= 1)
mtx_destroy(&context->mutex);
/* Free every list that was already created before the failing allocation:
`context` itself is freed below, so they would otherwise leak. */
if (context->would_delete)
array_list_delete(context->would_delete);
if (context->deleted_paths)
array_list_delete(context->deleted_paths);
free(context);
return NULL;
}
@@ -71,6 +88,8 @@ void pipeline_context_receiver_destroy(PipelineContextReceiver* context) {
dir_time_list_free(&context->dir_times);
if (context->would_delete)
array_list_delete(context->would_delete);
if (context->deleted_paths)
array_list_delete(context->deleted_paths);
mtx_destroy(&context->mutex);
cnd_destroy(&context->condition_not_full);
cnd_destroy(&context->condition_not_empty);
@@ -184,7 +203,8 @@ int receive_thread(void* pipeline_context) {
NULL,
receiver_pipeline_note_delete_limit,
&context->stats,
context->would_delete};
context->would_delete,
context->deleted_paths};
if (receiver_process_pending((Config*)config, file_descriptor, &sink, &context->deferred_manifest,
&context->deferred_plans) != 0) {
receiver_thread_fail(context);
@@ -228,12 +248,30 @@ int write_thread(void* pipeline_context) {
}
size_t file_bytes = file->data ? file->data->size : 0;
FileSaveResult result = FILE_SAVE_SKIPPED;
bool created = false;
unsigned created_dirs = 0;
/* Server-contacting --dry-run: never write. The receiver thread does not
enqueue anything on the dry-run path, but this keeps the writer thread
provably mutation-free if a data frame ever reached it. */
bool dry_run = context->config->dry_run;
if (save_to_disk && !dry_run) {
result = file_save_to_disk_full(root_directory, file, context->config);
result =
file_save_to_disk_full_ex(root_directory, file, context->config, &created, &created_dirs);
if (result == FILE_SAVE_WRITTEN) {
/* Protocol 2.28.0: fold the receiver-observed literal bytes and the
created-entry type into the shared stats block under its mutex (the
receive thread also writes stats.matched_data). */
mtx_lock(&context->mutex);
receiver_stats_note_saved(&context->stats, file, created, created_dirs);
mtx_unlock(&context->mutex);
}
/* --devices parity: a device node that could not be mknod'ed is counted
per-run but does NOT abort the transfer. */
if (result == FILE_SAVE_FAILED) {
mtx_lock(&context->mutex);
context->failed_entries++;
mtx_unlock(&context->mutex);
}
if (result == FILE_SAVE_ERROR) {
file_destroy(file);
pipeline_context_receiver_note_bytes_released(context, file_bytes);
+8
View File
@@ -61,6 +61,14 @@ typedef struct PipelineContextReceiver {
/* -n/--dry-run --delete would-delete path list, collected by receive_thread
and reported in the STATUS_STATS frame. */
struct ArrayList* would_delete;
/* --info=del actually-removed path list, collected by the deferred delete
commit in server.c and reported in the STATUS_STATS frame. */
struct ArrayList* deleted_paths;
/* Per-run count of entries that failed to materialize without aborting the
stream (currently ONLY a --devices mknod EPERM/EACCES). write_thread
increments it under `mutex`; server.c turns a nonzero count into a non-OK
terminal status so the client exits non-zero. */
size_t failed_entries;
} PipelineContextReceiver;
PipelineContextReceiver* pipeline_context_receiver_create(Config* config, Queue* queue_receiver,
+319 -200
View File
@@ -226,11 +226,6 @@ static void release_authorization(void) {
close(root_fd);
}
static bool path_is_within(const char* root, const char* path) {
size_t n = strlen(root);
return strncmp(root, path, n) == 0 && (path[n] == '\0' || path[n] == '/');
}
/* --mkpath contract: when the client's destination root directory does not
exist yet on the server side, --mkpath tells the server to create it (and
any missing leading components) below the authorized root at connection
@@ -710,69 +705,85 @@ static const char* server_module_gate(const Config* config, void* context) {
return module_gate_install_root(config, module);
}
void handler(int file_descriptor) {
SSL* ssl = io_get_ssl();
/* Per-connection state threaded through the handler phase helpers below. The
* fields are a faithful split of the former handler() locals: the protocol
* session, the config-frame gate context, the accepted config, the optional
* multithreaded pipeline context and the teardown bookkeeping all live here so
* the single `done` epilogue in handler() can release them exactly as before. */
typedef struct ServerSession {
int fd;
SSL* ssl;
ProtocolSession session;
protocol_session_init(&session, file_descriptor, file_descriptor);
protocol_session_set_ssl(&session, ssl);
protocol_session_bind(&session);
ModuleGateContext gate_ctx;
gate_ctx.ssl = ssl;
gate_ctx.fd = file_descriptor;
gate_ctx.super_mode_override = -1;
gate_ctx.has_peer_ip = false;
gate_ctx.peer_ip[0] = '\0';
gate_ctx.is_local = false;
/* All teardown state starts empty so the single `done` epilogue is safe to
* reach from any error path (including before the config frame arrives). */
Config* config = NULL;
PipelineContextReceiver* context = NULL;
char* joined_destination = NULL;
bool charset_ready = false;
config = config_receive_with_validate(file_descriptor, server_module_gate, &gate_ctx);
if (config == NULL) {
Config* config;
PipelineContextReceiver* context;
char* joined_destination;
bool charset_ready;
} ServerSession;
/* Phase 1 -- config receipt + validation. Receives the client config frame
* through the module gate, applies the super-mode override the gate recorded
* exactly once, and installs the per-connection protocol/compression state.
* Returns false when the config frame was refused (the gate has already
* answered the client); the caller jumps to the shared `done` epilogue. */
static bool server_accept_config(ServerSession* state) {
state->config = config_receive_with_validate(state->fd, server_module_gate, &state->gate_ctx);
if (state->config == NULL) {
log_message(LOG_LEVEL_ERROR, "Failed to receive config");
goto done;
return false;
}
/* Apply the super-mode veto the gate decided on (operator --no-super, or a
* daemon module without the `client owner = yes` opt-in) exactly once, so
* every downstream gate (identity_apply_ownership via privilege_super_permitted,
* device-node creation) sees SUPER_MODE_OFF. The gate never mutated the
* received config. */
if (gate_ctx.super_mode_override != -1)
config->super_mode = (SuperMode)gate_ctx.super_mode_override;
if (state->gate_ctx.super_mode_override != -1)
state->config->super_mode = (SuperMode)state->gate_ctx.super_mode_override;
/* Install the codec this connection negotiated before the receiver/writer
* threads start (the server forks per connection, so the process-global
* codec is private to this session). */
compression_set_algo((CompressionAlgo)config->compression_algo);
compression_set_algo((CompressionAlgo)state->config->compression_algo);
/* If the client requested ownership but the effective super mode forbids it
* (operator --no-super, a privileged standalone receiver's secure default, or
* a daemon module without `client owner = yes`), say so ONCE per connection so
* a successful -a/-o/-g transfer is not mistaken for preserved ownership. */
if (config->super_mode == SUPER_MODE_OFF && identity_ownership_requested(config))
if (state->config->super_mode == SUPER_MODE_OFF && identity_ownership_requested(state->config))
log_message(LOG_LEVEL_WARNING,
"requested ownership will NOT be applied: super-user activities are disabled "
"for this connection (operator veto, or module without `client owner = yes`)");
protocol_set_8_bit_output(config->eight_bit_output);
protocol_set_8_bit_output(state->config->eight_bit_output);
/* Server-side per-message protocol deadline for every frame from here on.
* `timeout` is not serialized, so this is the server's own config (the server
* has no --timeout CLI and defaults it to 0). A client's --timeout tightens
* only that client's own protocol I/O; the server floors its own deadline at
* SERVER_IO_TIMEOUT_SEC so a silent peer can never hold a session slot
* forever (the socket layer gets the same floor at startup). */
protocol_session_set_io_timeout(&session, protocol_server_io_timeout_sec(config->timeout));
protocol_session_set_io_timeout(&state->session,
protocol_server_io_timeout_sec(state->config->timeout));
return true;
}
/* Phase 2 -- security gates. The ORDER here is load-bearing and must not be
* merged or reordered: transport/authentication (plaintext refusal, TLS
* client-CN verification), then daemon-root confinement (absolute-destination
* rejection, traversal + within-authorized-root), then delete/force
* authorization -- exactly the sequence the former handler() used. Returns
* false after logging the matching rejection; the caller jumps to the shared
* `done` epilogue. */
static bool server_apply_security_gates(ServerSession* state) {
Config* config = state->config;
const char* authorized_root = utils_get_authorized_root_path();
if (!authorized_root) {
log_message(LOG_LEVEL_ERROR, "No server-side destination root configured");
goto done;
return false;
}
if (!allow_unauthenticated && ssl == NULL) {
if (!allow_unauthenticated && state->ssl == NULL) {
log_message(LOG_LEVEL_ERROR, "Rejected unauthenticated plaintext connection");
goto done;
return false;
}
if (ssl && required_client_cn && !tls_client_identity_allowed(ssl)) {
if (state->ssl && required_client_cn && !tls_client_identity_allowed(state->ssl)) {
log_message(LOG_LEVEL_ERROR, "Rejected TLS client with unauthorized identity");
goto done;
return false;
}
/* Daemon mode: the module's root is the authorized root (installed by
server_module_gate), and the client's destination is a MODULE-RELATIVE
@@ -782,27 +793,27 @@ void handler(int file_descriptor) {
if (g_daemon_conf && config->receive_root_directory && config->receive_root_directory[0] == '/') {
log_message(LOG_LEVEL_ERROR, "Rejected absolute daemon destination (must be relative to the "
"selected module root)");
goto done;
return false;
}
char* destination = config->receive_root_directory;
if (destination && destination[0] != '/')
joined_destination = path_cat(authorized_root, destination);
if (joined_destination)
destination = joined_destination;
state->joined_destination = path_cat(authorized_root, destination);
if (state->joined_destination)
destination = state->joined_destination;
if (!destination || has_path_traversal(destination) ||
!path_is_within(authorized_root, destination)) {
!path_is_within_root(authorized_root, destination)) {
log_message(LOG_LEVEL_ERROR, "Rejected destination outside authorized root");
free(joined_destination);
joined_destination = NULL;
goto done;
free(state->joined_destination);
state->joined_destination = NULL;
return false;
}
if (joined_destination) {
if (state->joined_destination) {
free(config->receive_root_directory);
config->receive_root_directory = joined_destination;
joined_destination = NULL;
config->receive_root_directory = state->joined_destination;
state->joined_destination = NULL;
}
if (!config->receive_root_directory) {
goto done;
return false;
}
config->use_delete = config->use_delete && allow_delete;
/* --force (receiver-side) is deletion authority too: it lets an incoming
@@ -812,6 +823,18 @@ void handler(int file_descriptor) {
* --delete-missing-args, so a client cannot use --force to bypass the delete
* policy. */
config->force_delete = config->force_delete && allow_delete;
return true;
}
/* Phase 3 -- session preparation. Installs the negotiated conversion, applies
* the remaining deletion policy, materializes the destination root (--mkpath),
* creates the --delay-updates staging tree, snapshots the identity policy, and
* publishes the --keep-dirlinks/--trust-sender globals and the daemon MOTD.
* All of it must happen before any receiver/writer thread is spawned. Returns
* false after logging the matching failure; the caller jumps to the shared
* `done` epilogue. */
static bool server_prepare_session(ServerSession* state) {
Config* config = state->config;
/* --iconv (protocol 2.16.0): install the receiver-side wire->local conversion
now that the client's full CONVERT_SPEC has been received and validated,
before any received file name is decoded. The server's own --iconv (if
@@ -822,14 +845,14 @@ void handler(int file_descriptor) {
if (!charset_wire_init_receiver(config->iconv_spec, server_iconv_spec)) {
log_message(LOG_LEVEL_ERROR,
"--iconv: unsupported charset conversion requested (LOCAL[,REMOTE])");
goto done;
return false;
}
charset_ready = true;
state->charset_ready = true;
}
/* --delete-missing-args deletes destination mirrors receiver-side, so it is
deletion and stays gated by the same --allow-delete server policy. When
the server policy is off the flag is inert (the missing entries are still
skipped via its implied --ignore-missing-args, but nothing is deleted). */
* deletion and stays gated by the same --allow-delete server policy. When
* the server policy is off the flag is inert (the missing entries are still
* skipped via its implied --ignore-missing-args, but nothing is deleted). */
config->delete_missing_args = config->delete_missing_args && allow_delete;
/* --mkpath: create the destination root (and its missing leading components)
* before anything else; without it the root must pre-exist. The precondition
@@ -844,7 +867,7 @@ void handler(int file_descriptor) {
log_message(LOG_LEVEL_ERROR, "destination root is not available: %s",
escaped_root ? escaped_root : "<allocation failed>");
free(escaped_root);
goto done;
return false;
}
/* A --delay-updates transfer stages under a private 0700 directory inside
the receive root. Create it up front (wiping leftovers of any previously
@@ -854,7 +877,7 @@ void handler(int file_descriptor) {
config->delay_context = delay_updates_context_create(config->receive_root_directory);
if (!config->delay_context || !delay_updates_prepare(config->delay_context)) {
log_message(LOG_LEVEL_ERROR, "Failed to initialize --delay-updates staging area");
goto done;
return false;
}
}
/* Preserve the negotiated identity policy for the fd-relative ownership
@@ -864,7 +887,7 @@ void handler(int file_descriptor) {
rather than silently applying the wrong ownership policy. */
if (!identity_set_active(config)) {
log_message(LOG_LEVEL_ERROR, "Failed to activate identity policy");
goto done;
return false;
}
/* Persist the negotiated --keep-dirlinks policy once, here at config-accept,
before any multithreaded receiver/writer threads are spawned, so the
@@ -892,136 +915,201 @@ void handler(int file_descriptor) {
Wave C note in config.h). */
if (g_daemon_conf) {
char* motd = motd_read_file(g_daemon_conf->global.motd_file);
if (!motd_send(file_descriptor, motd ? motd : "")) {
if (!motd_send(state->fd, motd ? motd : "")) {
free(motd);
log_message(LOG_LEVEL_ERROR, "Failed to send daemon MOTD");
goto done;
return false;
}
free(motd);
}
if (config->use_multithreading) {
Queue* q = queue_create(100, file_destroy);
if (q == NULL)
goto done;
context = pipeline_context_receiver_create(config, q, file_descriptor, ssl);
if (context == NULL) {
queue_destroy(q);
goto done;
}
protocol_session_set_max_alloc(&context->session, config->max_alloc);
protocol_session_set_io_timeout(&context->session,
protocol_server_io_timeout_sec(config->timeout));
atomic_store(&context->session.total_allocated_bytes,
atomic_load(&session.total_allocated_bytes));
pipeline_context_receiver_set_queue_byte_limit(context, RECEIVER_QUEUE_MAX_BYTES);
thrd_t receiver = {0};
thrd_t writer = {0};
bool receiver_created = thrd_create(&receiver, receive_thread, context) == thrd_success;
bool writer_created = false;
if (receiver_created)
writer_created = thrd_create(&writer, write_thread, context) == thrd_success;
if (!receiver_created || !writer_created) {
log_perror("Error creating Threads");
if (receiver_created) {
mtx_lock(&context->mutex);
atomic_store(&context->cancelled, true);
cnd_broadcast(&context->condition_not_full);
cnd_broadcast(&context->condition_not_empty);
mtx_unlock(&context->mutex);
/* Unblock a worker parked in socket I/O without closing the fd (the
* child owns the single close). shutdown() only affects sockets; for
* the --stdio pipe the receiver's per-message poll timeout still
* bounds the join, so do nothing there rather than close a descriptor
* another thread may still be using. */
struct stat fd_stat;
if (fstat(file_descriptor, &fd_stat) == 0 && S_ISSOCK(fd_stat.st_mode))
shutdown(file_descriptor, SHUT_RDWR);
thrd_join(receiver, NULL);
}
if (writer_created)
thrd_join(writer, NULL);
goto done;
}
int receiver_result;
int writer_result;
thrd_join(receiver, &receiver_result);
thrd_join(writer, &writer_result);
bool transfer_ok = receiver_result == thrd_success && writer_result == thrd_success;
if (transfer_ok && !config->dry_run) {
/* Commit-style (late) deletion: receive_thread handed the keep-set
manifest here instead of deleting while write_thread might still be
draining, so by now every file is on disk and the whole transfer is
known to have succeeded. Remove the extras before publishing a
--delay-updates run; the walker skips the staging directory. A
server-contacting --dry-run deletes nothing (no manifest is sent). */
if (context->deferred_manifest) {
size_t deleted = 0;
DeleteCommitResult deletion =
manifest_delete_all_counted(config, context->deferred_manifest, &deleted);
context->stats.deleted_files += deleted;
if (deletion == DELETE_COMMIT_ERROR) {
transfer_ok = false;
} else if (deletion == DELETE_COMMIT_LIMIT_REACHED) {
/* The transfer still succeeds; the terminal frame reports the capped
deletion so the sender exits 25 like rsync. */
context->delete_limit_reached = true;
}
delete_manifest_free(context->deferred_manifest);
context->deferred_manifest = NULL;
}
/* --delete-delay: receive_thread snapshotted each plan's extras as it
arrived; with the disk writer drained, commit the deferred removals.
--delete-during already applied its plans on the receive thread. */
if (context->deferred_plans) {
/* Defence in depth (the enclosing block already excludes dry-run): a
-n run never commits a deletion. */
DeleteCommitResult deletion =
config->dry_run ? DELETE_COMMIT_OK
: delete_plan_session_commit(context->deferred_plans, config);
context->stats.deleted_files += delete_plan_session_deleted(context->deferred_plans);
if (deletion == DELETE_COMMIT_ERROR) {
transfer_ok = false;
} else if (deletion == DELETE_COMMIT_LIMIT_REACHED) {
context->delete_limit_reached = true;
}
delete_plan_session_destroy(context->deferred_plans);
context->deferred_plans = NULL;
}
}
if (transfer_ok && !config->dry_run) {
/* --delay-updates: receive_thread has finished the whole protocol stream
(including manifest/delete handling) and write_thread has drained its
queue, so every staged file is complete. Publish atomically before the
success/outcome frame so a --remove-source-files sender only learns of
files that were actually installed. */
if (config->delay_updates && config->delay_context &&
!delay_updates_publish(config->delay_context, config)) {
transfer_ok = false;
}
/* P7 Wave D: all writers have joined and the late deletion (and
--delay-updates publication) has committed above, so it is finally safe
to stamp directory times; a directory's mtime must not be clobbered by
its children or by an extra removal. */
if (transfer_ok)
dir_metadata_list_apply(&context->dir_times, config->receive_root_directory, config);
}
if (transfer_ok) {
Status final_status = context->delete_limit_reached ? STATUS_DELETE_LIMIT : STATUS_OK;
/* Emit the optional wire-stats record first (protocol 2.25.0), then the
success/outcome frame, exactly like the single-threaded receiver. */
if (!receiver_send_stats_frame(file_descriptor, config, &context->stats,
context->would_delete) ||
!receiver_send_final_success(file_descriptor, config, &context->outcomes, final_status))
transfer_ok = false;
} else {
send_error_detail(file_descriptor, "transfer failed on receiver");
}
if (!transfer_ok)
log_message(LOG_LEVEL_ERROR, "Transfer failed");
} else {
if (receiver_receive_files(config, file_descriptor) != 0)
log_message(LOG_LEVEL_ERROR, "Transfer failed");
return true;
}
/* Phase 4a -- transfer via the multithreaded receiver. Spawns the receive/write
* thread pair, joins them, then commits the late deletion, --delay-updates
* publication and directory times before emitting the terminal stats/success
* frame. On any failure the helper just returns; the caller's `done` epilogue
* releases the pipeline context (which owns the config and queue) exactly as the
* former inline code did. */
static void server_run_mt_receiver(ServerSession* state) {
Config* config = state->config;
Queue* q = queue_create(100, file_destroy);
if (q == NULL)
return;
state->context = pipeline_context_receiver_create(config, q, state->fd, state->ssl);
if (state->context == NULL) {
queue_destroy(q);
return;
}
protocol_session_set_max_alloc(&state->context->session, config->max_alloc);
protocol_session_set_io_timeout(&state->context->session,
protocol_server_io_timeout_sec(config->timeout));
atomic_store(&state->context->session.total_allocated_bytes,
atomic_load(&state->session.total_allocated_bytes));
pipeline_context_receiver_set_queue_byte_limit(state->context, RECEIVER_QUEUE_MAX_BYTES);
thrd_t receiver = {0};
thrd_t writer = {0};
bool receiver_created = thrd_create(&receiver, receive_thread, state->context) == thrd_success;
bool writer_created = false;
if (receiver_created)
writer_created = thrd_create(&writer, write_thread, state->context) == thrd_success;
if (!receiver_created || !writer_created) {
log_perror("Error creating Threads");
if (receiver_created) {
mtx_lock(&state->context->mutex);
atomic_store(&state->context->cancelled, true);
cnd_broadcast(&state->context->condition_not_full);
cnd_broadcast(&state->context->condition_not_empty);
mtx_unlock(&state->context->mutex);
/* Unblock a worker parked in socket I/O without closing the fd (the
* child owns the single close). shutdown() only affects sockets; for
* the --stdio pipe the receiver's per-message poll timeout still
* bounds the join, so do nothing there rather than close a descriptor
* another thread may still be using. */
struct stat fd_stat;
if (fstat(state->fd, &fd_stat) == 0 && S_ISSOCK(fd_stat.st_mode))
shutdown(state->fd, SHUT_RDWR);
thrd_join(receiver, NULL);
}
if (writer_created)
thrd_join(writer, NULL);
return;
}
int receiver_result;
int writer_result;
thrd_join(receiver, &receiver_result);
thrd_join(writer, &writer_result);
bool transfer_ok = receiver_result == thrd_success && writer_result == thrd_success;
PipelineContextReceiver* context = state->context;
if (transfer_ok && !config->dry_run) {
/* Commit-style (late) deletion: receive_thread handed the keep-set
manifest here instead of deleting while write_thread might still be
draining, so by now every file is on disk and the whole transfer is
known to have succeeded. Remove the extras before publishing a
--delay-updates run; the walker skips the staging directory. A
server-contacting --dry-run deletes nothing (no manifest is sent). */
if (context->deferred_manifest) {
size_t deleted = 0;
DeletePathObserver observer = config->report_deletes ? receiver_record_deleted_path : NULL;
DeleteCommitResult deletion = manifest_delete_all_observed(
config, context->deferred_manifest, &deleted, observer, (void*)context->deleted_paths);
context->stats.deleted_files += deleted;
if (deletion == DELETE_COMMIT_ERROR) {
transfer_ok = false;
} else if (deletion == DELETE_COMMIT_LIMIT_REACHED) {
/* The transfer still succeeds; the terminal frame reports the capped
deletion so the sender exits 25 like rsync. */
context->delete_limit_reached = true;
}
delete_manifest_free(context->deferred_manifest);
context->deferred_manifest = NULL;
}
/* --delete-delay: receive_thread snapshotted each plan's extras as it
arrived; with the disk writer drained, commit the deferred removals.
--delete-during already applied its plans on the receive thread. */
if (context->deferred_plans) {
/* Defence in depth (the enclosing block already excludes dry-run): a
-n run never commits a deletion. */
if (config->report_deletes)
delete_plan_session_set_delete_observer(
context->deferred_plans, receiver_record_deleted_path, (void*)context->deleted_paths);
DeleteCommitResult deletion =
config->dry_run ? DELETE_COMMIT_OK
: delete_plan_session_commit(context->deferred_plans, config);
context->stats.deleted_files += delete_plan_session_deleted(context->deferred_plans);
if (deletion == DELETE_COMMIT_ERROR) {
transfer_ok = false;
} else if (deletion == DELETE_COMMIT_LIMIT_REACHED) {
context->delete_limit_reached = true;
}
delete_plan_session_destroy(context->deferred_plans);
context->deferred_plans = NULL;
}
}
if (transfer_ok && !config->dry_run) {
/* --delay-updates: receive_thread has finished the whole protocol stream
(including manifest/delete handling) and write_thread has drained its
queue, so every staged file is complete. Publish atomically before the
success/outcome frame so a --remove-source-files sender only learns of
files that were actually installed. */
if (config->delay_updates && config->delay_context &&
!delay_updates_publish(config->delay_context, config)) {
transfer_ok = false;
}
/* P7 Wave D: all writers have joined and the late deletion (and
--delay-updates publication) has committed above, so it is finally safe
to stamp directory times; a directory's mtime must not be clobbered by
its children or by an extra removal. */
if (transfer_ok)
dir_metadata_list_apply(&context->dir_times, config->receive_root_directory, config);
}
if (transfer_ok) {
if (context->failed_entries > 0)
log_message(LOG_LEVEL_WARNING,
"%zu entr%s failed to materialize; continuing (partial transfer)",
context->failed_entries, context->failed_entries == 1 ? "y" : "ies");
Status final_status = context->delete_limit_reached
? STATUS_DELETE_LIMIT
: (context->failed_entries > 0 ? STATUS_ERROR : STATUS_OK);
/* Emit the optional wire-stats record first (protocol 2.25.0), then the
success/outcome frame, exactly like the single-threaded receiver. */
if (!receiver_send_stats_frame(state->fd, config, &context->stats, context->would_delete,
context->deleted_paths) ||
!receiver_send_final_success(state->fd, config, &context->outcomes, final_status))
transfer_ok = false;
} else {
send_error_detail(state->fd, "transfer failed on receiver");
}
if (!transfer_ok)
log_message(LOG_LEVEL_ERROR, "Transfer failed");
}
/* Phase 4b -- transfer via the single-threaded receiver. Failure is logged
* exactly as before; the caller's `done` epilogue then releases the config. */
static void server_run_st_receiver(ServerSession* state) {
if (receiver_receive_files(state->config, state->fd) != 0)
log_message(LOG_LEVEL_ERROR, "Transfer failed");
}
/* Phase 4 dispatch -- choose the receiver implementation the config asks for.
* Both helpers own their success/failure logging; the caller falls through to
* the shared `done` epilogue either way. */
static void server_run_transfer(ServerSession* state) {
if (state->config->use_multithreading)
server_run_mt_receiver(state);
else
server_run_st_receiver(state);
}
void handler(int file_descriptor) {
/* Single per-connection state; every phase helper below advances it and
* returns false on a logged failure. All teardown state starts empty so the
* single `done` epilogue is safe to reach from any error path (including
* before the config frame arrives). */
ServerSession state;
state.fd = file_descriptor;
state.ssl = io_get_ssl();
protocol_session_init(&state.session, file_descriptor, file_descriptor);
protocol_session_set_ssl(&state.session, state.ssl);
protocol_session_bind(&state.session);
state.gate_ctx.ssl = state.ssl;
state.gate_ctx.fd = file_descriptor;
state.gate_ctx.super_mode_override = -1;
state.gate_ctx.has_peer_ip = false;
state.gate_ctx.peer_ip[0] = '\0';
state.gate_ctx.is_local = false;
state.config = NULL;
state.context = NULL;
state.joined_destination = NULL;
state.charset_ready = false;
if (!server_accept_config(&state))
goto done;
if (!server_apply_security_gates(&state))
goto done;
if (!server_prepare_session(&state))
goto done;
server_run_transfer(&state);
done:
/* Single cleanup epilogue: every error path jumps here, so the iconv
@@ -1030,38 +1118,63 @@ done:
* connection fd is deliberately NOT closed here -- the child functions own
* its single close (plain_child_fn / tls_child_fn), and the --stdio call
* site must leave stdin/stdout open. */
if (charset_ready)
if (state.charset_ready)
charset_wire_free();
/* The delay-updates staging tree is released by config_delete (which the
branch below always reaches), so it is cleaned exactly once. */
identity_clear_active();
protocol_session_unbind();
if (context != NULL) {
if (state.context != NULL) {
/* context owns both the config and the queue it was created with. */
pipeline_context_receiver_destroy(context);
context = NULL;
config = NULL;
pipeline_context_receiver_destroy(state.context);
state.context = NULL;
state.config = NULL;
} else {
config_delete(config);
config = NULL;
config_delete(state.config);
state.config = NULL;
}
free(joined_destination);
free(state.joined_destination);
}
#ifndef FASTSYNC_SERVER_AS_LIB
static Server* g_server = NULL;
/* Signal handler for the foreground daemon/standalone listener.
*
* Async-signal-safety: _exit(2) is on the POSIX async-signal-safe list and is
* the ONLY thing done here. The previous body called server_delete()
* (close/free/SSL_CTX_free), daemon_conf_free() and credentials_free(); none of
* those (free/malloc, and much of OpenSSL teardown) are async-signal-safe, so a
* signal delivered while the main thread was inside malloc/free could deadlock
* or corrupt the heap.
*
* Residual (documented, not hidden): the in-memory teardown is skipped on the
* signal path. That is safe because the parent daemon owns no persistent
* resource that survives process exit -- the listening socket is closed by the
* kernel, the connection registry is an anonymous MAP_SHARED mapping with no
* named backing object, and the daemon config/credential stores are plain heap
* allocations. Connection children are separate processes and handle their own
* temp files/locks. The normal (non-signal) shutdown path in main() still runs
* the full teardown, so no cleanup is dropped on the common path. Wiring the
* accept loop (transport_tcp.c, outside this change's scope) to a flag-based
* self-pipe shutdown would let the frees run context-safely; it is deliberately
* deferred rather than risk restructuring the daemon loop. */
static void cleanup(int sig) {
(void)sig;
if (g_server)
server_delete(&g_server);
daemon_conf_free(g_daemon_conf);
g_daemon_conf = NULL;
credentials_free(g_credentials);
g_credentials = NULL;
_exit(0);
}
/* Install a signal handler with sigaction(2) (the required async-signal-safe
* install primitive; signal(3) is not specified to be async-signal-safe). */
static void install_cleanup_handler(int signo) {
struct sigaction action;
memset(&action, 0, sizeof(action));
action.sa_handler = cleanup;
sigemptyset(&action.sa_mask);
action.sa_flags = 0;
sigaction(signo, &action, NULL);
}
static void print_server_usage(void) {
printf("FastSync Server\n");
printf("Usage: fastsync-server [options]\n\n");
@@ -1174,13 +1287,19 @@ static bool daemonize(void) {
close(devnull);
}
/* Do not pin the launch CWD (module-relative 'path' entries would resolve
* against an unstable working directory) and drop the restrictive host umask
* so modules can create files/dirs with the modes the config requests. */
* against an unstable working directory). Set a conservative daemon umask
* of 022 (the conventional service default): rsync never forces umask 0 --
* it reads and restores the inherited umask and creates new entries as
* 0777 & ~umask / source & ~umask without -p. Forcing 0 here made every
* implied parent directory world-writable (0777) whenever -p metadata was not
* applied. 022 gives 0755 directories and source&~022 files, matching rsync
* under a normal daemon umask; -p/-a still restore the exact source mode via
* fchmod, which is unaffected by the umask. */
if (chdir("/") != 0)
log_message(LOG_LEVEL_WARNING, "daemon: chdir to / failed: %s", strerror(errno));
umask(0);
umask(022);
/* Refresh the cached umask: main() captured the launch umask before this
* (single-threaded) umask(0), and file_mode_base() must see the daemon's
* (single-threaded) umask(022), and file_mode_base() must see the daemon's
* actual umask. */
file_umask_capture();
return true;
@@ -1249,8 +1368,8 @@ int main(int argc, char* argv[]) {
* this process-global policy cannot be re-enabled by a future caller. */
server_allow_super = opts.allow_super && !opts.stdio_mode;
server_iconv_spec = opts.iconv_spec;
signal(SIGINT, cleanup);
signal(SIGTERM, cleanup);
install_cleanup_handler(SIGINT);
install_cleanup_handler(SIGTERM);
/* Server-owned socket deadline floor: the client default --timeout=0 would
* otherwise leave accepted sockets without SO_RCVTIMEO/SO_SNDTIMEO and let a
* silent peer hold a connection (and its process slot) forever. */
+1 -9
View File
@@ -3,20 +3,12 @@
#include "credentials.h"
#include "utils.h"
#include <limits.h>
#include <stdarg.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/socket.h>
static void set_error(char* err, size_t err_size, const char* fmt, ...) {
if (!err || err_size == 0)
return;
va_list args;
va_start(args, fmt);
vsnprintf(err, err_size, fmt, args);
va_end(args);
}
#define set_error utils_set_error
void server_cli_options_default(ServerCliOptions* opts) {
if (!opts)
+3 -3
View File
@@ -9,7 +9,7 @@
ArrayList* array_list_create(void (*item_destroyer)(void* item)) {
ArrayList* list = (ArrayList*)protocol_alloc(sizeof(ArrayList));
if (list == NULL) {
log_perror("ERROR: Could not allocate memory for array list struct");
log_message(LOG_LEVEL_ERROR, "%s", "ERROR: Could not allocate memory for array list struct");
return NULL;
}
@@ -47,7 +47,7 @@ static bool array_list_extend(ArrayList* array_list) {
new_capacity = INITIAL_ARRAY_SIZE;
void* new_items = protocol_realloc(array_list->items, new_capacity * sizeof(void*));
if (new_items == NULL) {
log_perror("ERROR: Could not reallocate memory for array list items");
log_message(LOG_LEVEL_ERROR, "%s", "ERROR: Could not reallocate memory for array list items");
return false;
}
array_list->items = new_items;
@@ -73,7 +73,7 @@ void** array_list_to_array(const ArrayList* array_list) {
}
void** array = protocol_alloc(array_list->size * sizeof(void*));
if (array == NULL) {
log_perror("Could not malloc space for array from array list!");
log_message(LOG_LEVEL_ERROR, "%s", "Could not malloc space for array from array list!");
return NULL;
}
memcpy(array, array_list->items, array_list->size * sizeof(void*));
+15 -10
View File
@@ -168,9 +168,14 @@ bool charset_spec_valid_direction(const char* from_charset, const char* to_chars
return direction_probe_valid(from_charset, to_charset);
}
/* The receiver's real conversion is wire(client REMOTE) -> server-local (the
* server's own --iconv LOCAL half, or the client's LOCAL half when the server
* has no --iconv). A dedicated pre-ack check so an impossible direction is
/* The receiver's conversion is wire charset -> destination charset. rsync's
* CONVERT_SPEC is LOCAL,REMOTE and "stays the same whether you're pushing or
* pulling", so for a PUSH (FastSync's only direction) the destination end's
* charset is the spec's REMOTE half: the client converts LOCAL -> REMOTE on the
* sender and the receiver writes the wire bytes verbatim. Only a server that
* declares its OWN --iconv (the daemon "charset" analog) has a different local
* charset, and then it is that spec's LOCAL half and the receiver converts
* wire -> server-local. A dedicated pre-ack check so an impossible direction is
* rejected before the connection instead of refusing mid-transfer. */
bool charset_wire_receiver_spec_valid(const char* spec, const char* server_spec) {
if (!spec)
@@ -180,7 +185,7 @@ bool charset_wire_receiver_spec_valid(const char* spec, const char* server_spec)
if (charset_spec_parse(spec, &local, &remote) != 0)
return false;
const char* wire = remote;
const char* target_local = local;
const char* target_local = remote;
char* server_local = NULL;
char* server_remote = NULL;
if (server_spec) {
@@ -302,13 +307,13 @@ bool charset_wire_init_receiver(const char* spec, const char* server_spec) {
char* remote;
if (charset_spec_parse(spec, &local, &remote) != 0)
return false;
/* The wire charset is the client spec's REMOTE half; the local charset is
* the client spec's LOCAL half unless the server was itself started with
* --iconv naming a different local charset (the server halves above never
* travel, so the server's own flag is the only way its local charset can
* differ from what the client assumed). */
/* The wire charset is the client spec's REMOTE half (rsync's LOCAL,REMOTE
* spec stays the same push or pull, so on a push the destination end's
* charset is REMOTE and the receiver writes the wire bytes verbatim). Only a
* server started with its own --iconv declares a different local charset (the
* server halves above never travel), and then it is that spec's LOCAL half. */
const char* wire = remote;
const char* target_local = local;
const char* target_local = remote;
char* server_local = NULL;
char* server_remote = NULL;
if (server_spec) {
+6 -5
View File
@@ -57,8 +57,9 @@ void charset_conversion_close(void* conversion);
/* Process-wide wire conversion. charset_wire_init_sender (client side) opens
* LOCAL->REMOTE; charset_wire_init_receiver (server side) opens
* wire(REMOTE)->server-local. server_spec is the server's own --iconv, whose
* LOCAL half may override the local charset the client assumed; NULL reuses
* the client spec's LOCAL half. Both return false on an unsupported spec.
* LOCAL half overrides the destination charset; NULL means the destination
* charset is the client spec's REMOTE half (rsync's push semantics: the wire
* bytes are written verbatim). Both return false on an unsupported spec.
* The state is freed with charset_wire_free. */
bool charset_wire_init_sender(const char* spec);
bool charset_wire_init_receiver(const char* spec, const char* server_spec);
@@ -66,9 +67,9 @@ void charset_wire_free(void);
bool charset_wire_active(void);
/* Pre-ack receiver-direction sanity (see charset_wire_init_receiver): true
* when the exact wire->server-local conversion the receiver will use (client
* spec's REMOTE half into the server's own LOCAL half, or the client's LOCAL
* half when the server has no --iconv) opens and produces NUL-free output. */
* when the exact wire->destination conversion the receiver will use (client
* spec's REMOTE half into the server's own LOCAL half, or REMOTE->REMOTE when
* the server has no --iconv) opens and produces NUL-free output. */
bool charset_wire_receiver_spec_valid(const char* spec, const char* server_spec);
/* Convert a path across the wire in the process direction. Returns a malloc'd
+53 -12
View File
@@ -1,4 +1,5 @@
#include "checksum.h"
#include "utils.h"
#include <fcntl.h>
#include <openssl/evp.h>
#include <string.h>
@@ -223,17 +224,33 @@ bool checksum_digest_file(ChecksumAlgo algo, uint64_t seed, const char* path, ui
if (fd < 0)
return false;
bool ok = checksum_digest_fd(algo, seed, fd, out, out_capacity, out_len);
close(fd);
return ok;
}
bool checksum_digest_fd(ChecksumAlgo algo, uint64_t seed, int fd, uint8_t* out, size_t out_capacity,
size_t* out_len) {
if (fd < 0 || !out || !out_len || out_capacity < CHECKSUM_MAX_DIGEST_LEN)
return false;
if (algo == CHECKSUM_ALGO_NONE) {
/* No checksum requested: nothing to read; an empty digest succeeds. */
*out_len = 0;
return true;
}
uint8_t buffer[64 * 1024];
bool ok = false;
lseek(fd, 0, SEEK_SET);
if (algo == CHECKSUM_ALGO_MD5) {
if (algo == CHECKSUM_ALGO_MD5 || algo == CHECKSUM_ALGO_SHA1) {
const EVP_MD* md = algo == CHECKSUM_ALGO_MD5 ? EVP_md5() : EVP_sha1();
EVP_MD_CTX* ctx = EVP_MD_CTX_new();
if (!ctx) {
close(fd);
if (!ctx)
return false;
}
unsigned int digest_len = 0;
if (EVP_DigestInit_ex(ctx, EVP_md5(), NULL) == 1) {
if (EVP_DigestInit_ex(ctx, md, NULL) == 1) {
ok = true;
ssize_t got;
while ((got = read(fd, buffer, sizeof(buffer))) > 0) {
@@ -250,7 +267,22 @@ bool checksum_digest_file(ChecksumAlgo algo, uint64_t seed, const char* path, ui
ok = false;
}
EVP_MD_CTX_free(ctx);
close(fd);
return ok;
}
if (algo == CHECKSUM_ALGO_MD4) {
Md4Ctx ctx;
md4_init(&ctx);
ok = true;
ssize_t got;
while ((got = read(fd, buffer, sizeof(buffer))) > 0)
md4_update(&ctx, buffer, (size_t)got);
if (got < 0)
ok = false;
if (ok) {
md4_final(&ctx, out);
*out_len = 16;
}
return ok;
}
@@ -260,16 +292,13 @@ bool checksum_digest_file(ChecksumAlgo algo, uint64_t seed, const char* path, ui
XXH64_reset(&xxh64, seed);
} else if (algo == CHECKSUM_ALGO_XXH3 || algo == CHECKSUM_ALGO_XXH128) {
xxh3 = XXH3_createState();
if (!xxh3) {
close(fd);
if (!xxh3)
return false;
}
if (algo == CHECKSUM_ALGO_XXH3)
XXH3_64bits_reset_withSeed(xxh3, seed);
else
XXH3_128bits_reset_withSeed(xxh3, seed);
} else {
close(fd);
return false;
}
@@ -303,7 +332,6 @@ bool checksum_digest_file(ChecksumAlgo algo, uint64_t seed, const char* path, ui
}
if (xxh3)
XXH3_freeState(xxh3);
close(fd);
return ok;
}
@@ -371,7 +399,7 @@ uint8_t checksum_digest_len(ChecksumAlgo algo) {
return 0;
}
ChecksumAlgo checksum_negotiate_default(void) {
static ChecksumAlgo compiled_checksum_preference_first(void) {
/* rsync 3.4.1 default preference order; every entry is compiled in, so this
* resolves to xxh128. */
static const ChecksumAlgo preference[] = {
@@ -384,3 +412,16 @@ ChecksumAlgo checksum_negotiate_default(void) {
}
return CHECKSUM_ALGO_XXH64;
}
int checksum_choice_resolve(void) {
bool specified = false;
int env = env_choice_first("RSYNC_CHECKSUM_LIST", checksum_algo_from_name, &specified);
if (specified)
return env; /* -1 = the list named no supported checksum */
return (int)compiled_checksum_preference_first();
}
ChecksumAlgo checksum_negotiate_default(void) {
int resolved = checksum_choice_resolve();
return resolved >= 0 ? (ChecksumAlgo)resolved : compiled_checksum_preference_first();
}
+14
View File
@@ -51,6 +51,13 @@ bool checksum_digest(ChecksumAlgo algo, uint64_t seed, const void* data, size_t
bool checksum_digest_file(ChecksumAlgo algo, uint64_t seed, const char* path, uint8_t* out,
size_t out_capacity, size_t* out_len);
/* Descriptor form of the streaming digest: rewinds `fd` to the start and hashes
* to EOF without closing it. Used by the --verify-basis path to hash an
* already-open, root-confined basis descriptor. Same contract as
* checksum_digest_file. */
bool checksum_digest_fd(ChecksumAlgo algo, uint64_t seed, int fd, uint8_t* out, size_t out_capacity,
size_t* out_len);
/* Resolve a --checksum-choice string (case-insensitive) to an algorithm id.
* Accepts "xxh64"/"xxhash", "xxh3", "xxh128", "md5", "md4", "sha1", "none".
* "auto" is not an algorithm here; the caller resolves it to the negotiated
@@ -72,4 +79,11 @@ uint8_t checksum_digest_len(ChecksumAlgo algo);
* xxh128 xxh3 xxh64 md5 md4 sha1 none). Used to resolve "auto". */
ChecksumAlgo checksum_negotiate_default(void);
/* Resolve "auto" the way rsync does: the first supported name in
* RSYNC_CHECKSUM_LIST (whitespace-separated, client half ends at '&'), then the
* compiled-in preference order when the variable is unset/blank. Returns -1
* when the variable is set but names no supported checksum (rsync's failed
* negotiation), otherwise a valid ChecksumAlgo id. */
int checksum_choice_resolve(void);
#endif /* CHECKSUM_H */
+5 -4
View File
@@ -17,8 +17,9 @@
#include "protocol.h"
#include "utils.h"
/* Maximum individual file data size within a chunk (64 MB) */
#define MAX_FILE_DATA_SIZE (64ULL * 1024 * 1024)
/* Maximum individual file data size within a chunk (64 MB). Distinct from the
* receiver's whole-file MAX_FILE_DATA_SIZE (256 MB) in file_receive.c. */
#define MAX_CHUNK_FILE_DATA_SIZE (64ULL * 1024 * 1024)
#define MAX_FILES_PER_CHUNK 65536U
/* Reserve `charge` against `session`'s connection budget. This mirrors the
@@ -382,9 +383,9 @@ Chunk* chunk_deserialize(Data* data, bool use_metadata) {
}
// Reject individual file data larger than the maximum allowed size.
if (file_data_size > MAX_FILE_DATA_SIZE) {
if (file_data_size > MAX_CHUNK_FILE_DATA_SIZE) {
log_message(LOG_LEVEL_ERROR, "File data size %zu exceeds maximum %llu", file_data_size,
(unsigned long long)MAX_FILE_DATA_SIZE);
(unsigned long long)MAX_CHUNK_FILE_DATA_SIZE);
goto error;
}
+62 -4
View File
@@ -2,6 +2,7 @@
#include "data.h"
#include "log.h"
#include "protocol.h"
#include "utils.h"
#include <limits.h>
#include <lz4.h>
#include <stdatomic.h>
@@ -15,7 +16,14 @@
#include <zstd.h>
#define INITIAL_DECOMPRESS_BUF_SIZE (1024 * 1024)
#define MAX_DECOMPRESSED_SIZE (100ULL * 1024 * 1024) /* 100 MB hard ceiling */
/* Hard ceiling for a single decompression. The sender compresses whole files
* up to the protocol's whole-file receive bound, so the decompressor must
* accept payloads that large; referencing the protocol constant keeps the two
* bounds from drifting apart (they previously did: a 100 MB ceiling rejected
* 100-256 MB files). This remains a real bomb guard -- every allocation in the
* paths below is clamped to it -- so it must not exceed the protocol bound. */
#define MAX_DECOMPRESSED_SIZE MAX_RECEIVE_WHOLE_FILE_SIZE
/* rsync 3.4.1's built-in skip-compress suffix list (the `--skip-compress`
* defaults, in the man page's order). rsync stores it as space-separated
@@ -119,7 +127,7 @@ bool compression_algo_enabled(CompressionAlgo algo) {
return algo != COMPRESSION_ALGO_NONE;
}
CompressionAlgo compression_negotiate_default(void) {
static CompressionAlgo compiled_preference_first(void) {
/* rsync 3.4.1 default preference order; every entry is compiled in, so this
* resolves to zstd. */
static const CompressionAlgo preference[] = {
@@ -133,6 +141,57 @@ CompressionAlgo compression_negotiate_default(void) {
return COMPRESSION_ALGO_ZSTD;
}
int compression_choice_resolve(void) {
bool specified = false;
int env = env_choice_first("RSYNC_COMPRESS_LIST", compression_algo_from_name, &specified);
if (specified)
return env; /* -1 = the list named no supported codec */
return (int)compiled_preference_first();
}
CompressionAlgo compression_negotiate_default(void) {
int resolved = compression_choice_resolve();
return resolved >= 0 ? (CompressionAlgo)resolved : compiled_preference_first();
}
int compression_default_level(CompressionAlgo algo) {
switch (algo) {
case COMPRESSION_ALGO_ZSTD:
return ZSTD_CLEVEL_DEFAULT;
case COMPRESSION_ALGO_ZLIB:
case COMPRESSION_ALGO_ZLIBX:
return 6; /* rsync resolves zlib's Z_DEFAULT_COMPRESSION (-1) to 6 */
case COMPRESSION_ALGO_LZ4:
return 1; /* rsync lz4 level is 0/ignored; positive keeps the gate on */
case COMPRESSION_ALGO_NONE:
return 0;
}
return 0;
}
int compression_clamp_level(CompressionAlgo algo, int level) {
switch (algo) {
case COMPRESSION_ALGO_ZSTD:
if (level < 1)
return 1;
if (level > 22)
return 22;
return level;
case COMPRESSION_ALGO_ZLIB:
case COMPRESSION_ALGO_ZLIBX:
if (level < 1)
return 1;
if (level > 9)
return 9;
return level;
case COMPRESSION_ALGO_LZ4:
return 1; /* ignored by lz4_compress; keeps the "compress" gate on */
case COMPRESSION_ALGO_NONE:
return 0;
}
return level;
}
void compression_set_algo(CompressionAlgo algo) {
if (compression_algo_valid((int)algo))
atomic_store(&g_compression_algo, (int)algo);
@@ -585,8 +644,7 @@ static Data* zstd_decompress(Data* compressed_data, size_t maximum_size) {
}
if (ret > 0 && output.pos == output.size) {
if (buf_size >= hard_limit || buf_size > SIZE_MAX / 2) {
log_message(LOG_LEVEL_ERROR, "Decompressed data exceeds %llu bytes",
(unsigned long long)MAX_DECOMPRESSED_SIZE);
log_message(LOG_LEVEL_ERROR, "Decompressed data exceeds %llu bytes", hard_limit);
data_destroy(uncompressed_data);
uncompressed_data = NULL;
goto cleanup;
+20
View File
@@ -35,6 +35,26 @@ bool compression_algo_valid(int algo);
* "auto". */
CompressionAlgo compression_negotiate_default(void);
/* Resolve "auto" the way rsync does: the first supported name in
* RSYNC_COMPRESS_LIST (whitespace-separated, client half ends at '&'), then the
* compiled-in preference order when the variable is unset/blank. Returns -1
* when the variable is set but names no supported codec (rsync's failed
* negotiation), otherwise a valid CompressionAlgo id. */
int compression_choice_resolve(void);
/* rsync 3.4.1's per-codec default level, applied when the user did not pass
* --compress-level/--zl. zstd uses ZSTD_CLEVEL_DEFAULT (3) and zlib/zlibx the
* resolved Z_DEFAULT_COMPRESSION (6). lz4 has no tunable level in rsync
* (always the default acceleration); FastSync returns a positive placeholder so
* its "level > 0" compression gate stays engaged, and lz4_compress ignores the
* value, so the output is identical to rsync's. none is 0. */
int compression_default_level(CompressionAlgo algo);
/* Clamp an explicit --compress-level to the codec's accepted range the way
* rsync's init_compression_level() does: zstd 1..22, zlib/zlibx 1..9, lz4
* ignored (fixed positive placeholder), none 0. */
int compression_clamp_level(CompressionAlgo algo, int level);
/* True when the algorithm actually compresses (i.e. is not NONE). */
bool compression_algo_enabled(CompressionAlgo algo);
+169 -16
View File
@@ -20,9 +20,9 @@
static void config_set_defaults(Config* config) {
config->scanner_threads = 0;
config->metadata_explicitly_disabled = false;
config->preserve_perms_explicit_off = false;
config->preserve_times_explicit_off = false;
config->cli.preserve_perms_explicit_off = false;
config->cli.preserve_times_explicit_off = false;
config->cli.metadata_explicitly_disabled = false;
config->show_progress = false;
config->compression_threads = 0;
config->ssh_port = 22;
@@ -44,8 +44,8 @@ static void config_set_defaults(Config* config) {
config->tls_ca = NULL;
config->server_host = str_dup("127.0.0.1");
config->server_port = 8080;
config->server_port_set = false;
config->server_host_set = false;
config->cli.server_port_set = false;
config->cli.server_host_set = false;
/* rsync defaults: --timeout=0 (I/O timeouts disabled) and --contimeout=60.
* A value of 0 disables the client's own deadline on both the socket layer
* (tcp_set_timeouts) and the protocol layer
@@ -68,8 +68,10 @@ static void config_set_defaults(Config* config) {
config->human_readable = false;
config->ignore_errors = false;
config->ignore_missing_args = false;
config->checksum_transfer_algo = CHECKSUM_ALGO_DEFAULT;
config->cli_exit_code = 0;
config->cli.checksum_transfer_algo = CHECKSUM_ALGO_DEFAULT;
config->cli.cli_exit_code = 0;
config->cli.compression_level_set = false;
config->cli.checksum_choice_set = false;
config->filters = NULL;
config->files_from = NULL;
config->files_from_set = NULL;
@@ -77,13 +79,12 @@ static void config_set_defaults(Config* config) {
config->cvs_exclude = false;
config->per_dir_filter = false;
config->per_dir_filter_count = 0;
config->one_file_system = false;
config->one_file_system = 0;
config->no_implied_dirs = false;
config->dirs = false;
config->rsh_command = NULL;
config->blocking_io = false;
config->outbuf = OUTBUF_BLOCK;
config->old_args = false;
config->remote_options = NULL;
config->remote_option_count = 0;
config->address = NULL;
@@ -99,7 +100,7 @@ static void config_set_defaults(Config* config) {
config->trust_sender = false;
config->stop_after_mins = 0;
config->stop_at = 0;
config->stop_at_set = false;
config->cli.stop_at_set = false;
config->write_batch = NULL;
config->only_write_batch = NULL;
config->read_batch = NULL;
@@ -206,7 +207,8 @@ static bool validate_received_config(const Config* config) {
valid_wire_bool(config->preserve_perms) && valid_wire_bool(config->preserve_times) &&
valid_wire_bool(config->preserve_owner) && valid_wire_bool(config->preserve_group) &&
valid_wire_bool(config->munge_links) && valid_wire_bool(config->keep_dirlinks) &&
valid_wire_bool(config->fake_super) &&
valid_wire_bool(config->fake_super) && valid_wire_bool(config->report_dest_info) &&
valid_wire_bool(config->report_stats) && valid_wire_bool(config->report_deletes) &&
(!config->copy_as_set || (config->copy_as_uid >= 0 && config->copy_as_gid >= 0)) &&
(!config->use_compression ||
(config->compression_level >= 1 && config->compression_level <= 22)) &&
@@ -227,6 +229,13 @@ Config* config_create(void) {
if (!config)
return NULL;
config_set_defaults(config);
/* config_set_defaults() dups the default server host; a failure there leaves
* server_host NULL and would crash later consumers, so fail the whole create
* (every caller already handles a NULL return). */
if (!config->server_host) {
config_delete(config);
return NULL;
}
return config;
}
@@ -332,7 +341,8 @@ bool config_derived_use_metadata(const Config* config) {
config->chown_uid_set || config->chown_gid_set || config->usermap_count > 0 ||
config->groupmap_count > 0 || config->update)
return true;
return (config->use_incremental || config->use_delta) && !config->metadata_explicitly_disabled;
return (config->use_incremental || config->use_delta) &&
!config->cli.metadata_explicitly_disabled;
}
bool config_has_basis(const Config* config) {
@@ -683,9 +693,15 @@ int config_parse_ssh_dest(Config* config) {
return daemon_dest_parse_error("invalid remote destination user@host (must not be empty or "
"start with '-')",
dest);
config->transport = TRANSPORT_SSH;
config->ssh_destination = str_dup(dest);
char* ssh_destination = str_dup(dest);
char* path = str_dup(colon + 1);
if (!ssh_destination || !path) {
free(ssh_destination);
free(path);
return daemon_dest_parse_error("out of memory parsing remote destination", dest);
}
config->transport = TRANSPORT_SSH;
config->ssh_destination = ssh_destination;
free(config->receive_root_directory);
config->receive_root_directory = path;
return 0;
@@ -788,6 +804,8 @@ void config_delete(Config* config) {
if (config->filters) {
array_list_delete(config->filters);
}
filter_rule_list_free(config->protect_rules);
config->protect_rules = NULL;
/* A --delay-updates staging tree is transient receiver state: remove any
leftovers on every exit path (success already emptied it). */
if (config->delay_context)
@@ -1023,6 +1041,134 @@ static bool receive_basis_entries(int fd, Config* c, ConfigStringBudget* budget)
return true;
}
/* Receiver-side delete-protection rules (protocol 2.28.0). The sender compiles
* its command-line selection rules exactly as the scanner does and streams the
* result as one bounded, self-describing block (count + per-rule records); the
* receiver reconstructs a FilterRuleList for the --delete extras walk. owner
* and pattern are charged through the shared ConfigStringBudget and the block
* additionally enforces MAX_FILTER_RULES / MAX_FILTER_BYTES. */
static bool send_protect_entries(int fd, const Config* c) {
int count = c->filters ? c->filters->size : 0;
const char** texts = NULL;
if (count > 0) {
texts = malloc((size_t)count * sizeof(char*));
if (!texts)
return false;
for (int i = 0; i < count; i++)
texts[i] = (const char*)c->filters->items[i];
}
char err[160];
FilterRuleList* rules =
filter_base_build(texts, count, c->cvs_exclude, c->delete_excluded, err, sizeof(err));
free(texts);
if (!rules) {
log_message(LOG_LEVEL_ERROR, "invalid filter rule: %s", err);
return false;
}
/* The receiver rejects any block with more than MAX_FILTER_RULES entries as a
* protocol error; refuse to emit such a frame at all. filter_base_build()
* can expand the client rule set (cvs-exclude, merge files), so this is the
* authoritative bound, not config->filters->size. */
if (rules->count < 0 || rules->count > MAX_FILTER_RULES) {
log_message(LOG_LEVEL_ERROR, "too many filter rules: %d (maximum %d)", rules->count,
MAX_FILTER_RULES);
filter_rule_list_free(rules);
return false;
}
bool ok = send_int(fd, rules->count);
for (int i = 0; ok && i < rules->count; i++) {
const FilterRule* r = rules->items[i];
/* Mirror the receiver's limit so the peer never receives a rule it will
reject as a protocol error. */
if (r->pattern && strlen(r->pattern) > MAX_PROTECT_PATTERN_LEN) {
log_message(LOG_LEVEL_ERROR, "filter pattern exceeds %d bytes", MAX_PROTECT_PATTERN_LEN);
filter_rule_list_free(rules);
return false;
}
ok = send_int(fd, (int)r->action) && send_int(fd, (int)r->sides) &&
send_int(fd, r->anchored ? 1 : 0) && send_int(fd, r->dir_only ? 1 : 0) &&
send_int(fd, r->negate ? 1 : 0) && send_str(fd, r->owner ? r->owner : "") &&
send_str(fd, r->pattern ? r->pattern : "");
}
filter_rule_list_free(rules);
return ok;
}
static bool receive_protect_entries(int fd, Config* c, ConfigStringBudget* budget) {
int count;
if (!receive_int(fd, &count))
return false;
if (count < 0 || count > MAX_FILTER_RULES)
return false;
if (count == 0)
return true;
FilterRuleList* list = filter_rule_list_create();
if (!list)
return false;
size_t pattern_bytes = 0;
for (int i = 0; i < count; i++) {
int action;
int sides;
bool anchored;
bool dir_only;
bool negate;
if (!receive_int(fd, &action) ||
(action != FILTER_ACTION_EXCLUDE && action != FILTER_ACTION_INCLUDE) ||
!receive_int(fd, &sides) || sides < (int)FILTER_SIDE_SENDER ||
sides > (int)(FILTER_SIDE_SENDER | FILTER_SIDE_RECEIVER) ||
!receive_wire_bool(fd, &anchored) || !receive_wire_bool(fd, &dir_only) ||
!receive_wire_bool(fd, &negate))
goto fail;
char* owner = config_receive_str(fd, budget);
if (!owner)
goto fail;
char* pattern = config_receive_str(fd, budget);
if (!pattern || pattern[0] == '\0') {
free(owner);
free(pattern);
goto fail;
}
/* A pattern too long to be evaluated by glob_match against a PATH_MAX path
would silently fail to match and leave a protect rule inert (fail-open:
the entry is then deleted). Reject it up front as a protocol error
rather than accept a rule that can never shield anything. */
if (strlen(pattern) > MAX_PROTECT_PATTERN_LEN) {
free(owner);
free(pattern);
goto fail;
}
size_t bytes = strlen(owner) + strlen(pattern);
if (bytes > MAX_FILTER_BYTES - pattern_bytes) {
free(owner);
free(pattern);
goto fail;
}
pattern_bytes += bytes;
FilterRule* rule = calloc(1, sizeof(FilterRule));
if (!rule) {
free(owner);
free(pattern);
goto fail;
}
rule->action = (FilterAction)action;
rule->sides = (unsigned)sides;
rule->anchored = anchored;
rule->dir_only = dir_only;
rule->negate = negate;
rule->owner = owner;
rule->pattern = pattern;
if (!filter_rule_list_add(list, rule)) {
filter_rule_free(rule);
goto fail;
}
}
c->protect_rules = list;
return true;
fail:
filter_rule_list_free(list);
return false;
}
static bool send_identity_entries(int fd, const IdentityMap* map, int count) {
for (int i = 0; i < count; i++) {
if (!send_int(fd, map[i].from) || !send_int(fd, map[i].from_hi) || !send_int(fd, map[i].to) ||
@@ -1150,6 +1296,9 @@ fail:
#define CONFIG_RECV_BLOCK_IDMAP(name) \
receive_identity_entries(fd, budget, c->name##_count, &c->name)
#define CONFIG_SEND_BLOCK_PROTECT_RULES(name) send_protect_entries(fd, c)
#define CONFIG_RECV_BLOCK_PROTECT_RULES(name) receive_protect_entries(fd, c, budget)
/* One table entry, applied in sequence. XSEND/XRECV are statement macros so
* consecutive entries read as a plain sequence of assignments. */
#define XSEND(name, ctype, def, kind) ok = ok && (CONFIG_SEND_##kind(name));
@@ -1187,6 +1336,7 @@ CONFIG_DEFINE_SEND(send_privilege_options, CONFIG_WIRE_PRIVILEGE_FIELDS)
CONFIG_DEFINE_SEND(send_copy_as_options, CONFIG_WIRE_COPY_AS_FIELDS)
CONFIG_DEFINE_SEND(send_output_options, CONFIG_WIRE_OUTPUT_FIELDS)
CONFIG_DEFINE_SEND(send_codec_options, CONFIG_WIRE_CODEC_FIELDS)
CONFIG_DEFINE_SEND(send_protect_options, CONFIG_WIRE_PROTECT_FIELDS)
CONFIG_DEFINE_RECV(receive_core_fields, CONFIG_WIRE_CORE_FIELDS)
CONFIG_DEFINE_RECV(receive_delta_fields, CONFIG_WIRE_DELTA_FIELDS)
@@ -1207,6 +1357,7 @@ CONFIG_DEFINE_RECV(receive_privilege_options, CONFIG_WIRE_PRIVILEGE_FIELDS)
CONFIG_DEFINE_RECV(receive_copy_as_options, CONFIG_WIRE_COPY_AS_FIELDS)
CONFIG_DEFINE_RECV(receive_output_options, CONFIG_WIRE_OUTPUT_FIELDS)
CONFIG_DEFINE_RECV(receive_codec_options, CONFIG_WIRE_CODEC_FIELDS)
CONFIG_DEFINE_RECV(receive_protect_options, CONFIG_WIRE_PROTECT_FIELDS)
#undef XSEND
#undef XRECV
@@ -1325,7 +1476,8 @@ bool config_send_wire_block(int file_descriptor, const Config* config) {
send_privilege_options(file_descriptor, config) &&
send_copy_as_options(file_descriptor, config) &&
send_output_options(file_descriptor, config) &&
send_codec_options(file_descriptor, config);
send_codec_options(file_descriptor, config) &&
send_protect_options(file_descriptor, config);
}
bool config_send(int file_descriptor, const Config* config) {
@@ -1397,7 +1549,8 @@ Config* config_receive_with_validate(int file_descriptor, ConfigValidateFunc val
!receive_privilege_options(file_descriptor, config, &budget) ||
!receive_copy_as_options(file_descriptor, config, &budget) ||
!receive_output_options(file_descriptor, config, &budget) ||
!receive_codec_options(file_descriptor, config, &budget))
!receive_codec_options(file_descriptor, config, &budget) ||
!receive_protect_options(file_descriptor, config, &budget))
goto error;
/* Validate/normalize the negotiated codec. compress_choice is the human
* spelling (NULL or "" when -z was not given); compression_algo is the
+167 -53
View File
@@ -4,6 +4,7 @@
#include "array_list.h"
#include "checksum.h"
#include "compression.h"
#include "filter.h"
#include <stdbool.h>
#include <stdint.h>
#include <stdio.h>
@@ -82,7 +83,7 @@ typedef struct {
typedef enum SuperMode { SUPER_MODE_AUTO = 0, SUPER_MODE_ON = 1, SUPER_MODE_OFF = 2 } SuperMode;
/* ===========================================================================
* Config wire-field table (single source of truth for protocol 2.26.0).
* Config wire-field table (single source of truth for protocol 2.29.0).
*
* Every field below crosses the wire. The table is the ONLY place a
* serialized field is named: config.h expands CONFIG_WIRE_FIELDS() to declare
@@ -197,9 +198,18 @@ typedef enum SuperMode { SUPER_MODE_AUTO = 0, SUPER_MODE_ON = 1, SUPER_MODE_OFF
X(skip_compress_count, int, 0, INT_SKIPCOUNT) \
X(skip_compress_suffixes, char**, NULL, BLOCK_SKIP_SUFFIXES)
/* FastSync-only --verify-basis (protocol 2.28.0, no version bump by project
* decision): restores the stricter content equality on a basis hit. By
* default a basis hit is accepted on rsync's metadata quick-check alone (equal
* size plus equal mtime, or size alone under --size-only); with this flag the
* receiver ALSO requires the basis bytes' whole-file digest (the negotiated
* --checksum-choice algorithm) to equal the sender's, exactly FastSync's
* historical behavior. It is a receiver policy and crosses the wire so the
* receiver knows whether to read and hash the basis content. */
#define CONFIG_WIRE_BASIS_FIELDS(X) \
X(basis_count, int, 0, INT_BASISCOUNT) \
X(basis_dirs, BasisDest*, NULL, BLOCK_BASIS)
X(basis_dirs, BasisDest*, NULL, BLOCK_BASIS) \
X(verify_basis, bool, false, BOOL)
#define CONFIG_WIRE_FUZZY_FIELDS(X) X(fuzzy, bool, false, BOOL)
@@ -259,9 +269,18 @@ typedef enum SuperMode { SUPER_MODE_AUTO = 0, SUPER_MODE_ON = 1, SUPER_MODE_OFF
* for -n/--dry-run --delete, the destination-relative paths it WOULD have
* deleted. It is set by the client only when --stats, --progress/-P, an
* --out-format token needs a wire counter (%b/%c), or a dry-run carries
* --delete; the transfer decision itself is unchanged. */
* --delete; the transfer decision itself is unchanged.
*
* --info wave (protocol 2.27.0). report_deletes tells the receiver to include
* the destination-relative paths it ACTUALLY removed in its terminal
* STATUS_STATS record (the same path-list field the dry-run would-delete report
* uses), so the sender can print rsync's `deleting PATH`/`*deleting` lines for a
* real (non-dry-run) deletion. It is set when --delete is active and any of
* --info=del, -i/--itemize-changes or --out-format requests per-file change
* output; the transfer decision itself is unchanged. */
#define CONFIG_WIRE_OUTPUT_FIELDS(X) \
X(report_dest_info, bool, false, BOOL) X(report_stats, bool, false, BOOL)
X(report_dest_info, bool, false, BOOL) \
X(report_stats, bool, false, BOOL) X(report_deletes, bool, false, BOOL)
/* Codec-negotiation wave (protocol 2.26.0). compression_algo is the concrete
* codec the client selected for this transfer (a CompressionAlgo id) and is the
@@ -284,6 +303,20 @@ typedef enum SuperMode { SUPER_MODE_AUTO = 0, SUPER_MODE_ON = 1, SUPER_MODE_OFF
#define CONFIG_WIRE_CODEC_FIELDS(X) \
X(compression_algo, int, COMPRESSION_ALGO_ZSTD, INT_COMPRESSION_ALGO)
/* Receiver-side delete-protection filter rules (protocol 2.28.0). The sender
* compiles its root-level selection rules exactly as the scanner does
* (filter_base_build over --filter/-f/--exclude/--include/-C) and streams them
* as one self-describing, bounded block (count followed by per-rule records).
* The receiver reconstructs `protect_rules` and evaluates them against
* DESTINATION-ONLY entries during the --delete extras walk, so a
* `protect`/`P` rule protects an extra that never appeared on the sender
* (rsync re-derives deletion protection from the filter list; FastSync
* historically derived it only from the source scan). `protect_rules` is NULL
* on the sender and is owned/freed by the receiver Config. Bounded by
* MAX_FILTER_RULES and MAX_FILTER_BYTES; an unknown action/sides is a protocol
* error. */
#define CONFIG_WIRE_PROTECT_FIELDS(X) X(protect_rules, FilterRuleList*, NULL, BLOCK_PROTECT_RULES)
/* All serialized fields, in exact wire order. Concatenating the per-segment
* lists here is what keeps the declaration order = the wire order. */
#define CONFIG_WIRE_FIELDS(X) \
@@ -306,7 +339,54 @@ typedef enum SuperMode { SUPER_MODE_AUTO = 0, SUPER_MODE_ON = 1, SUPER_MODE_OFF
CONFIG_WIRE_PRIVILEGE_FIELDS(X) \
CONFIG_WIRE_COPY_AS_FIELDS(X) \
CONFIG_WIRE_OUTPUT_FIELDS(X) \
CONFIG_WIRE_CODEC_FIELDS(X)
CONFIG_WIRE_CODEC_FIELDS(X) \
CONFIG_WIRE_PROTECT_FIELDS(X)
/* Client-only, CLI-parse bookkeeping (never serialized). These members exist
* only so the client command-line parser can record HOW an option was
* specified (explicitly set, explicitly negated, or a parser-requested exit
* code); no other module and no wire peer ever needs them. Grouping them in
* one nested member keeps the public Config free of client-CLI-only state. */
typedef struct {
/* Set when the user explicitly turned an attribute off with --no-perms /
* --no-times (long or short form). --incremental/--delta historically
* auto-enabled mode and mtime preservation; these flags let
* cli_finalize_config restore that behavior while still honoring the
* explicit per-attribute negation. A later -p/-t re-enables the attribute
* directly, so the flag only prevents the incremental/delta implication,
* never a POSITIVE request. */
bool preserve_perms_explicit_off;
bool preserve_times_explicit_off;
/* Set by --no-preserve, the explicit opt-out of the whole preservation
* bundle, so the --incremental/--delta auto-preserve implication stays off. */
bool metadata_explicitly_disabled;
/* True when --server-port/--port was explicitly given. --dry-run uses it to
* decide whether a real server handshake was requested, so a plain local
* destination (no explicit port) keeps the existing client-side dry-run
* behavior instead of dialing the default 127.0.0.1:8080. */
bool server_port_set;
/* True when --server-host was explicitly given, and distinct from the
* "127.0.0.1" default: --dry-run uses it to route an explicit remote target
* to the server so it reports receiver state exactly like a real run,
* instead of silently running the client-side manifest. */
bool server_host_set;
/* Codec-negotiation CLI state. The effective pre-transfer checksum is
* Config->checksum_algo (serialized); checksum_transfer_algo is the rsync
* "transfer" half of a two-name --checksum-choice form (validated and used
* only to mirror rsync's whole-file forcing, since FastSync's per-block
* strong hash is fixed). cli_exit_code carries a parser-requested process
* exit status (rsync uses 4 for an unsupported checksum/compress algorithm)
* so main() can mirror it. */
int checksum_transfer_algo;
int cli_exit_code;
/* "The user explicitly chose" bits. They let the per-codec default level /
* checksum list be applied only when the corresponding rsync option was
* omitted (an explicit --compress-level / --checksum-choice always wins). */
bool compression_level_set;
bool checksum_choice_set;
/* True when --stop-at was given. */
bool stop_at_set;
} ConfigCliParse;
typedef struct Config {
/* -j/--threads=N: number of parallel scanner worker threads for the -m
@@ -314,16 +394,8 @@ typedef struct Config {
* scanner's built-in default" (4). CLIENT-ONLY: it is a local scheduling
* concern and is NEVER serialized into the wire config frame. */
int scanner_threads;
bool metadata_explicitly_disabled;
/* CLIENT-ONLY (never serialized; not in CONFIG_WIRE_FIELDS). Set when the
* user explicitly turned an attribute off with --no-perms / --no-times (long
* or short form). --incremental/--delta historically auto-enabled mode and
* mtime preservation; these flags let cli_finalize_config restore that
* behavior while still honoring the explicit per-attribute negation. A
* later -p/-t re-enables the attribute directly, so the flag only prevents
* the incremental/delta implication, never a POSITIVE request. */
bool preserve_perms_explicit_off;
bool preserve_times_explicit_off;
/* Client-only CLI-parse bookkeeping (never serialized). See ConfigCliParse. */
ConfigCliParse cli;
bool show_progress;
int compression_threads;
int ssh_port;
@@ -344,18 +416,6 @@ typedef struct Config {
bool use_tls;
char* server_host;
int server_port;
/* True when --server-port/--port was explicitly given. CLIENT-ONLY (never
* serialized): --dry-run uses it to decide whether a real server handshake
* was requested, so a plain local destination (no explicit port) keeps the
* existing client-side dry-run behavior instead of dialing the default
* 127.0.0.1:8080. */
bool server_port_set;
/* True when --server-host was explicitly given. CLIENT-ONLY (never
* serialized), and distinct from the "127.0.0.1" default: --dry-run uses it
* to route an explicit remote target to the server so it reports receiver
* state exactly like a real run, instead of silently running the client-side
* manifest. */
bool server_host_set;
char* tls_cert;
char* tls_key;
char* tls_ca;
@@ -401,16 +461,6 @@ typedef struct Config {
* enters the keep-set. Implied by --delete-missing-args. */
bool ignore_missing_args;
/* Codec-negotiation CLI state (all client-only, never serialized). The
* effective pre-transfer checksum is Config->checksum_algo (serialized);
* checksum_transfer_algo is the rsync "transfer" half of a two-name
* --checksum-choice form (validated and used only to mirror rsync's
* whole-file forcing, since FastSync's per-block strong hash is fixed).
* cli_exit_code carries a parser-requested process exit status (rsync uses 4
* for an unsupported checksum/compress algorithm) so main() can mirror it. */
int checksum_transfer_algo;
int cli_exit_code;
// Issue #129: Advanced file selection. These fields are CLIENT-ONLY: they are
// never serialized to the wire (the receiver must not learn them).
ArrayList* filters; /* --filter=RULE rule strings, in order */
@@ -423,9 +473,13 @@ typedef struct Config {
* /.rsync-filter' (the .rsync-filter files themselves are transferred); a
* repeated -F adds --filter='- .rsync-filter' so they are excluded too. */
int per_dir_filter_count;
bool one_file_system; /* -x/--one-file-system: do not cross filesystem boundaries */
/* --no-implied-dirs: client-only. With -R + --files-from, refuse to place a
* listed file whose ancestor directory is not itself explicitly listed. */
int one_file_system; /* -x/--one-file-system: do not cross filesystem boundaries.
Repeated -x (rsync's -xx) drops the mount-point
directory entirely instead of recreating it empty. */
/* --no-implied-dirs: client-only. With -R, do not transfer the source
* metadata of the parent directories implied by a listed path; an unlisted
* implied parent is still created (with default attributes) so the listed
* file can be placed, matching rsync. */
bool no_implied_dirs;
/* -d/--dirs: client-only. Transfer the directory entries named by the
* source argument / --files-from list without recursing into contents. */
@@ -442,7 +496,6 @@ typedef struct Config {
/* --outbuf mode (OutbufMode): stdout/stderr buffering. Client-only launch
* concern: NEVER crosses the wire. */
int outbuf;
bool old_args;
/* --remote-option=OPT (Phase 5, long form only): one or more extra command-line
* options to append to the REMOTE server invocation over SSH. CLIENT-ONLY:
* they are composed into the remote command line by ssh_build_remote_command()
@@ -512,7 +565,6 @@ typedef struct Config {
* process and are NEVER serialized into the config frame. */
int stop_after_mins; /* --stop-after=MINS minutes; 0 when unset */
time_t stop_at; /* --stop-at=... absolute wall-clock deadline */
bool stop_at_set; /* true when --stop-at was given */
/* Client-only residual-batch paths. A residual batch is a self-contained
* single-file record of the whole source tree (full file images using the
@@ -600,10 +652,14 @@ typedef struct Config {
source directory is streamed in directory order, and the receiver removes
each directory's extras when its plan arrives (during) or snapshots them
and removes them only after a successful transfer (delay). delete_after
(and plain --delete) keep the whole-tree commit mode: extras are removed
from a fresh end-of-transfer destination scan only after the whole transfer
succeeded. See config_delete_timing_early()/config_delete_timing_per_dir()
below. */
keeps the whole-tree commit mode: extras are removed from a fresh
end-of-transfer destination scan only after the whole transfer succeeded.
A plain --delete with no explicit timing flag defaults to delete_during on
the client (cli_finalize_config), matching rsync's --del default; the old
late-commit behavior is selected explicitly by --delete-after or the
FastSync-only long spelling --delete-commit (an exact alias for
--delete-after, mapped onto the same wire field). See
config_delete_timing_early()/config_delete_timing_per_dir() below. */
/* partial_dir */
// PR #174: Partial transfer resumption
/* suffix */
@@ -679,11 +735,13 @@ typedef struct Config {
* --copy-as) imply it. */
/* fake_super */
/* --fake-super: receiver-only. When set, each written file additionally gets
* a reserved user.fastsync.stat xattr recording the RESOLVED uid/gid (the
* rsync's reserved user.rsync.%stat xattr recording the RESOLVED uid/gid (the
* source's own when no ownership request is active, else the --chown/--usermap
* result) plus mode/mtime so a later privileged restore could re-apply them.
* It NEVER real-chowns: the point is to record the source ownership on an
* unprivileged receiver. Crosses the wire. */
* result) plus the full mode and rdev, in rsync 3.4.1's grammar, so the tree is
* interoperable and a later privileged restore could re-apply them. mtime is
* carried by the file's own timestamp, exactly as rsync does it. It NEVER
* real-chowns: the point is to record the source ownership on an unprivileged
* receiver. Crosses the wire. */
/* module */
/* Daemon module selection (Wave A, protocol 2.15.0). Client-composed from a
* host::module/path destination; NULL or "" means "no module" (the ordinary
@@ -988,11 +1046,64 @@ typedef struct Config {
* boundary, and the strict same-version handshake (config_receive rejects a
* mismatched version before parsing anything else) keeps mixed deployments from
* ever reaching that state. */
#define PROTOCOL_VERSION "2.26.0"
/* (7) --info=del report (protocol 2.27.0): the config frame gains one trailing
* bool, report_deletes, appended after report_stats. When set, the receiver
* lists the paths it actually removed in the terminal STATUS_STATS path list
* (the same count-delimited list the -n/--dry-run would-delete report uses), so
* the sender can print rsync's `deleting PATH` lines for a real deletion. No
* change to the fixed STATUS_STATS record itself; only a new trailing config
* bool, which still requires the version bump for the strict lockstep. */
/* (8) --stats receiver-observed counters (protocol 2.28.0): the fixed
* STATUS_STATS record grows from three counters to eight. The receiver now
* reports the bytes it literally stored (`literal_data`) and the count of
* destination entries it newly CREATED, split by type
* (reg/dir/link/special), so the sender can print rsync's exact
* `Number of created files: N (reg: X, dir: Y, link: Z, special: W)` line and
* an exact `Literal data` total even for delta transfers. The config-frame
* LAYOUT is unchanged (no new config field), but the STATUS_STATS body grows,
* so a 2.27 peer that does not consume the five new fixed-width counters would
* desynchronize on the trailing would-delete path list; the strict
* same-version handshake (config_receive rejects a mismatched version before
* parsing anything else) keeps mixed deployments from ever reaching that
* state. */
/* (9) Receiver-side delete protection (still protocol 2.28.0): the config frame
* gains one trailing self-describing block carrying the sender's compiled base
* filter rules so the receiver can protect DESTINATION-ONLY entries from
* --delete with `protect`/`risk` rules (rsync parity). The block appends after
* compression_algo; see CONFIG_WIRE_PROTECT_FIELDS. */
/* (10) Symlink xattrs/ACLs (protocol 2.29.0): the config-frame LAYOUT is
* unchanged (the derived use_xattrs bit already crosses the wire), but the
* STATUS_SYMLINK frame BODY grows a trailing bounded xattr block when -X/-A is
* negotiated -- exactly the block STATUS_MKDIR, STATUS_DIR_TIMES and regular
* files already carry. The sender captures the symlink's OWN xattrs with
* llistxattr/lgetxattr (so it can never attach the REFERENT's attributes to the
* link) and the receiver re-applies them to the link itself with lsetxattr on a
* confined /proc/self/fd/<parent>/<leaf> path (there is no *at xattr syscall and
* fsetxattr cannot target a symlink). A 2.28 peer that does not consume the new
* trailing block would desynchronize after every symlink, so the protocol
* version must bump; the strict same-version handshake (config_receive rejects a
* mismatched version before parsing anything else) keeps a 2.29 client and a
* 2.28 server from ever reaching that state. */
#define PROTOCOL_VERSION "2.29.0"
#define DEFAULT_CHUNK_SIZE (10 * 1024 * 1024)
/* Upper bound on total basis-dir entries (rsync caps --link-dest at 20). */
#define MAX_BASIS_DIRS 64
/* Bounds on the received receiver-side delete-protection rule block. The rule
* count and the aggregate pattern+owner bytes are each capped so a hostile
* peer cannot pin unbounded pre-auth memory; both are validated strictly on
* receive (alongside the per-string ConfigStringBudget). */
/* A peer may supply protect rules; cap the list so a crafted config cannot make
* the receiver's delete walk evaluate an unbounded number of glob patterns per
* destination entry (glob_match is O(pattern x path)). 1024 is far above any
* legitimate selection. */
#define MAX_FILTER_RULES 1024
#define MAX_FILTER_BYTES (256 * 1024)
/* glob_match's DP is capped at 64 Mi work units; a pattern longer than this
* could exceed the cap against a PATH_MAX path and silently stop matching,
* leaving a protect rule inert. Reject such a rule at receive time. */
#define MAX_PROTECT_PATTERN_LEN 8192
/* Upper bound on the number of --skip-compress suffixes accepted from the wire.
* Each suffix is an independent wire string (up to MAX_STRING_SIZE = 64 KiB), so
* without this a hostile pre-auth client could otherwise retain
@@ -1094,8 +1205,11 @@ bool config_delete_timing_early(const Config* config);
* commits them only after a fully-successful transfer (delay). */
bool config_delete_timing_per_dir(const Config* config);
/* Delete-timing sanity: with deletion enabled at most one timing flag may be
* set (none = the default delete-after commit timing); without deletion no
* timing flag may be set (each timing flag implies --delete). */
* set; without deletion no timing flag may be set (each timing flag implies
* --delete). A plain --delete is normalized to delete_during by
* cli_finalize_config on the client, so a transmitted use_delete config always
* carries exactly one timing; the zero-timing case remains valid only for a
* config that has not been through the CLI. */
bool config_has_valid_delete_timing(const Config* config);
/* Single source of truth for the cross-field ("combination") invariants a
+204 -18
View File
@@ -8,12 +8,13 @@
#include <openssl/evp.h>
#include <openssl/params.h>
#include <openssl/rand.h>
#include <stdarg.h>
#include <poll.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/stat.h>
#include <time.h>
#include <unistd.h>
/* One store entry: a username and its salted PBKDF2 verifier. The plaintext
@@ -58,19 +59,37 @@ struct CredentialStore {
static const uint8_t k_dummy_stored_key[CREDENTIAL_KEY_LEN] = {0};
static const uint8_t k_dummy_server_key[CREDENTIAL_KEY_LEN] = {0};
static void set_error(char* err, size_t err_size, const char* fmt, ...) {
if (!err || err_size == 0)
return;
va_list args;
va_start(args, fmt);
vsnprintf(err, err_size, fmt, args);
va_end(args);
}
#define set_error utils_set_error
static bool is_comment_char(char c) {
return c == '#' || c == ';';
}
/* True for a literal fd-backed store path: exactly "/dev/fd/<digits>" or
* "/proc/self/fd/<digits>", with no trailing component and no "..". These name
* the calling process's own open descriptors (e.g. a bash process substitution
* `<(...)`, which passes /dev/fd/N), and both prefixes are symlinks by
* construction. */
static bool is_fd_backed_path(const char* path) {
static const char* const prefixes[] = {"/dev/fd/", "/proc/self/fd/"};
if (!path)
return false;
for (size_t i = 0; i < sizeof(prefixes) / sizeof(prefixes[0]); i++) {
const char* prefix = prefixes[i];
size_t prefix_len = strlen(prefix);
if (strncmp(path, prefix, prefix_len) != 0)
continue;
const char* digits = path + prefix_len;
if (*digits < '0' || *digits > '9')
return false;
const char* p = digits;
while (*p >= '0' && *p <= '9')
p++;
return *p == '\0';
}
return false;
}
/* Open a --password-file / --early-input after verifying the EXACT inode we
* will read: it must be owned by the effective user and grant no group/other
* permission bit (so 0600 and stricter modes such as 0400 are accepted),
@@ -80,12 +99,31 @@ static bool is_comment_char(char c) {
* path and then fstat the resulting fd (rather than stat()ing the path first
* and reopening it), so the permission decision is made on the same inode that
* is read and cannot be raced by swapping the path between check and open.
* The path may be a process-substitution pipe (`<(...)` -> /dev/fd/N), so
* regular files and FIFOs are accepted when the ownership/mode checks pass.
* O_NOFOLLOW refuses a symlinked path outright (ELOOP fails closed) instead of
* following it before the owner/mode gate can run. The one exception is a
* literal fd-backed path (/dev/fd/N or /proc/self/fd/N, see
* is_fd_backed_path): those entries are symlinks to the CALLING process's own
* descriptors, so following them is not the untrusted-symlink hazard
* O_NOFOLLOW guards against, and requiring O_NOFOLLOW would break the
* documented process-substitution/FIFO usage. For them only, O_NOFOLLOW is
* omitted; the same fstat owner/mode gate still applies to the resolved inode.
* O_NONBLOCK keeps the OPEN itself from
* blocking forever on a writer-less FIFO (a blocking O_RDONLY open would wait
* for a writer). The fd is left nonblocking for FIFOs so a read never blocks
* either; the read loop (secret_read_line) absorbs the resulting EAGAIN by
* waiting, under a bounded deadline, for the writer -- this is what makes a
* slow process substitution (`--password-file <(sleep 1; ...)`) work while a
* writer-less FIFO still fails after the deadline instead of hanging. Only
* regular files and FIFOs pass the ownership/mode checks; O_NONBLOCK is
* cleared for regular files, where it is a no-op anyway and no EAGAIN can
* occur, so their stdio read path is byte-for-byte unchanged.
*
* Returns a FILE* the caller must fclose, or NULL with `err` filled. */
static FILE* secret_file_open(const char* path, char* err, size_t err_size) {
int fd = open(path, O_RDONLY | O_CLOEXEC);
int flags = O_RDONLY | O_NONBLOCK | O_CLOEXEC;
if (!is_fd_backed_path(path))
flags |= O_NOFOLLOW;
int fd = open(path, flags);
if (fd < 0) {
set_error(err, err_size, "cannot open secret file '%s': %s", path, strerror(errno));
return NULL;
@@ -105,6 +143,15 @@ static FILE* secret_file_open(const char* path, char* err, size_t err_size) {
close(fd);
return NULL;
}
/* O_NONBLOCK is only meaningful for the FIFO allowance. Restore blocking
* mode on a regular file so its read path is exactly as before; a no-op on
* most systems, but explicit. Failures here are ignored: O_NONBLOCK on a
* regular file does not affect reads either way. */
if (S_ISREG(st.st_mode)) {
int status_flags = fcntl(fd, F_GETFL);
if (status_flags >= 0)
(void)fcntl(fd, F_SETFL, status_flags & ~O_NONBLOCK);
}
FILE* fp = fdopen(fd, "r");
if (!fp) {
set_error(err, err_size, "cannot read secret file '%s': %s", path, strerror(errno));
@@ -114,6 +161,123 @@ static FILE* secret_file_open(const char* path, char* err, size_t err_size) {
return fp;
}
/* Overall bound on how long the reader waits for a process-substitution/FIFO
* writer to produce data before giving up. It must comfortably exceed a
* producer's startup delay (e.g. `--password-file <(sleep 1; ...)`) while still
* bounding a writer-less FIFO, so a stray or hostile FIFO cannot stall the
* daemon or client indefinitely. */
#define CREDENTIAL_FIFO_READ_TIMEOUT_MS 3000
/* Monotonic milliseconds, used only for the read deadline (wall-clock changes
* must not extend or shorten the wait). */
static int64_t credential_monotonic_ms(void) {
struct timespec ts;
if (clock_gettime(CLOCK_MONOTONIC, &ts) != 0)
return 0;
return (int64_t)ts.tv_sec * 1000 + (int64_t)(ts.tv_nsec / 1000000);
}
/* Wait until `fd` is readable or the deadline passes. Returns true when it is
* readable, false on timeout or a poll error (err filled). EINTR is retried
* against the same deadline, so signals cannot extend the wait. */
static bool credential_wait_readable(int fd, int64_t deadline, const char* label, const char* path,
char* err, size_t err_size) {
for (;;) {
int64_t remaining = deadline - credential_monotonic_ms();
if (remaining <= 0)
break;
if (remaining > INT_MAX)
remaining = INT_MAX;
struct pollfd pfd = {.fd = fd, .events = POLLIN, .revents = 0};
int rc = poll(&pfd, 1, (int)remaining);
if (rc > 0)
return true;
if (rc == 0)
break;
if (errno != EINTR) {
set_error(err, err_size, "error waiting for %s '%s': %s", label, path, strerror(errno));
return false;
}
}
set_error(err, err_size, "timed out after %d ms waiting for %s '%s'",
CREDENTIAL_FIFO_READ_TIMEOUT_MS, label, path);
return false;
}
typedef enum {
SECRET_READ_LINE,
SECRET_READ_EOF,
SECRET_READ_ERROR,
} SecretReadResult;
/* Read one complete line from `fp` into `line` (capacity `cap`), including the
* trailing newline when present and always NUL-terminating. `*out_len`
* receives strlen(line).
*
* A regular file is read exactly as before: secret_file_open leaves it
* blocking, so fgets never sees EAGAIN. A FIFO stays nonblocking, so fgets
* returns NULL (or a partial line) with EAGAIN while the writer is still
* starting up; instead of treating that as a fatal error the loop clearerr()s
* and polls for readability against one overall deadline. The `used`
* accumulator reassembles a line that arrived in several write()s into a single
* line, so a split write is not misparsed as two entries.
*
* Returns SECRET_READ_LINE, SECRET_READ_EOF, or SECRET_READ_ERROR (err filled)
* on timeout or a genuine read error. */
static SecretReadResult secret_read_line(char* line, size_t cap, FILE* fp, const char* label,
const char* path, size_t* out_len, char* err,
size_t err_size) {
int fd = fileno(fp);
int64_t deadline = credential_monotonic_ms() + CREDENTIAL_FIFO_READ_TIMEOUT_MS;
size_t used = 0;
line[0] = '\0';
for (;;) {
errno = 0;
if (fgets(line + used, (int)(cap - used), fp)) {
used += strlen(line + used);
if (used > 0 && line[used - 1] == '\n') {
*out_len = used;
return SECRET_READ_LINE;
}
if (feof(fp)) {
*out_len = used; /* final unterminated line */
return SECRET_READ_LINE;
}
/* No newline and not EOF. A full buffer is the caller's over-long-line
* case; otherwise the line is only partially available (a nonblocking
* FIFO under a slow writer), so any genuine read error fails and anything
* else waits for the rest. */
if (used >= cap - 1) {
*out_len = used;
return SECRET_READ_LINE;
}
int e = ferror(fp) ? errno : 0;
if (e != 0 && e != EAGAIN && e != EWOULDBLOCK) {
set_error(err, err_size, "error reading %s '%s': %s", label, path, strerror(e));
return SECRET_READ_ERROR;
}
clearerr(fp);
if (!credential_wait_readable(fd, deadline, label, path, err, err_size))
return SECRET_READ_ERROR;
continue;
}
/* fgets returned NULL: EOF, a not-yet-readable FIFO, or a real error. */
if (feof(fp)) {
*out_len = used;
return used > 0 ? SECRET_READ_LINE : SECRET_READ_EOF;
}
if (errno == EAGAIN || errno == EWOULDBLOCK) {
clearerr(fp);
if (!credential_wait_readable(fd, deadline, label, path, err, err_size))
return SECRET_READ_ERROR;
continue;
}
set_error(err, err_size, "error reading %s '%s': %s", label, path,
errno != 0 ? strerror(errno) : "read failed");
return SECRET_READ_ERROR;
}
}
/* Trim leading/trailing ASCII space and tab in place; returns the new start. */
static char* trim_space(char* s) {
while (*s == ' ' || *s == '\t')
@@ -509,9 +673,17 @@ static CredentialStore* load_store_file(const char* path, char* err, size_t err_
char line[CREDENTIAL_MAX_LINE + 2];
bool ok = true;
while (fgets(line, sizeof(line), fp)) {
for (;;) {
size_t len = 0;
SecretReadResult rr =
secret_read_line(line, sizeof(line), fp, "credential file", path, &len, err, err_size);
if (rr == SECRET_READ_EOF)
break;
if (rr == SECRET_READ_ERROR) {
ok = false;
break;
}
line_no++;
size_t len = strlen(line);
if (len == CREDENTIAL_MAX_LINE + 1 && line[len - 1] != '\n' && !feof(fp)) {
set_error(err, err_size, "credential file '%s' line %d exceeds the %d-byte limit", path,
line_no, CREDENTIAL_MAX_LINE);
@@ -1148,9 +1320,17 @@ int credentials_hash_file(const char* path, uint32_t iters, FILE* out, char* err
int line_no = 0;
int result = 0;
char line[CREDENTIAL_MAX_LINE + 2];
while (fgets(line, sizeof(line), fp)) {
for (;;) {
size_t len = 0;
SecretReadResult rr =
secret_read_line(line, sizeof(line), fp, "plaintext file", path, &len, err, err_size);
if (rr == SECRET_READ_EOF)
break;
if (rr == SECRET_READ_ERROR) {
result = -1;
break;
}
line_no++;
size_t len = strlen(line);
if (len == CREDENTIAL_MAX_LINE + 1 && line[len - 1] != '\n' && !feof(fp)) {
set_error(err, err_size, "plaintext file '%s' line %d exceeds the %d-byte limit", path,
line_no, CREDENTIAL_MAX_LINE);
@@ -1228,9 +1408,15 @@ int credentials_read_secret_file(const char* path, char** user_out, char** passw
char line[CREDENTIAL_MAX_LINE + 2];
int result = -1;
while (fgets(line, sizeof(line), fp)) {
for (;;) {
size_t len = 0;
SecretReadResult rr =
secret_read_line(line, sizeof(line), fp, "password file", path, &len, err, err_size);
if (rr == SECRET_READ_EOF)
break;
if (rr == SECRET_READ_ERROR)
goto done;
line_no++;
size_t len = strlen(line);
if (len == CREDENTIAL_MAX_LINE + 1 && line[len - 1] != '\n' && !feof(fp)) {
set_error(err, err_size, "password file '%s' line %d exceeds the %d-byte limit", path,
line_no, CREDENTIAL_MAX_LINE);
+217 -10
View File
@@ -1,12 +1,12 @@
#include "daemon_conf.h"
#include "credentials.h"
#include "log.h"
#include "utils.h"
#include <arpa/inet.h>
#include <ctype.h>
#include <errno.h>
#include <limits.h>
#include <netinet/in.h>
#include <stdarg.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
@@ -17,14 +17,7 @@
/* helpers */
/* ------------------------------------------------------------------ */
static void set_error(char* err, size_t err_size, const char* fmt, ...) {
if (!err || err_size == 0)
return;
va_list args;
va_start(args, fmt);
vsnprintf(err, err_size, fmt, args);
va_end(args);
}
#define set_error utils_set_error
/* Trim leading and trailing ASCII space/tab in place; returns the new start. */
static char* trim_ws(char* s) {
@@ -41,6 +34,158 @@ static bool key_equals(const char* key, const char* canonical) {
return strcasecmp(key, canonical) == 0;
}
/* True when `key` matches one of the NUL-terminated names in `list`. */
static bool key_in_list(const char* key, const char* const* list, size_t count) {
for (size_t i = 0; i < count; i++) {
if (strcasecmp(key, list[i]) == 0)
return true;
}
return false;
}
/* rsync 3.4.1 rsyncd.conf GLOBAL keys accepted in the pre-module section that
* have no FastSync equivalent. They are recognized and documented as inert:
* accepting a real rsync config must not fail on a logging/process key, but a
* silently-reinterpreted key is never invented. `pidfile`/`logfile` are the
* compact --dparam spellings rsync documents. The same list is used by the
* `--dparam` dispatch (apply_global_key), so there is a single impl. */
static const char* const kRsyncInertGlobalKeys[] = {
"pid file",
"pidfile",
"log file",
"logfile",
"socket options",
"sockopts",
"listen backlog",
"syslog facility",
"syslog tag",
"log format",
"use chroot",
"uid",
"gid",
"timeout",
"max verbosity",
"min verbosity",
"lock file",
"transfer logging",
"strict modes",
"reverse lookup",
"forward lookup",
"ignore errors",
"ignore nonreadable",
"dont compress",
};
/* rsync 3.4.1 rsyncd.conf MODULE keys accepted in a [module] section that have
* no FastSync equivalent (accepted-and-documented inert). Keys with a FastSync
* meaning (`path`, `read only`, `write only`, `auth users`, `max connections`,
* `hosts allow`/`hosts deny`, `client owner`) are handled by apply_module_key
* before this list is consulted. Security-relevant keys (`exclude`, `filter`,
* `secrets file`, `refuse options`, ...) are inert, so a daemon-side filter or
* rsync secrets file is NOT enforced: each is loudly warned about at load time
* (see kRsyncUnenforcedModuleSecurityKeys) and documented as a residual in
* RSYNC_COMPAT.md. */
static const char* const kRsyncInertModuleKeys[] = {
"comment",
"use chroot",
"daemon chroot",
"uid",
"gid",
"daemon uid",
"daemon gid",
"exclude",
"include",
"exclude from",
"include from",
"filter",
"max verbosity",
"min verbosity",
"lock file",
"transfer logging",
"log file",
"log format",
"syslog facility",
"syslog tag",
"timeout",
"secrets file",
"auth digest",
"strict modes",
"numeric ids",
"fake super",
"munge symlinks",
"list",
"dont compress",
"charset",
"refuse options",
"incoming chmod",
"outgoing chmod",
"open noatime",
"max size",
"min size",
"temp dir",
"pre-xfer exec",
"post-xfer exec",
"name converter",
"proxy protocol",
"proxy protocol hosts",
"reverse lookup",
"forward lookup",
"ignore errors",
"ignore nonreadable",
};
/* Subset of the inert rsync keys whose intent is access control (data
* visibility, credential source, transfer hooks, daemon privilege), plus the
* global keys that shape the daemon's privilege/identity. These load for
* rsync-config compatibility, but because FastSync ignores them an operator
* migrating a hardened rsyncd.conf must not believe the restriction applies.
* The loader emits one LOG_LEVEL_WARNING per occurrence naming the key (and the
* module, for a module key). `write only` is deliberately absent: it is mapped
* onto writability instead (FastSync is push-only, so a write-only module is
* simply writable). */
static const char* const kRsyncUnenforcedModuleSecurityKeys[] = {
"secrets file",
"auth digest",
"refuse options",
"exclude",
"include",
"exclude from",
"include from",
"filter",
"max size",
"min size",
"pre-xfer exec",
"post-xfer exec",
"incoming chmod",
"outgoing chmod",
"name converter",
"use chroot",
"daemon chroot",
"uid",
"gid",
"daemon uid",
"daemon gid",
"munge symlinks",
"fake super",
"strict modes",
"proxy protocol",
"proxy protocol hosts",
};
static const char* const kRsyncUnenforcedGlobalSecurityKeys[] = {
"use chroot",
"uid",
"gid",
"strict modes",
};
#define kRsyncInertGlobalCount (sizeof(kRsyncInertGlobalKeys) / sizeof(kRsyncInertGlobalKeys[0]))
#define kRsyncInertModuleCount (sizeof(kRsyncInertModuleKeys) / sizeof(kRsyncInertModuleKeys[0]))
#define kRsyncUnenforcedModuleSecurityCount \
(sizeof(kRsyncUnenforcedModuleSecurityKeys) / sizeof(kRsyncUnenforcedModuleSecurityKeys[0]))
#define kRsyncUnenforcedGlobalSecurityCount \
(sizeof(kRsyncUnenforcedGlobalSecurityKeys) / sizeof(kRsyncUnenforcedGlobalSecurityKeys[0]))
static bool parse_bool_value(const char* value, bool* out) {
if (strcasecmp(value, "yes") == 0 || strcasecmp(value, "true") == 0 || strcmp(value, "1") == 0) {
*out = true;
@@ -258,6 +403,10 @@ DaemonConf* daemon_conf_create(void) {
if (!conf)
return NULL;
conf->global.port = DAEMON_CONF_DEFAULT_PORT;
/* rsync modules are READ-ONLY unless `read only = no` (or `write only = yes`)
* is set, so FastSync must default the same way: a migrated rsyncd.conf that
* omits `read only` is served read-only, never writable. */
conf->global.read_only_default = true;
conf->global.max_connections = DAEMON_CONF_DEFAULT_MAX_CONNECTIONS;
conf->global.auth_failure_delay_ms = DAEMON_CONF_DEFAULT_AUTH_FAILURE_DELAY_MS;
conf->global.max_connections_per_host = DAEMON_CONF_DEFAULT_MAX_CONNECTIONS_PER_HOST;
@@ -332,7 +481,7 @@ static bool apply_global_key(DaemonConf* conf, char* key, const char* value, boo
char* err, size_t err_size) {
if (key_equals(key, "port"))
return store_port(&conf->global.port, value, err, err_size);
if (key_equals(key, "motd file")) {
if (key_equals(key, "motd file") || key_equals(key, "motdfile")) {
if (!store_string(&conf->global.motd_file, value)) {
set_error(err, err_size, "out of memory parsing 'motd file'");
return false;
@@ -346,6 +495,25 @@ static bool apply_global_key(DaemonConf* conf, char* key, const char* value, boo
}
return true;
}
/* rsync allows the `read only` module key in the global section as the
* default for modules defined after it. Map it to that default (a later
* --dparam re-applies it to modules that did not set their own value) so a
* global `read only = yes` cannot be silently dropped into a writable
* default. */
if (key_equals(key, "read only")) {
bool parsed;
if (!parse_bool_value(value, &parsed)) {
set_error(err, err_size, "global 'read only' must be yes/no (or true/false/1/0), got '%s'",
value);
return false;
}
conf->global.read_only_default = parsed;
for (int i = 0; i < conf->module_count; i++) {
if (!conf->modules[i].read_only_explicit)
conf->modules[i].read_only = parsed;
}
return true;
}
if (key_equals(key, "max connections"))
return store_max_connections(&conf->global.max_connections, value, NULL, err, err_size);
if (key_equals(key, "max connections per host"))
@@ -368,6 +536,15 @@ static bool apply_global_key(DaemonConf* conf, char* key, const char* value, boo
if (key_equals(key, "hosts deny"))
return store_host_list(&conf->global.hosts_deny, &conf->global.hosts_deny_count, value,
"hosts deny", NULL, replace_hosts, err, err_size);
/* A recognized rsync global key with no FastSync equivalent loads inert. */
if (key_in_list(key, kRsyncInertGlobalKeys, kRsyncInertGlobalCount)) {
if (key_in_list(key, kRsyncUnenforcedGlobalSecurityKeys, kRsyncUnenforcedGlobalSecurityCount))
log_message(LOG_LEVEL_WARNING,
"daemon config: global key '%s' is accepted for rsync compatibility but is NOT "
"enforced by FastSync; the restriction it expresses will not be applied",
key);
return true;
}
set_error(err, err_size, "unknown global key '%s'", key);
return false;
}
@@ -396,6 +573,26 @@ static bool apply_module_key(DaemonModule* module, char* key, char* value, char*
return false;
}
module->read_only = parsed;
module->read_only_explicit = true;
return true;
}
/* rsync's `write only = yes` makes the module client-writable. FastSync has
* no read/pull path, so mapping it to writability is the exact
* security-relevant effect; set `read_only_explicit` so a global default
* cannot override the module's explicit choice. `write only = no` is the
* rsync default and leaves the module's read-only state untouched. */
if (key_equals(key, "write only")) {
bool parsed;
if (!parse_bool_value(value, &parsed)) {
set_error(err, err_size,
"module '%s': 'write only' must be yes/no (or true/false/1/0), got '%s'",
module->name, value);
return false;
}
if (parsed) {
module->read_only = false;
module->read_only_explicit = true;
}
return true;
}
if (key_equals(key, "client owner")) {
@@ -465,6 +662,15 @@ static bool apply_module_key(DaemonModule* module, char* key, char* value, char*
if (key_equals(key, "hosts deny"))
return store_host_list(&module->hosts_deny, &module->hosts_deny_count, value, "hosts deny",
module->name, false, err, err_size);
/* A recognized rsync module key with no FastSync equivalent loads inert. */
if (key_in_list(key, kRsyncInertModuleKeys, kRsyncInertModuleCount)) {
if (key_in_list(key, kRsyncUnenforcedModuleSecurityKeys, kRsyncUnenforcedModuleSecurityCount))
log_message(LOG_LEVEL_WARNING,
"daemon config: module '%s' key '%s' is accepted for rsync compatibility but is "
"NOT enforced by FastSync; the restriction it expresses will not be applied",
module->name, key);
return true;
}
set_error(err, err_size, "unknown key '%s' in module '%s'", key, module->name);
return false;
}
@@ -515,6 +721,7 @@ static int open_module(DaemonConf* conf, int* current_module, const char* name,
}
conf->modules = grown;
memset(&conf->modules[conf->module_count], 0, sizeof(DaemonModule));
conf->modules[conf->module_count].read_only = conf->global.read_only_default;
conf->modules[conf->module_count].name = str_dup(name);
if (!conf->modules[conf->module_count].name) {
set_error(err, err_size, "out of memory adding module '%s'", name);
+39 -13
View File
@@ -17,7 +17,20 @@
* DAEMON_CONF_MAX_LINE all fail the whole load with a clear, line-numbered
* error instead of being silently ignored. This keeps a typo from silently
* changing what a module serves.
*/
*
* rsync compatibility: to reduce the divergence from rsync 3.4.1's rsyncd.conf
* grammar, the parser also ACCEPTS the common rsync GLOBAL and MODULE keys.
* Keys with a FastSync equivalent are mapped onto it (the native spellings are
* unchanged; `read only` defaults to yes like rsync, and `write only = yes`
* opts a module into writability). Keys with no FastSync equivalent are
* accepted and documented as inert (they load successfully but have no effect)
* rather than failing the whole config; the accepted inert set is listed in
* kRsyncInertGlobalKeys / kRsyncInertModuleKeys in daemon_conf.c and in
* RSYNC_COMPAT.md. Every inert key whose intent is access control is loudly
* warned about at load time (kRsyncUnenforced*SecurityKeys) so an operator
* migrating a hardened rsyncd.conf is never misled into believing the
* restriction is enforced. A key outside both the FastSync-native grammar and
* the recognized rsync subset is still rejected as unknown. */
/* A daemon module's configured root is used exactly like the standalone
* server's --destination-root: the daemon confines every connection that
@@ -42,15 +55,21 @@
* store refuses (fail closed) rather than falling open; see server.c. Auth is
* never bypassed by ignoring the list. */
typedef struct DaemonModule {
char* name; /* module name, as the client requests it */
char* path; /* module root (daemon-side authorized root) */
bool read_only; /* `read only = yes/no`; default no */
bool client_owner; /* `client owner = yes/no`; default no. Per-module opt-in
that lets this module's clients choose ownership
(--numeric-ids/--chown/--usermap/--groupmap/--fake-super/
--copy-as) and request explicit --super super-user
activities. Without it the daemon refuses all of them. */
char** auth_users; /* `auth users = a,b`; Wave B credential list */
char* name; /* module name, as the client requests it */
char* path; /* module root (daemon-side authorized root) */
bool read_only; /* `read only = yes/no`; defaults to the global `read only`
default (rsync allows it in the global section), which is
itself default YES (rsync modules are read-only unless
`read only = no` / `write only = yes` opts in) */
bool read_only_explicit; /* set when this module set its own `read only` or
`write only = yes`, so a later global default (from a
`--dparam read only=`) does not override it */
bool client_owner; /* `client owner = yes/no`; default no. Per-module opt-in
that lets this module's clients choose ownership
(--numeric-ids/--chown/--usermap/--groupmap/--fake-super/
--copy-as) and request explicit --super super-user
activities. Without it the daemon refuses all of them. */
char** auth_users; /* `auth users = a,b`; Wave B credential list */
int auth_user_count;
/* `max connections = N` (optional per-module cap). 0 means unlimited. The
* per-connection child records the selected module in the shared registry
@@ -69,6 +88,10 @@ typedef struct DaemonConfGlobals {
int port; /* `port`, default DAEMON_CONF_DEFAULT_PORT (873) */
char* motd_file; /* `motd file`, may be NULL */
char* address; /* `address` (optional bind address), may be NULL */
bool read_only_default; /* global `read only` default for modules defined
after it (rsync allows the module key in the
global section); default YES to match rsync's
read-only modules */
int max_connections; /* `max connections`, default
DAEMON_CONF_DEFAULT_MAX_CONNECTIONS (100) */
int auth_failure_delay_ms; /* `auth failure delay`, milliseconds; default
@@ -152,10 +175,13 @@ const DaemonModule* daemon_conf_find_module(const DaemonConf* conf, const char*
bool daemon_module_name_valid(const char* name);
/* Parse one --dparam=KEY=VALUE (or "--dparam KEY=VALUE") override string and
* apply it to the global keys only. Keys are case-insensitive and limited to
* the global keys defined by the grammar (port, motd file, address,
* apply it to the global keys only. Keys are case-insensitive and cover the
* global keys defined by the grammar (port, motd file, address, read only,
* max connections, max connections per host, auth failure delay,
* auth lockout threshold, auth lockout duration, hosts allow, hosts deny).
* auth lockout threshold, auth lockout duration, hosts allow, hosts deny) plus
* the recognized inert rsync global keys and the compact rsync spellings
* (`motdfile`, `pidfile`, `logfile`). Applying `read only` sets the global
* default and re-applies it to every module that did not set its own value.
* Returns 0 on success, -1 on error (err filled). */
int daemon_conf_apply_dparam(DaemonConf* conf, const char* assignment, char* err, size_t err_size);
+656
View File
@@ -0,0 +1,656 @@
#include "delete.h"
#include "delay_updates.h"
#include "filter.h"
#include "log.h"
#include "utils.h"
#include <dirent.h>
#include <errno.h>
#include <fcntl.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/stat.h>
#include <unistd.h>
/* Build the keep-set index from the exact manifest entries only. A lookup of
`rel` succeeds iff `rel` is a kept entry, a kept directory, or an ancestor
directory of kept content (the old is_dir_in_manifest predicate); the sorted
view answers "is an ancestor of kept content" without materializing any
per-component prefix copy, so the index is O(manifest size) memory. */
static bool build_keep_index(const ArrayList* manifest, PathIndex* index) {
if (!manifest || manifest->size <= 0)
return path_index_build(index, NULL, 0);
return path_index_build(index, (const char* const*)manifest->items, (size_t)manifest->size);
}
static bool keep_is_dir(const PathIndex* index, const char* rel_path) {
return path_index_contains(index, rel_path) || path_index_has_descendant(index, rel_path);
}
static bool keep_is_file(const PathIndex* index, const char* rel_path) {
return path_index_contains(index, rel_path);
}
/* True when child_rel is, or lies below, a protected entry. A prefix "a"
therefore protects "a" and "a/b/c" but not "ab". Entries with top_level_only
set only protect DIRECT children of the receive root (at_root); nested
directories that share such a name stay ordinary destination content. */
bool path_under_skip_prefix(const char* child_rel, bool at_root, const DeleteSkipEntry* skips,
int skip_count) {
for (int i = 0; i < skip_count; i++) {
if (skips[i].top_level_only && !at_root)
continue;
size_t prefix_len = strlen(skips[i].prefix);
if (strncmp(child_rel, skips[i].prefix, prefix_len) == 0 &&
(child_rel[prefix_len] == '\0' || child_rel[prefix_len] == '/'))
return true;
}
return false;
}
/* Per-run deletion budget and tallies. `max_delete` is the cap on the number
of entries the walker may remove (SIZE_MAX = unlimited); once it is reached
the remaining extras are counted in `skipped` and left in place, matching
rsync's partial --max-delete behavior. */
typedef struct {
size_t max_delete;
size_t deleted;
size_t skipped;
bool limit_hit;
} DeleteBudget;
/* True when direct children of the directory named by `rel` may be removed.
With no synchronization info (dirs == NULL) the whole tree is deletable; when
a dirs index is supplied only its exact entries are (the receive root is the
"." sentinel). */
static bool is_synced_dir(const PathIndex* dirs, const char* rel) {
if (!dirs)
return true;
return path_index_contains(dirs, rel[0] == '\0' ? "." : rel);
}
/* Unsigned byte-wise string compare, matching rsync's u_strcmp (a signed
strcmp would order bytes >= 0x80 differently). */
static int delete_name_cmp(const char* a, const char* b) {
const unsigned char* pa = (const unsigned char*)a;
const unsigned char* pb = (const unsigned char*)b;
while (*pa != '\0' && *pa == *pb) {
pa++;
pb++;
}
return (int)*pa - (int)*pb;
}
bool delete_dir_entries_collect(int dirfd, DeleteDirEntry** out, size_t* count,
bool* operation_ok) {
*out = NULL;
*count = 0;
if (operation_ok)
*operation_ok = true;
int scanfd = openat(dirfd, ".", O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
if (scanfd < 0)
return false;
DIR* dir = fdopendir(scanfd);
if (!dir) {
close(scanfd);
return false;
}
DeleteDirEntry* entries = NULL;
size_t used = 0;
size_t capacity = 0;
bool ok = true;
const struct dirent* entry;
while ((entry = readdir(dir)) != NULL) {
if (strcmp(entry->d_name, ".") == 0 || strcmp(entry->d_name, "..") == 0)
continue;
struct stat st;
if (fstatat(dirfd, entry->d_name, &st, AT_SYMLINK_NOFOLLOW) != 0) {
if (errno != ENOENT && operation_ok)
*operation_ok = false;
continue;
}
if (used == capacity) {
size_t next = capacity == 0 ? 16 : capacity * 2;
DeleteDirEntry* grown = realloc(entries, next * sizeof(*grown));
if (!grown) {
ok = false;
break;
}
entries = grown;
capacity = next;
}
entries[used].name = str_dup(entry->d_name);
if (!entries[used].name) {
ok = false;
break;
}
entries[used].is_dir = S_ISDIR(st.st_mode);
used++;
}
closedir(dir);
if (!ok) {
delete_dir_entries_free(entries, used);
return false;
}
*out = entries;
*count = used;
return true;
}
void delete_dir_entries_free(DeleteDirEntry* entries, size_t count) {
if (!entries)
return;
for (size_t i = 0; i < count; i++)
free(entries[i].name);
free(entries);
}
/* rsync's extraneous-entry order: subdirectories before files, each group in
descending name order. */
int delete_dir_entry_cmp_desc(const void* a, const void* b) {
const DeleteDirEntry* ea = a;
const DeleteDirEntry* eb = b;
if (ea->is_dir != eb->is_dir)
return ea->is_dir ? -1 : 1;
return -delete_name_cmp(ea->name, eb->name);
}
/* rsync's kept-subdirectory order: plain ascending name. */
int delete_dir_entry_cmp_asc(const void* a, const void* b) {
const DeleteDirEntry* ea = a;
const DeleteDirEntry* eb = b;
return delete_name_cmp(ea->name, eb->name);
}
/* How the shared classification/descent walk disposes of an extra it has
identified. LIST records the destination-relative path without touching disk
(the -n/--dry-run would-delete enumeration); DELETE unlinks/rmdirs it, charges
the shared --max-delete budget and notifies the observer. Both modes classify
and traverse identically, so the dry-run enumeration and the real deletion
cannot drift. */
typedef enum { DELETE_WALK_MODE_DELETE, DELETE_WALK_MODE_LIST } DeleteWalkMode;
typedef struct {
DeleteWalkMode mode;
DeleteBudget* budget; /* DELETE mode */
ArrayList* out; /* LIST mode: receives strdup'd relative paths */
size_t* recorded; /* LIST mode */
DeletePathObserver observer; /* DELETE mode */
void* observer_context; /* DELETE mode */
} DeleteWalkState;
/* Remove the extras directly inside the directory open on `dirfd` (DELETE mode)
or record the paths that WOULD be removed (LIST mode), recursing into every
child directory so kept content below a synchronized prefix is reached.
`all_removed` reports whether every child entry was removed (so the caller may
rmdir this directory). A child directory is never removed when it is itself a
synchronized directory or holds kept content; with a dirs index supplied,
direct children of a non-synchronized directory are never extras at all (they
are left in place but still descended into). Symlinks are unlinked like any
other non-directory extra (never followed).
Entries are processed in rsync's order (extraneous subdirectories in
descending name order, then extraneous files, then kept subdirectories in
ascending order) rather than readdir() order, so `--max-delete` leaves the
same survivors and the `--info=del`/dry-run line order matches rsync. */
static bool delete_walk_fd(int dirfd, const char* rel_path, const PathIndex* keep,
const PathIndex* dirs, DeleteWalkState* state,
const DeleteSkipEntry* skips, int skip_count,
const FilterRuleList* protect_rules, bool parent_deletable,
bool* all_removed) {
DeleteDirEntry* entries = NULL;
size_t count = 0;
bool collect_ok = true;
if (!delete_dir_entries_collect(dirfd, &entries, &count, &collect_ok))
return false;
bool operation_ok = collect_ok;
bool local_survives = false;
bool* shielded = calloc(count ? count : 1, sizeof(bool));
bool* is_extra = calloc(count ? count : 1, sizeof(bool));
if (!shielded || !is_extra) {
free(shielded);
free(is_extra);
delete_dir_entries_free(entries, count);
return false;
}
/* A directory is deletable when it or ANY ancestor is synchronized; the
`parent_deletable` flag carries that down the recursion so dest-only
directories below a synchronized root are removed wholesale. */
bool deletable = parent_deletable || is_synced_dir(dirs, rel_path);
bool at_root = rel_path[0] == '\0';
/* Reproduce rsync's traversal order: extraneous subdirectories in descending
name order, then extraneous files in descending name order, and kept
subdirectories only afterwards (ascending). Sorting up front also fixes the
identity of the survivors under a partial --max-delete. */
if (count > 1)
qsort(entries, count, sizeof(*entries), delete_dir_entry_cmp_desc);
size_t dir_count = 0;
while (dir_count < count && entries[dir_count].is_dir)
dir_count++;
/* Classify every entry up front (the verdict does not depend on processing
order) so the ordered passes below can act on it. */
for (size_t i = 0; i < count; i++) {
char* child_rel = path_cat((char*)rel_path, entries[i].name);
if (!child_rel) {
operation_ok = false;
continue;
}
/* A --delay-updates run keeps its staging directory as a direct child of
the receive root, and basis-dir snapshots live below it too. Their
contents are not manifest entries, so descending into them would delete
every staged / basis file as an "extra". Only the staging name (a
top-level-only prefix) and the basis prefixes are protected: a nested
destination directory that happens to be called .fastsync-stage is
ordinary content. */
if (path_under_skip_prefix(child_rel, at_root, skips, skip_count)) {
shielded[i] = true;
local_survives = true;
} else if (protect_rules &&
filter_rules_apply_side(protect_rules, child_rel, entries[i].name, entries[i].is_dir,
FILTER_SIDE_RECEIVER) == FILTER_ACTION_PROTECT) {
/* A first-match protect rule shields the extra; for a directory the whole
subtree is shielded (rsync prunes an excluded directory), so do not
descend. */
shielded[i] = true;
local_survives = true;
} else if (entries[i].is_dir) {
bool child_synced = dirs && path_index_contains(dirs, child_rel);
is_extra[i] = deletable && !child_synced && !keep_is_dir(keep, child_rel);
if (!is_extra[i])
local_survives = true;
} else {
is_extra[i] = deletable && !keep_is_file(keep, child_rel);
if (!is_extra[i])
local_survives = true;
}
free(child_rel);
}
/* Pass 1: extraneous subdirectories, descending. */
for (size_t i = 0; i < dir_count; i++) {
if (!is_extra[i])
continue;
char* child_rel = path_cat((char*)rel_path, entries[i].name);
if (!child_rel) {
operation_ok = false;
continue;
}
int childfd = openat(dirfd, entries[i].name, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
bool child_all_removed = false;
if (childfd >= 0) {
if (!delete_walk_fd(childfd, child_rel, keep, dirs, state, skips, skip_count, protect_rules,
deletable, &child_all_removed))
operation_ok = false;
close(childfd);
} else if (errno != ENOENT) {
operation_ok = false;
}
if (child_all_removed && deletable) {
if (state->mode == DELETE_WALK_MODE_LIST) {
/* Record the directory with rsync's trailing slash. */
size_t len = strlen(child_rel);
char* copy = malloc(len + 2);
if (!copy) {
operation_ok = false;
} else {
memcpy(copy, child_rel, len);
copy[len] = '/';
copy[len + 1] = '\0';
if (!array_list_add(state->out, copy)) {
free(copy);
operation_ok = false;
} else {
(*state->recorded)++;
}
}
} else if (state->budget->deleted >= state->budget->max_delete) {
state->budget->limit_hit = true;
state->budget->skipped++;
local_survives = true;
} else if (unlinkat(dirfd, entries[i].name, AT_REMOVEDIR) != 0) {
/* ENOENT: already gone (fine). ENOTEMPTY/EEXIST: the directory still
holds entries the walker leaves in place (a protected excluded
prefix, a kept file the manifest protects, a symlink); rsync leaves
such a directory behind, so this is not an error. Only genuine I/O
failures abort the deletion. */
if (errno != ENOENT && errno != ENOTEMPTY && errno != EEXIST)
operation_ok = false;
local_survives = true;
} else {
state->budget->deleted++;
/* rsync reports a removed directory with a trailing slash. */
if (state->observer) {
size_t len = strlen(child_rel);
char* with_slash = malloc(len + 2);
if (with_slash) {
memcpy(with_slash, child_rel, len);
with_slash[len] = '/';
with_slash[len + 1] = '\0';
state->observer(state->observer_context, with_slash);
free(with_slash);
} else {
state->observer(state->observer_context, child_rel);
}
}
}
} else {
local_survives = true;
}
free(child_rel);
}
/* Pass 2: extraneous files, descending. */
for (size_t i = dir_count; i < count; i++) {
if (!is_extra[i])
continue;
if (state->mode == DELETE_WALK_MODE_LIST) {
char* child_rel = path_cat((char*)rel_path, entries[i].name);
if (!child_rel) {
operation_ok = false;
continue;
}
char* copy = str_dup(child_rel);
if (!copy || !array_list_add(state->out, copy)) {
free(copy);
operation_ok = false;
} else {
(*state->recorded)++;
}
free(child_rel);
} else if (state->budget->deleted >= state->budget->max_delete) {
state->budget->limit_hit = true;
state->budget->skipped++;
local_survives = true;
} else if (unlinkat(dirfd, entries[i].name, 0) != 0) {
if (errno != ENOENT)
operation_ok = false;
local_survives = true;
} else {
state->budget->deleted++;
char* child_rel = path_cat((char*)rel_path, entries[i].name);
if (child_rel) {
if (state->observer)
state->observer(state->observer_context, child_rel);
char* escaped_path = output_escape(child_rel, log_get_8_bit_output());
fprintf(stderr, " Deleted: %s\n", escaped_path ? escaped_path : "<allocation failed>");
free(escaped_path);
}
free(child_rel);
}
}
/* Pass 3: kept subdirectories, ascending (rsync descends into these only
after the parent's own extras have been handled). */
for (size_t i = dir_count; i-- > 0;) {
if (is_extra[i] || shielded[i])
continue;
char* child_rel = path_cat((char*)rel_path, entries[i].name);
if (!child_rel) {
operation_ok = false;
continue;
}
int childfd = openat(dirfd, entries[i].name, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
bool child_all_removed = false;
if (childfd >= 0) {
if (!delete_walk_fd(childfd, child_rel, keep, dirs, state, skips, skip_count, protect_rules,
deletable, &child_all_removed))
operation_ok = false;
close(childfd);
} else if (errno != ENOENT) {
operation_ok = false;
}
/* A kept/synchronized directory is never removed. */
local_survives = true;
free(child_rel);
}
free(shielded);
free(is_extra);
delete_dir_entries_free(entries, count);
*all_removed = !local_survives;
return operation_ok;
}
/* Open the receive root following the same authorized-root confinement the
walker uses, or dest_root directly when no authorized root is installed. */
static int open_destination_root(const char* dest_root) {
int root_fd = utils_get_authorized_root_fd();
if (root_fd >= 0) {
if (utils_get_authorized_root_path())
return utils_open_authorized_destination(dest_root);
if (dest_root == NULL)
return dup(root_fd);
return -1;
}
return open(dest_root, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
}
bool delete_extras_list(const char* dest_root, const ArrayList* manifest,
const ArrayList* synced_dirs, const DeleteSkipEntry* skips, int skip_count,
const FilterRuleList* protect_rules, ArrayList* out, size_t* count_out) {
if (count_out)
*count_out = 0;
if (!manifest || !out)
return false;
PathIndex keep;
if (!build_keep_index(manifest, &keep))
return false;
PathIndex dirs;
bool have_dirs = synced_dirs != NULL;
if (have_dirs &&
!path_index_build(&dirs, (const char* const*)synced_dirs->items, (size_t)synced_dirs->size)) {
path_index_free(&keep);
return false;
}
int rootfd = open_destination_root(dest_root);
if (rootfd < 0) {
path_index_free(&keep);
if (have_dirs)
path_index_free(&dirs);
return false;
}
bool all_removed = false;
size_t recorded = 0;
DeleteWalkState state = {.mode = DELETE_WALK_MODE_LIST,
.budget = NULL,
.out = out,
.recorded = &recorded,
.observer = NULL,
.observer_context = NULL};
bool ok = delete_walk_fd(rootfd, "", &keep, have_dirs ? &dirs : NULL, &state, skips, skip_count,
protect_rules, false, &all_removed);
if (close(rootfd) != 0)
ok = false;
path_index_free(&keep);
if (have_dirs)
path_index_free(&dirs);
if (count_out)
*count_out = recorded;
return ok;
}
DeleteWalkResult delete_extras_limited_observed(const char* dest_root, const ArrayList* manifest,
const ArrayList* synced_dirs, size_t max_delete,
const DeleteSkipEntry* skips, int skip_count,
const FilterRuleList* protect_rules,
size_t* deleted_out, size_t* skipped_out,
DeletePathObserver observer,
void* observer_context) {
if (deleted_out)
*deleted_out = 0;
if (skipped_out)
*skipped_out = 0;
if (!manifest)
return DELETE_WALK_ERROR;
/* Index the keep-set (and the synchronized-dir set, when supplied) once so
membership is answered in O(path length) instead of scanning every entry
for every destination entry. */
PathIndex keep;
if (!build_keep_index(manifest, &keep))
return DELETE_WALK_ERROR;
PathIndex dirs;
bool have_dirs = synced_dirs != NULL;
if (have_dirs &&
!path_index_build(&dirs, (const char* const*)synced_dirs->items, (size_t)synced_dirs->size)) {
path_index_free(&keep);
return DELETE_WALK_ERROR;
}
int rootfd = open_destination_root(dest_root);
if (rootfd < 0) {
path_index_free(&keep);
if (have_dirs)
path_index_free(&dirs);
return DELETE_WALK_ERROR;
}
DeleteBudget budget = {.max_delete = max_delete, .deleted = 0, .skipped = 0, .limit_hit = false};
bool all_removed = false;
DeleteWalkState state = {.mode = DELETE_WALK_MODE_DELETE,
.budget = &budget,
.out = NULL,
.recorded = NULL,
.observer = observer,
.observer_context = observer_context};
bool ok = delete_walk_fd(rootfd, "", &keep, have_dirs ? &dirs : NULL, &state, skips, skip_count,
protect_rules, false, &all_removed);
if (close(rootfd) != 0)
ok = false;
path_index_free(&keep);
if (have_dirs)
path_index_free(&dirs);
if (deleted_out)
*deleted_out = budget.deleted;
if (skipped_out)
*skipped_out = budget.skipped;
if (!ok)
return DELETE_WALK_ERROR;
return budget.limit_hit ? DELETE_WALK_LIMIT_REACHED : DELETE_WALK_OK;
}
DeleteWalkResult delete_extras_limited(const char* dest_root, const ArrayList* manifest,
const ArrayList* synced_dirs, size_t max_delete,
const DeleteSkipEntry* skips, int skip_count,
const FilterRuleList* protect_rules, size_t* deleted_out,
size_t* skipped_out) {
return delete_extras_limited_observed(dest_root, manifest, synced_dirs, max_delete, skips,
skip_count, protect_rules, deleted_out, skipped_out, NULL,
NULL);
}
bool delete_extras(const char* dest_root, const ArrayList* manifest) {
return delete_extras_limited(dest_root, manifest, NULL, SIZE_MAX, NULL, 0, NULL, NULL, NULL) ==
DELETE_WALK_OK;
}
/* Build the delete-walk protection prefix for one basis directory. The walker
compares paths relative to the receive root, so a relative entry is already
in the right form; an absolute entry that lies below the root is converted to
its root-relative form, and one outside the root returns NULL (the walk
cannot reach it, and it is not protected data beneath the root). Exposed so
tests can exercise the root-of-"/" child mapping directly. */
char* delete_basis_relative(const Config* config, const char* path) {
if (!path)
return NULL;
if (path[0] != '/')
return str_dup(path);
const char* root = config->receive_root_directory;
if (!root || root[0] != '/')
return NULL;
size_t root_len = strlen(root);
while (root_len > 1 && root[root_len - 1] == '/')
root_len--;
if (strncmp(path, root, root_len) != 0)
return NULL;
if (root_len == 1) {
/* `root` is "/" (the only single-character absolute root): every absolute
path is below it, and the child relative form is everything after the
leading '/'. */
if (path[1] == '\0')
return NULL; /* identical to the root, not a child */
return str_dup(path + 1);
}
if (path[root_len] != '/')
return NULL; /* identical or a sibling sharing a name prefix */
return str_dup(path + root_len + 1);
}
bool delete_skips_build(const Config* config, const ArrayList* protected_paths,
const ArrayList* size_skipped, bool basis_root_relative,
DeleteSkipSet* out) {
if (!out)
return false;
out->entries = NULL;
out->owned_prefixes = NULL;
out->count = 0;
out->owned_count = 0;
if (!config)
return false;
int protected_count = protected_paths ? protected_paths->size : 0;
int size_skipped_count = size_skipped ? size_skipped->size : 0;
int count =
(config->delay_updates ? 1 : 0) + config->basis_count + protected_count + size_skipped_count;
if (count == 0)
return true;
out->entries = calloc((size_t)count, sizeof(DeleteSkipEntry));
if (!out->entries)
return false;
if (basis_root_relative && config->basis_count > 0) {
out->owned_prefixes = calloc((size_t)config->basis_count, sizeof(char*));
if (!out->owned_prefixes) {
free(out->entries);
out->entries = NULL;
return false;
}
out->owned_count = config->basis_count;
}
int idx = 0;
if (config->delay_updates) {
out->entries[idx].prefix = DELAY_UPDATES_STAGING_DIR;
out->entries[idx].top_level_only = true;
idx++;
}
for (int i = 0; i < config->basis_count; i++) {
const char* prefix = config->basis_dirs[i].path;
if (basis_root_relative) {
/* An absolute basis outside the receive root is unreachable by this walk,
so it contributes no protection prefix (and no slot). */
char* relative = delete_basis_relative(config, config->basis_dirs[i].path);
if (!relative)
continue;
out->owned_prefixes[i] = relative;
prefix = relative;
}
out->entries[idx].prefix = prefix;
out->entries[idx].top_level_only = false;
idx++;
}
for (int i = 0; i < protected_count; i++) {
out->entries[idx].prefix = (const char*)protected_paths->items[i];
out->entries[idx].top_level_only = false;
idx++;
}
for (int i = 0; i < size_skipped_count; i++) {
out->entries[idx].prefix = (const char*)size_skipped->items[i];
out->entries[idx].top_level_only = false;
idx++;
}
out->count = idx;
return true;
}
void delete_skips_free(DeleteSkipSet* set) {
if (!set)
return;
if (set->owned_prefixes) {
for (int i = 0; i < set->owned_count; i++)
free(set->owned_prefixes[i]);
}
free(set->owned_prefixes);
free(set->entries);
set->entries = NULL;
set->owned_prefixes = NULL;
set->count = 0;
set->owned_count = 0;
}
+151
View File
@@ -0,0 +1,151 @@
#ifndef DELETE_H
#define DELETE_H
#include "array_list.h"
#include "config.h"
#include <stdbool.h>
#include <stddef.h>
/* Delete engine.
*
* This module owns destination-relative delete traversal: the ordered directory
* walker that reproduces rsync's extraneous-entry order, the skip-prefix
* protection set shared by every delete pass, and the read-only enumeration
* that mirrors the walker for -n/--dry-run. The budgeted manifest commit
* (delete_commit.c) and the per-directory delete plans (delete_plan.c) are
* built on the primitives exported here. */
/* Result of a bounded extra-file deletion run. */
typedef enum {
/* Every extra entry was removed (or there were none). */
DELETE_WALK_OK = 0,
/* The numeric cap for this run was reached before every extra was removed.
The walker removed exactly the entries the cap allowed and skipped (without
removing) the rest, matching rsync's partial --max-delete behavior. */
DELETE_WALK_LIMIT_REACHED,
/* A traversal or unlink failure aborted the deletion (partial removal is
possible, mirroring the delete pass). */
DELETE_WALK_ERROR
} DeleteWalkResult;
/* One protected entry for the delete walker. When top_level_only is true the
prefix is skipped only as a DIRECT child of dest_root (the --delay-updates
staging directory, which must not hide genuine extras inside a nested
destination directory that happens to share the staging name); otherwise the
prefix is skipped at any depth (the --compare-dest/--copy-dest/--link-dest
basis trees, and the sender-side protected filter-excluded prefixes, which
are never destination content). */
typedef struct {
const char* prefix;
bool top_level_only;
} DeleteSkipEntry;
/* A built skip-prefix set. `entries`/`count` are what path_under_skip_prefix()
consumes. `owned_prefixes` holds any prefix strings the builder had to
allocate (root-relative basis-dir conversions); it is NULL when every prefix
is borrowed from the config or the caller's lists. Release with
delete_skips_free(). */
typedef struct {
DeleteSkipEntry* entries;
char** owned_prefixes;
int count;
int owned_count;
} DeleteSkipSet;
/* True when child_rel is, or lies below, one of the protected entries (a prefix
"a" protects "a" and "a/b/c" but not "ab"; top_level_only entries protect
only DIRECT children of the destination root, i.e. child_rel has no '/'). */
bool path_under_skip_prefix(const char* child_rel, bool at_root, const DeleteSkipEntry* skips,
int skip_count);
/* One destination-directory entry collected up front so the delete walkers can
reproduce rsync's traversal order instead of readdir() order. rsync processes
a directory's extraneous subdirectories first (descending name, depth-first),
then its extraneous files (descending name), and only afterwards descends into
its kept subdirectories (ascending name). */
typedef struct {
char* name;
bool is_dir;
} DeleteDirEntry;
/* Collect the entries of the directory open on `dirfd` (excluding "." and ".."),
stat'ing each with AT_SYMLINK_NOFOLLOW. On success *out is a malloc'd array of
*count entries whose names the caller frees with delete_dir_entries_free().
Returns false on an allocation/readdir failure; a vanished entry (ENOENT) is
skipped, any other stat failure is reported through *operation_ok while the
walk continues. */
bool delete_dir_entries_collect(int dirfd, DeleteDirEntry** out, size_t* count, bool* operation_ok);
void delete_dir_entries_free(DeleteDirEntry* entries, size_t count);
/* Sort comparators: `_desc` orders subdirectories before files and each group by
descending name (rsync's extraneous-entry order); `_asc` orders plain ascending
name (rsync's kept-subdirectory order). */
int delete_dir_entry_cmp_desc(const void* a, const void* b);
int delete_dir_entry_cmp_asc(const void* a, const void* b);
/* Remove files/dirs/symlinks under dest_root that are not listed in manifest
without ever descending into a protected prefix (see DeleteSkipEntry). When
`synced_dirs` is non-NULL, extras are only removed directly inside a directory
whose destination-relative path is an exact entry in that list (the receive
root is the "." sentinel); directories outside the synchronized set are still
descended into so kept content below a listed directory is preserved, but
nothing in them is removed. A NULL `synced_dirs` keeps the legacy behavior of
treating the whole destination tree as deletable. `max_delete` caps the
number of removed entries (SIZE_MAX = unlimited): the walker removes up to the
cap and returns DELETE_WALK_LIMIT_REACHED when more extras remained.
`deleted_out`/`skipped_out` optionally receive the number of entries removed
and the number skipped because of the cap. */
DeleteWalkResult delete_extras_limited(const char* dest_root, const ArrayList* manifest,
const ArrayList* synced_dirs, size_t max_delete,
const DeleteSkipEntry* skips, int skip_count,
const FilterRuleList* protect_rules, size_t* deleted_out,
size_t* skipped_out);
/* Optional per-deletion observer: called for each destination-relative path
actually removed (a file, symlink, or directory), in removal order, so the
receiver can stream rsync's `--info=del`/`--info=remove` lines. */
typedef void (*DeletePathObserver)(void* context, const char* rel_path);
/* `delete_extras_limited_observed` is delete_extras_limited with an optional
* observer; the observer is invoked only for entries truly removed. When
* `protect_rules` is non-NULL its receiver-side verdict is evaluated for every
* candidate extra: a first-match PROTECT leaves the entry (and, for a
* directory, its whole subtree) in place, while RISK/NONE fall through to the
* ordinary skip-prefix/keep-set logic. */
DeleteWalkResult delete_extras_limited_observed(const char* dest_root, const ArrayList* manifest,
const ArrayList* synced_dirs, size_t max_delete,
const DeleteSkipEntry* skips, int skip_count,
const FilterRuleList* protect_rules,
size_t* deleted_out, size_t* skipped_out,
DeletePathObserver observer,
void* observer_context);
/* Read-only companion to delete_extras_limited: walk the destination exactly as
the delete pass would and APPEND (strdup'd) destination-relative paths that
WOULD be removed, without touching disk. Used for -n/--dry-run --delete
would-delete reporting. Returns true on a clean walk; the caller owns the
strings appended to `out` and receives their count in *count_out. */
bool delete_extras_list(const char* dest_root, const ArrayList* manifest,
const ArrayList* synced_dirs, const DeleteSkipEntry* skips, int skip_count,
const FilterRuleList* protect_rules, ArrayList* out, size_t* count_out);
bool delete_extras(const char* dest_root, const ArrayList* manifest);
/* Build the delete walk's skip-prefix set from the config's --delay-updates
staging directory, its --compare-dest/--copy-dest/--link-dest basis dirs, and
the caller-supplied protection lists, in that order. `protected_paths` and
`size_skipped` are borrowed (may be NULL); every entry in them is protected at
any depth. The staging directory is protected only as a DIRECT child of the
receive root. `basis_root_relative` selects how a basis path becomes a
prefix: true converts an absolute path under the receive root to its
root-relative form (the whole-tree commit walk; an unreachable path
contributes no slot), false keeps the configured path verbatim (the
per-directory plan walk). On success the caller releases `*out` with
delete_skips_free(); returns false on allocation failure. */
bool delete_skips_build(const Config* config, const ArrayList* protected_paths,
const ArrayList* size_skipped, bool basis_root_relative,
DeleteSkipSet* out);
void delete_skips_free(DeleteSkipSet* set);
/* Convert one basis-directory path to the receive-root-relative protection
prefix the delete walker uses (NULL when it lies outside the root). Exposed
for unit tests of the root-of-"/" and normalization edge cases. */
char* delete_basis_relative(const Config* config, const char* path);
#endif
+488
View File
@@ -0,0 +1,488 @@
#include <errno.h>
#include <ctype.h>
#include <dirent.h>
#include <fcntl.h>
#include <libgen.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/stat.h>
#include <sys/sysmacros.h>
#include <unistd.h>
#include "array_list.h"
#include "charset.h"
#include "chmod.h"
#include "chunk.h"
#include "compression.h"
#include "config.h"
#include "data.h"
#include "delay_updates.h"
#include "delete_commit.h"
#include "delta.h"
#include "file.h"
#include "format.h"
#include "identity.h"
#include "log.h"
#include "metadata.h"
#include "protocol.h"
#include "utils.h"
#include "xattr.h"
#define MAX_SERVER_DELETE_COUNT 100000U
/* Retained cost of one delete-manifest entry beyond its path bytes: the
ArrayList pointer slot plus an approximate malloc header/rounding for the
heap copy. Charged against MAX_MANIFEST_BYTES so a frame full of tiny paths
cannot retain far more than the byte budget (B5). */
#define MANIFEST_ENTRY_OVERHEAD (sizeof(char*) + 16)
/* Read a delete-manifest frame (the STATUS_MANIFEST leading code has already
been consumed): a keep-set entry count followed by that many
destination-relative paths, then a protected-prefix count followed by that
many destination-relative prefixes, then a missing-args count followed by that
many destination-relative delete paths, then (protocol 2.23.0) a
synchronized-directory count followed by that many destination-relative
directory paths (the receive root is the "." sentinel). The frame is
self-delimiting (the counts are authoritative), so the caller decides what to
do next and continues reading the following STATUS_* frame. Every section is
validated identically: an entry must be non-empty, relative and traversal-free
and the aggregate length across ALL sections is capped by MAX_MANIFEST_BYTES
(so the missing-args deletion requests are confined like the rest of the
manifest). Returns an owned DeleteManifest, or NULL after sending STATUS_ERROR
when the frame is malformed (bad count, empty/absolute path, path traversal,
or an aggregate size beyond MAX_MANIFEST_BYTES). */
static bool receive_manifest_section(int fd, ArrayList* list, size_t* manifest_bytes,
size_t* manifest_entries) {
int count;
if (!receive_int(fd, &count)) {
send_status(fd, STATUS_ERROR);
return false;
}
if (count < 0 || count > MAX_MANIFEST_ENTRIES ||
(size_t)count > MAX_MANIFEST_ENTRIES - *manifest_entries) {
send_status(fd, STATUS_ERROR);
return false;
}
for (int i = 0; i < count; i++) {
char* s = receive_wire_str(fd);
size_t entry_size = s ? strlen(s) + MANIFEST_ENTRY_OVERHEAD : 0;
if (!s || s[0] == '\0' || s[0] == '/' || has_path_traversal(s) ||
entry_size > MAX_MANIFEST_BYTES - *manifest_bytes ||
(*manifest_bytes += entry_size) > MAX_MANIFEST_BYTES || !array_list_add(list, s)) {
free(s);
send_status(fd, STATUS_ERROR);
return false;
}
}
*manifest_entries += (size_t)count;
return true;
}
DeleteManifest* receive_manifest_entries(int fd) {
DeleteManifest* manifest = calloc(1, sizeof(DeleteManifest));
if (!manifest) {
send_status(fd, STATUS_ERROR);
return NULL;
}
manifest->keeps = array_list_create(free);
manifest->protected = array_list_create(free);
manifest->missing = array_list_create(free);
manifest->dirs = array_list_create(free);
if (!manifest->keeps || !manifest->protected || !manifest->missing || !manifest->dirs) {
delete_manifest_free(manifest);
send_status(fd, STATUS_ERROR);
return NULL;
}
size_t manifest_bytes = 0;
size_t manifest_entries = 0;
if (!receive_manifest_section(fd, manifest->keeps, &manifest_bytes, &manifest_entries) ||
!receive_manifest_section(fd, manifest->protected, &manifest_bytes, &manifest_entries) ||
!receive_manifest_section(fd, manifest->missing, &manifest_bytes, &manifest_entries) ||
!receive_manifest_section(fd, manifest->dirs, &manifest_bytes, &manifest_entries)) {
delete_manifest_free(manifest);
return NULL;
}
return manifest;
}
void delete_manifest_free(DeleteManifest* manifest) {
if (!manifest)
return;
array_list_delete(manifest->keeps);
array_list_delete(manifest->protected);
array_list_delete(manifest->missing);
array_list_delete(manifest->dirs);
free(manifest);
}
/* Shared --max-delete budget for one receiver-side deletion commit. Both the
--delete-missing-args exact-path removals and the ordinary extras walk draw
from the same tally, matching rsync (whose --max-delete counts every deleted
file or directory). `max_delete` is SIZE_MAX for an unlimited budget. */
typedef struct {
size_t max_delete;
size_t deleted;
size_t skipped;
bool limit_hit;
} DeleteBudgetState;
/* Remove every destination entry under the receive root that is not in the
keep-set, bounded by the shared budget (a smaller client --max-delete=NUM
replaces the server hard bound; rsync deletes up to the bound and skips the
rest). With --delay-updates the not-yet-published staging directory is a
direct child of the receive root and must not be treated as a set of extras;
the manifest's protected prefixes (paths excluded on the source), the
size-pruned prefixes (--max-size/--min-size, always protected) and the
alternate basis directories are never destination content and are skipped at
any depth. Returns true unless a traversal/unlink error aborted the walk;
the budget's limit_hit/skipped fields report a cap-stopped run. */
static bool delete_extras_budgeted_observed(const Config* config, const DeleteManifest* manifest,
DeleteBudgetState* budget, DeletePathObserver observer,
void* observer_context) {
if (!config || !manifest || !manifest->keeps)
return false;
fprintf(stderr, "Deleting files not in manifest...\n");
/* Protected entries: the --delay-updates staging name (only as a DIRECT child
of the receive root), the alternate basis directories and the sender-side
protected prefixes (filter-excluded and size-pruned source mirrors), all at
any depth. See delete_skips_build(). */
DeleteSkipSet skips;
if (!delete_skips_build(config, manifest->protected, NULL, true, &skips))
return false;
/* Clamp rather than subtract: an accounting bug where deleted already exceeds
max_delete must never underflow into an effectively unlimited budget. */
size_t remaining;
if (budget->max_delete == SIZE_MAX)
remaining = SIZE_MAX;
else if (budget->deleted >= budget->max_delete)
remaining = 0;
else
remaining = budget->max_delete - budget->deleted;
size_t deleted = 0;
size_t skipped = 0;
DeleteWalkResult result = delete_extras_limited_observed(
config->receive_root_directory, manifest->keeps, manifest->dirs, remaining, skips.entries,
skips.count, config->protect_rules, &deleted, &skipped, observer, observer_context);
delete_skips_free(&skips);
budget->deleted += deleted;
budget->skipped += skipped;
if (result == DELETE_WALK_LIMIT_REACHED) {
budget->limit_hit = true;
return true;
}
if (result != DELETE_WALK_OK) {
log_message(LOG_LEVEL_ERROR, "deletion failed while removing extraneous files");
return false;
}
return true;
}
static bool delete_extras_budgeted(const Config* config, const DeleteManifest* manifest,
DeleteBudgetState* budget) {
return delete_extras_budgeted_observed(config, manifest, budget, NULL, NULL);
}
/* Prefixes every observed path with a fixed subtree root, so a nested walk
(a recursively removed missing-arg directory) reports receive-root-relative
names like the rest of the delete output. */
typedef struct {
DeletePathObserver inner;
void* inner_context;
const char* prefix;
} PrefixedDeleteObserver;
static void prefixed_delete_observer(void* context, const char* rel) {
PrefixedDeleteObserver* prefixed = context;
if (!prefixed->inner || !rel)
return;
char* joined = path_cat((char*)prefixed->prefix, rel);
if (joined) {
prefixed->inner(prefixed->inner_context, joined);
free(joined);
}
}
/* --delete-missing-args exact-path deletions: each destination mirror in
manifest->missing is an explicit user request, so it is removed even when the
ordinary extras walk (with its protected prefixes) would leave it alone. The
--delay-updates staging directory and basis snapshots are receiver artifacts
and stay protected exactly as in the extras walker. A regular file or
symlink is unlinked, an empty directory removed, and a NON-empty directory is
removed recursively only when --delete or --force is in effect (rsync parity:
the man page says a non-empty directory mirror is only deleted with --force
or --delete); otherwise it is left with a warning and the run continues. A
mirror that does not exist is a no-op. Each removal draws from the shared
--max-delete budget: once it is exhausted the remaining requests are skipped
and counted. Returns false only on a genuine error (a confinement failure on
a validated path or an I/O error), which fails the run. */
static bool delete_missing_args_budgeted_observed(const Config* config,
const DeleteManifest* manifest,
DeleteBudgetState* budget,
DeletePathObserver observer,
void* observer_context) {
if (!config || !manifest)
return false;
if (!manifest->missing || manifest->missing->size == 0)
return true;
fprintf(stderr, "Deleting destination mirrors of missing source arguments...\n");
/* The staging directory and basis snapshots stay protected exactly as in the
extras walker (the missing-args path overrides the ordinary protected
prefixes, so those are not passed here). */
DeleteSkipSet skips;
if (!delete_skips_build(config, NULL, NULL, true, &skips))
return false;
bool ok = true;
for (int i = 0; i < manifest->missing->size; i++) {
const char* rel = (const char*)manifest->missing->items[i];
if (!rel || *rel == '\0' || *rel == '/' || has_path_traversal(rel)) {
/* Defensive only: receive_manifest_entries already validated every
section identically, so a controlled peer never reaches this branch. */
log_message(LOG_LEVEL_ERROR, "invalid missing-args delete path");
ok = false;
continue;
}
bool at_root = strchr(rel, '/') == NULL;
if (path_under_skip_prefix(rel, at_root, skips.entries, skips.count)) {
char* escaped = output_escape(rel, log_get_8_bit_output());
log_message(LOG_LEVEL_WARNING,
"missing-args path '%s' is protected (staging directory or basis snapshot); "
"not deleting",
escaped ? escaped : "<allocation failed>");
free(escaped);
continue;
}
char* full = path_cat(config->receive_root_directory, rel);
if (!full) {
ok = false;
continue;
}
char* leaf = NULL;
int parent_fd = file_open_secure_parent(full, &leaf, false);
if (parent_fd < 0) {
/* The mirror's parent directory may itself not exist on the destination
(a deeper missing entry whose leading directories were never created).
That is a no-op -- there is nothing to delete -- matching
file_remove_tree_secure's absent-path handling; only a genuine I/O
error (EACCES, a symlink loop, ...) fails the run. */
bool absent = errno == ENOENT || errno == ENOTDIR;
free(full);
free(leaf);
if (!absent)
ok = false;
continue;
}
struct stat st;
if (fstatat(parent_fd, leaf, &st, AT_SYMLINK_NOFOLLOW) != 0) {
/* Already absent: nothing to delete (a no-op, not a deletion). */
if (errno != ENOENT)
ok = false;
close(parent_fd);
free(leaf);
free(full);
continue;
}
/* An entry that exists is one deletion: skip it (and count it) when the
shared --max-delete budget is already exhausted. */
if (budget->deleted >= budget->max_delete) {
budget->limit_hit = true;
budget->skipped++;
close(parent_fd);
free(leaf);
free(full);
continue;
}
bool removed = false;
if (S_ISDIR(st.st_mode)) {
if (unlinkat(parent_fd, leaf, AT_REMOVEDIR) == 0) {
removed = true;
} else if (errno == ENOTEMPTY || errno == EEXIST) {
close(parent_fd);
parent_fd = -1;
free(leaf);
leaf = NULL;
if (config->use_delete || config->force_delete) {
/* Remove the contents entry-by-entry through the budgeted extras
walker so every deleted file/dir counts toward --max-delete (rsync
parity); the now-empty directory itself costs one more. A run that
hits the cap leaves the remaining entries in place. */
ArrayList* no_keeps = array_list_create(free);
/* Never let an accounting slip (deleted > max_delete) underflow the
remaining budget into SIZE_MAX, which would grant unlimited
deletions. */
size_t remaining =
budget->deleted >= budget->max_delete ? 0 : budget->max_delete - budget->deleted;
size_t contents_deleted = 0;
size_t contents_skipped = 0;
PrefixedDeleteObserver nested = {observer, observer_context, rel};
DeleteWalkResult walk =
no_keeps ? delete_extras_limited_observed(full, no_keeps, NULL, remaining, NULL, 0,
NULL, &contents_deleted, &contents_skipped,
observer ? prefixed_delete_observer : NULL,
observer ? &nested : NULL)
: DELETE_WALK_ERROR;
if (no_keeps)
array_list_delete(no_keeps);
budget->deleted += contents_deleted;
budget->skipped += contents_skipped;
if (walk == DELETE_WALK_LIMIT_REACHED) {
budget->limit_hit = true;
} else if (walk != DELETE_WALK_OK) {
ok = false;
} else if (budget->deleted >= budget->max_delete) {
budget->limit_hit = true;
budget->skipped++;
} else if (file_remove_tree_secure(full)) {
/* The shared `if (removed)` tail charges this directory exactly
once; counting it here too would consume two budget units. */
removed = true;
} else {
ok = false;
}
} else {
char* escaped = output_escape(rel, log_get_8_bit_output());
log_message(LOG_LEVEL_WARNING,
"missing-args destination '%s' is a non-empty directory; use --force or "
"--delete to remove it",
escaped ? escaped : "<allocation failed>");
free(escaped);
}
} else if (errno != ENOENT) {
ok = false;
}
} else {
if (unlinkat(parent_fd, leaf, 0) == 0) {
removed = true;
} else if (errno != ENOENT) {
ok = false;
}
}
if (removed) {
budget->deleted++;
if (observer)
observer(observer_context, rel);
char* escaped = output_escape(rel, log_get_8_bit_output());
fprintf(stderr, " Deleted: %s\n", escaped ? escaped : "<allocation failed>");
free(escaped);
}
if (parent_fd >= 0)
close(parent_fd);
free(leaf);
free(full);
if (!ok)
break;
}
delete_skips_free(&skips);
return ok;
}
/* Public wrappers used outside the commit path (and by unit tests): no
--max-delete budget. */
bool manifest_would_delete_list(const Config* config, const DeleteManifest* manifest,
ArrayList* out, size_t* count_out) {
if (count_out)
*count_out = 0;
if (!config || !manifest || !manifest->keeps || !out)
return false;
DeleteSkipSet skips;
if (!delete_skips_build(config, manifest->protected, NULL, true, &skips))
return false;
bool ok = delete_extras_list(config->receive_root_directory, manifest->keeps, manifest->dirs,
skips.entries, skips.count, config->protect_rules, out, count_out);
delete_skips_free(&skips);
return ok;
}
bool manifest_delete_extras(const Config* config, const DeleteManifest* manifest) {
DeleteBudgetState budget = {
.max_delete = SIZE_MAX, .deleted = 0, .skipped = 0, .limit_hit = false};
return delete_extras_budgeted(config, manifest, &budget);
}
bool manifest_delete_missing_args(const Config* config, const DeleteManifest* manifest) {
DeleteBudgetState budget = {
.max_delete = SIZE_MAX, .deleted = 0, .skipped = 0, .limit_hit = false};
return delete_missing_args_budgeted_observed(config, manifest, &budget, NULL, NULL);
}
bool manifest_delete_missing_args_limited(const Config* config, const DeleteManifest* manifest,
size_t max_delete, size_t* deleted, size_t* skipped,
bool* limit_hit) {
return manifest_delete_missing_args_limited_observed(config, manifest, max_delete, deleted,
skipped, limit_hit, NULL, NULL);
}
bool manifest_delete_missing_args_limited_observed(
const Config* config, const DeleteManifest* manifest, size_t max_delete, size_t* deleted,
size_t* skipped, bool* limit_hit, DeletePathObserver observer, void* observer_context) {
DeleteBudgetState budget = {
.max_delete = max_delete, .deleted = 0, .skipped = 0, .limit_hit = false};
bool ok =
delete_missing_args_budgeted_observed(config, manifest, &budget, observer, observer_context);
if (deleted)
*deleted = budget.deleted;
if (skipped)
*skipped = budget.skipped;
if (limit_hit)
*limit_hit = budget.limit_hit;
return ok;
}
/* Commit every deletion family the manifest carries. The --delete-missing-args
exact-path deletions run FIRST: they are explicit user requests and must not
be blocked by the extras walker's filter-exclusion protection (a protected
leftover inside a missing-argument directory must not make that user-requested
removal fail). The ordinary extras walk then runs when --delete is active.
Both draw from one --max-delete budget; the result reports a cap-stopped
(partial) commit distinctly so the client can exit 25 like rsync. */
DeleteCommitResult manifest_delete_all(const Config* config, const DeleteManifest* manifest) {
return manifest_delete_all_counted(config, manifest, NULL);
}
DeleteCommitResult manifest_delete_all_counted(const Config* config, const DeleteManifest* manifest,
size_t* deleted) {
return manifest_delete_all_observed(config, manifest, deleted, NULL, NULL);
}
DeleteCommitResult manifest_delete_all_observed(const Config* config,
const DeleteManifest* manifest, size_t* deleted,
DeletePathObserver observer,
void* observer_context) {
if (deleted)
*deleted = 0;
if (!config || !manifest)
return DELETE_COMMIT_ERROR;
/* Central no-mutation guard: a dry-run never deletes. No manifest is sent on
the dry-run path, but a hostile/buggy peer could; treat it as a no-op so
the receiver can never remove anything. */
if (config->dry_run)
return DELETE_COMMIT_OK;
/* A client --max-delete=NUM smaller than the server's hard bound replaces it
for this run; both still bound the commit. */
bool user_limited =
config->max_delete >= 0 && (size_t)config->max_delete < MAX_SERVER_DELETE_COUNT;
DeleteBudgetState budget = {.max_delete = user_limited ? (size_t)config->max_delete
: MAX_SERVER_DELETE_COUNT,
.deleted = 0,
.skipped = 0,
.limit_hit = false};
if (config->delete_missing_args &&
!delete_missing_args_budgeted_observed(config, manifest, &budget, observer, observer_context))
return DELETE_COMMIT_ERROR;
if (config->use_delete &&
!delete_extras_budgeted_observed(config, manifest, &budget, observer, observer_context))
return DELETE_COMMIT_ERROR;
if (deleted)
*deleted = budget.deleted;
if (budget.limit_hit) {
if (user_limited) {
log_message(LOG_LEVEL_ERROR, "Deletions stopped due to --max-delete limit (%zu skipped)",
budget.skipped);
} else {
log_message(LOG_LEVEL_ERROR,
"Deletions stopped due to the server deletion limit of %u (%zu skipped)",
(unsigned)MAX_SERVER_DELETE_COUNT, budget.skipped);
}
return DELETE_COMMIT_LIMIT_REACHED;
}
return DELETE_COMMIT_OK;
}
+105
View File
@@ -0,0 +1,105 @@
#ifndef DELETE_COMMIT_H
#define DELETE_COMMIT_H
#include "array_list.h"
#include "config.h"
#include "delete.h"
#include <stdbool.h>
/* Delete-commit module: delete-manifest receive plus the budgeted extras and
* --delete-missing-args walkers. These declarations are re-exported by the
* file_receive.h facade. */
/* A received delete-manifest frame: the keep-set (`keeps`, destination-relative
paths the sender transferred/keeps) plus `protected`, destination-relative
prefixes the sender asks the receiver never to delete (paths excluded on the
source, protected at any depth). When --delete-excluded is given the sender
transmits an empty protected list so excluded destination mirrors are treated
as ordinary extras. With --delete-missing-args a third section (`missing`)
carries the destination mirrors of explicitly-listed source entries that do
not exist: each is an exact deletion request, independent of the ordinary
extras walk (never blocked by the protected prefixes) and processed when the
manifest is committed. */
typedef struct DeleteManifest {
ArrayList* keeps;
ArrayList* protected;
ArrayList* missing;
/* Destination-relative paths of the directories the sender synchronized for
this run. The extras walker only removes entries directly inside one of
these (the receive root is the "." sentinel); `--files-from` runs therefore
leave untransmitted directories and the unlisted parts of listed ones
alone, matching rsync's "delete only in synchronized directories". */
ArrayList* dirs;
} DeleteManifest;
void delete_manifest_free(DeleteManifest* manifest);
/* Read a delete-manifest frame (protocol 2.23.0): keep count + keeps, then
protected count + protected prefixes, then missing count + missing paths,
then synchronized-directory count + directory paths (self-delimiting; the
leading STATUS_MANIFEST code has been consumed). Returns an owned
DeleteManifest, or NULL after signalling STATUS_ERROR on a malformed frame. */
DeleteManifest* receive_manifest_entries(int fd);
/* Remove destination entries under config->receive_root_directory that are not
in `manifest` (bounded, all-or-nothing walk; staging-dir, basis-dir and
protected-prefix skips). `--max-delete` and `--force` are honored here. The
caller decides WHEN to run it based on the negotiated delete timing. Returns
false (and the transfer fails) when the deletion cannot be committed. */
bool manifest_delete_extras(const Config* config, const DeleteManifest* manifest);
/* --delete-missing-args exact-path deletions: remove each destination mirror
in `manifest->missing` (never blocked by the protected prefixes, staging dir
and basis dirs excluded). A regular file/symlink is unlinked; an empty
directory is removed; a NON-empty directory is removed recursively only when
--delete or --force is in effect, otherwise it is left with a warning (rsync
parity). A missing path is a no-op. Returns false only on a genuine
confinement or I/O error (the run then fails); tolerated per-path cases are
reported and skipped. */
bool manifest_delete_missing_args(const Config* config, const DeleteManifest* manifest);
/* Budgeted form of manifest_delete_missing_args for the per-directory delete
session: each removed mirror draws from `max_delete` (SIZE_MAX = unlimited)
and the tallies are accumulated into `*deleted`/`*skipped`. `*limit_hit` is set
when the budget stopped the pass with entries left over. Returns false only
on a genuine deletion error. */
bool manifest_delete_missing_args_limited(const Config* config, const DeleteManifest* manifest,
size_t max_delete, size_t* deleted, size_t* skipped,
bool* limit_hit);
/* Observer-aware form of manifest_delete_missing_args_limited: `observer` (may
be NULL) is invoked for every destination-relative path truly removed. */
bool manifest_delete_missing_args_limited_observed(
const Config* config, const DeleteManifest* manifest, size_t max_delete, size_t* deleted,
size_t* skipped, bool* limit_hit, DeletePathObserver observer, void* observer_context);
/* Outcome of committing a delete manifest. LIMIT_REACHED reports rsync's
partial --max-delete result: the budget allowed some deletions and the rest
were skipped (the run still stores all file data but the client exits 25). */
typedef enum {
DELETE_COMMIT_OK = 0,
DELETE_COMMIT_LIMIT_REACHED,
DELETE_COMMIT_ERROR
} DeleteCommitResult;
/* Run every deletion family the manifest carries: the --delete-missing-args
exact-path deletions first (user requests are not blocked by exclusion
protection), then the ordinary extras walk when --delete is active. Both
share one --max-delete budget. Returns DELETE_COMMIT_OK when nothing was to
do or everything committed, DELETE_COMMIT_LIMIT_REACHED when the budget
stopped part of the work, or DELETE_COMMIT_ERROR on a genuine failure. */
DeleteCommitResult manifest_delete_all(const Config* config, const DeleteManifest* manifest);
/* Like manifest_delete_all, but reports how many destination entries the commit
removed (for the end-of-transfer wire stats). `deleted` may be NULL. */
DeleteCommitResult manifest_delete_all_counted(const Config* config, const DeleteManifest* manifest,
size_t* deleted);
/* Observer-aware form of manifest_delete_all_counted: `observer` (may be NULL)
is invoked for every destination-relative path truly removed. */
DeleteCommitResult manifest_delete_all_observed(const Config* config,
const DeleteManifest* manifest, size_t* deleted,
DeletePathObserver observer,
void* observer_context);
/* -n/--dry-run --delete would-delete reporting: walk the destination exactly as
the delete pass would and append (strdup'd) destination-relative paths that
WOULD be removed to `out`, without touching disk. Uses the same staging-dir,
basis-dir and protected-prefix skips as the real commit. Returns true on a
clean walk; `*count_out` receives the number of paths appended. */
bool manifest_would_delete_list(const Config* config, const DeleteManifest* manifest,
ArrayList* out, size_t* count_out);
#endif
+289 -108
View File
@@ -2,6 +2,7 @@
#include "charset.h"
#include "delay_updates.h"
#include "delete.h"
#include "file.h"
#include "log.h"
#include "utils.h"
@@ -331,6 +332,8 @@ static int send_plan_node(int fd, DeletePlanSender* sender, PlanNode* node) {
return -1;
sender->config_sent = true;
}
if (!send_int(fd, 1)) /* apply = true */
return -1;
if (!send_wire_str(fd, node->dir))
return -1;
if (send_str_section(fd, node->dirs) != 0 || send_str_section(fd, node->files) != 0)
@@ -339,6 +342,29 @@ static int send_plan_node(int fd, DeletePlanSender* sender, PlanNode* node) {
return 0;
}
/* Transmit the one-shot per-run config block (protected prefixes, size-pruned
* mirrors, --delete-missing-args exact paths) on its own carrier frame, with
* apply=false so the receiver consumes the config but walks nothing. This is
* how the config still reaches the receiver when the scope allows no directory
* plan at all (a --files-from list of bare files synchronizes no directory):
* without it, the missing-args exact deletions would be lost. Idempotent. */
static int send_config_only(int fd, DeletePlanSender* sender) {
if (!sender || sender->config_sent)
return 0;
if (!send_status(fd, STATUS_DELETE_PLAN) || !send_int(fd, 1))
return -1;
if (send_str_section(fd, sender->protected_prefixes) != 0 ||
send_str_section(fd, sender->size_skipped) != 0 ||
send_str_section(fd, sender->missing_args) != 0)
return -1;
sender->config_sent = true;
if (!send_int(fd, 0)) /* apply = false */
return -1;
if (!send_wire_str(fd, ".") || !send_int(fd, 0) || !send_int(fd, 0))
return -1;
return 0;
}
static int send_prefix_plan(int fd, DeletePlanSender* sender, const char* dir) {
PlanNode* node = plan_find(sender, dir);
if (!node || node->sent)
@@ -354,6 +380,10 @@ int delete_plan_send_root(int fd, DeletePlanSender* sender) {
const char* root = sender->walk_root ? sender->walk_root : ".";
if (!plan_ensure(sender, root))
return -1;
/* Put the config block on the wire first, on its own carrier frame, so the
receiver always sees it even when the scope permits no directory plan. */
if (send_config_only(fd, sender) != 0)
return -1;
return send_prefix_plan(fd, sender, root);
}
@@ -403,6 +433,17 @@ int delete_plan_send_remaining(int fd, DeletePlanSender* sender, const ArrayList
return 0;
}
int delete_plan_send_all(int fd, DeletePlanSender* sender, const ArrayList* dirs) {
if (!sender)
return -1;
/* Root first: this also transmits the one-shot per-run config block on its
own carrier frame (see send_config_only), so it reaches the receiver even
when the scope permits no directory plan at all. */
if (delete_plan_send_root(fd, sender) != 0)
return -1;
return delete_plan_send_remaining(fd, sender, dirs);
}
/* ------------------------------------------------------------------ */
/* Receiver: delete session */
/* ------------------------------------------------------------------ */
@@ -412,6 +453,17 @@ struct DeletePlanSession {
bool dry_run;
size_t max_delete;
size_t deleted;
/* Removals charged against --max-delete. The budget is charged on ACTUAL
removals (an unlink/rmdir that succeeded), matching rsync: a snapshotted
entry that fails removal consumes nothing, so a later extra is still
deleted. `planned` and `deleted` advance together for the inline paths and
`apply_missing`; `deleted` is the reported count. */
size_t planned;
/* Hard bound on the deferred snapshot list. Because the budget is no longer
charged at snapshot time, this independent cap keeps a huge destination
from growing the list without limit (it matches the receiver's overall
deletion bound). */
size_t defer_cap;
size_t skipped;
bool limit_hit;
bool limit_logged;
@@ -421,8 +473,34 @@ struct DeletePlanSession {
ArrayList* size_skipped;
ArrayList* missing;
ArrayList* deferred;
DeletePathObserver observer;
void* observer_context;
};
/* Report one path the session truly removed (no-op without an observer). */
static void notify_deleted(DeletePlanSession* session, const char* rel) {
if (session && session->observer && rel)
session->observer(session->observer_context, rel);
}
/* A removed directory is reported with rsync's trailing slash (`deleting dir/`)
while files keep their bare path. */
static void notify_deleted_dir(DeletePlanSession* session, const char* rel) {
if (!session || !session->observer || !rel)
return;
size_t len = strlen(rel);
char* with_slash = malloc(len + 2);
if (!with_slash) {
session->observer(session->observer_context, rel);
return;
}
memcpy(with_slash, rel, len);
with_slash[len] = '/';
with_slash[len + 1] = '\0';
session->observer(session->observer_context, with_slash);
free(with_slash);
}
DeletePlanSession* delete_plan_session_create(const Config* config) {
if (!config)
return NULL;
@@ -435,6 +513,7 @@ DeletePlanSession* delete_plan_session_create(const Config* config) {
config->max_delete >= 0 && (size_t)config->max_delete < DELETE_PLAN_SERVER_LIMIT;
session->max_delete =
user_limited ? (size_t)config->max_delete : (size_t)DELETE_PLAN_SERVER_LIMIT;
session->defer_cap = DELETE_PLAN_SERVER_LIMIT;
session->protected_prefixes = array_list_create(free);
session->size_skipped = array_list_create(free);
session->missing = array_list_create(free);
@@ -523,49 +602,27 @@ static int open_plan_dir(const Config* config, const char* dir) {
return fd;
}
typedef struct PlanSkips {
DeleteSkipEntry* entries;
int count;
typedef struct {
DeleteSkipSet set;
/* Receiver-side delete-protection rules received on the config frame (NULL
when the sender sent none). Evaluated per extra so a protect/risk rule is
honored under --delete-during/--delete-delay exactly like the whole-tree
commit walker. */
const FilterRuleList* protect_rules;
} PlanSkips;
static bool build_plan_skips(const Config* config, const DeletePlanSession* session,
PlanSkips* out) {
out->entries = NULL;
out->count = 0;
int count = (config->delay_updates ? 1 : 0) + config->basis_count +
session->protected_prefixes->size + session->size_skipped->size;
if (count == 0)
return true;
out->entries = calloc((size_t)count, sizeof(DeleteSkipEntry));
if (!out->entries)
return false;
int idx = 0;
if (config->delay_updates) {
out->entries[idx].prefix = DELAY_UPDATES_STAGING_DIR;
out->entries[idx].top_level_only = true;
idx++;
}
for (int i = 0; i < config->basis_count; i++) {
out->entries[idx].prefix = config->basis_dirs[i].path;
out->entries[idx].top_level_only = false;
idx++;
}
for (int i = 0; i < session->protected_prefixes->size; i++) {
out->entries[idx].prefix = (const char*)session->protected_prefixes->items[i];
out->entries[idx].top_level_only = false;
idx++;
}
for (int i = 0; i < session->size_skipped->size; i++) {
out->entries[idx].prefix = (const char*)session->size_skipped->items[i];
out->entries[idx].top_level_only = false;
idx++;
}
out->count = idx;
return true;
out->protect_rules = config->protect_rules;
/* The per-directory plan walk keeps each basis path verbatim (it does not
convert an absolute under-root path to its root-relative form, unlike the
whole-tree commit walk). */
return delete_skips_build(config, session->protected_prefixes, session->size_skipped, false,
&out->set);
}
static bool budget_available(const DeletePlanSession* session) {
return session->deleted < session->max_delete;
return session->planned < session->max_delete;
}
static void note_skipped(DeletePlanSession* session) {
@@ -579,8 +636,15 @@ static void log_deleted(const char* rel) {
free(escaped);
}
/* Append a snapshot path for --delete-delay. */
/* Append a snapshot path for --delete-delay. The budget is NOT charged here:
* the remover charges --max-delete only when a path is actually unlinked (see
* apply_deferred_path), so a snapshotted entry that survives ENOTEMPTY cannot
* deny budget to a later extra. The independent `defer_cap` bounds the list. */
static bool defer_add(DeletePlanSession* session, const char* rel) {
if ((size_t)session->deferred->size >= session->defer_cap) {
note_skipped(session);
return true;
}
char* copy = str_dup(rel);
if (!copy)
return false;
@@ -588,7 +652,6 @@ static bool defer_add(DeletePlanSession* session, const char* rel) {
free(copy);
return false;
}
session->deleted++;
return true;
}
@@ -621,19 +684,21 @@ static bool process_extra_dir(int dirfd, const char* name, const char* child_rel
return false;
if (survives)
return true;
if (!budget_available(session)) {
note_skipped(session);
return true;
}
if (session->defer && !force_now) {
if (!defer_add(session, child_rel))
return false;
*removed = true;
return true;
}
if (!budget_available(session)) {
note_skipped(session);
return true;
}
if (unlinkat(dirfd, name, AT_REMOVEDIR) == 0) {
session->deleted++;
session->planned++;
log_deleted(child_rel);
notify_deleted_dir(session, child_rel);
*removed = true;
return true;
}
@@ -648,16 +713,18 @@ static bool process_extra_dir(int dirfd, const char* name, const char* child_rel
static bool process_extra_file(int dirfd, const char* name, const char* child_rel, bool force_now,
DeletePlanSession* session) {
if (session->defer && !force_now) {
return defer_add(session, child_rel);
}
if (!budget_available(session)) {
note_skipped(session);
return true;
}
if (session->defer && !force_now) {
return defer_add(session, child_rel);
}
if (unlinkat(dirfd, name, 0) == 0) {
session->deleted++;
session->planned++;
log_deleted(child_rel);
notify_deleted(session, child_rel);
} else if (errno != ENOENT) {
return false;
}
@@ -668,70 +735,108 @@ static bool process_children(int dirfd, const char* dir_rel, const ArrayList* ke
const ArrayList* keep_files, bool at_root, bool force_now,
const PlanSkips* skips, DeletePlanSession* session, bool* survives) {
*survives = false;
int scanfd = openat(dirfd, ".", O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
if (scanfd < 0)
DeleteDirEntry* entries = NULL;
size_t count = 0;
bool collect_ok = true;
if (!delete_dir_entries_collect(dirfd, &entries, &count, &collect_ok))
return false;
DIR* dir = fdopendir(scanfd);
if (!dir) {
close(scanfd);
bool operation_ok = collect_ok;
bool local_survives = false;
bool* shielded = calloc(count ? count : 1, sizeof(bool));
bool* is_extra = calloc(count ? count : 1, sizeof(bool));
bool* force = calloc(count ? count : 1, sizeof(bool));
if (!shielded || !is_extra || !force) {
free(shielded);
free(is_extra);
free(force);
delete_dir_entries_free(entries, count);
return false;
}
bool operation_ok = true;
bool local_survives = false;
const struct dirent* entry;
while ((entry = readdir(dir)) != NULL) {
if (strcmp(entry->d_name, ".") == 0 || strcmp(entry->d_name, "..") == 0)
continue;
/* rsync's order: extraneous subdirectories in descending name order, then
extraneous files in descending name order (kept entries survive and are not
touched here — a kept subdirectory gets its own per-directory plan). */
if (count > 1)
qsort(entries, count, sizeof(*entries), delete_dir_entry_cmp_desc);
size_t dir_count = 0;
while (dir_count < count && entries[dir_count].is_dir)
dir_count++;
for (size_t i = 0; i < count; i++) {
char* child_rel =
(strcmp(dir_rel, ".") == 0) ? str_dup(entry->d_name) : path_cat(dir_rel, entry->d_name);
(strcmp(dir_rel, ".") == 0) ? str_dup(entries[i].name) : path_cat(dir_rel, entries[i].name);
if (!child_rel) {
operation_ok = false;
continue;
}
if (path_under_skip_prefix(child_rel, at_root, skips->entries, skips->count)) {
if (path_under_skip_prefix(child_rel, at_root, skips->set.entries, skips->set.count)) {
shielded[i] = true;
local_survives = true;
free(child_rel);
continue;
}
struct stat st;
if (fstatat(dirfd, entry->d_name, &st, AT_SYMLINK_NOFOLLOW) != 0) {
if (errno != ENOENT)
operation_ok = false;
free(child_rel);
continue;
}
bool is_dir = S_ISDIR(st.st_mode);
bool in_keep_dirs = is_dir && list_contains_str(keep_dirs, entry->d_name);
bool in_keep_files = !is_dir && list_contains_str(keep_files, entry->d_name);
if (in_keep_dirs) {
bool is_dir = entries[i].is_dir;
bool in_keep_dirs = is_dir && list_contains_str(keep_dirs, entries[i].name);
bool in_keep_files = !is_dir && list_contains_str(keep_files, entries[i].name);
bool rule_protected =
skips->protect_rules &&
filter_rules_apply_side(skips->protect_rules, child_rel, entries[i].name, is_dir,
FILTER_SIDE_RECEIVER) == FILTER_ACTION_PROTECT;
if (in_keep_dirs || in_keep_files || rule_protected) {
shielded[i] = true;
local_survives = true;
} else if (keep_dirs && !is_dir && list_contains_str(keep_dirs, entry->d_name)) {
/* Destination file blocks a source directory: clear it now, whatever the
delete timing, so the directory can be created. */
if (!process_extra_file(dirfd, entry->d_name, child_rel, true, session))
operation_ok = false;
} else if (in_keep_files) {
local_survives = true;
} else if (keep_files && is_dir && list_contains_str(keep_files, entry->d_name)) {
/* Destination directory blocks a source file: remove it now. */
bool removed = false;
if (!process_extra_dir(dirfd, entry->d_name, child_rel, true, skips, session, &removed))
operation_ok = false;
else if (!removed)
local_survives = true;
} else if (is_dir) {
bool removed = false;
if (!process_extra_dir(dirfd, entry->d_name, child_rel, force_now, skips, session, &removed))
operation_ok = false;
else if (!removed)
local_survives = true;
/* A destination directory blocks a source file of the same name: remove
it now, whatever the delete timing, so the file can be created. */
is_extra[i] = true;
force[i] = keep_files && list_contains_str(keep_files, entries[i].name);
} else {
if (!process_extra_file(dirfd, entry->d_name, child_rel, force_now, session))
operation_ok = false;
/* A destination file blocks a source directory of the same name: clear it
now so the directory can be created. */
is_extra[i] = true;
force[i] = keep_dirs && list_contains_str(keep_dirs, entries[i].name);
}
free(child_rel);
}
closedir(dir);
/* Pass 1: extraneous subdirectories, descending. */
for (size_t i = 0; i < dir_count; i++) {
if (!is_extra[i])
continue;
char* child_rel =
(strcmp(dir_rel, ".") == 0) ? str_dup(entries[i].name) : path_cat(dir_rel, entries[i].name);
if (!child_rel) {
operation_ok = false;
continue;
}
bool removed = false;
if (!process_extra_dir(dirfd, entries[i].name, child_rel, force[i] || force_now, skips, session,
&removed))
operation_ok = false;
else if (!removed)
local_survives = true;
free(child_rel);
}
/* Pass 2: extraneous files, descending. */
for (size_t i = dir_count; i < count; i++) {
if (!is_extra[i])
continue;
char* child_rel =
(strcmp(dir_rel, ".") == 0) ? str_dup(entries[i].name) : path_cat(dir_rel, entries[i].name);
if (!child_rel) {
operation_ok = false;
continue;
}
if (!process_extra_file(dirfd, entries[i].name, child_rel, force[i] || force_now, session))
operation_ok = false;
free(child_rel);
}
free(shielded);
free(is_extra);
free(force);
delete_dir_entries_free(entries, count);
*survives = local_survives;
return operation_ok;
}
@@ -751,7 +856,7 @@ static bool apply_plan_dir(DeletePlanSession* session, const Config* config, con
bool survives = false;
bool ok = process_children(dirfd, dir, dirs, files, strcmp(dir, ".") == 0, false, &skips, session,
&survives);
free(skips.entries);
delete_skips_free(&skips.set);
close(dirfd);
if (!ok)
log_message(LOG_LEVEL_ERROR, "deletion failed while removing extraneous files");
@@ -768,13 +873,15 @@ static bool apply_missing(DeletePlanSession* session, const Config* config) {
return true;
DeleteManifest manifest = {
.keeps = NULL, .protected = NULL, .missing = session->missing, .dirs = NULL};
size_t remaining = budget_available(session) ? session->max_delete - session->deleted : 0;
size_t remaining = budget_available(session) ? session->max_delete - session->planned : 0;
size_t deleted = 0;
size_t skipped = 0;
bool limit = false;
bool ok = manifest_delete_missing_args_limited(config, &manifest, remaining, &deleted, &skipped,
&limit);
bool ok = manifest_delete_missing_args_limited_observed(config, &manifest, remaining, &deleted,
&skipped, &limit, session->observer,
session->observer_context);
session->deleted += deleted;
session->planned += deleted;
session->skipped += skipped;
if (limit)
session->limit_hit = true;
@@ -801,6 +908,14 @@ int delete_plan_session_receive(DeletePlanSession* session, const Config* config
}
session->config_seen = true;
}
/* apply=false is the config-only carrier frame: the receiver consumes the
config (and the missing-args exact deletions) but must not walk any
directory. Every real plan carries apply=true. */
int apply;
if (!receive_int(fd, &apply) || (apply != 0 && apply != 1)) {
send_status(fd, STATUS_ERROR);
return -1;
}
char* dir = receive_wire_str(fd);
ArrayList* dirs = array_list_create(free);
ArrayList* files = array_list_create(free);
@@ -818,7 +933,7 @@ int delete_plan_session_receive(DeletePlanSession* session, const Config* config
if (!session->dry_run && enabled) {
if (!session->defer && !apply_missing(session, config))
ok = false;
if (ok && !apply_plan_dir(session, config, dir, dirs, files))
if (ok && apply && !apply_plan_dir(session, config, dir, dirs, files))
ok = false;
}
free(dir);
@@ -837,37 +952,103 @@ int delete_plan_session_receive(DeletePlanSession* session, const Config* config
}
/* Apply one snapshotted --delete-delay path (post-order: children precede their
* parent directory). */
* parent directory). A directory that is still present is re-scanned so content
* created after the plan is removed too; every actual removal charges
* --max-delete. */
static bool apply_deferred_path(DeletePlanSession* session, const Config* config, const char* rel) {
(void)session;
char* full = path_cat(config->receive_root_directory, rel);
if (!full)
return false;
char* leaf = NULL;
int parent_fd = file_open_secure_parent(full, &leaf, false);
int open_errno = errno;
free(full);
if (parent_fd < 0) {
free(leaf);
return errno == ENOENT || errno == ENOTDIR;
return open_errno == ENOENT || open_errno == ENOTDIR;
}
struct stat st;
if (fstatat(parent_fd, leaf, &st, AT_SYMLINK_NOFOLLOW) != 0) {
bool absent = errno == ENOENT;
bool absent = errno == ENOENT || errno == ENOTDIR;
close(parent_fd);
free(leaf);
return absent;
}
int rc;
if (S_ISDIR(st.st_mode))
rc = unlinkat(parent_fd, leaf, AT_REMOVEDIR);
else
rc = unlinkat(parent_fd, leaf, 0);
bool ok = rc == 0 || errno == ENOENT || errno == ENOTEMPTY || errno == EEXIST;
if (rc == 0)
if (S_ISDIR(st.st_mode)) {
if (!budget_available(session)) {
note_skipped(session);
close(parent_fd);
free(leaf);
return true;
}
int dirfd = openat(parent_fd, leaf, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
if (dirfd < 0) {
bool absent = errno == ENOENT || errno == ENOTDIR;
close(parent_fd);
free(leaf);
return absent;
}
PlanSkips skips;
if (!build_plan_skips(config, session, &skips)) {
close(dirfd);
close(parent_fd);
free(leaf);
return false;
}
bool survives = false;
bool ok = process_children(dirfd, rel, NULL, NULL, false, true, &skips, session, &survives);
delete_skips_free(&skips.set);
close(dirfd);
if (!ok) {
close(parent_fd);
free(leaf);
return false;
}
if (!survives) {
if (!budget_available(session)) {
note_skipped(session);
} else if (unlinkat(parent_fd, leaf, AT_REMOVEDIR) == 0) {
session->deleted++;
session->planned++;
log_deleted(rel);
notify_deleted_dir(session, rel);
} else if (errno != ENOENT && errno != ENOTEMPTY && errno != EEXIST) {
close(parent_fd);
free(leaf);
return false;
}
}
close(parent_fd);
free(leaf);
return true;
}
if (!budget_available(session)) {
note_skipped(session);
close(parent_fd);
free(leaf);
return true;
}
if (unlinkat(parent_fd, leaf, 0) == 0) {
session->deleted++;
session->planned++;
log_deleted(rel);
notify_deleted(session, rel);
} else if (errno != ENOENT) {
close(parent_fd);
free(leaf);
return false;
}
close(parent_fd);
free(leaf);
return ok;
return true;
}
void delete_plan_session_set_delete_observer(DeletePlanSession* session,
DeletePathObserver observer, void* context) {
if (!session)
return;
session->observer = observer;
session->observer_context = context;
}
DeleteCommitResult delete_plan_session_commit(DeletePlanSession* session, const Config* config) {
+29 -6
View File
@@ -3,8 +3,10 @@
#include "array_list.h"
#include "config.h"
#include "delete.h"
#include "file_receive.h"
#include "protocol.h"
#include "utils.h"
#include <stdbool.h>
/* Per-directory delete plans (protocol 2.24.0).
@@ -47,19 +49,32 @@ void delete_plan_sender_finalize(DeletePlanSender* sender, const ArrayList* sync
Directory keep entries do not count, so an I/O error that hid every file
still refuses to delete. */
bool delete_plan_sender_empty(const DeletePlanSender* sender);
/* Attach the global config sections advertised on the first plan frame. */
/* Attach the global config sections advertised on the first plan frame. The
* block is always transmitted by delete_plan_send_root(), on a config-only
* carrier frame when the scope allows no directory plan. */
void delete_plan_sender_set_config(DeletePlanSender* sender, const ArrayList* protected_prefixes,
const ArrayList* size_skipped, const ArrayList* missing_args);
/* Send the root plan (even before any data, so root extras are handled like
* rsync's first generator directory). Returns -1 on I/O error. */
* rsync's first generator directory), after transmitting the per-run config
* block on its own carrier frame. Returns -1 on I/O error. */
int delete_plan_send_root(int fd, DeletePlanSender* sender);
/* Send the plans for every ancestor of `path` (root-first) and, when is_dir,
* for `path` itself; already-sent plans are skipped. */
int delete_plan_send_for_path(int fd, DeletePlanSender* sender, const char* path, bool is_dir);
/* Send the plan for every directory in `dirs` that has not been transmitted
* yet. Called after the data stream so an empty source directory's plan still
* clears its destination extras even though no file frame triggered it. */
* yet. */
int delete_plan_send_remaining(int fd, DeletePlanSender* sender, const ArrayList* dirs);
/* Transmit the COMPLETE per-directory plan set in one pass, before any data
* frame: the root plan (with the one-shot per-run config block on its carrier
* frame) followed by every directory in `dirs`. Because the whole plan set is
* known from the path-only pre-scan, sending it all up front means a
* mid-transfer abort has already applied every planned removal, matching
* rsync's generator (which runs ahead of its throttled sender). A completed
* run is unaffected. `dirs` is the set of directories whose direct children
* were enumerated (the scanner's plan_dirs sink), so a merely listed but
* untraversed directory never gets a plan and its mirror is left intact.
* Returns -1 on I/O error. */
int delete_plan_send_all(int fd, DeletePlanSender* sender, const ArrayList* dirs);
/* ---- Receiver: delete session ---- */
@@ -76,8 +91,16 @@ int delete_plan_session_receive(DeletePlanSession* session, const Config* config
DeleteCommitResult delete_plan_session_commit(DeletePlanSession* session, const Config* config);
/* True once the shared --max-delete budget stopped part of a deletion. */
bool delete_plan_session_limit_reached(const DeletePlanSession* session);
/* Number of destination entries the session's plans removed (or, for
--delete-delay, snapshotted for removal), for the end-of-transfer stats. */
/* Number of destination entries the session actually removed, for the
end-of-transfer stats. For --delete-delay this excludes a snapshotted entry
that survived (e.g. a refilled directory that failed ENOTEMPTY), even though
that entry already consumed --max-delete budget at snapshot time. */
size_t delete_plan_session_deleted(const DeletePlanSession* session);
/* Install an observer invoked for every destination-relative path the session
truly removes (including the deferred --delete-delay commit), so the receiver
can report rsync's `deleting PATH` lines through the terminal STATUS_STATS
record. Pass NULL/0 to clear. */
void delete_plan_session_set_delete_observer(DeletePlanSession* session,
DeletePathObserver observer, void* context);
#endif
+434 -38
View File
@@ -15,6 +15,7 @@
#include <unistd.h>
#include "data.h"
#include "checksum.h"
#include "delta.h"
#include "file.h"
#include "file_store.h"
@@ -39,6 +40,31 @@ static bool write_all(int fd, const void* data, unsigned long long size) {
return true;
}
/* Streaming copy of an open source descriptor into the just-created destination
`fd` (already at offset 0). Used by the --copy-dest basis install so a basis
larger than any in-memory whole-file bound still materializes without
buffering the entire file. `expected_size` is the caller-verified basis
size; the copy must produce exactly that many bytes (a short source is a hard
error, never a silently truncated destination). The final ftruncate drops
any residual tail a raced-in longer source might have left. */
static bool copy_fd_all(int dst_fd, int src_fd, unsigned long long expected_size) {
unsigned char buf[1 << 20];
unsigned long long done = 0;
while (done < expected_size) {
unsigned long long remaining = expected_size - done;
size_t want = remaining < sizeof(buf) ? (size_t)remaining : sizeof(buf);
ssize_t n = read(src_fd, buf, want);
if (n < 0 && errno == EINTR)
continue;
if (n <= 0)
return false;
if (!write_all(dst_fd, buf, (unsigned long long)n))
return false;
done += (unsigned long long)n;
}
return ftruncate(dst_fd, (off_t)expected_size) == 0;
}
/* Preallocate `size` bytes on `fd` before any data is written (--preallocate).
* fallocate(2) reserves real disk blocks, so an out-of-space condition
* (ENOSPC/EDQUOT) surfaces up front instead of partway through a transfer;
@@ -140,6 +166,12 @@ bool file_checksum(File* file, ChecksumAlgo algo, uint64_t seed, uint8_t* out, s
if (file->data->size == 0) {
return checksum_digest(algo, seed, "", 0, out, out_capacity, out_len);
}
/* A streamed source (data not loaded) may exceed any in-memory whole-file
bound; hash it from the file path in bounded buffers instead of forcing a
full load. This is the same digest the receiver recomputes on the basis. */
if (!file->data->data && file->path && file->data->size > STREAM_THRESHOLD &&
checksum_digest_file(algo, seed, file->path, out, out_capacity, out_len))
return true;
if (!file->data->data && !file_load_data(file))
return false;
return checksum_digest(algo, seed, file->data->data, file->data->size, out, out_capacity,
@@ -176,6 +208,7 @@ File* file_create(const char* path) {
file->is_dir = false;
file->dir_time_only = false;
file->basis_link = NULL;
file->basis_copy = NULL;
file->link_group = 0;
file->link_first = false;
file->hardlink_target = NULL;
@@ -187,6 +220,7 @@ File* file_create(const char* path) {
file->xattrs = NULL;
file->dest_state = (OutputDestState){0};
file->matched_bytes = 0;
file->literal_bytes = 0;
return file;
}
@@ -204,6 +238,8 @@ void file_destroy(void* item) {
file->send_path = NULL;
free(file->basis_link);
file->basis_link = NULL;
free(file->basis_copy);
file->basis_copy = NULL;
free(file->hardlink_target);
file->hardlink_target = NULL;
free(file->symlink_target);
@@ -614,6 +650,103 @@ static int open_dir_beneath_root(const char* resolved, const char* root) {
}
int file_open_secure_parent(const char* path, char** leaf_out, bool create_dirs) {
return file_open_secure_parent_counted(path, leaf_out, create_dirs, NULL, NULL);
}
/* The logical transfer root expressed in the same coordinate as the secure
* parent walk's `rel_buf` (relative to the authorized root, with a leading
* '/'), used as the floor at or below which a created directory is a real
* file-list entry. The on-disk transfer root is the receive root joined to the
* wire path; the mirror scaffolding above it (the absolute source path below
* the destination root) is not an rsync entry. Returns an allocated string or
* NULL (count every created component). */
static char* transfer_root_floor(const Config* config) {
if (!config || !config->send_directory || config->send_directory[0] == '\0')
return NULL;
const char* spec = config->send_directory;
const char* after = spec;
if (spec[0] == '.' && spec[1] == '/') {
after = spec + 2;
} else {
const char* cut = strstr(spec, "/./");
if (cut)
after = cut + 3;
}
while (*after == '/')
after++;
char* wire_root = str_dup(after);
if (!wire_root)
return NULL;
size_t wlen = strlen(wire_root);
while (wlen > 0 && wire_root[wlen - 1] == '/')
wire_root[--wlen] = '\0';
if (wlen == 0) {
free(wire_root);
return NULL;
}
char* disk_root = config->receive_root_directory
? path_cat(config->receive_root_directory, wire_root)
: str_dup(wire_root);
free(wire_root);
if (!disk_root)
return NULL;
const char* root_path = utils_get_authorized_root_path();
const char* floor = disk_root;
if (root_path && root_path[0] == '/') {
size_t rl = strlen(root_path);
while (rl > 0 && root_path[rl - 1] == '/')
rl--;
if (strncmp(disk_root, root_path, rl) == 0 && (disk_root[rl] == '/' || disk_root[rl] == '\0'))
floor = disk_root + rl;
}
while (*floor == '/')
floor++;
char* out = str_dup(floor);
free(disk_root);
if (!out)
return NULL;
if (out[0] == '\0') {
free(out);
return NULL;
}
return out;
}
/* A created parent component counts toward `Number of created files` only when
* its receive-root-relative path is at or below the logical transfer root
* (`count_floor`). The transfer root itself corresponds to rsync's `.` entry
* (created on a fresh destination, pre-existing otherwise); the mirror
* scaffolding above it is FastSync's absolute-path layout, not an rsync entry. */
static bool created_dir_counts(const char* count_floor, const char* rel_buf,
const char* component) {
if (!count_floor)
return true;
char candidate[PATH_MAX];
int n = snprintf(candidate, sizeof(candidate), "%s/%s", rel_buf, component);
if (n < 0 || (size_t)n >= sizeof(candidate))
return false;
const char* cand = candidate;
while (*cand == '/')
cand++;
size_t fl = strlen(count_floor);
if (strncmp(cand, count_floor, fl) != 0)
return false;
return cand[fl] == '\0' || cand[fl] == '/';
}
/* Public wrapper for the receiver's created-directory accounting: the logical
* transfer root expressed receive-root-relative, or NULL when the wire paths
* carry no mirror scaffolding above it (--relative and --files-from, whose
* paths are already relative to the transfer root). The caller frees a
* non-NULL result. */
char* file_transfer_root_floor(const Config* config) {
if (!config || config->relative || config->files_from_set != NULL)
return NULL;
return transfer_root_floor(config);
}
int file_open_secure_parent_counted(const char* path, char** leaf_out, bool create_dirs,
unsigned* dirs_created, const char* count_floor) {
char* copy = str_dup(path);
if (!copy)
return -1;
@@ -674,6 +807,14 @@ int file_open_secure_parent(const char* path, char** leaf_out, bool create_dirs)
if (next < 0 && create_dirs && errno == ENOENT) {
bool created = mkdirat(fd, component, (mode_t)(0777 & ~(mode_t)file_process_umask())) == 0;
if (created || errno == EEXIST) {
/* Protocol 2.28.0: only directories the logical file list would
create count toward `Number of created files`; the mirror
scaffolding above the transfer root (e.g. the absolute source path
under the destination root) is not an rsync entry. `count_floor`
is a receive-root-relative prefix that must be reached before a
created component is counted. */
if (created && dirs_created && created_dir_counts(count_floor, rel_buf, component))
(*dirs_created)++;
/* P7 Wave E: --copy-as owns EVERY entry, including the intermediate
directories this walk creates implicitly. Its target ids are a
global policy, so they are available here without per-entry source
@@ -807,6 +948,25 @@ bool file_ensure_directory_secure(const char* path) {
} else if (errno == EEXIST) {
dir_fd = openat(parent_fd, leaf, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
}
} else if (dir_fd < 0 && errno == ENOTDIR) {
/* rsync replaces a destination non-directory (regular file) with an
incoming directory. Confined to the already-opened secure parent fd:
the leaf is unlinked by name (never followed) and only a non-directory
is ever removed, so this cannot escape the authorized root or remove a
pre-existing directory tree. A symlink is left alone (openat with
O_NOFOLLOW reports ELOOP, which takes no branch here), since replacing
it is not required for FastSync's transferred directories and keeps
--keep-dirlinks semantics untouched. */
struct stat leaf_st;
if (fstatat(parent_fd, leaf, &leaf_st, AT_SYMLINK_NOFOLLOW) == 0 && !S_ISDIR(leaf_st.st_mode) &&
!S_ISLNK(leaf_st.st_mode)) {
if (unlinkat(parent_fd, leaf, 0) == 0) {
if (mkdirat(parent_fd, leaf, (mode_t)(0777 & ~(mode_t)file_process_umask())) == 0)
created = true;
/* On failure dir_fd stays < 0 below, so the caller still sees it. */
dir_fd = openat(parent_fd, leaf, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
}
}
}
bool ok = dir_fd >= 0;
/* --copy-as owns a directory this call just created (the final component;
@@ -969,17 +1129,51 @@ int file_open_private_dir(const char* dir_path) {
return fd;
}
/* Open a --temp-dir scratch directory exactly as rsync does: the directory must
* already exist and is used as given (an absolute path is used verbatim, a
* relative one was already resolved against the destination root by the
* caller). Unlike file_open_private_dir this neither creates it nor confines
* it below the receive root, because rsync accepts any temp dir -- including
* one outside the destination tree or on another filesystem. Returns an
* O_DIRECTORY|O_CLOEXEC fd, or -1 on error. */
/* Open a --temp-dir scratch directory. The directory must already exist (rsync
* never creates it); a relative path was already resolved against the
* destination root by the caller. Unlike file_open_private_dir this neither
* creates it nor requires it to be a direct child of the receive root, because
* rsync permits a scratch dir that (via a symlink) lands on another filesystem
* -- but it MUST resolve inside the authorized receive root. The directory is
* opened following symlinks and then judged by the REAL path of the opened fd
* (through /proc/self/fd), so a client-planted symlink under the receive root
* can never redirect receiver scratch files outside the sandbox while an
* in-root link to another filesystem (the EXDEV fallback case) still works.
* Returns an O_DIRECTORY|O_CLOEXEC fd, or -1 on error (errno set; an escaping
* target is reported as EACCES with a logged reason). */
int file_open_temp_dir(const char* dir_path) {
if (!dir_path)
return -1;
return open(dir_path, O_RDONLY | O_DIRECTORY | O_CLOEXEC);
int fd = open(dir_path, O_RDONLY | O_DIRECTORY | O_CLOEXEC);
if (fd < 0)
return -1;
const char* root = utils_get_authorized_root_path();
if (!root) {
/* No authorized root (e.g. a local batch apply): nothing to confine
against, so preserve the historical open-as-given behavior. */
return fd;
}
char fd_path[64];
int fd_path_length = snprintf(fd_path, sizeof(fd_path), "/proc/self/fd/%d", fd);
char resolved[PATH_MAX];
if (fd_path_length < 0 || (size_t)fd_path_length >= sizeof(fd_path) ||
!realpath(fd_path, resolved)) {
int saved_errno = errno;
close(fd);
errno = saved_errno;
return -1;
}
if (!path_is_within_root(root, resolved)) {
char* escaped = output_escape(dir_path, log_get_8_bit_output());
log_message(LOG_LEVEL_ERROR,
"--temp-dir '%s' resolves outside the authorized receive root; refusing",
escaped ? escaped : "<allocation failed>");
free(escaped);
close(fd);
errno = EACCES;
return -1;
}
return fd;
}
/* After the content and mode/times are restored on the just-written file, apply
@@ -988,34 +1182,37 @@ int file_open_temp_dir(const char* dir_path) {
* destination file) and best-effort: a per-attribute or privilege failure is
* logged and skipped, never fatal. */
static void restore_extra_fd(int fd, const FileMetadata* metadata, const FileXattrList* xattrs,
bool fake_super, FileAttrPolicy policy) {
bool fake_super, FileAttrPolicy policy, uint32_t fake_super_rdev_major,
uint32_t fake_super_rdev_minor) {
xattr_apply_fd(fd, xattrs);
if (fake_super && metadata) {
/* Record the ownership that WOULD have been applied: when an explicit
ownership request (--chown/--usermap/--groupmap/--copy-as or -o/-g) is
active, the resolved mapping; otherwise the source's own id. The real
chown is suppressed (identity_apply_ownership early-returns under
--fake-super) so recording never defeats the flag. Mode/mtime are still
replayed (policy-gated) so unprivileged --fake-super keeps working. */
--fake-super) so recording never defeats the flag. The recorded stat is
rsync's format; the permission bits are replayed (policy-gated) so
unprivileged --fake-super keeps working while mtime comes from the
normal metadata path above. */
uint32_t store_uid;
uint32_t store_gid;
identity_resolve_storage_ids((int32_t)metadata->uid, (int32_t)metadata->gid, &store_uid,
&store_gid);
fake_super_store_fd(fd, store_uid, store_gid, (uint32_t)metadata->mode, metadata->mtime_sec,
metadata->mtime_nsec);
fake_super_store_fd(fd, store_uid, store_gid, (uint32_t)metadata->mode, fake_super_rdev_major,
fake_super_rdev_minor);
fake_super_restore_fd(fd, policy);
}
}
static bool file_to_disk_secure_impl(const char* path, const void* data,
unsigned long long data_size, bool inplace, bool sparse,
bool preallocate, const FileMetadata* metadata,
FileAttrPolicy policy, bool update, bool no_replace,
bool use_fsync, const char* temp_dir,
const FileXattrList* xattrs, bool fake_super,
bool keep_partial) {
static bool
file_to_disk_secure_impl(const char* path, const void* data, unsigned long long data_size,
bool inplace, bool sparse, bool preallocate, const FileMetadata* metadata,
FileAttrPolicy policy, bool update, bool no_replace, bool use_fsync,
const char* temp_dir, const FileXattrList* xattrs, bool fake_super,
bool keep_partial, unsigned* dirs_created, const char* count_floor,
uint32_t fake_super_rdev_major, uint32_t fake_super_rdev_minor) {
char* leaf = NULL;
int dirfd = file_open_secure_parent(path, &leaf, true);
int dirfd = file_open_secure_parent_counted(path, &leaf, true, dirs_created, count_floor);
if (dirfd < 0)
return false;
int fd = -1;
@@ -1120,7 +1317,8 @@ static bool file_to_disk_secure_impl(const char* path, const void* data,
}
}
if (ok)
restore_extra_fd(fd, metadata, xattrs, fake_super, policy);
restore_extra_fd(fd, metadata, xattrs, fake_super, policy, fake_super_rdev_major,
fake_super_rdev_minor);
if (ok && use_fsync)
ok = fsync(fd) == 0;
}
@@ -1238,7 +1436,8 @@ static bool file_to_disk_secure_impl(const char* path, const void* data,
}
}
if (ok)
restore_extra_fd(fd, metadata, xattrs, fake_super, policy);
restore_extra_fd(fd, metadata, xattrs, fake_super, policy, fake_super_rdev_major,
fake_super_rdev_minor);
if (ok && use_fsync)
ok = fsync(fd) == 0;
}
@@ -1306,7 +1505,8 @@ static bool file_to_disk_secure_impl(const char* path, const void* data,
"non-atomic copy into the destination directory");
return file_to_disk_secure_impl(path, data, data_size, inplace, sparse, preallocate, metadata,
policy, update, no_replace, use_fsync, NULL, xattrs, fake_super,
keep_partial);
keep_partial, dirs_created, count_floor, fake_super_rdev_major,
fake_super_rdev_minor);
}
return ok;
}
@@ -1315,7 +1515,8 @@ bool file_to_disk_secure(const char* path, const void* data, unsigned long long
bool inplace, bool sparse, bool preallocate, const FileMetadata* metadata,
FileAttrPolicy policy, const char* temp_dir) {
return file_to_disk_secure_impl(path, data, data_size, inplace, sparse, preallocate, metadata,
policy, false, false, false, temp_dir, NULL, false, false);
policy, false, false, false, temp_dir, NULL, false, false, NULL,
NULL, 0, 0);
}
bool file_to_disk_secure_update(const char* path, const void* data, unsigned long long data_size,
@@ -1323,7 +1524,8 @@ bool file_to_disk_secure_update(const char* path, const void* data, unsigned lon
const FileMetadata* metadata, FileAttrPolicy policy,
const char* temp_dir) {
return file_to_disk_secure_impl(path, data, data_size, inplace, sparse, preallocate, metadata,
policy, true, false, false, temp_dir, NULL, false, false);
policy, true, false, false, temp_dir, NULL, false, false, NULL,
NULL, 0, 0);
}
bool file_to_disk_secure_with_fsync(const char* path, const void* data,
@@ -1331,7 +1533,8 @@ bool file_to_disk_secure_with_fsync(const char* path, const void* data,
bool preallocate, const FileMetadata* metadata,
FileAttrPolicy policy, bool use_fsync, const char* temp_dir) {
return file_to_disk_secure_impl(path, data, data_size, inplace, sparse, preallocate, metadata,
policy, false, false, use_fsync, temp_dir, NULL, false, false);
policy, false, false, use_fsync, temp_dir, NULL, false, false,
NULL, NULL, 0, 0);
}
bool file_to_disk_secure_no_replace(const char* path, const void* data,
@@ -1339,7 +1542,8 @@ bool file_to_disk_secure_no_replace(const char* path, const void* data,
const FileMetadata* metadata, FileAttrPolicy policy,
const char* temp_dir) {
return file_to_disk_secure_impl(path, data, data_size, false, sparse, preallocate, metadata,
policy, false, true, false, temp_dir, NULL, false, false);
policy, false, true, false, temp_dir, NULL, false, false, NULL,
NULL, 0, 0);
}
/* Receiver write-path variant that also applies the per-file xattrs (-X/-A)
@@ -1352,9 +1556,21 @@ bool file_to_disk_secure_attrs(const char* path, const void* data, unsigned long
const FileMetadata* metadata, FileAttrPolicy policy, bool update,
bool no_replace, bool use_fsync, const FileXattrList* xattrs,
bool fake_super, bool keep_partial, const char* temp_dir) {
return file_to_disk_secure_attrs_counted(path, data, data_size, inplace, sparse, preallocate,
metadata, policy, update, no_replace, use_fsync, xattrs,
fake_super, keep_partial, temp_dir, NULL, NULL, 0, 0);
}
bool file_to_disk_secure_attrs_counted(
const char* path, const void* data, unsigned long long data_size, bool inplace, bool sparse,
bool preallocate, const FileMetadata* metadata, FileAttrPolicy policy, bool update,
bool no_replace, bool use_fsync, const FileXattrList* xattrs, bool fake_super,
bool keep_partial, const char* temp_dir, unsigned* dirs_created, const char* count_floor,
uint32_t fake_super_rdev_major, uint32_t fake_super_rdev_minor) {
return file_to_disk_secure_impl(path, data, data_size, inplace, sparse, preallocate, metadata,
policy, update, no_replace, use_fsync, temp_dir, xattrs,
fake_super, keep_partial);
fake_super, keep_partial, dirs_created, count_floor,
fake_super_rdev_major, fake_super_rdev_minor);
}
/* Atomic --link-dest install. The destination is replaced (via a temporary
@@ -1372,16 +1588,176 @@ bool file_to_disk_secure_attrs(const char* path, const void* data, unsigned long
* basis). Likewise `xattrs`/`fake_super` are applied only on the copy
* fallback, so a fallback copy preserves the per-file attributes instead of
* silently dropping them. */
/* Streaming --copy-dest basis install: atomically materialize `path` from the
* bytes of `basis_path` without holding the file in memory, so a basis larger
* than any whole-file bound still works. Mirrors the ordinary secure store
* path (confined parent walk, temp + rename, --update/--ignore-existing/
* --preallocate/--temp-dir) but sources the data from the basis descriptor
* rather than a caller buffer, and applies the SOURCE metadata (rsync copies
* then fixes attributes). A hard-link install that falls back to a byte copy
* also routes through here when the caller supplies the basis path. */
static bool file_copy_basis_stream_impl(const char* path, const char* basis_path,
unsigned long long expected_size, bool preallocate,
const FileMetadata* metadata, FileAttrPolicy policy,
bool update, bool no_replace, bool use_fsync,
const FileXattrList* xattrs, bool fake_super,
const char* temp_dir, unsigned* dirs_created,
const char* count_floor) {
if (!path || !basis_path)
return false;
char* leaf = NULL;
int dirfd = file_open_secure_parent_counted(path, &leaf, true, dirs_created, count_floor);
if (dirfd < 0)
return false;
char* basis_leaf = NULL;
int basis_dirfd = file_open_secure_parent(basis_path, &basis_leaf, false);
int src_fd = -1;
if (basis_dirfd >= 0 && basis_leaf != NULL) {
/* O_NONBLOCK rejects a raced-in FIFO without blocking; the S_ISREG gate
below is the real type check. */
src_fd = openat(basis_dirfd, basis_leaf, O_RDONLY | O_CLOEXEC | O_NOFOLLOW | O_NONBLOCK);
struct stat src_st;
if (src_fd >= 0 && (fstat(src_fd, &src_st) != 0 || !S_ISREG(src_st.st_mode))) {
close(src_fd);
src_fd = -1;
}
}
if (basis_dirfd >= 0)
close(basis_dirfd);
free(basis_leaf);
if (src_fd < 0) {
close(dirfd);
free(leaf);
return false;
}
struct stat destination_stat;
bool destination_is_regular = fstatat(dirfd, leaf, &destination_stat, AT_SYMLINK_NOFOLLOW) == 0 &&
S_ISREG(destination_stat.st_mode);
if (update && metadata && destination_is_regular && stat_is_newer(&destination_stat, metadata)) {
close(src_fd);
close(dirfd);
free(leaf);
return true;
}
if (no_replace && file_path_exists_secure(path)) {
close(src_fd);
close(dirfd);
free(leaf);
return true;
}
int scratch_dirfd = -1;
if (temp_dir) {
scratch_dirfd = file_open_temp_dir(temp_dir);
if (scratch_dirfd < 0) {
int saved_errno = errno;
log_message(LOG_LEVEL_ERROR,
"--temp-dir '%s' could not be opened (rsync requires it to already exist): %s",
temp_dir, strerror(saved_errno));
close(src_fd);
close(dirfd);
free(leaf);
return false;
}
}
int target_dirfd = scratch_dirfd >= 0 ? scratch_dirfd : dirfd;
int tmp_size = snprintf(NULL, 0, ".%s.tmp.%ld.%llu", leaf, (long)getpid(), ~0ULL);
char* tmp = NULL;
bool ok = false;
if (tmp_size >= 0)
tmp = malloc((size_t)tmp_size + 1);
if (tmp) {
for (unsigned int i = 0; i < 100 && !ok; ++i) {
if (scratch_dirfd >= 0)
snprintf(tmp, (size_t)tmp_size + 1, ".%s.tmp.%ld.%llu", leaf, (long)getpid(),
next_temp_sequence());
else
snprintf(tmp, (size_t)tmp_size + 1, ".%s.tmp.%ld.%u", leaf, (long)getpid(), i);
int fd =
openat(target_dirfd, tmp, O_WRONLY | O_CREAT | O_EXCL | O_CLOEXEC | O_NOFOLLOW, 0600);
if (fd < 0) {
if (errno != EEXIST)
break;
continue;
}
bool wrote = true;
if (preallocate && expected_size > 0 && preallocate_fd(fd, expected_size) != 0)
wrote = false;
if (wrote)
wrote = copy_fd_all(fd, src_fd, expected_size);
if (wrote && metadata) {
if (!policy.perms &&
fchmod(fd, file_mode_base(metadata, destination_is_regular,
destination_is_regular ? destination_stat.st_mode & 0777
: 0)) != 0)
wrote = false;
if (wrote)
wrote = file_restore_metadata_fd(fd, metadata, policy);
} else if (wrote && fchmod(fd, S_IRUSR | S_IWUSR | S_IRGRP | S_IROTH) != 0) {
wrote = false;
}
if (wrote)
restore_extra_fd(fd, metadata, xattrs, fake_super, policy, 0, 0);
if (wrote && use_fsync)
wrote = fsync(fd) == 0;
if (close(fd) != 0)
wrote = false;
if (wrote && renameat(target_dirfd, tmp, dirfd, leaf) != 0)
wrote = false;
if (!wrote)
unlinkat(target_dirfd, tmp, 0);
ok = wrote;
}
free(tmp);
}
if (!ok && scratch_dirfd >= 0) {
/* Retry once with no scratch dir (rsync's EXDEV fallback). */
close(scratch_dirfd);
close(src_fd);
close(dirfd);
free(leaf);
return file_copy_basis_stream_impl(path, basis_path, expected_size, preallocate, metadata,
policy, update, no_replace, use_fsync, xattrs, fake_super,
NULL, dirs_created, count_floor);
}
if (scratch_dirfd >= 0)
close(scratch_dirfd);
close(src_fd);
close(dirfd);
free(leaf);
return ok;
}
/* --copy-dest basis install (streaming). Applies the source metadata and the
per-file xattrs / --fake-super record. */
bool file_copy_basis_stream_attrs(const char* path, const char* basis_path,
unsigned long long expected_size, bool preallocate,
const FileMetadata* metadata, FileAttrPolicy policy, bool update,
bool use_fsync, const FileXattrList* xattrs, bool fake_super,
const char* temp_dir) {
return file_copy_basis_stream_impl(path, basis_path, expected_size, preallocate, metadata, policy,
update, false, use_fsync, xattrs, fake_super, temp_dir, NULL,
NULL);
}
static bool file_to_disk_secure_link_impl(const char* path, const char* basis_path,
const void* data, unsigned long long data_size,
bool preallocate, const FileMetadata* metadata,
FileAttrPolicy policy, bool use_fsync,
const FileXattrList* xattrs, bool fake_super,
const char* temp_dir) {
const char* temp_dir, unsigned* dirs_created,
const char* count_floor) {
if (!path || !basis_path)
return false;
/* The caller-supplied buffer is no longer used: the copy fallback streams
from the basis path (which may hold an over-limit file). Kept in the
signature for the existing API. */
(void)data;
char* leaf = NULL;
int dirfd = file_open_secure_parent(path, &leaf, true);
int dirfd = file_open_secure_parent_counted(path, &leaf, true, dirs_created, count_floor);
if (dirfd < 0)
return false;
@@ -1463,10 +1839,17 @@ static bool file_to_disk_secure_link_impl(const char* path, const char* basis_pa
close(dirfd);
free(leaf);
/* The basis file could not be linked in (missing, cross-device, refused
by the filesystem). Write a byte-identical local copy instead. */
return file_to_disk_secure_attrs(path, data, data_size, false, false, preallocate, metadata,
policy, false, false, use_fsync, xattrs, fake_super, false,
temp_dir);
by the filesystem). Stream a byte-identical local copy from the basis
itself (never the possibly-absent caller buffer) so an over-limit basis
still materializes. When the basis path is not a readable regular file
(e.g. a directory raced in), fall back to the caller-supplied bytes. */
if (file_copy_basis_stream_impl(path, basis_path, data_size, preallocate, metadata, policy,
false, false, use_fsync, xattrs, fake_super, temp_dir,
dirs_created, count_floor))
return true;
return file_to_disk_secure_attrs_counted(
path, data, data_size, false, false, preallocate, metadata, policy, false, false, use_fsync,
xattrs, fake_super, false, temp_dir, dirs_created, count_floor, 0, 0);
}
if (scratch_dirfd >= 0)
@@ -1481,7 +1864,7 @@ bool file_to_disk_secure_link(const char* path, const char* basis_path, const vo
const FileMetadata* metadata, FileAttrPolicy policy, bool use_fsync,
const char* temp_dir) {
return file_to_disk_secure_link_impl(path, basis_path, data, data_size, preallocate, metadata,
policy, use_fsync, NULL, false, temp_dir);
policy, use_fsync, NULL, false, temp_dir, NULL, NULL);
}
bool file_to_disk_secure_link_attrs(const char* path, const char* basis_path, const void* data,
@@ -1489,14 +1872,27 @@ bool file_to_disk_secure_link_attrs(const char* path, const char* basis_path, co
const FileMetadata* metadata, FileAttrPolicy policy,
bool use_fsync, const FileXattrList* xattrs, bool fake_super,
const char* temp_dir) {
return file_to_disk_secure_link_attrs_counted(path, basis_path, data, data_size, preallocate,
metadata, policy, use_fsync, xattrs, fake_super,
temp_dir, NULL, NULL);
}
bool file_to_disk_secure_link_attrs_counted(const char* path, const char* basis_path,
const void* data, unsigned long long data_size,
bool preallocate, const FileMetadata* metadata,
FileAttrPolicy policy, bool use_fsync,
const FileXattrList* xattrs, bool fake_super,
const char* temp_dir, unsigned* dirs_created,
const char* count_floor) {
return file_to_disk_secure_link_impl(path, basis_path, data, data_size, preallocate, metadata,
policy, use_fsync, xattrs, fake_super, temp_dir);
policy, use_fsync, xattrs, fake_super, temp_dir,
dirs_created, count_floor);
}
bool file_write_to_disk(const char* path, const void* data, unsigned long long data_size,
bool inplace, bool sparse) {
if (!path || (!data && data_size != 0) || has_path_traversal(path))
return false;
FileAttrPolicy policy = {false, false, false, false};
FileAttrPolicy policy = {0};
return file_to_disk_secure(path, data, data_size, inplace, sparse, false, NULL, policy, NULL);
}
+42 -2
View File
@@ -87,6 +87,11 @@ bool file_path_exists_secure(const char* path);
bool file_stat_secure(const char* path, struct stat* st);
bool file_destination_is_newer_secure(const char* path, const FileMetadata* metadata);
int file_open_secure_parent(const char* path, char** leaf_out, bool create_dirs);
/* Protocol 2.28.0 variant: also increments *dirs_created for every missing
* parent directory this walk creates that lies strictly below `count_floor`
* (a receive-root-relative path, or NULL to count all of them). */
int file_open_secure_parent_counted(const char* path, char** leaf_out, bool create_dirs,
unsigned* dirs_created, const char* count_floor);
bool file_ensure_directory_secure(const char* path);
bool file_directory_exists_secure(const char* path);
bool file_rename_secure(const char* old_path, const char* new_path);
@@ -98,8 +103,12 @@ bool file_remove_tree_secure(const char* path);
the authorized root. Used for the --delay-updates staging directory. */
int file_open_private_dir(const char* dir_path);
/* Open an existing --temp-dir scratch directory as-is (absolute or relative;
no creation, no root confinement), matching rsync's --temp-dir handling. */
/* Open an existing --temp-dir scratch directory (relative or absolute; no
creation). When an authorized receive root is configured the directory's
REAL path (symlinks resolved) must lie within it, so a client-planted
symlink cannot redirect receiver scratch files outside the sandbox; an
in-root symlink to another filesystem is still allowed for rsync's EXDEV
fallback. */
int file_open_temp_dir(const char* dir_path);
/* The file_to_disk_secure* variants write a temporary copy in the destination
@@ -158,5 +167,36 @@ bool file_to_disk_secure_link_attrs(const char* path, const char* basis_path, co
const FileMetadata* metadata, FileAttrPolicy policy,
bool use_fsync, const FileXattrList* xattrs, bool fake_super,
const char* temp_dir);
/* Streaming --copy-dest install: atomically materialize `path` by copying the
* bytes of `basis_path` through a bounded buffer (no whole-file buffering, so
* an arbitrarily large basis works), applying the SOURCE metadata and the
* per-file xattrs / --fake-super record. `update` honors a newer destination;
* a --temp-dir scratch location falls back to a direct write on EXDEV. */
bool file_copy_basis_stream_attrs(const char* path, const char* basis_path,
unsigned long long expected_size, bool preallocate,
const FileMetadata* metadata, FileAttrPolicy policy, bool update,
bool use_fsync, const FileXattrList* xattrs, bool fake_super,
const char* temp_dir);
/* Protocol 2.28.0 receiver-stat variants: like the two above but additionally
* report through `dirs_created` (when non-NULL) how many parent directories the
* confined secure walk had to create that lie strictly below `count_floor` (a
* receive-root-relative prefix, or NULL for all). Used to reproduce rsync's
* `Number of created files` directory count on a fresh destination. */
bool file_to_disk_secure_attrs_counted(
const char* path, const void* data, unsigned long long data_size, bool inplace, bool sparse,
bool preallocate, const FileMetadata* metadata, FileAttrPolicy policy, bool update,
bool no_replace, bool use_fsync, const FileXattrList* xattrs, bool fake_super,
bool keep_partial, const char* temp_dir, unsigned* dirs_created, const char* count_floor,
uint32_t fake_super_rdev_major, uint32_t fake_super_rdev_minor);
bool file_to_disk_secure_link_attrs_counted(const char* path, const char* basis_path,
const void* data, unsigned long long data_size,
bool preallocate, const FileMetadata* metadata,
FileAttrPolicy policy, bool use_fsync,
const FileXattrList* xattrs, bool fake_super,
const char* temp_dir, unsigned* dirs_created,
const char* count_floor);
/* The logical transfer root expressed receive-root-relative, or NULL when the
* wire paths carry no mirror scaffolding above it. Caller frees non-NULL. */
char* file_transfer_root_floor(const Config* config);
#endif
+6
View File
@@ -29,6 +29,12 @@ typedef struct FileAttrPolicy {
bool times; /* config->preserve_times: apply the source mtime */
bool atimes; /* config->preserve_atimes (-U): apply the source atime */
bool executability; /* config->use_executability (-E): exec-bits-only mode */
/* privilege_super_mode_permitted(): when false (SUPER_MODE_OFF / --no-super,
or a daemon that did not grant `client owner = yes`), the setuid/setgid/
sticky bits are stripped from every applied mode (source mode and any
--chmod result) even under --perms. When true, rsync's exact semantics are
preserved: -p copies the special bits and the kernel decides. */
bool super_permitted;
} FileAttrPolicy;
/* Build the per-attribute policy from a connection's Config. A NULL config
+15 -3066
View File
File diff suppressed because it is too large Load Diff
+10 -104
View File
@@ -2,10 +2,19 @@
#define FILE_RECEIVE_H
#include "config.h"
#include "delete_commit.h"
#include "file_save.h"
#include "file_types.h"
#include "incremental_check.h"
#include "utils.h"
#include <stdbool.h>
/* Server-side file receive/save path. */
/* Server-side file receive/save path.
*
* This header is the public facade for the file_receive module family: the
* wire receive dispatch (this file) plus the save-to-disk (file_save.h), the
* incremental check (incremental_check.h) and the delete-commit
* (delete_commit.h) modules. */
/* Cumulative caps for the deferred directory-time accumulator. The sender may
* legitimately split a large tree across repeated STATUS_DIR_TIMES frames, so a
@@ -22,15 +31,6 @@ File* file_receive_dir_time(int file_descriptor, const Config* config);
File* file_receive_hardlink(int file_descriptor);
File* file_receive_symlink(int file_descriptor, const Config* config);
File* file_receive_special(int file_descriptor);
bool file_special_rdev_valid(int32_t major, int32_t minor, mode_t mode);
File* receive_incremental_check(int fd, const Config* config, bool* skipped);
/* Extended variant used by the receiver. `would_transfer` (may be NULL) is set
* true only on the server-contacting --dry-run path when the file is not up to
* date: the receiver has already sent STATUS_DRY_RUN_TRANSFER and returns NULL
* without storing anything. On that path `*skipped` is true for an up-to-date
* (STATUS_OK) file and both flags are false for a genuine error. */
File* receive_incremental_check_ex(int fd, const Config* config, bool* skipped,
bool* would_transfer);
/* P7 Wave D directory-time accumulator. The receiver collects the metadata of
* every directory it creates/receives (STATUS_MKDIR with metadata and/or the
@@ -74,98 +74,4 @@ bool dir_time_list_add(DirTimeList* list, const char* wire_path, const FileMetad
void dir_metadata_list_apply(const DirTimeList* list, const char* root_directory,
const Config* config);
/* A received delete-manifest frame: the keep-set (`keeps`, destination-relative
paths the sender transferred/keeps) plus `protected`, destination-relative
prefixes the sender asks the receiver never to delete (paths excluded on the
source, protected at any depth). When --delete-excluded is given the sender
transmits an empty protected list so excluded destination mirrors are treated
as ordinary extras. With --delete-missing-args a third section (`missing`)
carries the destination mirrors of explicitly-listed source entries that do
not exist: each is an exact deletion request, independent of the ordinary
extras walk (never blocked by the protected prefixes) and processed when the
manifest is committed. */
typedef struct DeleteManifest {
ArrayList* keeps;
ArrayList* protected;
ArrayList* missing;
/* Destination-relative paths of the directories the sender synchronized for
this run. The extras walker only removes entries directly inside one of
these (the receive root is the "." sentinel); `--files-from` runs therefore
leave untransmitted directories and the unlisted parts of listed ones
alone, matching rsync's "delete only in synchronized directories". */
ArrayList* dirs;
} DeleteManifest;
void delete_manifest_free(DeleteManifest* manifest);
/* Read a delete-manifest frame (protocol 2.23.0): keep count + keeps, then
protected count + protected prefixes, then missing count + missing paths,
then synchronized-directory count + directory paths (self-delimiting; the
leading STATUS_MANIFEST code has been consumed). Returns an owned
DeleteManifest, or NULL after signalling STATUS_ERROR on a malformed frame. */
DeleteManifest* receive_manifest_entries(int fd);
/* Remove destination entries under config->receive_root_directory that are not
in `manifest` (bounded, all-or-nothing walk; staging-dir, basis-dir and
protected-prefix skips). `--max-delete` and `--force` are honored here. The
caller decides WHEN to run it based on the negotiated delete timing. Returns
false (and the transfer fails) when the deletion cannot be committed. */
bool manifest_delete_extras(const Config* config, DeleteManifest* manifest);
/* --delete-missing-args exact-path deletions: remove each destination mirror
in `manifest->missing` (never blocked by the protected prefixes, staging dir
and basis dirs excluded). A regular file/symlink is unlinked; an empty
directory is removed; a NON-empty directory is removed recursively only when
--delete or --force is in effect, otherwise it is left with a warning (rsync
parity). A missing path is a no-op. Returns false only on a genuine
confinement or I/O error (the run then fails); tolerated per-path cases are
reported and skipped. */
bool manifest_delete_missing_args(const Config* config, DeleteManifest* manifest);
/* Budgeted form of manifest_delete_missing_args for the per-directory delete
session: each removed mirror draws from `max_delete` (SIZE_MAX = unlimited)
and the tallies are accumulated into `*deleted`/`*skipped`. `*limit_hit` is set
when the budget stopped the pass with entries left over. Returns false only
on a genuine deletion error. */
bool manifest_delete_missing_args_limited(const Config* config, DeleteManifest* manifest,
size_t max_delete, size_t* deleted, size_t* skipped,
bool* limit_hit);
/* Outcome of committing a delete manifest. LIMIT_REACHED reports rsync's
partial --max-delete result: the budget allowed some deletions and the rest
were skipped (the run still stores all file data but the client exits 25). */
typedef enum {
DELETE_COMMIT_OK = 0,
DELETE_COMMIT_LIMIT_REACHED,
DELETE_COMMIT_ERROR
} DeleteCommitResult;
/* Run every deletion family the manifest carries: the --delete-missing-args
exact-path deletions first (user requests are not blocked by exclusion
protection), then the ordinary extras walk when --delete is active. Both
share one --max-delete budget. Returns DELETE_COMMIT_OK when nothing was to
do or everything committed, DELETE_COMMIT_LIMIT_REACHED when the budget
stopped part of the work, or DELETE_COMMIT_ERROR on a genuine failure. */
DeleteCommitResult manifest_delete_all(const Config* config, DeleteManifest* manifest);
/* Like manifest_delete_all, but reports how many destination entries the commit
removed (for the end-of-transfer wire stats). `deleted` may be NULL. */
DeleteCommitResult manifest_delete_all_counted(const Config* config, DeleteManifest* manifest,
size_t* deleted);
/* -n/--dry-run --delete would-delete reporting: walk the destination exactly as
the delete pass would and append (strdup'd) destination-relative paths that
WOULD be removed to `out`, without touching disk. Uses the same staging-dir,
basis-dir and protected-prefix skips as the real commit. Returns true on a
clean walk; `*count_out` receives the number of paths appended. */
bool manifest_would_delete_list(const Config* config, DeleteManifest* manifest, ArrayList* out,
size_t* count_out);
/* Convert one basis-directory path to the receive-root-relative protection
prefix the delete walker uses (NULL when it lies outside the root). Exposed
for unit tests of the root-of-"/" and normalization edge cases. */
char* file_receive_basis_delete_relative(const Config* config, const char* path);
/* Outcome of a single file_save_to_disk operation. The receiver needs to
distinguish "written" from "skipped" so --remove-source-files can be told
which sources were actually stored. */
typedef enum { FILE_SAVE_ERROR = 0, FILE_SAVE_WRITTEN = 1, FILE_SAVE_SKIPPED = 2 } FileSaveResult;
FileSaveResult file_save_to_disk_full(const char* root_directory, const File* file,
const Config* config);
bool file_save_to_disk(const char* root_directory, const File* file, const Config* config);
#endif
File diff suppressed because it is too large Load Diff
+48
View File
@@ -0,0 +1,48 @@
#ifndef FILE_SAVE_H
#define FILE_SAVE_H
#include "config.h"
#include "file_types.h"
#include "format.h"
#include <stdbool.h>
/* Save-to-disk module: regular-file/symlink/hardlink/special install, xattr
* application, --fake-super and the --delay-updates staging path. These
* declarations are re-exported by the file_receive.h facade. */
/* Outcome of a single file_save_to_disk operation. The receiver needs to
distinguish "written" from "skipped" so --remove-source-files can be told
which sources were actually stored. FILE_SAVE_FAILED is a per-entry failure
(for example a device node that mknodat() refused with EPERM/EACCES): it is
logged and counted by the receiver but does NOT abort the transfer, matching
rsync's continue-and-exit-partial behavior. */
typedef enum {
FILE_SAVE_ERROR = 0,
FILE_SAVE_WRITTEN = 1,
FILE_SAVE_SKIPPED = 2,
FILE_SAVE_FAILED = 3
} FileSaveResult;
bool file_special_rdev_valid(int32_t major, int32_t minor, mode_t mode);
FileSaveResult file_save_to_disk_full(const char* root_directory, const File* file,
const Config* config);
/* Protocol 2.28.0 variant: also reports through `created` (when non-NULL)
* whether the destination entry did not exist before this save, and through
* `created_dirs` how many parent directories the confined walk created, so the
* receiver can build rsync's `Number of created files` breakdown. The plain
* file_save_to_disk_full() is this with both out-params NULL. */
FileSaveResult file_save_to_disk_full_ex(const char* root_directory, const File* file,
const Config* config, bool* created,
unsigned* created_dirs);
bool file_save_to_disk(const char* root_directory, const File* file, const Config* config);
/* Protocol 2.28.0 receiver counter accumulator: fold one successfully saved
* entry into `stats`, adding its receiver-observed literal bytes and, when
* `created`, the matching created-by-type counter (regular file / symlink /
* special) plus `created_dirs` implicitly-created parent directories.
* Non-first hardlink siblings contribute no literal bytes. */
void receiver_stats_note_saved(ReceiverStats* stats, const File* file, bool created,
unsigned created_dirs);
#endif
+9 -2
View File
@@ -124,8 +124,12 @@ bool file_send_sendfile_with_skip(File* file, int file_descriptor, bool use_meta
}
/* sendfile cannot encrypt TLS records. Keep the framing identical but
route encrypted transfers through the deadline-aware IO layer. */
if (io_get_ssl() != NULL) {
route encrypted transfers through the deadline-aware IO layer. Resolve
the transport from the bound session, not the thread-local io_ssl: a
worker thread running a TLS transfer has its SSL only on the session it
bound, so io_get_ssl() would be NULL there and the raw sendfile() path
would be taken on an encrypted socket. */
if (protocol_current_ssl() != NULL) {
unsigned char buffer[64 * 1024];
unsigned long long remaining = file_size;
bool ok = true;
@@ -166,6 +170,8 @@ bool file_send_sendfile_with_skip(File* file, int file_descriptor, bool use_meta
}
struct pollfd pfd = {.fd = file_descriptor, .events = POLLOUT};
int polled = poll(&pfd, 1, timeout);
if (polled < 0 && errno == EINTR)
continue;
if (polled <= 0 || (pfd.revents & (POLLERR | POLLHUP | POLLNVAL))) {
close(fd);
return false;
@@ -183,6 +189,7 @@ bool file_send_sendfile_with_skip(File* file, int file_descriptor, bool use_meta
return false;
}
protocol_note_bytes_written((unsigned long long)sent);
protocol_throttle_bytes(file_descriptor, (size_t)sent);
}
close(fd);
+10
View File
@@ -57,6 +57,11 @@ typedef struct {
* equals the incoming file, and `data` is kept as the cross-filesystem
* fallback (a local copy) if the hard link cannot be created. */
char* basis_link;
/* Receiver-only, --copy-dest: when set (and basis_link is NULL), stream the
* basis file's bytes into the destination instead of `data`/`data->size`.
* This lets a basis larger than any whole-file bound materialize without
* buffering it; the source metadata on `metadata` is applied afterwards. */
char* basis_copy;
/* --hard-links (-H), sender + receiver wire state. link_group is a run-local
* id shared by every member of one source inode (0 = not part of a group).
* The FIRST member (link_first == true) carries its data on the wire and is
@@ -97,6 +102,11 @@ typedef struct {
* basis file (matched delta blocks) for this entry. 0 when the file was sent
* whole. Accumulated into ReceiverStats.matched_data by the receiver sink. */
unsigned long long matched_bytes;
/* Receiver-only (protocol 2.28.0) wire-stats tally: the literal delta fragment
* bytes this entry carried (DELTA_INSTR_LITERAL). 0 when the file was sent
* whole; the sink then falls back to the whole payload size. Accumulated
* into ReceiverStats.literal_bytes. */
unsigned long long literal_bytes;
} File;
/* The path that should be sent on the wire and used for the receiver-side
+95 -27
View File
@@ -4,22 +4,12 @@
#include <ctype.h>
#include <errno.h>
#include <limits.h>
#include <stdarg.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
/* Write a diagnostic message into the caller's optional buffer. A NULL `err`
* (or a zero size) is a no-op, so a caller that only needs the boolean status
* may pass NULL without the snprintf-on-NULL undefined behaviour. */
static void filter_set_error(char* err, size_t err_size, const char* fmt, ...) {
if (!err || err_size == 0)
return;
va_list ap;
va_start(ap, fmt);
vsnprintf(err, err_size, fmt, ap);
va_end(ap);
}
/* Write a diagnostic message into the caller's optional buffer. */
#define filter_set_error utils_set_error
/* ---- Ordered rule lists ---- */
@@ -172,14 +162,74 @@ static bool is_modifier_char(char c) {
return c == 's' || c == 'r' || c == 'p' || c == 'x' || c == '/' || c == '!' || c == 'C';
}
/* merge/dir-merge rules are the only rules rsync accepts the merge-file
* modifiers on. */
static bool is_merge_rule(RuleKind kind) {
return kind == RULE_KIND_MERGE || kind == RULE_KIND_DIR_MERGE;
}
/* Merge-file modifiers rsync defines but FastSync does not implement:
* 'e' exclude the merge file itself, 'n' do not inherit the merge file, 'w'
* word-split the merge file. They are recognized as part of a modifier run on
* every rule (so a pure e/n/w token is rejected rather than folded into the
* pattern), but are accepted (and ignored) only on merge/dir-merge rules. */
static bool is_unsupported_modifier_char(char c) {
return c == 'e' || c == 'n' || c == 'w';
}
/* Merge-file modifiers rsync accepts on merge/dir-merge rules: 'e', 'n', 'w'
* and '-' (do not transfer the merge file). */
static bool is_merge_modifier_char(char c) {
return c == 'e' || c == 'n' || c == 'w' || c == '-';
}
/* Characters that count as part of a modifier run for `kind` when deciding
* whether a token is a pure modifier run. e/n/w count on every rule so that a
* pure e/n/w token is rejected on non-merge rules; '-' only on merge rules. */
static bool is_modifier_scan_char(char c, RuleKind kind) {
return is_modifier_char(c) || is_unsupported_modifier_char(c) ||
(is_merge_rule(kind) && is_merge_modifier_char(c));
}
/* Characters actually consumed as modifiers for `kind`. The merge-file
* modifiers are consumed only on merge/dir-merge rules; elsewhere e/n/w fall
* through to the pattern (so mixed tokens such as "H,!secret" keep their
* historical "ecret" pattern). */
static bool is_consumed_modifier_char(char c, RuleKind kind) {
return is_modifier_char(c) || (is_merge_rule(kind) && is_merge_modifier_char(c));
}
/* Inspect the token that follows a rule name (up to the first space/underscore
* or the end). If the token is composed *solely* of modifier characters and
* includes one that is invalid for `kind`, it is unambiguously a modifier run:
* return that character so the caller can reject it. A token that contains any
* non-modifier character is a pattern (e.g. "-newfile") and returns '\0', which
* keeps the historical parsing of mixed tokens such as "H,!secret" intact. */
static char unsupported_modifier_in_token(const char* tok, RuleKind kind) {
if (*tok == '\0' || *tok == ' ' || *tok == '_')
return '\0';
char bad = '\0';
for (const char* q = tok; *q != '\0' && *q != ' ' && *q != '_'; q++) {
if (!is_modifier_scan_char(*q, kind))
return '\0';
if (!is_merge_rule(kind) && is_unsupported_modifier_char(*q))
bad = *q;
}
return bad;
}
/* Parse "RULE[,MODIFIERS] [PATTERN]". On success `kind`, `sides`,
* `sides_explicit`, `negate`, `anchored_mod`, `perishable`, `xattr`,
* `cvs_inject` and the pattern span (`pat_start`/`pat_len`, possibly 0 for
* merge/clear) are filled. Returns true on success. */
* merge/clear) are filled. Returns true on success.
*
* On failure `*bad_mod` is set to the offending modifier character when the
* rule carried a modifier FastSync does not implement, and left '\0' for a
* generic syntax error so callers can emit a precise diagnostic. */
static bool parse_rule_syntax(const char* text, RuleKind* kind, unsigned* sides,
bool* sides_explicit, bool* negate, bool* anchored_mod,
bool* perishable, bool* xattr, bool* cvs_inject,
const char** pat_start, size_t* pat_len) {
const char** pat_start, size_t* pat_len, char* bad_mod) {
const char* p = text;
*sides = FILTER_SIDE_SENDER | FILTER_SIDE_RECEIVER;
*sides_explicit = false;
@@ -190,6 +240,7 @@ static bool parse_rule_syntax(const char* text, RuleKind* kind, unsigned* sides,
*cvs_inject = false;
*pat_start = NULL;
*pat_len = 0;
*bad_mod = '\0';
bool is_short = false;
if (short_rule_char(*p, kind)) {
@@ -210,17 +261,25 @@ static bool parse_rule_syntax(const char* text, RuleKind* kind, unsigned* sides,
/* Modifiers: long names require a comma; short names may attach directly.
Only commit a modifier run that terminates at a separator or the end, so a
pattern such as "*.tmp" written as "-*.tmp" is not mistaken for modifiers. */
if (*p == ',') {
*bad_mod = unsupported_modifier_in_token(p + 1, *kind);
} else if (is_short) {
*bad_mod = unsupported_modifier_in_token(p, *kind);
}
if (*bad_mod != '\0')
return false;
const char* mod_start = p;
const char* mod_end = p;
if (*p == ',') {
p++;
mod_start = p;
while (is_modifier_char(*p))
while (is_consumed_modifier_char(*p, *kind))
p++;
mod_end = p;
} else if (is_short) {
const char* scan = p;
while (is_modifier_char(*scan))
while (is_consumed_modifier_char(*scan, *kind))
scan++;
if (*scan == '\0' || *scan == ' ' || *scan == '_') {
mod_start = p;
@@ -290,9 +349,13 @@ FilterRule* filter_rule_parse(const char* line, const FilterParseOptions* opts,
bool sides_explicit, negate, anchored_mod, perishable, xattr, cvs_inject;
const char* pat;
size_t pat_len;
char bad_mod;
if (!parse_rule_syntax(p, &kind, &sides, &sides_explicit, &negate, &anchored_mod, &perishable,
&xattr, &cvs_inject, &pat, &pat_len)) {
filter_set_error(err, err_size, "unrecognized filter rule syntax");
&xattr, &cvs_inject, &pat, &pat_len, &bad_mod)) {
if (bad_mod != '\0')
filter_set_error(err, err_size, "unsupported filter modifier '%c'", bad_mod);
else
filter_set_error(err, err_size, "unrecognized filter rule syntax");
return NULL;
}
if (cvs_inject) {
@@ -401,7 +464,6 @@ FilterRule* filter_rule_parse(const char* line, const FilterParseOptions* opts,
rule->dir_only = dir_only;
rule->negate = negate;
rule->perishable = perishable;
(void)xattr; /* xattr-name rules never match file/dir names; accepted/ignored */
return rule;
}
@@ -530,16 +592,27 @@ static bool filter_list_parse_append_depth(FilterRuleList* list, const char* lin
bool sides_explicit, negate, anchored_mod, perishable, xattr, cvs_inject;
const char* pat;
size_t pat_len;
char bad_mod;
if (!parse_rule_syntax(p, &kind, &sides, &sides_explicit, &negate, &anchored_mod, &perishable,
&xattr, &cvs_inject, &pat, &pat_len)) {
filter_set_error(err, err_size, "unrecognized filter rule syntax: %s", p);
&xattr, &cvs_inject, &pat, &pat_len, &bad_mod)) {
if (bad_mod != '\0')
filter_set_error(err, err_size, "unsupported filter modifier '%c': %s", bad_mod, p);
else
filter_set_error(err, err_size, "unrecognized filter rule syntax: %s", p);
return false;
}
(void)sides_explicit;
(void)negate;
(void)anchored_mod;
(void)perishable;
(void)xattr;
/* xattr-name rules are not implemented; reject them everywhere (including on
* merge/dir-merge, where the flag would otherwise be silently dropped) with
* the same diagnostic the standalone parser gives. */
if (xattr) {
filter_set_error(err, err_size, "xattr-name filter rules (the x modifier) are not supported");
return false;
}
if (cvs_inject) {
/* "C" injects the CVS defaults in place; no pattern is expected. */
@@ -811,8 +884,3 @@ FilterAction filter_rules_apply_side(const FilterRuleList* list, const char* rel
}
return FILTER_ACTION_NONE;
}
FilterAction filter_rules_apply(const FilterRuleList* list, const char* rel_path, const char* leaf,
bool is_dir) {
return filter_rules_apply_side(list, rel_path, leaf, is_dir, FILTER_SIDE_SENDER);
}
+8 -7
View File
@@ -23,7 +23,13 @@
* dir-merge/: per-directory merge file (registered for the scanner)
* clear/! clear the current rule list (takes no argument)
* Modifiers: '/' absolute anchor, '!' negate match, 'C' inject CVS defaults,
* 's' sender side, 'r' receiver side, 'p' perishable, 'x' xattr name rule.
* 's' sender side, 'r' receiver side, 'p' perishable. The rsync 'x'
* (xattr-name) modifier is not implemented and is rejected explicitly
* everywhere. The merge-file modifiers 'e' (exclude the merge file itself),
* 'n' (do not inherit the merge file), 'w' (word-split the merge file) and '-'
* (do not transfer the merge file) are accepted and consumed only on merge/
* dir-merge rules (rejected on every other rule, matching rsync); their
* semantics are not implemented and they are otherwise ignored.
* A trailing '/' makes a pattern match directories only. A leading '/' anchors
* the pattern to its owner directory.
*/
@@ -53,7 +59,7 @@ typedef struct {
char* pattern; /* cleaned glob pattern (no leading '/', no trailing '/') */
} FilterRule;
typedef struct {
typedef struct FilterRuleList {
FilterRule** items; /* owned array of rule pointers */
int count;
int capacity;
@@ -129,9 +135,4 @@ FilterRuleList* filter_file_read(const char* dir_path, const char* owner_rel, bo
FilterAction filter_rules_apply_side(const FilterRuleList* list, const char* rel_path,
const char* leaf, bool is_dir, unsigned side);
/* Sender-side convenience wrapper (kept for callers/tests that only need the
* transfer decision). */
FilterAction filter_rules_apply(const FilterRuleList* list, const char* rel_path, const char* leaf,
bool is_dir);
#endif
+15 -13
View File
@@ -105,25 +105,27 @@ bool format_dest_state_receive(int fd, OutputDestState* state) {
bool format_stats_send(int fd, const ReceiverStats* stats) {
if (!stats)
return false;
unsigned long long matched = stats->matched_data;
unsigned long long deleted = stats->deleted_files;
unsigned long long would = stats->would_delete_count;
return send_n_data(fd, &matched, sizeof(matched)) && send_n_data(fd, &deleted, sizeof(deleted)) &&
send_n_data(fd, &would, sizeof(would));
unsigned long long fields[8] = {
stats->matched_data, stats->deleted_files, stats->would_delete_count, stats->literal_bytes,
stats->created_reg, stats->created_dir, stats->created_link, stats->created_special,
};
return send_n_data(fd, fields, sizeof(fields));
}
bool format_stats_receive(int fd, ReceiverStats* stats) {
if (!stats)
return false;
unsigned long long matched = 0;
unsigned long long deleted = 0;
unsigned long long would = 0;
if (!receive_n_data(fd, &matched, sizeof(matched)) ||
!receive_n_data(fd, &deleted, sizeof(deleted)) || !receive_n_data(fd, &would, sizeof(would)))
unsigned long long fields[8] = {0};
if (!receive_n_data(fd, fields, sizeof(fields)))
return false;
memset(stats, 0, sizeof(*stats));
stats->matched_data = matched;
stats->deleted_files = deleted;
stats->would_delete_count = would;
stats->matched_data = fields[0];
stats->deleted_files = fields[1];
stats->would_delete_count = fields[2];
stats->literal_bytes = fields[3];
stats->created_reg = fields[4];
stats->created_dir = fields[5];
stats->created_link = fields[6];
stats->created_special = fields[7];
return true;
}
+39 -4
View File
@@ -57,14 +57,26 @@ bool format_dest_state_send(int fd, const OutputDestState* state);
bool format_dest_state_receive(int fd, OutputDestState* state);
/* End-of-transfer receiver counters reported through STATUS_STATS (protocol
* 2.25.0) when the wire config carries report_stats. `would_delete_count` is
* the number of destination-relative paths the receiver would have deleted in a
* -n/--dry-run --delete run; that many wire strings immediately follow the
* fixed record (sent/read by the caller). */
* 2.25.0, extended in 2.28.0) when the wire config carries report_stats.
* `would_delete_count` is the number of destination-relative paths the receiver
* would have deleted in a -n/--dry-run --delete run; that many wire strings
* immediately follow the fixed record (sent/read by the caller).
*
* Protocol 2.28.0 adds the receiver-observed counters the sender cannot see:
* `literal_bytes` is the file data the receiver actually stored literally
* (whole files plus the literal fragments of a delta) and the four `created_*`
* counters split the destination entries the receiver newly created by type,
* reproducing rsync's `Number of created files` breakdown and an exact
* `Literal data` for a delta run. */
typedef struct {
unsigned long long matched_data;
unsigned long long deleted_files;
unsigned long long would_delete_count;
unsigned long long literal_bytes;
unsigned long long created_reg;
unsigned long long created_dir;
unsigned long long created_link;
unsigned long long created_special;
} ReceiverStats;
/* Fixed-width STATUS_STATS counter record. The status frame and the optional
@@ -73,4 +85,27 @@ typedef struct {
bool format_stats_send(int fd, const ReceiverStats* stats);
bool format_stats_receive(int fd, ReceiverStats* stats);
/* Sender-side file-list accounting for rsync's `--stats` block. Filled while
* the scan/send loops walk each entry: the flist counters describe every
* scanned source entry (transferred or skipped), while the transferred/literal
* counters describe only the regular files the receiver actually stored. The
* type split lets the client print rsync's `Number of files` breakdown; the
* receiver-only counters (matched data, deleted, created) come from
* STATUS_STATS. */
typedef struct {
unsigned long long flist_reg;
unsigned long long flist_dir;
unsigned long long flist_link;
unsigned long long flist_special;
unsigned long long total_file_size; /* sum of entry sizes (link target len) */
unsigned long long transferred_regular; /* regular files actually stored */
unsigned long long transferred_file_size; /* source size of those files */
/* Whole-file accuracy: the `--stats` "Literal data" row. The sender counts
* the source size of every stored file, so a whole-file transfer matches
* rsync. A delta run actually ships only the literal fragments of the diff
* (the rest is matched/copied), so here the value is an upper bound, not
* rsync's literal-byte total; see RSYNC_COMPAT.md's `--stats` row. */
unsigned long long literal_data;
} TransferStats;
#endif
+163 -11
View File
@@ -3,6 +3,7 @@
#include "utils.h"
#include <errno.h>
#include <fcntl.h>
#include <fnmatch.h>
#include <grp.h>
#include <limits.h>
#include <pwd.h>
@@ -12,6 +13,8 @@
#include <sys/stat.h>
#include <unistd.h>
static bool identity_id_fits_int32(unsigned long id);
/* The active identity snapshot lives in a per-process global. The TCP server
* forks one child process per connection, so a connection never shares this
* with another; within a connection the multithreaded receiver reads it without
@@ -419,17 +422,10 @@ static int identity_parse_from(const char* token, bool is_group, int32_t* out_fr
/* Not a numeric LOW-HIGH range: fall through and treat as a name (a
* hyphenated account name like "wayne-smith" must still resolve). */
}
/* A sender-side name. A wildcard other than the bare '*' is matched by rsync
* against the sender's names; because FastSync transmits numeric ids only, the
* receiver cannot evaluate it, so reject rather than silently mis-match. */
if (identity_token_has_glob(token)) {
log_message(LOG_LEVEL_ERROR,
"%smap FROM '%s': name wildcards other than '*' are not supported "
"(FastSync transmits numeric ids, so sender names are unavailable on the "
"receiver)",
is_group ? "--group" : "--user", token);
return -1;
}
/* A sender-side name. A FROM name wildcard other than the bare '*' is handled
* by identity_expand_from_glob() in the caller (it expands against the
* sender's account database at CLI-parse time), so this function only sees the
* bare '*' or a literal name here. */
int32_t id;
if (identity_resolve_token(token, is_group, &id) != 0)
return -1;
@@ -486,6 +482,138 @@ static int identity_append_rule(IdentityMap** map, int* count, const IdentityMap
return 0;
}
/* True when `lo` and `hi` are adjacent ids (no overflow at INT32_MAX). */
static bool identity_ids_adjacent(int32_t lo, int32_t hi) {
return lo < INT32_MAX && hi == lo + 1;
}
static int identity_id_cmp(const void* a, const void* b) {
int32_t x = *(const int32_t*)a;
int32_t y = *(const int32_t*)b;
return (x > y) - (x < y);
}
static bool identity_ids_push(int32_t** ids, size_t* count, size_t* cap, int32_t id) {
if (*count == *cap) {
size_t grown_cap = *cap ? *cap * 2 : 16;
int32_t* grown = realloc(*ids, grown_cap * sizeof(int32_t));
if (!grown)
return false;
*ids = grown;
*cap = grown_cap;
}
(*ids)[(*count)++] = id;
return true;
}
/* Expand a FROM name wildcard (rsync's match against sender-side account names)
* into one rule per contiguous run of matching numeric ids, all sharing the same
* TO side. FastSync transmits numeric ids only, so the wildcard must be
* resolved here -- at CLI-parse time -- against the SENDER's passwd/group
* database; the receiver has no sender names to match. Contiguous matched ids
* are collapsed into a single LOW-HIGH range (a range of adjacent ids contains
* exactly the ids it spans, so this is semantically exact). Returns 0 on
* success, -1 on an allocation failure, a wildcard that matches no sender
* account, or an expansion that would push the map past MAX_IDENTITY_MAP. */
static int identity_expand_from_glob(Config* config, const char* glob, bool is_group,
const IdentityMap* to_rule) {
const char* optname = is_group ? "--groupmap" : "--usermap";
size_t cap = 0;
size_t n = 0;
int32_t* ids = NULL;
bool alloc_failed = false;
if (is_group) {
setgrent();
struct group* gr;
while ((gr = getgrent()) != NULL) {
if (fnmatch(glob, gr->gr_name, 0) != 0)
continue;
if (!identity_id_fits_int32((unsigned long)gr->gr_gid))
continue;
if (!identity_ids_push(&ids, &n, &cap, (int32_t)gr->gr_gid)) {
alloc_failed = true;
break;
}
}
endgrent();
} else {
setpwent();
struct passwd* pw;
while ((pw = getpwent()) != NULL) {
if (fnmatch(glob, pw->pw_name, 0) != 0)
continue;
if (!identity_id_fits_int32((unsigned long)pw->pw_uid))
continue;
if (!identity_ids_push(&ids, &n, &cap, (int32_t)pw->pw_uid)) {
alloc_failed = true;
break;
}
}
endpwent();
}
if (alloc_failed) {
free(ids);
log_message(LOG_LEVEL_ERROR, "%s: memory allocation failed expanding FROM '%s'", optname, glob);
return -1;
}
if (n == 0) {
free(ids);
log_message(LOG_LEVEL_ERROR, "%s FROM '%s': no source account name matches the wildcard",
optname, glob);
return -1;
}
qsort(ids, n, sizeof(int32_t), identity_id_cmp);
size_t unique = 0;
for (size_t i = 0; i < n; i++) {
if (unique == 0 || ids[unique - 1] != ids[i])
ids[unique++] = ids[i];
}
n = unique;
int runs = 0;
for (size_t i = 0; i < n; i++) {
if (i == 0 || !identity_ids_adjacent(ids[i - 1], ids[i]))
runs++;
}
IdentityMap** map = is_group ? &config->groupmap : &config->usermap;
int* count = is_group ? &config->groupmap_count : &config->usermap_count;
if (*count > MAX_IDENTITY_MAP - runs) {
log_message(LOG_LEVEL_ERROR,
"%s FROM '%s': the name wildcard expands to %d rule(s), which would exceed "
"the maximum of %d map rules",
optname, glob, runs, MAX_IDENTITY_MAP);
free(ids);
return -1;
}
for (size_t i = 0; i < n;) {
size_t j = i;
while (j + 1 < n && identity_ids_adjacent(ids[j], ids[j + 1]))
j++;
IdentityMap rule;
rule.from = ids[i];
rule.from_hi = ids[j];
rule.to = to_rule->to;
rule.to_name = to_rule->to_name ? str_dup(to_rule->to_name) : NULL;
if (to_rule->to_name && !rule.to_name) {
free(ids);
return -1;
}
if (identity_append_rule(map, count, &rule) != 0) {
free(rule.to_name);
free(ids);
return -1;
}
i = j + 1;
}
free(ids);
return 0;
}
int identity_parse_map(Config* config, const char* value, bool is_group) {
if (!config || !value || *value == '\0') {
log_message(LOG_LEVEL_ERROR, "%smap requires a value", is_group ? "--group" : "--user");
@@ -509,6 +637,30 @@ int identity_parse_map(Config* config, const char* value, bool is_group) {
char* to_token = colon + 1;
IdentityMap parsed;
memset(&parsed, 0, sizeof(parsed));
/* A FROM name wildcard (anything with a glob metacharacter other than the
* bare '*') is expanded against the sender's account database here, while
* the sender's passwd/group DB is still available; the resulting numeric
* rules travel on the wire like an explicit list. The TO side is parsed
* first so every expanded rule shares it. */
if (strcmp(from_token, "*") != 0 && identity_token_has_glob(from_token)) {
if (identity_parse_to(to_token, is_group, &parsed.to, &parsed.to_name) != 0) {
log_message(LOG_LEVEL_ERROR, "%s could not parse TO '%s' in '%s'", optname, to_token,
value);
free(list);
return -1;
}
if (identity_expand_from_glob(config, from_token, is_group, &parsed) != 0) {
free(parsed.to_name);
free(list);
return -1;
}
/* Every rule emitted by the expansion took its own str_dup of the name,
* so the parse-time copy is unreachable on success: release it here (the
* failure path above already does). `parsed.to_name` is NULL for a
* numeric TO. */
free(parsed.to_name);
continue;
}
if (identity_parse_from(from_token, is_group, &parsed.from, &parsed.from_hi) != 0) {
log_message(LOG_LEVEL_ERROR,
"%s could not resolve FROM '%s' in '%s' (a name must exist on the "
+8 -2
View File
@@ -25,8 +25,14 @@
/* Parse one --usermap= / --groupmap= value (comma-separated FROM:TO rules,
* first match wins) into config->usermap / config->groupmap. is_group selects
* the group tables and name databases. Returns 0 on success, -1 on a
* malformed spec or an unresolvable name (never a silent no-op). */
* the group tables and name databases. A FROM name wildcard (containing `*`,
* `?` or `[...]`, but not the bare `*`) is expanded against the SENDER's
* account database at parse time into one or more numeric id/range rules
* (contiguous ids collapse to a range) sharing the same TO, because only
* numeric ids cross the wire; the expansion is capped at MAX_IDENTITY_MAP and a
* wildcard matching no account is an error. Returns 0 on success, -1 on a
* malformed spec, an unresolvable name, an unmatched wildcard, or a map that
* would exceed MAX_IDENTITY_MAP (never a silent no-op). */
int identity_parse_map(Config* config, const char* value, bool is_group);
/* Parse --chown=USER:GROUP. Supports USER:GROUP, USER (owner only), :GROUP
File diff suppressed because it is too large Load Diff
+41
View File
@@ -0,0 +1,41 @@
#ifndef INCREMENTAL_CHECK_H
#define INCREMENTAL_CHECK_H
#include "config.h"
#include "file_types.h"
#include "protocol.h"
#include <stdbool.h>
/* Incremental-check module: the per-file STATUS_CHECK state machine, the
* incremental delta / alternate-basis / fuzzy matching helpers and the shared
* xattr receive helper. These declarations are re-exported by the
* file_receive.h facade. */
/* Whole-file payload bound shared by the plain receive path and the
* incremental check paths. */
#define MAX_FILE_DATA_SIZE MAX_RECEIVE_WHOLE_FILE_SIZE
/* Receive a file's xattr block (when the config enables xattr transport) and
* attach it to `file`. Returns false on a malformed/oversized frame. */
bool receive_file_xattrs(File* file, int fd, const Config* config);
File* receive_incremental_check(int fd, const Config* config, bool* skipped);
/* Extended variant used by the receiver. `would_transfer` (may be NULL) is set
* true only on the server-contacting --dry-run path when the file is not up to
* date: the receiver has already sent STATUS_DRY_RUN_TRANSFER and returns NULL
* without storing anything. On that path `*skipped` is true for an up-to-date
* (STATUS_OK) file and both flags are false for a genuine error. */
File* receive_incremental_check_ex(int fd, const Config* config, bool* skipped,
bool* would_transfer);
/* Testable basis quick-check / verification policy. file_basis_quick_match is
* rsync's metadata quick-check for a basis candidate (equal size is required
* separately by the caller; this adds the --size-only / mtime / --modify-window
* leg). file_basis_content_required reports whether a hit must ALSO be
* confirmed by a whole-file content digest (--verify-basis; false is the
* default rsync-parity behavior). */
bool file_basis_quick_match(const Config* config, const struct stat* st, time_t check_mtime,
long check_mtime_nsec);
bool file_basis_content_required(const Config* config);
#endif
+49 -5
View File
@@ -5,6 +5,15 @@
#include <stdbool.h>
#include <stdint.h>
/* Ask the compiler to type-check the printf-style arguments of the variadic
* logging helpers. Only enabled for GNU-compatible compilers (gcc/clang). */
#if defined(__GNUC__)
#define LOG_PRINTF_ATTR(fmt_idx, first_vararg_idx) \
__attribute__((format(printf, fmt_idx, first_vararg_idx)))
#else
#define LOG_PRINTF_ATTR(fmt_idx, first_vararg_idx)
#endif
typedef enum { LOG_LEVEL_DEBUG, LOG_LEVEL_INFO, LOG_LEVEL_WARNING, LOG_LEVEL_ERROR } LogLevel;
typedef enum { LOG_STDERR_ERRORS, LOG_STDERR_ALL } LogStderrMode;
@@ -13,7 +22,19 @@ typedef enum {
LOG_DEBUG_PROTO = 1u << 1,
LOG_DEBUG_PACK = 1u << 2,
LOG_DEBUG_UTIL = 1u << 3,
LOG_DEBUG_ALL = (1u << 4) - 1,
/* rsync --debug categories that now map to a natural FastSync event:
* flist (file-list scan progress), del (deletions), hash/deltasum
* (whole-file hashing and delta-sum generation), recv (receiver
* responses/signatures), filter (selection/exclusion decisions) and send
* (files handed to the sender). Only emitted when the category is
* explicitly enabled; a normal run stays silent. */
LOG_DEBUG_FLIST = 1u << 4,
LOG_DEBUG_DEL = 1u << 5,
LOG_DEBUG_HASH = 1u << 6,
LOG_DEBUG_RECV = 1u << 7,
LOG_DEBUG_FILTER = 1u << 8,
LOG_DEBUG_SEND = 1u << 9,
LOG_DEBUG_ALL = (1u << 10) - 1,
} LogDebugFlag;
typedef enum {
@@ -21,10 +42,33 @@ typedef enum {
LOG_INFO_MISC = 1u << 1,
LOG_INFO_SKIP = 1u << 2,
LOG_INFO_STATS = 1u << 3,
LOG_INFO_ALL = LOG_INFO_COPY | LOG_INFO_MISC | LOG_INFO_SKIP | LOG_INFO_STATS,
/* rsync categories that map to a FastSync event (emitted in rsync's line
* format): del (deletions), remove (sender-side source removal), name
* (transferred entry names), flist (file-list header), nonreg (skipped
* non-regular files), progress (per-file progress). rsync's `backup`
* category is accepted for CLI parity but stays silent: the receiver does the
* backing-up and FastSync has no backup event to report from the sender. */
LOG_INFO_DEL = 1u << 4,
LOG_INFO_REMOVE = 1u << 5,
LOG_INFO_NAME = 1u << 6,
LOG_INFO_FLIST = 1u << 7,
LOG_INFO_NONREG = 1u << 8,
LOG_INFO_PROGRESS = 1u << 9,
/* Marker for `--info=name2` and higher: also print rsync's
"NAME is uptodate" line for entries the receiver already has. It rides in
the info_level bitset (there is no separate Config field) and is never set
by --info=all (which selects level 1). */
LOG_INFO_NAME_UPTODATE = 1u << 10,
/* --info=mount: print rsync's `[sender] skipping mount-point dir NAME` when
* -xx/--one-file-system drops a mount-point directory (FastSync's client is
* the sender). */
LOG_INFO_MOUNT = 1u << 11,
LOG_INFO_ALL = LOG_INFO_COPY | LOG_INFO_MISC | LOG_INFO_SKIP | LOG_INFO_STATS | LOG_INFO_DEL |
LOG_INFO_REMOVE | LOG_INFO_NAME | LOG_INFO_FLIST | LOG_INFO_NONREG |
LOG_INFO_PROGRESS | LOG_INFO_MOUNT,
} LogInfoFlag;
void log_message(LogLevel log_level, const char* message, ...);
void log_message(LogLevel log_level, const char* message, ...) LOG_PRINTF_ATTR(2, 3);
void log_perror(const char* context);
void set_log_level(LogLevel level);
void set_log_debug_flags(uint32_t flags);
@@ -33,10 +77,10 @@ uint32_t get_log_debug_flags(void);
* the debug log level is enabled AND the flag is selected. Hot paths use this
* to skip expensive message formatting/escaping when the line is filtered. */
bool log_debug_enabled(LogDebugFlag flag);
void log_debug_message(LogDebugFlag flag, const char* message, ...);
void log_debug_message(LogDebugFlag flag, const char* message, ...) LOG_PRINTF_ATTR(2, 3);
void set_log_info_flags(uint32_t flags);
uint32_t get_log_info_flags(void);
void log_info_message(LogInfoFlag flag, const char* message, ...);
void log_info_message(LogInfoFlag flag, const char* message, ...) LOG_PRINTF_ATTR(2, 3);
void log_set_file(FILE* fp);
void log_set_8_bit_output(bool enabled);
bool log_get_8_bit_output(void);
+20 -5
View File
@@ -211,13 +211,22 @@ FileMetadata* metadata_receive(int file_descriptor, int* ok) {
bool metadata_mode_for_policy(mode_t source_mode, mode_t current_mode, FileAttrPolicy policy,
mode_t* out_mode) {
const mode_t special_bits = (mode_t)(S_ISUID | S_ISGID | S_ISVTX);
const mode_t execute_bits = S_IXUSR | S_IXGRP | S_IXOTH;
if (policy.perms) {
/* rsync --perms copies the source's permission and special bits exactly,
* including group/other write and setuid/setgid/sticky. The kernel may
* still clear setgid when the receiver is not in the file's group; the
* caller logs a failed chmod rather than silently masking the bits here. */
*out_mode = source_mode & (mode_t)(S_ISUID | S_ISGID | S_ISVTX | 0777);
* caller logs a failed chmod rather than silently masking the bits here.
* Setuid/setgid/sticky are super-user activities: when the connection did
* not permit them (SUPER_MODE_OFF / --no-super) they are stripped, so a
* client can never install a privileged bit on a receiver that forbade
* super-user activities. This also covers bits introduced by --chmod,
* whose result is fed in as source_mode. */
mode_t bits = source_mode & (mode_t)(special_bits | 0777);
if (!policy.super_permitted)
bits &= ~special_bits;
*out_mode = bits;
return true;
}
if (policy.executability) {
@@ -227,8 +236,11 @@ bool metadata_mode_for_policy(mode_t source_mode, mode_t current_mode, FileAttrP
* execute); otherwise clear every execute bit. This runs on the
* destination-derived base (pre-existing dest mode, or source&~umask for a
* new file), and leaves the special bits untouched. --perms wins when both
* are set (handled above). */
mode_t base = current_mode & (mode_t)(S_ISUID | S_ISGID | S_ISVTX | 0777);
* are set (handled above). The destination's own special bits survive
* unless super-user activities are forbidden. */
mode_t base = current_mode & (mode_t)(special_bits | 0777);
if (!policy.super_permitted)
base &= ~special_bits;
if (source_mode & 0111)
*out_mode = base | ((base & 0444) >> 2);
else
@@ -240,12 +252,13 @@ bool metadata_mode_for_policy(mode_t source_mode, mode_t current_mode, FileAttrP
}
FileAttrPolicy file_attr_policy_from_config(const Config* config) {
FileAttrPolicy policy = {false, false, false, false};
FileAttrPolicy policy = {0};
if (config) {
policy.perms = config->preserve_perms;
policy.times = config->preserve_times;
policy.atimes = config->preserve_atimes;
policy.executability = config->use_executability;
policy.super_permitted = privilege_super_mode_permitted(config->super_mode);
}
return policy;
}
@@ -314,6 +327,8 @@ bool file_restore_symlink_metadata(const char* path, const FileMetadata* metadat
transfer never fails over it. */
if (policy.perms) {
mode_t link_mode = metadata->mode & (mode_t)(S_ISUID | S_ISGID | S_ISVTX | 0777);
if (!policy.super_permitted)
link_mode &= ~(mode_t)(S_ISUID | S_ISGID | S_ISVTX);
if (fchmodat(parent_fd, leaf, link_mode, AT_SYMLINK_NOFOLLOW) != 0 && errno != EOPNOTSUPP &&
errno != ENOTSUP && errno != ENOSYS) {
log_message(LOG_LEVEL_DEBUG, "Could not set symlink mode on %s: %s", path, strerror(errno));
+9 -1
View File
@@ -37,17 +37,21 @@ PipelineContextSender* pipeline_context_sender_create(Config* config, Queue* que
context->scan_had_io_error = false;
context->remove_source_files = NULL;
context->early_delete = false;
context->prescan_chunks = NULL;
context->delete_plans = NULL;
context->delete_suppressed = false;
context->scan_stopped_early = false;
context->total_files = 0;
context->progress_bytes = 0;
context->total_bytes = 0;
memset(&context->stats, 0, sizeof(context->stats));
context->sender_done = false;
atomic_init(&context->cancelled, false);
protocol_session_init(&context->allocation_session, -1, -1);
protocol_session_set_max_alloc(&context->allocation_session, config->max_alloc);
context->dir_entries = NULL;
context->dir_entries_mutex_init = false;
atomic_init(&context->dir_count, 0);
context->delete_limit = false;
int init = 0;
if (config->use_metadata) {
@@ -83,11 +87,13 @@ PipelineContextSender* pipeline_context_sender_create(Config* config, Queue* que
return context;
fail:
log_perror("Error initializing synchronization objects");
log_message(LOG_LEVEL_ERROR, "%s", "Error initializing synchronization objects");
if (context->dir_entries_mutex_init)
mtx_destroy(&context->dir_entries_mutex);
if (context->dir_entries)
array_list_delete(context->dir_entries);
if (init >= 7)
mtx_destroy(&context->mutex_progress);
if (init >= 6)
cnd_destroy(&context->condition_not_empty_loader);
if (init >= 5)
@@ -189,6 +195,8 @@ void pipeline_context_sender_destroy(PipelineContextSender* context) {
if (context->manifest) {
array_list_delete(context->manifest);
}
if (context->prescan_chunks)
array_list_delete(context->prescan_chunks);
if (context->delete_plans)
delete_plan_sender_destroy(context->delete_plans);
if (context->excluded_paths)
+24
View File
@@ -9,6 +9,7 @@
#include "config.h"
#include "delete_plan.h"
#include "file.h"
#include "format.h"
#include "protocol.h"
#include "queue.h"
#include "stop_condition.h"
@@ -77,15 +78,34 @@ typedef struct {
path-only pre-scan on the calling thread and the pipeline scanner must not
append to it. Set once before the worker threads start. */
bool early_delete;
/* --delete-before: the path-only pre-scan that built the early keep-set,
retained as the pipeline's file list (owning Chunk*; consumed and NULLed by
the scanner thread) so the data pass replays rsync's single file list
instead of re-reading the source. NULL in every other mode, where the
scanner thread scans normally. Set once before the worker threads start
and freed with the context. */
ArrayList* prescan_chunks;
/* Non-NULL for --delete-during/--delete-delay: the per-directory plan set
prebuilt by the path-only pre-scan on the calling thread. The sender
thread transmits the root plan before any data and the remaining plans
alongside the chunks. Set once before the worker threads start. */
DeletePlanSender* delete_plans;
/* A scan I/O error without --ignore-errors suppressed deletion: the prebuilt
keep-set/plans were dropped, and the streaming scanner must not build a
fresh manifest or re-send the per-directory plans. Set once before the
worker threads start. */
bool delete_suppressed;
mtx_t mutex_progress;
int total_files;
unsigned long long progress_bytes;
unsigned long long total_bytes;
/* Per-type flist / transferred accounting for the rsync --stats breakdown and
the progress `to-chk` denominator. Owned by the sender thread: it is the
only writer (the entry/transfer notes in send_chunks_multithreaded) and it
reads the totals in its completion tail, so no lock is needed. This is NOT
guarded by mutex_progress (which covers total_files/progress_bytes/
total_bytes/sender_done). */
TransferStats stats;
bool sender_done;
atomic_bool cancelled;
ProtocolSession allocation_session;
@@ -106,6 +126,10 @@ typedef struct {
ArrayList* dir_entries;
mtx_t dir_entries_mutex;
bool dir_entries_mutex_init;
/* --stats directory accounting for a `-r` scan (no directory metadata):
shared by the parallel scanner workers, read by the sender thread once the
scanner is done. See ScannerOptions.dir_count. */
atomic_ullong dir_count;
/* Set by the sender thread when the receiver reported a --max-delete-capped
deletion (STATUS_DELETE_LIMIT): the transfer succeeded and the process must
exit 25 like rsync. Read by the caller after the sender thread is joined. */
+232 -78
View File
@@ -38,6 +38,129 @@ static atomic_ullong io_bytes_read = 0;
static unsigned long long global_bwlimit(void);
/* ------------------------------------------------------------------------- *
* Transport vtable implementations.
*
* Each op performs exactly one transfer attempt. WANT_READ/WANT_WRITE and an
* EINTR-interrupted syscall are reported as PROTOCOL_IO_RETRY (with
* *wait_events set to the poll event the caller must wait on); a clean peer
* close is PROTOCOL_IO_CLOSED and anything else is PROTOCOL_IO_ERROR. This
* keeps every WANT_READ/WANT_WRITE and EINTR retry exactly where it was before
* the vtable was introduced, just moved behind the function pointer.
* ------------------------------------------------------------------------- */
static ssize_t plain_io_send(ProtocolSession* session, const void* data, size_t size,
short* wait_events) {
ssize_t written = write(session->write_fd, data, size);
if (written < 0) {
if (errno == EINTR)
return PROTOCOL_IO_RETRY;
return PROTOCOL_IO_ERROR;
}
if (written == 0)
return PROTOCOL_IO_ERROR;
*wait_events = POLLOUT;
return written;
}
static ssize_t plain_io_recv(ProtocolSession* session, void* data, size_t size,
short* wait_events) {
ssize_t received = read(session->read_fd, data, size);
if (received < 0) {
if (errno == EINTR)
return PROTOCOL_IO_RETRY;
return PROTOCOL_IO_ERROR;
}
if (received == 0)
return PROTOCOL_IO_CLOSED;
*wait_events = POLLIN;
return received;
}
static bool plain_io_has_pending(const ProtocolSession* session) {
(void)session;
return false;
}
static ssize_t tls_io_send(ProtocolSession* session, const void* data, size_t size,
short* wait_events) {
/* SSL_write takes an int length; clamp a >INT_MAX request into chunks so the
* size_t downcast can never truncate into a negative/partial write. */
size_t chunk = size > (size_t)INT_MAX ? (size_t)INT_MAX : size;
ssize_t written = SSL_write(session->ssl, data, (int)chunk);
if (written <= 0) {
int ssl_err = SSL_get_error(session->ssl, (int)written);
if (ssl_err == SSL_ERROR_WANT_WRITE) {
*wait_events = POLLOUT;
return PROTOCOL_IO_RETRY;
}
if (ssl_err == SSL_ERROR_WANT_READ) {
*wait_events = POLLIN;
return PROTOCOL_IO_RETRY;
}
/* A signal (e.g. Ctrl-C) interrupts the blocking TLS write: retry so the
* send loop can observe the abort flag at the next checkpoint. Only an
* actual negative return is an interrupted syscall; a 0-byte SSL_write is
* not a valid EINTR retry. */
if (written < 0 && ssl_err == SSL_ERROR_SYSCALL && errno == EINTR)
return PROTOCOL_IO_RETRY;
return PROTOCOL_IO_ERROR;
}
*wait_events = POLLOUT;
return written;
}
static ssize_t tls_io_recv(ProtocolSession* session, void* data, size_t size, short* wait_events) {
/* SSL_read takes an int length; clamp a >INT_MAX request into chunks
* (mirrors the send path) so the size_t downcast can never truncate into a
* negative/partial read. */
size_t chunk = size > (size_t)INT_MAX ? (size_t)INT_MAX : size;
ssize_t received = SSL_read(session->ssl, data, (int)chunk);
if (received <= 0) {
int ssl_err = SSL_get_error(session->ssl, (int)received);
if (ssl_err == SSL_ERROR_WANT_WRITE) {
*wait_events = POLLOUT;
return PROTOCOL_IO_RETRY;
}
if (ssl_err == SSL_ERROR_WANT_READ) {
*wait_events = POLLIN;
return PROTOCOL_IO_RETRY;
}
/* A signal interrupts the blocking TLS read: retry (mirrors the send path)
* so the loop reaches its next abort/deadline checkpoint. Only an actual
* negative return is an interrupted syscall: a 0-byte SSL_read is an
* unexpected EOF (the peer closed without close_notify), which OpenSSL also
* reports as SSL_ERROR_SYSCALL with errno possibly still EINTR from an
* earlier interrupted poll/read. Retrying that would busy-spin the
* status-read loop until its deadline, so classify it as closed instead. */
if (received < 0 && ssl_err == SSL_ERROR_SYSCALL && errno == EINTR)
return PROTOCOL_IO_RETRY;
/* A zero-length SSL_read is the peer's clean close_notify (or EOF without
* one); report it distinctly so the caller can log it as a close. */
if (received == 0)
return PROTOCOL_IO_CLOSED;
return PROTOCOL_IO_ERROR;
}
*wait_events = POLLIN;
return received;
}
static bool tls_io_has_pending(const ProtocolSession* session) {
return session->ssl != NULL && SSL_pending(session->ssl) > 0;
}
static const ProtocolIoOps plain_io_ops = {
.send = plain_io_send,
.recv = plain_io_recv,
.has_pending = plain_io_has_pending,
};
static const ProtocolIoOps tls_io_ops = {
.send = tls_io_send,
.recv = tls_io_recv,
.has_pending = tls_io_has_pending,
};
static bool protocol_reserve_memory(ProtocolSession* session, size_t charge) {
unsigned long long allocated = atomic_load(&session->total_allocated_bytes);
while (true) {
@@ -79,6 +202,7 @@ void io_set_fds(int read_fd, int write_fd) {
legacy_io_session.read_fd = read_fd;
legacy_io_session.write_fd = write_fd;
legacy_io_session.ssl = NULL;
legacy_io_session.ops = &plain_io_ops;
legacy_io_session.eight_bit_output = false;
atomic_store(&legacy_io_session.total_allocated_bytes, 0);
legacy_io_session.max_alloc = DEFAULT_MAX_ALLOC;
@@ -91,6 +215,7 @@ void protocol_session_init(ProtocolSession* session, int read_fd, int write_fd)
memset(session, 0, sizeof(*session));
session->read_fd = read_fd;
session->write_fd = write_fd;
session->ops = &plain_io_ops;
session->max_alloc = DEFAULT_MAX_ALLOC;
session->io_timeout_sec = RECEIVE_TIMEOUT_SEC;
atomic_init(&session->total_allocated_bytes, 0);
@@ -158,8 +283,12 @@ void protocol_session_unbind(void) {
}
void protocol_session_set_ssl(ProtocolSession* session, SSL* ssl) {
if (session)
session->ssl = ssl;
if (!session)
return;
session->ssl = ssl;
/* Select the transport dispatch once, here, instead of branching on the SSL
* pointer inside every I/O loop. */
session->ops = ssl ? &tls_io_ops : &plain_io_ops;
}
static void bw_mutex_init(void) {
@@ -183,12 +312,28 @@ void io_set_bwlimit(unsigned long long bytes_per_sec) {
mtx_unlock(&bw_mutex);
}
unsigned long long io_get_bwlimit(void) {
return global_bwlimit();
}
/* rsync's throttle (io.c sleep_for_bwlimit) sleeps once its unslept debt
* reaches ~100 ms of bandwidth, so its effective initial burst is about 0.1 s
* worth of bytes, not a full second. FastSync models the same with a token
* bucket whose capacity is bwlimit/10, so a throttled run paces like rsync
* instead of sending a full second's worth up front. */
static long long bw_burst_capacity(unsigned long long bwlimit) {
if (bwlimit == 0)
return 0;
long long burst = (long long)(bwlimit / 10);
return burst > 0 ? burst : 1;
}
void protocol_session_set_bwlimit(ProtocolSession* session, unsigned long long bytes_per_sec) {
if (!session)
return;
session->bwlimit =
bytes_per_sec > (unsigned long long)LLONG_MAX ? (unsigned long long)LLONG_MAX : bytes_per_sec;
session->bw_tokens = (long long)session->bwlimit;
session->bw_tokens = bw_burst_capacity(session->bwlimit);
struct timespec now;
clock_gettime(CLOCK_MONOTONIC, &now);
session->bw_last_refill_sec = now.tv_sec;
@@ -222,8 +367,9 @@ static void bw_throttle_session(ProtocolSession* session, size_t bytes_written)
long long tokens_to_add = (long long)((double)session->bwlimit * elapsed_ns / 1000000000.0);
session->bw_tokens += tokens_to_add;
if (session->bw_tokens > (long long)session->bwlimit)
session->bw_tokens = (long long)session->bwlimit;
long long burst = bw_burst_capacity(session->bwlimit);
if (session->bw_tokens > burst)
session->bw_tokens = burst;
session->bw_tokens -= bytes_written;
@@ -234,7 +380,11 @@ static void bw_throttle_session(ProtocolSession* session, size_t bytes_written)
poll(NULL, 0, (int)(deficit_us / 1000));
else
usleep((useconds_t)deficit_us);
/* Reset the bucket AFTER the sleep: crediting the sleep duration as elapsed
refill time would cancel half the throttle (the next call would see the
whole sleep as refill and immediately grant a fresh burst). */
session->bw_tokens = 0;
clock_gettime(CLOCK_MONOTONIC, &now);
session->bw_last_refill_sec = now.tv_sec;
session->bw_last_refill_nsec = now.tv_nsec;
}
@@ -249,6 +399,20 @@ SSL* io_get_ssl(void) {
return io_ssl;
}
SSL* protocol_current_ssl(void) {
/* The bound session is the authoritative transport for a worker thread: it
* was explicitly handed to protocol_session_bind() and carries its own SSL,
* whereas io_ssl is thread-local and NULL in a thread that never performed
* the handshake. Only a session whose selected dispatch is TLS may supply
* the SSL: a bound plaintext session has ssl == NULL and must not shadow a
* live thread-local io_ssl, or file_send.c would take the raw sendfile(2)
* path on a socket this thread is driving with TLS. With no TLS session
* bound (plaintext session, or the fd-shim path), fall back to io_ssl. */
if (bound_session && bound_session->ops == &tls_io_ops && bound_session->ssl)
return bound_session->ssl;
return io_ssl;
}
unsigned long long protocol_bytes_written(void) {
return atomic_load(&io_bytes_written);
}
@@ -277,9 +441,21 @@ static ProtocolSession* legacy_session(int read_fd, int write_fd) {
protocol_session_set_bwlimit(&legacy_io_session, global_bwlimit());
}
legacy_io_session.ssl = io_ssl;
legacy_io_session.ops = io_ssl ? &tls_io_ops : &plain_io_ops;
return &legacy_io_session;
}
/* Pace an out-of-band write that bypassed protocol_send_n_data (the plaintext
* sendfile fast path). The bound/legacy session is resolved exactly as the
* preceding send_n_data(fd, ...) resolved it, so the same token-bucket state is
* throttled and the TLS and plaintext transports share identical --bwlimit
* semantics. Passing the wire fd (rather than -1) is essential: the sendfile
* send left legacy_io_session.write_fd bound to it, so resolving with -1 would
* mismatch, re-initialize the session and hand out a second first-call burst. */
void protocol_throttle_bytes(int file_descriptor, size_t bytes) {
bw_throttle_session(legacy_session(-1, file_descriptor), bytes);
}
bool send_n_data(int file_descriptor, const void* data, size_t data_size) {
return protocol_send_n_data(legacy_session(-1, file_descriptor), data, data_size);
}
@@ -303,8 +479,9 @@ bool protocol_send_n_data(ProtocolSession* session, const void* data, size_t dat
if (!data && data_size != 0)
return false;
log_debug_message(LOG_DEBUG_IO, " Sending n Data: %zu", data_size);
if (!session)
if (!session || !session->ops)
return false;
const ProtocolIoOps* ops = session->ops;
/* A non-positive session timeout disables the deadline entirely (rsync's
* --timeout=0 default); poll then blocks until the socket becomes writable. */
int timeout_sec = session->io_timeout_sec > 0 ? session->io_timeout_sec : 0;
@@ -330,38 +507,18 @@ bool protocol_send_n_data(ProtocolSession* session, const void* data, size_t dat
continue;
if (pfd.revents & (POLLERR | POLLNVAL))
return false;
ssize_t bytes_send;
if (session->ssl) {
/* SSL_write takes an int length; clamp a >INT_MAX request into chunks so
* the size_t downcast can never truncate into a negative/partial write. */
size_t ssl_chunk = chunk > (size_t)INT_MAX ? (size_t)INT_MAX : chunk;
bytes_send = SSL_write(session->ssl, (const char*)data + total_bytes_send, (int)ssl_chunk);
} else {
bytes_send = write(fd, (const char*)data + total_bytes_send, chunk);
}
ssize_t bytes_send =
ops->send(session, (const char*)data + total_bytes_send, chunk, &wait_events);
if (bytes_send == PROTOCOL_IO_RETRY)
continue;
if (bytes_send <= 0) {
if (session->ssl) {
int ssl_err = SSL_get_error(session->ssl, (int)bytes_send);
if (ssl_err == SSL_ERROR_WANT_WRITE || ssl_err == SSL_ERROR_WANT_READ) {
wait_events = ssl_err == SSL_ERROR_WANT_WRITE ? POLLOUT : POLLIN;
continue;
}
/* A signal (e.g. Ctrl-C) interrupts the blocking TLS write: retry so
the send loop can observe the abort flag at the next checkpoint. */
if (ssl_err == SSL_ERROR_SYSCALL && errno == EINTR)
continue;
} else if (errno == EINTR) {
continue;
}
log_message(LOG_LEVEL_ERROR, "Could not send data");
return false;
}
bw_throttle_session(session, (size_t)bytes_send);
total_bytes_send += bytes_send;
if (session->ssl)
wait_events = POLLOUT;
}
log_debug_message(LOG_DEBUG_IO, " Send n Data: %zu", total_bytes_send);
log_debug_message(LOG_DEBUG_IO, " Send n Data: %zd", total_bytes_send);
atomic_fetch_add(&io_bytes_written, (unsigned long long)total_bytes_send);
return true;
}
@@ -392,14 +549,15 @@ bool protocol_receive_n_data(ProtocolSession* session, void* data, size_t data_s
static bool protocol_receive_n_data_until(ProtocolSession* session, void* data, size_t data_size,
const struct timespec* deadline) {
log_debug_message(LOG_DEBUG_IO, " Receiving n Data: %zu", data_size);
if (!session)
if (!session || !session->ops)
return false;
const ProtocolIoOps* ops = session->ops;
int fd = session->read_fd;
size_t total_bytes_received = 0;
short wait_events = POLLIN;
while (total_bytes_received < data_size) {
if (!session->ssl || SSL_pending(session->ssl) == 0) {
if (!ops->has_pending(session)) {
struct pollfd pfd = {.fd = fd, .events = wait_events};
/* A NULL deadline means "wait indefinitely" (timeout disabled). */
int poll_result = poll(&pfd, 1, deadline ? deadline_remaining_ms(deadline) : -1);
@@ -417,37 +575,19 @@ static bool protocol_receive_n_data_until(ProtocolSession* session, void* data,
return false;
}
ssize_t bytes_received;
if (session->ssl)
bytes_received = SSL_read(session->ssl, (char*)data + total_bytes_received,
data_size - total_bytes_received);
else
bytes_received =
read(fd, (char*)data + total_bytes_received, data_size - total_bytes_received);
ssize_t bytes_received = ops->recv(session, (char*)data + total_bytes_received,
data_size - total_bytes_received, &wait_events);
if (bytes_received == PROTOCOL_IO_RETRY)
continue;
if (bytes_received == PROTOCOL_IO_CLOSED) {
log_message(LOG_LEVEL_ERROR, "Connection closed while receiving data");
return false;
}
if (bytes_received <= 0) {
if (session->ssl) {
int ssl_err = SSL_get_error(session->ssl, (int)bytes_received);
if (ssl_err == SSL_ERROR_WANT_WRITE || ssl_err == SSL_ERROR_WANT_READ) {
wait_events = ssl_err == SSL_ERROR_WANT_WRITE ? POLLOUT : POLLIN;
continue;
}
/* A signal interrupts the blocking TLS read: retry (mirrors the send
path and protocol_read_status_until) so the loop reaches its next
abort/deadline checkpoint instead of failing spuriously. */
if (ssl_err == SSL_ERROR_SYSCALL && errno == EINTR)
continue;
} else if (errno == EINTR) {
continue;
}
if (bytes_received == 0)
log_message(LOG_LEVEL_ERROR, "Connection closed while receiving data");
else
log_message(LOG_LEVEL_ERROR, "Could not receive bytes");
log_message(LOG_LEVEL_ERROR, "Could not receive bytes");
return false;
}
total_bytes_received += (size_t)bytes_received;
if (session->ssl)
wait_events = POLLIN;
}
log_debug_message(LOG_DEBUG_IO, " Received n Data: %zu", total_bytes_received);
atomic_fetch_add(&io_bytes_read, (unsigned long long)total_bytes_received);
@@ -529,6 +669,15 @@ static const char* status_to_string(Status status) {
}
}
/* Reject a raw wire status outside the known enum range before it is handed to
* callers, so an unknown/corrupt frame fails as a protocol error instead of
* being silently interpreted as an unexpected-but-valid verdict. STATUS_OK is
* the first enumerator and STATUS_STATS the last, so the range check accepts
* every status the protocol defines. */
static bool status_is_valid(Status status) {
return status >= STATUS_OK && status <= STATUS_STATS;
}
/* Shared string send/receive implementation. `redact` selects whether the
* payload body is written to the LOG_DEBUG_PROTO debug log: daemon auth material
* (the username and the proof/signature fields) sets it so a --verbose log never
@@ -612,7 +761,7 @@ bool protocol_send_data(ProtocolSession* session, const Data* data) {
return false;
if (!protocol_send_n_data(session, data->data, data_size))
return false;
log_debug_message(LOG_DEBUG_PROTO, "Send %lld data", data_size);
log_debug_message(LOG_DEBUG_PROTO, "Send %llu data", data_size);
return true;
}
@@ -646,7 +795,7 @@ Data* protocol_receive_data_limited(ProtocolSession* session, unsigned long long
protocol_release_memory_for_session(session, allocation_size);
return NULL;
}
log_debug_message(LOG_DEBUG_PROTO, "Received %lld data", size);
log_debug_message(LOG_DEBUG_PROTO, "Received %llu data", size);
Data* result = data_create(data, (size_t)size);
if (!result) {
protocol_release_memory_for_session(session, allocation_size);
@@ -756,6 +905,10 @@ bool protocol_receive_status(ProtocolSession* session, Status* status) {
}
if (!protocol_receive_n_data_until(session, status, sizeof(Status), deadline_ptr))
return false;
if (!status_is_valid(*status)) {
log_message(LOG_LEVEL_ERROR, "Received unknown protocol status %d", *status);
return false;
}
if (!protocol_capture_error_detail(session, status, deadline_ptr, NULL))
return false;
log_debug_message(LOG_DEBUG_PROTO, "Received Status: %s", status_to_string(*status));
@@ -777,6 +930,10 @@ bool protocol_receive_status_timed(ProtocolSession* session, Status* status, int
deadline.tv_sec += timeout_sec;
if (!protocol_receive_n_data_until(session, status, sizeof(Status), &deadline))
return false;
if (!status_is_valid(*status)) {
log_message(LOG_LEVEL_ERROR, "Received unknown protocol status %d", *status);
return false;
}
if (!protocol_capture_error_detail(session, status, &deadline, NULL))
return false;
log_debug_message(LOG_DEBUG_PROTO, "Received Status: %s", status_to_string(*status));
@@ -790,11 +947,14 @@ bool protocol_receive_status_timed(ProtocolSession* session, Status* status, int
* reply across a frame boundary. Returns false on timeout/EOF/error. */
static bool protocol_read_status_until(ProtocolSession* session, Status* status,
const struct timespec* deadline) {
if (!session || !session->ops)
return false;
const ProtocolIoOps* ops = session->ops;
Status received = STATUS_ERROR;
size_t got = 0;
short wait_events = POLLIN;
while (got < sizeof(Status)) {
if (!session->ssl || SSL_pending(session->ssl) == 0) {
if (!ops->has_pending(session)) {
int remaining_ms = deadline ? deadline_remaining_ms(deadline) : -1;
if (remaining_ms == 0) {
log_message(LOG_LEVEL_ERROR, "Receive timeout while reading status");
@@ -814,21 +974,11 @@ static bool protocol_read_status_until(ProtocolSession* session, Status* status,
if (pfd.revents & (POLLERR | POLLNVAL))
return false;
}
ssize_t bytes_received;
if (session->ssl)
bytes_received = SSL_read(session->ssl, (char*)&received + got, sizeof(Status) - got);
else
bytes_received = read(session->read_fd, (char*)&received + got, sizeof(Status) - got);
ssize_t bytes_received =
ops->recv(session, (char*)&received + got, sizeof(Status) - got, &wait_events);
if (bytes_received == PROTOCOL_IO_RETRY)
continue;
if (bytes_received <= 0) {
if (session->ssl) {
int ssl_err = SSL_get_error(session->ssl, (int)bytes_received);
if (ssl_err == SSL_ERROR_WANT_READ || ssl_err == SSL_ERROR_WANT_WRITE) {
wait_events = ssl_err == SSL_ERROR_WANT_WRITE ? POLLOUT : POLLIN;
continue;
}
}
if (bytes_received < 0 && errno == EINTR)
continue;
log_message(LOG_LEVEL_ERROR, "Connection closed while receiving status");
return false;
}
@@ -857,7 +1007,7 @@ bool protocol_receive_status_keepalive(ProtocolSession* session, Status* status,
while (true) {
if (abort_check && abort_check())
return false;
if (!session->ssl || SSL_pending(session->ssl) == 0) {
if (!session->ops || !session->ops->has_pending(session)) {
int remaining_ms = deadline_remaining_ms(&deadline);
if (remaining_ms <= 0) {
log_message(LOG_LEVEL_ERROR, "Receive timeout after %ds", timeout_sec);
@@ -890,6 +1040,10 @@ bool protocol_receive_status_keepalive(ProtocolSession* session, Status* status,
Status received;
if (!protocol_read_status_until(session, &received, &deadline))
return false;
if (!status_is_valid(received)) {
log_message(LOG_LEVEL_ERROR, "Received unknown protocol status %d", received);
return false;
}
if (!protocol_capture_error_detail(session, &received, &deadline, abort_check))
return false;
if (received == STATUS_KEEPALIVE) {
+65 -3
View File
@@ -28,6 +28,10 @@
/* Maximum chunk size (64 MB) — prevents unbounded allocation from the wire */
#define MAX_CHUNK_SIZE (64ULL * 1024 * 1024)
/* Files larger than this are not kept fully in memory while loading: the
* loader skips them so the sender streams from the path, and file_checksum
* hashes them from disk in bounded buffers instead of forcing a full load. */
#define STREAM_THRESHOLD (64ULL * 1024 * 1024)
#define MAX_MANIFEST_ENTRIES (1024 * 1024)
/* Aggregate bytes retained by one received deletion manifest. */
#define MAX_MANIFEST_BYTES (16ULL * 1024 * 1024)
@@ -46,16 +50,48 @@
typedef struct ssl_st SSL;
typedef struct ProtocolSession ProtocolSession;
/*
* Transport vtable: the per-session set of I/O primitives the three protocol
* loops (send, receive, status-read) dispatch through. The ops are selected
* once, when the session is initialized or its SSL is installed, so the loops
* never branch on the transport at runtime. A plaintext session uses the
* read()/write() ops; a TLS session uses the SSL_read()/SSL_write() ops.
*
* `send`/`recv` attempt exactly one transfer and return:
* > 0 bytes transferred,
* PROTOCOL_IO_RETRY no progress; poll on *wait_events and retry,
* PROTOCOL_IO_CLOSED peer closed the stream,
* PROTOCOL_IO_ERROR fatal transport error.
* `has_pending` reports bytes already buffered by the transport (a TLS record
* residue); the receive loops skip the poll() gate when it is true.
*/
typedef struct ProtocolIoOps {
ssize_t (*send)(ProtocolSession* session, const void* data, size_t size, short* wait_events);
ssize_t (*recv)(ProtocolSession* session, void* data, size_t size, short* wait_events);
bool (*has_pending)(const ProtocolSession* session);
} ProtocolIoOps;
/* Negative sentinels returned by ProtocolIoOps.send/recv (see above). */
enum {
PROTOCOL_IO_RETRY = -1,
PROTOCOL_IO_CLOSED = -2,
PROTOCOL_IO_ERROR = -3,
};
/*
* Explicit owner of protocol I/O. A session does not own the descriptors or
* SSL object; it only describes the transport used by a transfer. This makes
* it safe to pass the transport to a worker without relying on inherited
* thread-local state.
*/
typedef struct ProtocolSession {
struct ProtocolSession {
int read_fd;
int write_fd;
SSL* ssl;
/* Transport dispatch selected by protocol_session_init()/set_ssl(). */
const ProtocolIoOps* ops;
unsigned long long bwlimit;
long long bw_tokens;
long long bw_last_refill_sec;
@@ -71,7 +107,7 @@ typedef struct ProtocolSession {
* SO_RCVTIMEO/SO_SNDTIMEO. The server does not propagate a client 0 here: it
* installs protocol_server_io_timeout_sec() so its sessions keep a floor. */
int io_timeout_sec;
} ProtocolSession;
};
typedef int Status;
enum NET_STATUS {
@@ -191,7 +227,9 @@ enum NET_STATUS {
* (--delete-delay). Payload: an int32 has_config flag (1 on the first plan
* of the run, 0 afterwards); when set, the three global config sections
* (protected-prefix count+paths, size-skipped count+paths, missing-args
* count+paths); then the destination-relative directory path wire string
* count+paths); then an int32 apply flag (1 for a real plan, 0 for a
* config-only carrier frame that must not walk a directory); then the
* destination-relative directory path wire string
* ("." for the receive root); then the child-directory count + names and the
* child-file count + names that must be kept. Appended after
* STATUS_DEST_INFO so no existing status is renumbered. */
@@ -207,8 +245,20 @@ enum NET_STATUS {
void io_set_fds(int read_fd, int write_fd);
void io_set_bwlimit(unsigned long long bytes_per_sec);
unsigned long long io_get_bwlimit(void);
void io_set_ssl(SSL* ssl);
SSL* io_get_ssl(void);
/* SSL object of the transport in effect on this thread: the currently bound
* session's SSL when a TLS session is bound, otherwise the legacy thread-local
* io_ssl. NULL for a plaintext transport. Unlike io_get_ssl(), this resolves
* worker threads that bound a TLS session via protocol_session_set_ssl()/
* protocol_session_bind() but never called io_set_ssl() themselves (C11
* _Thread_local state is not inherited by a new thread). A bound session only
* wins when its selected dispatch is TLS; a bound plaintext session (ssl ==
* NULL) falls back to io_ssl so it can never mask a live encrypted transport.
* Callers that must choose a TLS-only code path (e.g. file_send.c's sendfile
* fallback) must use this instead of io_get_ssl(). */
SSL* protocol_current_ssl(void);
/* Process-wide wire byte counters. protocol_send_n_data/protocol_receive_n_data
* update them; the zero-copy sendfile path reports through
@@ -217,6 +267,15 @@ SSL* io_get_ssl(void);
unsigned long long protocol_bytes_written(void);
unsigned long long protocol_bytes_read(void);
void protocol_note_bytes_written(unsigned long long bytes);
/* Apply --bwlimit pacing to bytes written outside protocol_send_n_data (the
* plaintext zero-copy sendfile fast path). `file_descriptor` is the wire fd
* the bytes were written to, so the legacy session is resolved exactly as the
* preceding send_n_data call resolved it (the bound TLS session still wins when
* set); resolving with the same fd avoids re-initializing the legacy session
* and granting a second first-call burst. Runs the same token-bucket throttle,
* so the sendfile transport is paced identically to the buffered/TLS paths. A
* no-op when the effective session has no bandwidth limit. */
void protocol_throttle_bytes(int file_descriptor, size_t bytes);
void protocol_session_init(ProtocolSession* session, int read_fd, int write_fd);
/* Transitional bridge for helpers whose signatures still carry only an fd. */
@@ -260,6 +319,9 @@ bool protocol_send_int(ProtocolSession* session, int data);
bool protocol_receive_int(ProtocolSession* session, int* data);
bool protocol_send_status(ProtocolSession* session, Status status);
bool protocol_receive_status(ProtocolSession* session, Status* status);
/* As protocol_receive_status, but with an explicit per-message deadline
* (seconds) instead of the session's configured io_timeout_sec. */
bool protocol_receive_status_timed(ProtocolSession* session, Status* status, int timeout_sec);
bool send_n_data(int file_descriptor, const void* data, size_t data_size);
bool receive_n_data(int file_descriptor, void* data, size_t data_size);
+18 -1
View File
@@ -128,7 +128,7 @@ bool queue_enqueue_multithreaded_cancel(Queue* queue, void* item, mtx_t* mutex,
void* queue_dequeue(Queue* queue) {
if (queue == NULL || queue_is_empty(queue)) {
log_perror("ERROR: Could not dequeue from null or empty queue.");
log_message(LOG_LEVEL_ERROR, "%s", "ERROR: Could not dequeue from null or empty queue.");
return NULL;
}
@@ -139,6 +139,23 @@ void* queue_dequeue(Queue* queue) {
return item;
}
bool queue_push(Queue* queue, void* item) {
return queue_enqueue(queue, item);
}
void* queue_pop(Queue* queue) {
if (queue == NULL || queue_is_empty(queue)) {
log_message(LOG_LEVEL_ERROR, "%s", "ERROR: Could not pop from null or empty queue.");
return NULL;
}
queue->rear = (queue->rear - 1 + queue->capacity) % queue->capacity;
void* item = queue->items[queue->rear];
queue->items[queue->rear] = NULL;
queue->size--;
return item;
}
void* queue_dequeue_multithreaded(Queue* queue, mtx_t* mutex, cnd_t* condition_not_empty,
cnd_t* condition_not_full, const bool* other_thread_done) {
mtx_lock(mutex);
+7
View File
@@ -28,4 +28,11 @@ void* queue_dequeue(Queue* queue);
void* queue_dequeue_multithreaded(Queue* queue, mtx_t* mutex, cnd_t* condition_not_empty,
cnd_t* condition_not_full, const bool* other_thread_done);
/* LIFO stack operations over the same ring buffer. queue_push() is the enqueue
primitive; queue_pop() removes from the rear, so a sequence of pushes is
returned in reverse order. Used by the sequential scanner's depth-first
traversal. */
bool queue_push(Queue* queue, void* item);
void* queue_pop(Queue* queue);
#endif
+5 -7
View File
@@ -87,7 +87,7 @@ static int parse_remote_dest(const char* dest, RemoteDest* r) {
return 0;
}
char* ssh_build_remote_command(const char* server_path, bool old_args, char* const* remote_options,
char* ssh_build_remote_command(const char* server_path, char* const* remote_options,
int remote_option_count) {
const char* path = server_path ? server_path : "fastsync-server";
const char* suffix = " --stdio";
@@ -105,9 +105,7 @@ char* ssh_build_remote_command(const char* server_path, bool old_args, char* con
shell word (remote options below reuse the same escaping), then
" --stdio". Quoting the path is the only injection-safe construction: an
unquoted path would carry shell metacharacters straight into the remote
shell command. --old-args is kept for CLI/ABI compatibility but no longer
disables that protection. */
(void)old_args;
shell command. (rsync's --old-args no longer disables that protection.) */
size_t quote_count = 0;
for (const char* p = path; *p; p++)
if (*p == '\'')
@@ -296,8 +294,8 @@ void ssh_free_client_argv(char** argv) {
}
Client* client_connect_ssh(const char* destination, int port, const char* server_path,
bool old_args, const char* rsh_command, bool blocking_io,
char* const* remote_options, int remote_option_count) {
const char* rsh_command, bool blocking_io, char* const* remote_options,
int remote_option_count) {
RemoteDest r;
if (parse_remote_dest(destination, &r) != 0) {
char* escaped = output_escape(destination, false);
@@ -376,7 +374,7 @@ Client* client_connect_ssh(const char* destination, int port, const char* server
snprintf(ssh_user, ssh_user_len, "%s", r.host);
char* remote_command =
ssh_build_remote_command(server_path, old_args, remote_options, remote_option_count);
ssh_build_remote_command(server_path, remote_options, remote_option_count);
if (!remote_command)
ssh_child_setup_failed(exec_pipe[1]);
char** ssh_argv = ssh_build_client_argv(rsh_command, port, ssh_user, remote_command);
+8 -8
View File
@@ -4,17 +4,17 @@
#include "transport_tcp.h"
Client* client_connect_ssh(const char* destination, int port, const char* server_path,
bool old_args, const char* rsh_command, bool blocking_io,
char* const* remote_options, int remote_option_count);
const char* rsh_command, bool blocking_io, char* const* remote_options,
int remote_option_count);
/* Build the escaped remote-shell command string (the server program path always
* quoted as one remote-shell word, followed by ` --stdio` and each
* --remote-option value appended as an individually single-quoted shell word).
* `old_args` is accepted for CLI/ABI compatibility but no longer disables
* quoting: the path is always escaped so a metacharacter-bearing
* --rsync-path can never be interpreted by the remote shell. Every
* --remote-option value is individually escaped with the '\'' sequence and
* values with empty/control characters are rejected at the CLI parse layer. */
char* ssh_build_remote_command(const char* server_path, bool old_args, char* const* remote_options,
* The path is always escaped so a metacharacter-bearing --rsync-path can never
* be interpreted by the remote shell (the --old-args no-op does not disable
* quoting). Every --remote-option value is individually escaped with the '\''
* sequence and values with empty/control characters are rejected at the CLI
* parse layer. */
char* ssh_build_remote_command(const char* server_path, char* const* remote_options,
int remote_option_count);
/* Build the NULL-terminated child argv for the remote-shell client (argv[0] is
* the exec/execvp program). rsh_command is whitespace-split into leading argv
+77 -383
View File
@@ -2,10 +2,12 @@
#include "array_list.h"
#include "log.h"
#include <arpa/inet.h>
#include <ctype.h>
#include <dirent.h>
#include <errno.h>
#include <fcntl.h>
#include <netinet/in.h>
#include <stdarg.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
@@ -47,11 +49,45 @@ const char* utils_get_authorized_root_path(void) {
return authorized_root_path;
}
void utils_set_error(char* err, size_t err_size, const char* fmt, ...) {
if (!err || err_size == 0)
return;
va_list args;
va_start(args, fmt);
vsnprintf(err, err_size, fmt, args);
va_end(args);
}
bool path_is_within_root(const char* root, const char* path) {
size_t root_len = strlen(root);
return strncmp(root, path, root_len) == 0 && (path[root_len] == '\0' || path[root_len] == '/');
}
/* Borrowed transfer-relative view of `path`: strip any leading '/' and then a
* `root` prefix (its own leading/trailing slashes tolerated), returning a
* pointer into `path`. Non-allocating, so it is safe on the hot scan/print
* paths. A NULL/empty root, or a path not under `root`, leaves only the
* leading-slash strip. `path` must be NUL-terminated and live in the caller. */
const char* utils_strip_transfer_root(const char* path, const char* root) {
if (path == NULL)
return NULL;
const char* rel = path;
while (*rel == '/')
rel++;
if (root == NULL)
return rel;
while (*root == '/')
root++;
size_t root_len = strlen(root);
while (root_len > 0 && root[root_len - 1] == '/')
root_len--;
if (root_len == 0)
return rel;
if (strncmp(rel, root, root_len) == 0 && (rel[root_len] == '/' || rel[root_len] == '\0'))
return rel + root_len + (rel[root_len] == '/' ? 1 : 0);
return rel;
}
/* Open the destination root directory itself, confined to the authorized root.
* NOTE (do not merge with file_open_secure_parent): this walk opens dest_root
* (a directory that must already exist) and returns its fd, whereas
@@ -117,6 +153,47 @@ char* str_dup(const char* string) {
return new_string;
}
int env_choice_first(const char* env_name, int (*resolve)(const char*), bool* specified) {
if (specified)
*specified = false;
if (!env_name || !resolve)
return -1;
const char* env = getenv(env_name);
if (!env)
return -1;
bool saw_nonblank = false;
const char* p = env;
while (*p) {
if (*p == '&')
break;
if (isspace((unsigned char)*p)) {
p++;
continue;
}
saw_nonblank = true;
char token[64];
size_t len = 0;
while (*p && *p != '&' && !isspace((unsigned char)*p)) {
if (len < sizeof(token) - 1)
token[len++] = *p;
p++;
}
token[len] = '\0';
if (len > 0) {
int id = resolve(token);
if (id >= 0) {
if (specified)
*specified = true;
return id;
}
}
}
if (specified)
*specified = saw_nonblank;
return -1;
}
#define STR_HASH_SET_MIN_CAPACITY 16
static size_t str_hash_set_hash(const char* key, size_t len) {
@@ -548,389 +625,6 @@ bool format_human_bytes(unsigned long long bytes, char* buffer, size_t buffer_si
return written >= 0 && (size_t)written < buffer_size;
}
/* Build the keep-set index from the exact manifest entries only. A lookup of
`rel` succeeds iff `rel` is a kept entry, a kept directory, or an ancestor
directory of kept content (the old is_dir_in_manifest predicate); the sorted
view answers "is an ancestor of kept content" without materializing any
per-component prefix copy, so the index is O(manifest size) memory. */
static bool build_keep_index(const ArrayList* manifest, PathIndex* index) {
if (!manifest || manifest->size <= 0)
return path_index_build(index, NULL, 0);
return path_index_build(index, (const char* const*)manifest->items, (size_t)manifest->size);
}
static bool keep_is_dir(const PathIndex* index, const char* rel_path) {
return path_index_contains(index, rel_path) || path_index_has_descendant(index, rel_path);
}
static bool keep_is_file(const PathIndex* index, const char* rel_path) {
return path_index_contains(index, rel_path);
}
/* True when child_rel is, or lies below, a protected entry. A prefix "a"
therefore protects "a" and "a/b/c" but not "ab". Entries with top_level_only
set only protect DIRECT children of the receive root (at_root); nested
directories that share such a name stay ordinary destination content. */
bool path_under_skip_prefix(const char* child_rel, bool at_root, const DeleteSkipEntry* skips,
int skip_count) {
for (int i = 0; i < skip_count; i++) {
if (skips[i].top_level_only && !at_root)
continue;
size_t prefix_len = strlen(skips[i].prefix);
if (strncmp(child_rel, skips[i].prefix, prefix_len) == 0 &&
(child_rel[prefix_len] == '\0' || child_rel[prefix_len] == '/'))
return true;
}
return false;
}
/* Per-run deletion budget and tallies. `max_delete` is the cap on the number
of entries the walker may remove (SIZE_MAX = unlimited); once it is reached
the remaining extras are counted in `skipped` and left in place, matching
rsync's partial --max-delete behavior. */
typedef struct {
size_t max_delete;
size_t deleted;
size_t skipped;
bool limit_hit;
} DeleteBudget;
/* True when direct children of the directory named by `rel` may be removed.
With no synchronization info (dirs == NULL) the whole tree is deletable; when
a dirs index is supplied only its exact entries are (the receive root is the
"." sentinel). */
static bool is_synced_dir(const PathIndex* dirs, const char* rel) {
if (!dirs)
return true;
return path_index_contains(dirs, rel[0] == '\0' ? "." : rel);
}
/* Remove the extras directly inside the directory open on `dirfd`, recursing
into every child directory so kept content below a synchronized prefix is
reached. `all_removed` reports whether every child entry was removed (so the
caller may rmdir this directory). A child directory is never removed when it
is itself a synchronized directory or holds kept content; with a dirs index
supplied, direct children of a non-synchronized directory are never extras at
all (they are left in place but still descended into). Symlinks are unlinked
like any other non-directory extra (never followed). */
static bool delete_extras_fd(int dirfd, const char* rel_path, const PathIndex* keep,
const PathIndex* dirs, DeleteBudget* budget,
const DeleteSkipEntry* skips, int skip_count, bool parent_deletable,
bool* all_removed) {
/* openat(dirfd, ".") opens an independent file description: a dup() would
share dirfd's file offset and a prior pass could leave the stream drained. */
int scanfd = openat(dirfd, ".", O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
if (scanfd < 0)
return false;
DIR* dir = fdopendir(scanfd);
if (!dir) {
close(scanfd);
return false;
}
bool operation_ok = true;
bool local_survives = false;
/* A directory is deletable when it or ANY ancestor is synchronized; the
`parent_deletable` flag carries that down the recursion so dest-only
directories below a synchronized root are removed wholesale. */
bool deletable = parent_deletable || is_synced_dir(dirs, rel_path);
const struct dirent* entry;
while ((entry = readdir(dir)) != NULL) {
if (strcmp(entry->d_name, ".") == 0 || strcmp(entry->d_name, "..") == 0)
continue;
char* child_rel = path_cat((char*)rel_path, entry->d_name);
if (!child_rel) {
operation_ok = false;
continue;
}
/* A --delay-updates run keeps its staging directory as a direct child of
the receive root, and basis-dir snapshots live below it too. Their
contents are not manifest entries, so descending into them would delete
every staged / basis file as an "extra". Only the staging name (a
top-level-only prefix) and the basis prefixes are protected: a nested
destination directory that happens to be called .fastsync-stage is
ordinary content. */
if (path_under_skip_prefix(child_rel, rel_path[0] == '\0', skips, skip_count)) {
local_survives = true;
free(child_rel);
continue;
}
struct stat st;
if (fstatat(dirfd, entry->d_name, &st, AT_SYMLINK_NOFOLLOW) != 0) {
if (errno != ENOENT)
operation_ok = false;
free(child_rel);
continue;
}
if (S_ISDIR(st.st_mode)) {
int childfd = openat(dirfd, entry->d_name, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
bool child_all_removed = false;
if (childfd >= 0) {
if (!delete_extras_fd(childfd, child_rel, keep, dirs, budget, skips, skip_count, deletable,
&child_all_removed))
operation_ok = false;
close(childfd);
} else if (errno != ENOENT) {
operation_ok = false;
}
bool child_synced = dirs && path_index_contains(dirs, child_rel);
if (child_synced || keep_is_dir(keep, child_rel)) {
/* A synchronized directory and a directory holding kept content are
never removed. */
local_survives = true;
} else if (child_all_removed && deletable) {
if (budget->deleted >= budget->max_delete) {
budget->limit_hit = true;
budget->skipped++;
local_survives = true;
} else if (unlinkat(dirfd, entry->d_name, AT_REMOVEDIR) != 0) {
/* ENOENT: already gone (fine). ENOTEMPTY/EEXIST: the directory
still holds entries the walker leaves in place (a protected
excluded prefix, a kept file the manifest protects, a symlink);
rsync leaves such a directory behind, so this is not an error.
Only genuine I/O failures abort the deletion. */
if (errno != ENOENT && errno != ENOTEMPTY && errno != EEXIST)
operation_ok = false;
local_survives = true;
} else {
budget->deleted++;
}
} else {
local_survives = true;
}
} else {
bool found = keep_is_file(keep, child_rel);
if (found || !deletable) {
/* Kept file, or a child of a directory that is not synchronized: never
an extra for this run. */
local_survives = true;
} else if (budget->deleted >= budget->max_delete) {
budget->limit_hit = true;
budget->skipped++;
local_survives = true;
} else if (unlinkat(dirfd, entry->d_name, 0) != 0) {
if (errno != ENOENT)
operation_ok = false;
local_survives = true;
} else {
budget->deleted++;
char* escaped_path = output_escape(child_rel, log_get_8_bit_output());
fprintf(stderr, " Deleted: %s\n", escaped_path ? escaped_path : "<allocation failed>");
free(escaped_path);
}
}
free(child_rel);
}
closedir(dir);
*all_removed = !local_survives;
return operation_ok;
}
/* Read-only mirror of delete_extras_fd: records the paths that WOULD be removed
without unlinking anything. A child directory is reported after its own
reportable children (depth-first), matching the delete pass's ordering. */
static bool list_extras_fd(int dirfd, const char* rel_path, const PathIndex* keep,
const PathIndex* dirs, ArrayList* out, size_t* recorded,
const DeleteSkipEntry* skips, int skip_count, bool parent_deletable,
bool* all_removed) {
int scanfd = openat(dirfd, ".", O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
if (scanfd < 0)
return false;
DIR* dir = fdopendir(scanfd);
if (!dir) {
close(scanfd);
return false;
}
bool operation_ok = true;
bool local_survives = false;
bool deletable = parent_deletable || is_synced_dir(dirs, rel_path);
const struct dirent* entry;
while ((entry = readdir(dir)) != NULL) {
if (strcmp(entry->d_name, ".") == 0 || strcmp(entry->d_name, "..") == 0)
continue;
char* child_rel = path_cat((char*)rel_path, entry->d_name);
if (!child_rel) {
operation_ok = false;
continue;
}
if (path_under_skip_prefix(child_rel, rel_path[0] == '\0', skips, skip_count)) {
local_survives = true;
free(child_rel);
continue;
}
struct stat st;
if (fstatat(dirfd, entry->d_name, &st, AT_SYMLINK_NOFOLLOW) != 0) {
if (errno != ENOENT)
operation_ok = false;
free(child_rel);
continue;
}
if (S_ISDIR(st.st_mode)) {
int childfd = openat(dirfd, entry->d_name, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
bool child_all_removed = false;
if (childfd >= 0) {
if (!list_extras_fd(childfd, child_rel, keep, dirs, out, recorded, skips, skip_count,
deletable, &child_all_removed))
operation_ok = false;
close(childfd);
} else if (errno != ENOENT) {
operation_ok = false;
}
bool child_synced = dirs && path_index_contains(dirs, child_rel);
if (child_synced || keep_is_dir(keep, child_rel)) {
local_survives = true;
} else if (child_all_removed && deletable) {
size_t len = strlen(child_rel);
char* copy = malloc(len + 2);
if (!copy) {
operation_ok = false;
} else {
memcpy(copy, child_rel, len);
copy[len] = '/';
copy[len + 1] = '\0';
if (!array_list_add(out, copy)) {
free(copy);
operation_ok = false;
} else {
(*recorded)++;
}
}
} else {
local_survives = true;
}
} else {
bool found = keep_is_file(keep, child_rel);
if (found || !deletable) {
local_survives = true;
} else {
char* copy = str_dup(child_rel);
if (!copy || !array_list_add(out, copy)) {
free(copy);
operation_ok = false;
} else {
(*recorded)++;
}
}
}
free(child_rel);
}
closedir(dir);
*all_removed = !local_survives;
return operation_ok;
}
bool delete_extras_list(const char* dest_root, const ArrayList* manifest,
const ArrayList* synced_dirs, const DeleteSkipEntry* skips, int skip_count,
ArrayList* out, size_t* count_out) {
if (count_out)
*count_out = 0;
if (!manifest || !out)
return false;
PathIndex keep;
if (!build_keep_index(manifest, &keep))
return false;
PathIndex dirs;
bool have_dirs = synced_dirs != NULL;
if (have_dirs &&
!path_index_build(&dirs, (const char* const*)synced_dirs->items, (size_t)synced_dirs->size)) {
path_index_free(&keep);
return false;
}
int rootfd;
int root_fd = utils_get_authorized_root_fd();
if (root_fd >= 0) {
if (utils_get_authorized_root_path())
rootfd = utils_open_authorized_destination(dest_root);
else if (dest_root == NULL)
rootfd = dup(root_fd);
else
rootfd = -1;
} else {
rootfd = open(dest_root, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
}
if (rootfd < 0) {
path_index_free(&keep);
if (have_dirs)
path_index_free(&dirs);
return false;
}
bool all_removed = false;
size_t recorded = 0;
bool ok = list_extras_fd(rootfd, "", &keep, have_dirs ? &dirs : NULL, out, &recorded, skips,
skip_count, false, &all_removed);
if (close(rootfd) != 0)
ok = false;
path_index_free(&keep);
if (have_dirs)
path_index_free(&dirs);
if (count_out)
*count_out = recorded;
return ok;
}
DeleteWalkResult delete_extras_limited(const char* dest_root, const ArrayList* manifest,
const ArrayList* synced_dirs, size_t max_delete,
const DeleteSkipEntry* skips, int skip_count,
size_t* deleted_out, size_t* skipped_out) {
if (deleted_out)
*deleted_out = 0;
if (skipped_out)
*skipped_out = 0;
if (!manifest)
return DELETE_WALK_ERROR;
/* Index the keep-set (and the synchronized-dir set, when supplied) once so
membership is answered in O(path length) instead of scanning every entry
for every destination entry. */
PathIndex keep;
if (!build_keep_index(manifest, &keep))
return DELETE_WALK_ERROR;
PathIndex dirs;
bool have_dirs = synced_dirs != NULL;
if (have_dirs &&
!path_index_build(&dirs, (const char* const*)synced_dirs->items, (size_t)synced_dirs->size)) {
path_index_free(&keep);
return DELETE_WALK_ERROR;
}
int rootfd;
int root_fd = utils_get_authorized_root_fd();
if (root_fd >= 0) {
if (utils_get_authorized_root_path())
rootfd = utils_open_authorized_destination(dest_root);
else if (dest_root == NULL)
rootfd = dup(root_fd);
else
rootfd = -1;
} else {
rootfd = open(dest_root, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
}
if (rootfd < 0) {
path_index_free(&keep);
if (have_dirs)
path_index_free(&dirs);
return DELETE_WALK_ERROR;
}
DeleteBudget budget = {.max_delete = max_delete, .deleted = 0, .skipped = 0, .limit_hit = false};
bool all_removed = false;
bool ok = delete_extras_fd(rootfd, "", &keep, have_dirs ? &dirs : NULL, &budget, skips,
skip_count, false, &all_removed);
if (close(rootfd) != 0)
ok = false;
path_index_free(&keep);
if (have_dirs)
path_index_free(&dirs);
if (deleted_out)
*deleted_out = budget.deleted;
if (skipped_out)
*skipped_out = budget.skipped;
if (!ok)
return DELETE_WALK_ERROR;
return budget.limit_hit ? DELETE_WALK_LIMIT_REACHED : DELETE_WALK_OK;
}
bool delete_extras(const char* dest_root, const ArrayList* manifest) {
return delete_extras_limited(dest_root, manifest, NULL, SIZE_MAX, NULL, 0, NULL, NULL) ==
DELETE_WALK_OK;
}
bool has_path_traversal(const char* path) {
if (!path)
return true;
+23 -53
View File
@@ -2,6 +2,7 @@
#define UTILS_H
#include "array_list.h"
#include "filter.h"
#include <stddef.h>
#include <stdbool.h>
#include <stdio.h>
@@ -81,6 +82,17 @@ bool path_index_has_descendant(const PathIndex* index, const char* path);
char* str_dup(const char* string);
char* output_escape(const char* string, bool eight_bit_output);
/* Resolve the first supported name from a rsync algorithm-preference
* environment variable (RSYNC_COMPRESS_LIST / RSYNC_CHECKSUM_LIST). `resolve`
* maps a case-insensitive name to an algorithm id (>= 0) or -1 for an unknown
* name. rsync's syntax is a whitespace-separated list (comma/colon are NOT
* separators); the client-side half ends at '&'. Unknown entries are skipped
* and the first resolvable one wins. *specified is set true when the variable
* holds at least one non-blank character. Returns the first resolvable id, or
* -1 when the variable is unset/blank or names no supported algorithm. */
int env_choice_first(const char* env_name, int (*resolve)(const char*), bool* specified);
/* Upper bound on one line/token read from a local list file (--files-from,
* --exclude-from/--include-from, .rsync-filter). Mirrors MAX_STRING_SIZE and
* stops a hostile multi-gigabyte line from forcing unbounded allocation. */
@@ -94,59 +106,7 @@ char* output_escape(const char* string, bool eight_bit_output);
ssize_t utils_getdelim_bounded(FILE* stream, char** line, size_t* cap, int delim, size_t max_len);
char* path_cat(const char* path1, const char* path2);
bool glob_match(const char* pattern, const char* str);
/* Result of a bounded extra-file deletion run. */
typedef enum {
/* Every extra entry was removed (or there were none). */
DELETE_WALK_OK = 0,
/* The numeric cap for this run was reached before every extra was removed.
The walker removed exactly the entries the cap allowed and skipped (without
removing) the rest, matching rsync's partial --max-delete behavior. */
DELETE_WALK_LIMIT_REACHED,
/* A traversal or unlink failure aborted the deletion (partial removal is
possible, mirroring the delete pass). */
DELETE_WALK_ERROR
} DeleteWalkResult;
/* One protected entry for the delete walker. When top_level_only is true the
prefix is skipped only as a DIRECT child of dest_root (the --delay-updates
staging directory, which must not hide genuine extras inside a nested
destination directory that happens to share the staging name); otherwise the
prefix is skipped at any depth (the --compare-dest/--copy-dest/--link-dest
basis trees, and the sender-side protected filter-excluded prefixes, which
are never destination content). */
typedef struct {
const char* prefix;
bool top_level_only;
} DeleteSkipEntry;
/* True when child_rel is, or lies below, one of the protected entries (a prefix
"a" protects "a" and "a/b/c" but not "ab"; top_level_only entries protect
only DIRECT children of the destination root, i.e. child_rel has no '/'). */
bool path_under_skip_prefix(const char* child_rel, bool at_root, const DeleteSkipEntry* skips,
int skip_count);
/* Remove files/dirs/symlinks under dest_root that are not listed in manifest
without ever descending into a protected prefix (see DeleteSkipEntry). When
`synced_dirs` is non-NULL, extras are only removed directly inside a directory
whose destination-relative path is an exact entry in that list (the receive
root is the "." sentinel); directories outside the synchronized set are still
descended into so kept content below a listed directory is preserved, but
nothing in them is removed. A NULL `synced_dirs` keeps the legacy behavior of
treating the whole destination tree as deletable. `max_delete` caps the
number of removed entries (SIZE_MAX = unlimited): the walker removes up to the
cap and returns DELETE_WALK_LIMIT_REACHED when more extras remained.
`deleted_out`/`skipped_out` optionally receive the number of entries removed
and the number skipped because of the cap. */
DeleteWalkResult delete_extras_limited(const char* dest_root, const ArrayList* manifest,
const ArrayList* synced_dirs, size_t max_delete,
const DeleteSkipEntry* skips, int skip_count,
size_t* deleted_out, size_t* skipped_out);
/* Read-only companion to delete_extras_limited: walk the destination exactly as
the delete pass would and APPEND (strdup'd) destination-relative paths that
WOULD be removed, without touching disk. Used for -n/--dry-run --delete
would-delete reporting. Returns true on a clean walk; the caller owns the
strings appended to `out` and receives their count in *count_out. */
bool delete_extras_list(const char* dest_root, const ArrayList* manifest,
const ArrayList* synced_dirs, const DeleteSkipEntry* skips, int skip_count,
ArrayList* out, size_t* count_out);
bool delete_extras(const char* dest_root, const ArrayList* manifest);
/* Open the existing destination directory at `dest_root`, confined to the
authorized root with an O_NOFOLLOW component walk (the same confinement the
deletion walker uses for its root). Returns a new fd the caller owns, or -1
@@ -171,12 +131,22 @@ void utils_set_authorized_root_fd(int fd);
* threads spawn; see utils.c). */
int utils_get_authorized_root_fd(void);
const char* utils_get_authorized_root_path(void);
/* Write a diagnostic message into a caller-supplied buffer, mirroring
* vsnprintf. A NULL `err` or a zero `err_size` is a no-op, so a caller that
* only needs the boolean status may safely pass NULL. Returns nothing; the
* buffer is always NUL-terminated by vsnprintf when err_size > 0. */
void utils_set_error(char* err, size_t err_size, const char* fmt, ...);
/* True when `path` is `root` itself or lies directly beneath it: a lexical
* prefix test requiring the byte after `root` to be '\0' or '/'. Both `root`
* and `path` must be absolute canonical paths free of "."/".." components (the
* callers guarantee this); this is containment by string, not by resolved
* symlinks. Shared by the utils and file secure-walk root confinement. */
bool path_is_within_root(const char* root, const char* path);
/* Non-allocating transfer-relative view of `path`: strip any leading '/' and
* then a `root` prefix (leading/trailing slashes tolerated), returning a
* borrowed pointer into `path`. A NULL/empty root, or a path not under
* `root`, yields just the leading-slash strip. `path`/`root` must stay alive. */
const char* utils_strip_transfer_root(const char* path, const char* root);
/* True when `path` contains a ".." component. This is a purely lexical
* dot-dot check: an absolute path is NOT rejected here, because default
* (non-relative) transfers legitimately put the sender's absolute source path
+159 -38
View File
@@ -7,6 +7,7 @@
#include "utils.h"
#include "file_types.h"
#include <errno.h>
#include <limits.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
@@ -133,16 +134,23 @@ static bool xattr_name_is_posix_acl(const char* name) {
/* ---- SENDER: capture ---- */
FileXattrList* xattr_capture_path(const char* path, bool preserve_acls) {
/* The two syscall families differ only in whether the FINAL component is
* followed (`listxattr`/`getxattr` follow; `llistxattr`/`lgetxattr` do not), so
* one common implementation backs both public entry points. */
typedef ssize_t (*XattrListFn)(const char* path, char* list, size_t size);
typedef ssize_t (*XattrGetFn)(const char* path, const char* name, void* value, size_t size);
static FileXattrList* xattr_capture_common(const char* path, bool preserve_acls,
XattrListFn list_fn, XattrGetFn get_fn) {
if (!path)
return NULL;
ssize_t list_size = listxattr(path, NULL, 0);
ssize_t list_size = list_fn(path, NULL, 0);
if (list_size <= 0)
return NULL; /* no xattrs, ENOTSUP, or error: nothing appliable */
char* names = malloc((size_t)list_size);
if (!names)
return NULL;
ssize_t got = listxattr(path, names, (size_t)list_size);
ssize_t got = list_fn(path, names, (size_t)list_size);
if (got < 0) {
free(names);
return NULL;
@@ -165,7 +173,7 @@ FileXattrList* xattr_capture_path(const char* path, bool preserve_acls) {
negotiated. Without it a plain -X capture never carries an ACL. */
if (!xattr_name_appliable(name, preserve_acls))
continue;
ssize_t value_size = getxattr(path, name, NULL, 0);
ssize_t value_size = get_fn(path, name, NULL, 0);
if (value_size < 0)
continue;
if (value_size > XATTR_VALUE_MAX)
@@ -178,7 +186,7 @@ FileXattrList* xattr_capture_path(const char* path, bool preserve_acls) {
free(names);
return NULL;
}
ssize_t read_len = getxattr(path, name, buffer, (size_t)value_size);
ssize_t read_len = get_fn(path, name, buffer, (size_t)value_size);
if (read_len < 0 || read_len != value_size) {
free(buffer);
continue;
@@ -200,6 +208,14 @@ FileXattrList* xattr_capture_path(const char* path, bool preserve_acls) {
return list;
}
FileXattrList* xattr_capture_path(const char* path, bool preserve_acls) {
return xattr_capture_common(path, preserve_acls, listxattr, getxattr);
}
FileXattrList* xattr_capture_path_nofollow(const char* path, bool preserve_acls) {
return xattr_capture_common(path, preserve_acls, llistxattr, lgetxattr);
}
/* ---- WIRE ---- */
bool xattr_send(int fd, const FileXattrList* list) {
@@ -363,16 +379,74 @@ bool xattr_apply_fd(int fd, const FileXattrList* list) {
return true;
}
/* ---- --fake-super: park ownership/mode/mtime in a reserved xattr ---- */
/* Symlink counterpart of xattr_apply_fd(): target the link ITSELF, never its
* referent. fsetxattr cannot be used (no *at xattr syscall exists, and the
* kernel rejects xattr syscalls on an O_PATH descriptor), so the already-open,
* confinement-checked parent directory is addressed through /proc/self/fd and
* the final component is applied with lsetxattr, which does not follow it.
*
* The list is trusted to come from xattr_receive() (already whitelisted), but
* every name is re-validated here so this path-based primitive is confined on
* its own -- this is the only apply primitive that addresses a path, and the
* header promises a whitelisted apply. The apply is best-effort: if /proc is
* not mounted (the anchor cannot be formed) or the kernel refuses the set, the
* failure is skipped and never fails the transfer. See xattr.h for the bounded
* residual TOCTOU between link creation and lsetxattr. */
bool xattr_apply_path_nofollow(int parent_fd, const char* leaf, const FileXattrList* list,
bool preserve_acls) {
if (parent_fd < 0 || !leaf || leaf[0] == '\0' || strchr(leaf, '/') != NULL || !list)
return false;
if (list->count == 0)
return true;
char prefix[64];
int prefix_len = snprintf(prefix, sizeof(prefix), "/proc/self/fd/%d/", parent_fd);
if (prefix_len < 0 || (size_t)prefix_len >= sizeof(prefix))
return false;
size_t leaf_len = strlen(leaf);
char* path = malloc((size_t)prefix_len + leaf_len + 1);
if (!path)
return false;
memcpy(path, prefix, (size_t)prefix_len);
memcpy(path + prefix_len, leaf, leaf_len + 1);
bool warned = false;
int first_errno = 0;
for (int i = 0; i < list->count; i++) {
const FileXattr* xa = &list->items[i];
/* Defense in depth: re-validate against the receiver's full whitelist, so a
hand-crafted list can never apply a privileged namespace or the reserved
--fake-super key through this path-based primitive. */
if (!xattr_name_appliable(xa->name, preserve_acls))
continue;
if (lsetxattr(path, xa->name, xa->value, xa->value_len, 0) != 0) {
if (!warned) {
warned = true;
first_errno = errno;
}
}
}
if (warned)
log_message(LOG_LEVEL_WARNING,
"could not set one or more xattrs on the destination symlink: %s",
strerror(first_errno));
free(path);
return true;
}
void fake_super_store_fd(int fd, uint32_t uid, uint32_t gid, uint32_t mode, int64_t mtime_sec,
int64_t mtime_nsec) {
/* ---- --fake-super: park ownership/mode/rdev in a reserved xattr ---- */
void fake_super_store_fd(int fd, uint32_t uid, uint32_t gid, uint32_t mode, uint32_t rdev_major,
uint32_t rdev_minor) {
if (fd < 0)
return;
char record[128];
int len =
snprintf(record, sizeof(record), "%lu:%lu:%03o:%lld:%ld", (unsigned long)uid,
(unsigned long)gid, (unsigned)mode & 0777U, (long long)mtime_sec, (long)mtime_nsec);
/* rsync 3.4.1's exact grammar: "<octal full st_mode> <rdev_major>,<rdev_minor>
* <uid>:<gid>". The octal mode carries the S_IFMT bits (e.g. 0104711 for a
* setuid regular file, 020644 for a char device); the rdev pair is 0,0 for a
* non-device. No mtime field: rsync leaves the file's own timestamp in
* charge of mtime. This value is what rsync reads back to restore a
* fake-super tree, so the field order and separators must not change. */
char record[96];
int len = snprintf(record, sizeof(record), "%o %u,%u %u:%u", (unsigned)mode, (unsigned)rdev_major,
(unsigned)rdev_minor, (unsigned)uid, (unsigned)gid);
if (len <= 0 || (size_t)len >= sizeof(record))
return;
if (fsetxattr(fd, FAKESUPER_XATTR, record, (size_t)len, 0) != 0) {
@@ -381,11 +455,64 @@ void fake_super_store_fd(int fd, uint32_t uid, uint32_t gid, uint32_t mode, int6
}
}
/* --fake-super replay: read the freshly-stored record and re-apply mode/mtime
* fd-relative. The recorded uid/gid are retained for a later privileged
* restore but are NEVER chowned here: --fake-super only RECORDS ownership, it
* must not real-chown the recorded (resolved) owner. Mode/mtime still apply so
* unprivileged --fake-super keeps working. */
/* Parse rsync's `user.rsync.%stat` grammar strictly:
* "<octal st_mode> <rdev_major>,<rdev_minor> <uid>:<gid>"
* Every field is parsed with strtoul() so an out-of-range value is a clean
* rejection rather than the undefined behavior sscanf("%u") exhibited, each
* field is range-checked against the same bounds the wire validator uses, and
* the whole record must be consumed (only trailing whitespace is tolerated) so
* trailing garbage is refused. Returns false on any malformed input. */
static bool fake_super_parse_stat(const char* record, unsigned* mode_out, unsigned* rdev_major_out,
unsigned* rdev_minor_out, unsigned* uid_out, unsigned* gid_out) {
if (!record)
return false;
char* end = NULL;
const char* p = record;
errno = 0;
unsigned long mode = strtoul(p, &end, 8);
if (errno != 0 || end == p || mode > (unsigned long)UINT_MAX || *end != ' ')
return false;
p = end + 1;
errno = 0;
unsigned long rdev_major = strtoul(p, &end, 10);
if (errno != 0 || end == p || rdev_major > 0xffffUL || *end != ',')
return false;
p = end + 1;
errno = 0;
unsigned long rdev_minor = strtoul(p, &end, 10);
if (errno != 0 || end == p || rdev_minor > 0x00ffffffUL || *end != ' ')
return false;
p = end + 1;
errno = 0;
unsigned long uid = strtoul(p, &end, 10);
if (errno != 0 || end == p || uid > (unsigned long)UINT_MAX || *end != ':')
return false;
p = end + 1;
errno = 0;
unsigned long gid = strtoul(p, &end, 10);
if (errno != 0 || end == p || gid > (unsigned long)UINT_MAX)
return false;
p = end;
while (*p == ' ' || *p == '\t' || *p == '\n' || *p == '\r')
p++;
if (*p != '\0')
return false;
*mode_out = (unsigned)mode;
*rdev_major_out = (unsigned)rdev_major;
*rdev_minor_out = (unsigned)rdev_minor;
*uid_out = (unsigned)uid;
*gid_out = (unsigned)gid;
return true;
}
/* --fake-super replay: read the freshly-stored record and re-apply its
* permission bits fd-relative. The recorded uid/gid are retained for a later
* privileged restore but are NEVER chowned here: --fake-super only RECORDS
* ownership, it must not real-chown the recorded (resolved) owner. The
* recorded rdev is likewise parsed for grammar compatibility but is not acted
* on (device recreation is a separate, privilege-gated path). mtime is not in
* the record: the normal metadata path applies it (policy.times), exactly as
* rsync relies on the file's own timestamp. */
bool fake_super_restore_fd(int fd, FileAttrPolicy policy) {
if (fd < 0)
return false;
@@ -394,24 +521,25 @@ bool fake_super_restore_fd(int fd, FileAttrPolicy policy) {
if (len < 0)
return false; /* absent or filesystem without xattrs: silent no-op */
record[len] = '\0';
unsigned long ul_uid, ul_gid, ul_mode;
long long mtime_sec;
long mtime_nsec;
if (sscanf(record, "%lu:%lu:%lo:%lld:%ld", &ul_uid, &ul_gid, &ul_mode, &mtime_sec, &mtime_nsec) !=
5)
unsigned ul_mode, rdev_major, rdev_minor, ul_uid, ul_gid;
if (!fake_super_parse_stat(record, &ul_mode, &rdev_major, &rdev_minor, &ul_uid, &ul_gid))
return false; /* malformed record: skip, never fatal */
/* --fake-super NEVER performs a real chown: that would defeat the whole
point of the flag (record privileged ownership on an unprivileged receiver
for a later privileged restore). The uid/gid parsed above are retained in
the record for that later restore, but no ownership change happens here. */
/* --fake-super NEVER performs a real chown: that would defeat the whole point
of the flag (record privileged ownership on an unprivileged receiver for a
later privileged restore). The uid/gid parsed above are retained in the
record for that later restore, but no ownership change happens here. The
rdev is retained for the same reason. */
(void)rdev_major;
(void)rdev_minor;
(void)ul_uid;
(void)ul_gid;
/* Mode is applied only when the per-attribute policy asks for it, through the
SAME shared helper the normal metadata path uses (metadata_mode_for_policy):
under --perms the recorded source mode is copied exactly, including
group/other write and setuid/setgid/sticky bits (rsync parity), and the -E
rule derives exec bits from the destination's read bits exactly like
SAME shared helper the normal metadata path uses (metadata_mode_for_policy).
The recorded special bits are stripped first: rsync's fake-super receiver
stores the full mode in the xattr but never installs setuid/setgid/sticky on
the real file, so only the 0777 permission bits may be replayed. The -E
rule then derives exec bits from the destination's read bits exactly like
file_restore_metadata_fd. */
if (policy.perms || policy.executability) {
struct stat cur;
@@ -419,19 +547,12 @@ bool fake_super_restore_fd(int fd, FileAttrPolicy policy) {
if (fstat(fd, &cur) != 0) {
log_message(LOG_LEVEL_WARNING, "--fake-super: could not read destination mode: %s",
strerror(errno));
} else if (metadata_mode_for_policy((mode_t)ul_mode, cur.st_mode, policy, &want)) {
} else if (metadata_mode_for_policy((mode_t)(ul_mode & 0777U), cur.st_mode, policy, &want)) {
if (fchmod(fd, want) != 0)
log_message(LOG_LEVEL_WARNING,
"--fake-super: could not restore mode on destination file: %s",
strerror(errno));
}
}
if (policy.times) {
struct timespec times[2] = {{.tv_sec = 0, .tv_nsec = UTIME_OMIT},
{.tv_sec = (time_t)mtime_sec, .tv_nsec = mtime_nsec}};
if (futimens(fd, times) != 0)
log_message(LOG_LEVEL_WARNING,
"--fake-super: could not restore mtime on destination file: %s", strerror(errno));
}
return true;
}
+73 -23
View File
@@ -26,16 +26,22 @@
* and total bytes) on BOTH ends to prevent OOM/memory abuse; an oversized
* or malformed frame is a clean protocol rejection, never an allocation
* blowup.
* * Application is confined to the exact destination file descriptor
* (fsetxattr on the just-written fd), never a caller-controlled path.
* * Application is confined to the exact destination entry: fsetxattr on the
* just-written fd for regular files/directories, and for a symlink an
* lsetxattr on "/proc/self/fd/<parent_fd>/<leaf>" reached through the
* already-opened, confinement-checked parent directory -- never a
* caller-controlled path, and never following the link.
*/
/* Reserved key used by --fake-super to park the source's privileged ownership
* / mode / mtime on the destination file as an unprivileged user.* xattr, so a
* later privileged restore could re-apply them. Exact documented format:
* uid:gid:mode:mtime_sec:mtime_nsec (decimal, decimal, octal, dec, dec)
* e.g. "1000:1000:644:1765238400:0". */
#define FAKESUPER_XATTR "user.fastsync.stat"
* / mode / rdev on the destination file as an unprivileged user.* xattr, so the
* tree is interoperable with rsync 3.4.1 and a later privileged restore can
* re-apply them. This is rsync's own key and value grammar exactly:
* <octal st_mode with S_IFMT> <rdev_major>,<rdev_minor> <uid>:<gid>
* e.g. "104711 0,0 1234:5678" for a setuid regular file owned by 1234:5678,
* or "20644 1,3 111:222" for a char device. mtime is deliberately NOT part of
* the record: exactly like rsync, the file's own timestamp carries it. */
#define FAKESUPER_XATTR "user.rsync.%stat"
/* --- bounds --- */
#define XATTR_NAME_MAX 255 /* xattr names are limited to 255 bytes */
@@ -76,6 +82,17 @@ bool xattr_name_appliable(const char* name, bool preserve_acls);
* distinct from NULL. */
FileXattrList* xattr_capture_path(const char* path, bool preserve_acls);
/* Sender: like xattr_capture_path() but reads the xattrs of `path` ITSELF,
* never following a final symlink (llistxattr/lgetxattr). A symlink entry must
* use this so the scanner never captures the REFERENT's attributes onto the
* link (the path-following variant would). On Linux the VFS refuses to
* associate xattrs with symlinks at all, so this normally returns NULL; it is
* still correct and portable for a filesystem/platform that supports them.
* The same whitelist/bounds as xattr_capture_path() apply. Returns NULL when
* the link has no appliable xattrs (or the filesystem does not support them);
* an empty-but-valid list is never returned distinct from NULL. */
FileXattrList* xattr_capture_path_nofollow(const char* path, bool preserve_acls);
/* Wire: bounded serialization. xattr_send returns false on write failure; an
* empty/NULL list transmits a zero-count block. xattr_receive returns NULL and
* sets *ok = 0 on any malformed / oversized / non-whitelisted entry. When
@@ -91,24 +108,57 @@ FileXattrList* xattr_receive(int fd, int* ok, bool preserve_acls);
* true when apply was attempted (allowing callers to treat it as best-effort). */
bool xattr_apply_fd(int fd, const FileXattrList* list);
/* --fake-super: write the source uid/gid/mode/mtime record into the reserved
* FAKESUPER_XATTR on `fd`. Best-effort (logged, never fatal). Only meaningful
* when metadata was transmitted so the values exist. */
void fake_super_store_fd(int fd, uint32_t uid, uint32_t gid, uint32_t mode, int64_t mtime_sec,
int64_t mtime_nsec);
/* Receiver: apply every entry to the symlink named by (parent_fd, leaf) WITHOUT
* following it, via lsetxattr() on the confined path
* "/proc/self/fd/<parent_fd>/<leaf>". Every incoming name is independently
* re-validated against xattr_name_appliable() with `preserve_acls`, exactly like
* xattr_apply_fd(): a non-whitelisted namespace (including the reserved
* --fake-super key) is skipped, so this primitive stays confined even if handed
* a hand-crafted list. A symlink cannot be targeted by the fd-relative
* fsetxattr() path: there is no *at() xattr syscall and the kernel rejects
* xattr syscalls on an O_PATH descriptor, so the already-opened,
* confinement-checked parent directory is the anchor and only the final
* component is the (no-follow) link. `leaf` must be a single path component.
*
* Portability: the "/proc/self/fd/<parent_fd>" anchor requires a mounted /proc.
* Where /proc is unavailable (or the fd cannot be addressed that way) the
* lsetxattr simply fails and is skipped -- the apply is best-effort exactly like
* xattr_apply_fd(), so no error is propagated and the transfer continues. A
* per-attribute failure (on Linux every set on a symlink fails with EPERM) is
* logged once and skipped, never fatal. Returns false only for an invalid
* anchor/list; true when an apply was attempted.
*
* Residual TOCTOU: `leaf` is a caller-supplied name resolved by path in the
* parent, so a local writer could replace the just-created symlink between its
* creation and lsetxattr(). This is bounded: it requires write access to the
* confinement-checked destination directory (already trusted), can only install
* a whitelisted user namespace or POSIX-ACL name, and never follows the link (a
* replacement symlink is still applied to as the final, no-follow component). */
bool xattr_apply_path_nofollow(int parent_fd, const char* leaf, const FileXattrList* list,
bool preserve_acls);
/* --fake-super: write the source uid/gid/mode/rdev record into the reserved
* FAKESUPER_XATTR on `fd`, using rsync 3.4.1's exact grammar (see the key
* comment above). `mode` is the full st_mode including its S_IFMT bits.
* Best-effort (logged, never fatal). Only meaningful when metadata was
* transmitted so the values exist. */
void fake_super_store_fd(int fd, uint32_t uid, uint32_t gid, uint32_t mode, uint32_t rdev_major,
uint32_t rdev_minor);
/* --fake-super replay: parse the FAKESUPER_XATTR record previously written on
* `fd` by fake_super_store_fd and re-apply mode/mtime fd-relative. The
* recorded uid/gid are deliberately NOT chowned for real: --fake-super only
* RECORDS ownership (the caller stores the resolved mapping via
* identity_resolve_storage_ids), it never performs a real chown. Best-effort:
* absence of the xattr or a malformed record is a silent no-op that never fails
* the transfer. The MODE leg is applied only when policy.perms||policy.
* executability and the MTIME leg only when policy.times, so the fake-super
* replay cannot bypass the per-attribute split; the mode follows the normal
* metadata path exactly (under --perms the source mode is copied verbatim,
* special and group/other write bits included).
* Returns true when the xattr was present and parsed. */
* `fd` by fake_super_store_fd and re-apply the recorded permission bits
* fd-relative. The recorded uid/gid are deliberately NOT chowned for real:
* --fake-super only RECORDS ownership (the caller stores the resolved mapping
* via identity_resolve_storage_ids), it never performs a real chown. The
* recorded rdev is retained for a later privileged restore but is not acted on
* here. Best-effort: absence of the xattr or a malformed record is a silent
* no-op that never fails the transfer. The MODE leg is applied only when
* policy.perms||policy.executability, and the recorded special bits
* (setuid/setgid/sticky) are NOT applied to the real file -- exactly like
* rsync's fake-super receiver, which stores the full mode in the xattr but
* strips the special bits on disk. mtime is not part of the record; the normal
* metadata path carries it (policy.times) exactly as rsync sets the file's own
* timestamp. Returns true when the xattr was present and parsed. */
bool fake_super_restore_fd(int fd, FileAttrPolicy policy);
#endif
+7
View File
@@ -119,6 +119,13 @@ static void build_canonical_frame(void) {
cfg->usermap[0].to = MAP_TO;
cfg->usermap[0].to_name = NULL;
}
/* Force a non-empty receiver delete-protection block so the fuzzer mutates
* its rule count, action/sides codes and pattern strings. */
cfg->filters = array_list_create(free);
if (cfg->filters) {
array_list_add(cfg->filters, str_dup("P *.log"));
array_list_add(cfg->filters, str_dup("+r keep/**"));
}
if (!cfg->send_directory || !cfg->receive_root_directory || !cfg->usermap) {
config_delete(cfg);
return;
+80
View File
@@ -0,0 +1,80 @@
# Integration tests
The integration suite drives the built `build/server` and `build/client`
against local corpora. Unit tests live in `tests/` (the custom C framework);
the Python suite here covers the full transfer pipeline, transports, features,
and rsync parity.
## Running
```bash
# Full suite (excludes privilege-dependent tests on CI runners)
python3 -m pytest tests/integration/ -n 4 --dist=load -m "not setpriv"
# Fast PR subset only
python3 -m pytest tests/integration/ -n 4 --dist=load -m ci
```
The tests expect `build/server` and `build/client` (configure/build with CMake
first); `common.py` derives `BUILD_DIR` from the repository root.
## Differential rsync-parity gate
`test_differential_parity.py` runs the **same** transfer with real
`rsync 3.4.1` and with FastSync over separate destinations, then compares:
- the destination trees — relative paths, file content hashes, symlink
targets, modes (where the case is about perms), and hard-link grouping;
- the normalized stdout for output-oriented flags (`-i`,
`--out-format=...`, `--stats`), after stripping volatile fields
(timings, rates, wire byte counts) and directory-only itemize lines that
FastSync's recursive scanner documents as absent.
FastSync mirrors the absolute source path under its receive root (see
`get_dest_received_dir`); the harness normalizes that layout (and the
`-R`/`--files-from` layouts) before comparing.
```bash
# Fast subset that guards the ✅ surface on pull requests
python3 -m pytest tests/integration/test_differential_parity.py -n 4 --dist=load -m parity_ci
# Full differential case table (`_CASES`): every row is marked `parity`, and a
# case with an allowlisted residual in `parity_caveats.py` is included too.
python3 -m pytest tests/integration/test_differential_parity.py -n 4 --dist=load -m parity
```
`-m parity` selects only the `_CASES` table in this module. Differential
coverage for options outside that table (`--temp-dir`, `--delay-updates`,
`--dry-run`, `--fuzzy`, the basis-dir options, `-M` over daemon/TCP, and
receiver filter-protect) lives in dedicated modules (`test_option_parity.py`,
`test_parity_blockers.py`, `test_parity_quickwins.py`, ...) and is not part of
this gate. The suite skips cleanly when `rsync` is not installed.
## Allowlist (`parity_caveats.py`)
`parity_caveats.py` is the single data-driven allowlist of known differences.
Each entry maps a case id to the aspects that may differ (`tree`, `stdout`,
`extra`, `rc`) and cites the governing row in `RSYNC_COMPAT.md`:
```python
CAVEATS = {
# no known residuals at present -- the burn-down reached zero
# "some_case_id": {"tree": "documented residual ... ref: RSYNC_COMPAT.md ..."},
}
```
A differential mismatch in an aspect that is **not** listed fails the gate with
a readable tree/stdout diff.
If a case is allowlisted but now matches rsync, the gate emits a loud warning
naming the stale entry — that is the parity burn-down signal. Run with
`FASTSYNC_PARITY_STRICT=1` to make stale entries fail instead (the full CI
parity job sets this). To add a residual:
1. Reproduce it with `-m parity` and read the failure's tree/stdout diff.
2. Confirm it is a documented `⚠️`/`❌` residual (or get the `✅` row
reclassified) and cite the row.
3. Add the case id and aspect(s) to `CAVEATS`, keeping the reason concise.
Do not allowlist an undocumented divergence from a `✅` row — fix it or get the
row reclassified first.
+55 -12
View File
@@ -20,6 +20,12 @@ CLIENT_CMD = [os.path.join(BUILD_DIR, "client")]
_WORKER = os.environ.get("PYTEST_XDIST_WORKER")
TEST_DATA_DIR = os.path.join(PROJECT_ROOT, f"test_data-{_WORKER}" if _WORKER else "test_data")
# Default wall-clock budget for a short-lived client invocation. Every client
# is expected to finish well within this; the bound exists so a hung client
# fails the test instead of stalling the whole CI run indefinitely. Callers
# that legitimately need longer can pass an explicit ``timeout``.
CLIENT_TIMEOUT = 180
class ServerManager:
"""Manages a long-lived server process. Reuses across test cases."""
@@ -137,7 +143,40 @@ class CountingProxy:
return result
def run_client(source_dir, dest_dir, flags=None, port=None, extra_args=None):
def _run_client_cmd(cmd, timeout):
"""Run one client command, returning ``(result, duration)``.
On timeout the client is killed and a result-like ``CompletedProcess`` with
a non-zero returncode is returned instead of raising, so callers keep the
established ``(result, duration)`` contract and the failure carries the
command plus whatever output was captured for diagnosis.
"""
start = time.monotonic()
try:
result = subprocess.run(cmd, text=True, capture_output=True, timeout=timeout)
except subprocess.TimeoutExpired as exc:
duration = time.monotonic() - start
stdout = exc.stdout or ""
stderr = exc.stderr or ""
if isinstance(stdout, bytes):
stdout = stdout.decode(errors="replace")
if isinstance(stderr, bytes):
stderr = stderr.decode(errors="replace")
diagnostic = (
f"client timed out after {timeout}s\n"
f"command: {cmd!r}\n"
f"--- captured stdout ---\n{stdout}\n"
f"--- captured stderr ---\n{stderr}"
)
result = subprocess.CompletedProcess(cmd, returncode=-1,
stdout=stdout, stderr=diagnostic)
return result, duration
duration = time.monotonic() - start
return result, duration
def run_client(source_dir, dest_dir, flags=None, port=None, extra_args=None,
timeout=CLIENT_TIMEOUT):
"""Run the client and return (result, duration)."""
cmd = CLIENT_CMD + ["--source-dir", source_dir, "--dest-dir", dest_dir, "--save-to-disk"]
if port:
@@ -146,23 +185,18 @@ def run_client(source_dir, dest_dir, flags=None, port=None, extra_args=None):
cmd += flags
if extra_args:
cmd += extra_args
start = time.monotonic()
result = subprocess.run(cmd, text=True, capture_output=True)
duration = time.monotonic() - start
return result, duration
return _run_client_cmd(cmd, timeout)
def run_client_posix(source_dir, dest_dir, flags=None, port=None):
def run_client_posix(source_dir, dest_dir, flags=None, port=None,
timeout=CLIENT_TIMEOUT):
"""Run the client with positional args (rsync-style)."""
cmd = CLIENT_CMD + [source_dir, dest_dir, "--save-to-disk"]
if port:
cmd += ["--server-port", str(port)]
if flags:
cmd += flags
start = time.monotonic()
result = subprocess.run(cmd, text=True, capture_output=True)
duration = time.monotonic() - start
return result, duration
return _run_client_cmd(cmd, timeout)
def generate_test_files(source_dir, full=False):
@@ -238,8 +272,17 @@ def make_result(name, success, duration=None, error=""):
def get_dest_received_dir(dest_dir, source_dir):
"""Get the path where received files land inside dest_dir."""
return os.path.join(dest_dir, os.path.abspath(source_dir).lstrip(os.sep))
"""Get the path where received files land inside dest_dir.
FastSync mirrors the absolute source path below the receive root with the
leading root separator removed. Strip that separator explicitly rather
than with ``str.lstrip(os.sep)``: ``lstrip`` removes a *set* of characters
rather than a path prefix, which is not the same operation.
"""
abs_source = os.path.abspath(source_dir)
if abs_source.startswith(os.sep):
abs_source = abs_source[len(os.sep):]
return os.path.join(dest_dir, abs_source)
def _find_free_port():
+34
View File
@@ -0,0 +1,34 @@
"""Data-driven allowlist for the differential rsync-parity gate.
Every entry maps a case id (see ``test_differential_parity.py``) to the aspects
that are *known* to differ from ``rsync 3.4.1`` and the documented reason. A
differential mismatch in an aspect that is **not** listed here fails the gate.
Aspect keys
-----------
``tree`` destination tree differs (paths, file hashes, symlink targets,
modes, hardlink grouping)
``stdout`` normalized output for ``-i`` / ``--stats`` / ``--out-format``
``extra`` a case-specific assertion differs (basis/inode checks, ...)
``rc`` exit status differs
Burn-down
---------
If a case is listed here but now matches rsync, the gate emits a loud
``pytest`` warning naming the stale entry: delete the entry (and, when the
underlying row in ``RSYNC_COMPAT.md`` is now parity, update that row). Set
``FASTSYNC_PARITY_STRICT=1`` to turn stale entries into failures in CI.
Keep the values concise but cite the governing row so the entry can be
re-triaged when the row moves.
"""
# case id -> {aspect: "reason (ref: RSYNC_COMPAT.md ...)"}
CAVEATS = {}
# Accepted aspect names (guards against typos in this file).
ASPECTS = ("tree", "stdout", "extra", "rc")
def caveat_for(case_id: str) -> dict:
return CAVEATS.get(case_id, {})
+584
View File
@@ -0,0 +1,584 @@
"""Differential rsync-parity harness.
Runs the SAME transfer with real ``rsync`` and with FastSync over separate
destinations and compares the resulting trees and (optionally) normalized
stdout. ``test_differential_parity.py`` drives this module with a table of
cases; ``parity_caveats.py`` is the data-driven allowlist of documented
residuals.
Design notes
------------
FastSync mirrors the *absolute* source path below its receive root, while
rsync copies the source contents directly into the destination. ``Case.layout``
tells the harness which pair of directory roots to compare:
* ``MIRROR`` -- rsync ``DEST/`` vs FastSync ``DEST/<abs-src>/`` (the common
case; matches ``common.get_dest_received_dir``).
* ``MIRROR_ABS`` -- ``rsync -R`` without a cut lays the full absolute path
under the destination, so rsync ``DEST/<abs-src>/`` is compared against the
same FastSync mirror path.
* ``RELATIVE`` -- ``rsync -R --files-from`` lays bare relative paths under the
destination and FastSync does the same, so both destination roots compare
directly.
Only ``tests/integration/common.py`` is used to reach the build products and the
server manager; the harness never duplicates that plumbing.
"""
import difflib
import hashlib
import os
import re
import shutil
import subprocess
import sys
from dataclasses import dataclass
from typing import Callable, Dict, List, Optional, Tuple
sys.path.insert(0, os.path.dirname(__file__))
from common import ( # noqa: E402 (path bootstrap above)
TEST_DATA_DIR,
clean_dir,
get_dest_received_dir,
run_client,
)
RSYNC = shutil.which("rsync")
# Comparison layouts (see module docstring).
MIRROR = "mirror"
MIRROR_ABS = "mirror_abs"
RELATIVE = "relative"
# stdout comparators.
STDOUT_NONE = None
STDOUT_ITEMIZE = "itemize"
STDOUT_OUTFMT = "outfmt"
STDOUT_STATS = "stats"
STDOUT_PROGRESS = "progress"
# rsync --stats lines that are protocol-independent and must match exactly.
# `Number of files` and `Number of created files` carry rsync's per-type
# breakdown; protocol 2.28.0 reports the receiver-created split over
# STATUS_STATS. Deliberately excluded: Total bytes sent/received (protocol
# framing differs, see the `--stats` row in RSYNC_COMPAT.md).
STATS_KEYS = (
"Number of files",
"Number of created files",
"Number of deleted files",
"Number of regular files transferred",
"Total file size",
"Total transferred file size",
"Literal data",
"Matched data",
"File list size",
)
_ITEMIZE_RE = re.compile(r"^(<|>|c|h|\.|\*)[fdLDS][.+\-][.+\-][.+\-][.+\-]")
_PROGRESS_TOTAL_RE = re.compile(r"to-chk=\d+/(\d+)")
_PROGRESS_XFR_RE = re.compile(r"xfr#(\d+)")
@dataclass
class Case:
"""One differential scenario: a corpus, a flag set, and how to compare."""
id: str
corpus: str
flags: List[str]
fastsync_flags: Optional[List[str]] = None
layout: str = MIRROR
server_args: Tuple[str, ...] = ("--allow-super",)
seed: Optional[Callable] = None
stdout: Optional[str] = STDOUT_NONE
compare_modes: bool = False
compare_hardlinks: bool = False
ignore_paths: Tuple[str, ...] = ()
extra_check: Optional[Callable] = None
files_from: Optional[Tuple[str, ...]] = None
# rsync receives ``src + "/"``; FastSync mirrors the path it is given, so a
# trailing-slash-sensitive case must hand FastSync the same form.
fs_src_suffix: str = ""
# Some cases have an unspecified result (e.g. which extras survive a
# partial --max-delete abort): assert the case-specific invariants via
# extra_check and skip the exact-tree comparison.
compare_tree: bool = True
ci: bool = False
ref: str = ""
def fs_flags(self) -> List[str]:
return list(self.flags if self.fastsync_flags is None else self.fastsync_flags)
# ---------------------------------------------------------------------------
# Corpora
# ---------------------------------------------------------------------------
# Deterministic mtimes so quick-check decisions are reproducible.
_SRC_MTIME = 1_600_000_000
def _write(path: str, data: bytes, mode: Optional[int] = None) -> None:
os.makedirs(os.path.dirname(path), exist_ok=True)
with open(path, "wb") as fh:
fh.write(data)
os.utime(path, (_SRC_MTIME, _SRC_MTIME))
if mode is not None:
os.chmod(path, mode)
def _set_mode(path: str, mode: int) -> None:
os.chmod(path, mode)
def corpus_basic(root: str) -> None:
"""Regular files + nested dirs (dirs are implied by their files)."""
clean_dir(root)
_write(os.path.join(root, "a.txt"), b"hello world\n")
_write(os.path.join(root, "sub", "b.bin"),
bytes((i * 7) & 0xFF for i in range(5000)))
_write(os.path.join(root, "sub", "deep", "c.txt"), "w\u00f6rld\n".encode())
def corpus_unicode(root: str) -> None:
clean_dir(root)
_write(os.path.join(root, "uni \u00f1\u6587.txt"), b"unicode\n")
_write(os.path.join(root, "sub", "sp ace \u00e9.dat"), b"spaced\n")
_set_mode(os.path.join(root, "sub"), 0o750)
def corpus_links(root: str) -> None:
corpus_basic(root)
os.symlink("a.txt", os.path.join(root, "rel_link"))
os.symlink("/etc/hostname", os.path.join(root, "abs_link"))
os.symlink("nowhere/target", os.path.join(root, "broken_link"))
def corpus_hardlinks(root: str) -> None:
clean_dir(root)
_write(os.path.join(root, "h1.txt"), b"hardlinked payload\n")
os.link(os.path.join(root, "h1.txt"), os.path.join(root, "h2.txt"))
_write(os.path.join(root, "other.txt"), b"other\n")
def corpus_sparse(root: str) -> None:
clean_dir(root)
_write(os.path.join(root, "small.txt"), b"small\n")
sparse = os.path.join(root, "sparse.bin")
with open(sparse, "wb") as fh:
fh.seek(1024 * 1024 - 1)
fh.write(b"\0")
os.utime(sparse, (_SRC_MTIME, _SRC_MTIME))
def corpus_filters(root: str) -> None:
clean_dir(root)
_write(os.path.join(root, "keep.txt"), b"keep\n")
_write(os.path.join(root, "drop.log"), b"log\n")
_write(os.path.join(root, "sub", "keep2.txt"), b"keep2\n")
_write(os.path.join(root, "sub", "drop2.log"), b"log2\n")
_write(os.path.join(root, "sub", "data.bin"), b"bin\n")
def corpus_empty_dir(root: str) -> None:
clean_dir(root)
_write(os.path.join(root, "keep.txt"), b"keep\n")
os.makedirs(os.path.join(root, "emptydir"), exist_ok=True)
os.utime(os.path.join(root, "emptydir"), (_SRC_MTIME, _SRC_MTIME))
_write(os.path.join(root, "nonempty", "f.txt"), b"f\n")
def corpus_multidir(root: str) -> None:
"""Multi-directory tree for the --progress file-list naming/denominator.
Nested files, a directory-only branch, an empty directory and a symlink
exercise every file-list entry type rsync counts in `to-chk` but FastSync's
streaming scanner never emits as a transfer entry.
"""
clean_dir(root)
_write(os.path.join(root, "a.txt"), b"alpha\n")
_write(os.path.join(root, "b.txt"), b"bravo\n")
_write(os.path.join(root, "sub1", "c.txt"), b"charlie\n")
_write(os.path.join(root, "sub1", "deep", "d.txt"), b"delta\n")
_write(os.path.join(root, "sub2", "e.txt"), b"echo\n")
os.symlink("a.txt", os.path.join(root, "link1"))
os.makedirs(os.path.join(root, "emptydir"), exist_ok=True)
os.utime(os.path.join(root, "emptydir"), (_SRC_MTIME, _SRC_MTIME))
def corpus_relative(root: str) -> None:
"""Tree for the -R/--files-from cases."""
clean_dir(root)
_write(os.path.join(root, "a.txt"), b"a\n")
_write(os.path.join(root, "b.txt"), b"b\n")
_write(os.path.join(root, "sub", "x.txt"), b"x\n")
_write(os.path.join(root, "sub", "y.txt"), b"y\n")
os.makedirs(os.path.join(root, "dir1"), exist_ok=True)
os.utime(os.path.join(root, "dir1"), (_SRC_MTIME, _SRC_MTIME))
_write(os.path.join(root, "dir1", "keep.txt"), b"keep\n")
def corpus_iconv(root: str) -> None:
"""Latin-1 (ISO-8859-1) encoded filenames, matching the --iconv direction."""
clean_dir(root)
for rel, data in ((b"caf\xe9.txt", b"caf\xe9\n"),
(os.path.join(b"sub", b"\xfcber.txt"), b"\xfcber\n")):
full = os.path.join(os.fsencode(root), rel)
os.makedirs(os.path.dirname(full), exist_ok=True)
with open(full, "wb") as fh:
fh.write(data)
os.utime(full, (_SRC_MTIME, _SRC_MTIME))
# Payload for the --fuzzy basis corpus: large enough for the delta engine's
# 16 KiB minimum and with repeated content so a coinciding basis yields a
# non-zero (and identical) Matched data count in both tools.
FUZZY_PAYLOAD = (b"the quick brown fox jumps over the lazy dog\n" * 2000)[:65536]
def corpus_fuzzy(root: str) -> None:
"""A named regular file; the `fuzzy` seed adds the similar-suffix sibling."""
clean_dir(root)
_write(os.path.join(root, "report_v2.txt"), FUZZY_PAYLOAD)
CORPORA: Dict[str, Callable[[str], None]] = {
"basic": corpus_basic,
"unicode": corpus_unicode,
"links": corpus_links,
"hardlinks": corpus_hardlinks,
"sparse": corpus_sparse,
"filters": corpus_filters,
"empty_dir": corpus_empty_dir,
"multidir": corpus_multidir,
"relative": corpus_relative,
"iconv": corpus_iconv,
"fuzzy": corpus_fuzzy,
}
# ---------------------------------------------------------------------------
# Tree snapshotting / comparison
# ---------------------------------------------------------------------------
def snapshot(root: str, compare_modes: bool = False) -> Dict[str, tuple]:
"""Map relative path -> descriptor for every entry below ``root``.
Files hash their contents with SHA-256 (structural comparison, so differing
quick-check metadata cannot mask a payload difference). Symlinks record
their target. Empty directories are included (as ``("dir", ...)``) so the
recursive-empty-directory residual is observable.
"""
out: Dict[str, tuple] = {}
if not os.path.isdir(root):
return out
def describe(path: str) -> Optional[tuple]:
st = os.lstat(path)
if os.path.islink(path):
return ("link", os.readlink(path))
if os.path.isdir(path):
mode = oct(st.st_mode & 0o7777) if compare_modes else None
return ("dir", mode)
h = hashlib.sha256()
with open(path, "rb") as fh:
for chunk in iter(lambda: fh.read(65536), b""):
h.update(chunk)
mode = oct(st.st_mode & 0o7777) if compare_modes else None
return ("file", h.hexdigest()[:16], mode)
# The comparison root itself is not part of the tree diff: a no-transfer
# result legitimately leaves FastSync's mirror directory absent while rsync
# leaves an existing (empty) destination root.
for dirpath, dirnames, filenames in os.walk(root, followlinks=False):
dirnames.sort()
for name in sorted(dirnames):
p = os.path.join(dirpath, name)
rel = os.path.relpath(p, root)
if os.path.islink(p):
out[rel] = ("link", os.readlink(p))
dirnames.remove(name)
else:
out[rel] = describe(p)
for name in sorted(filenames):
p = os.path.join(dirpath, name)
out[os.path.relpath(p, root)] = describe(p)
return out
def _hardlink_groups(root: str) -> Dict[str, str]:
"""Assign a stable group letter to each inode shared by >1 regular file."""
inodes: Dict[tuple, List[str]] = {}
for dirpath, _dirs, filenames in os.walk(root, followlinks=False):
for name in filenames:
p = os.path.join(dirpath, name)
if os.path.islink(p):
continue
st = os.lstat(p)
if st.st_nlink > 1:
inodes.setdefault((st.st_dev, st.st_ino), []).append(
os.path.relpath(p, root))
groups: Dict[str, str] = {}
for i, (_key, members) in enumerate(sorted(inodes.items())):
for rel in members:
groups[rel] = chr(ord("A") + i)
return groups
def _drop_ignored(tree: Dict[str, tuple], ignore_paths) -> Dict[str, tuple]:
if not ignore_paths:
return tree
out = {}
for rel, desc in tree.items():
if any(rel == ig or rel.startswith(ig.rstrip("/") + "/") for ig in ignore_paths):
continue
out[rel] = desc
return out
def tree_diff(rsync_root: str, fs_root: str, case: Case) -> List[str]:
"""Return a list of human-readable differences (empty when identical)."""
rtree = _drop_ignored(snapshot(rsync_root, case.compare_modes), case.ignore_paths)
ftree = _drop_ignored(snapshot(fs_root, case.compare_modes), case.ignore_paths)
if case.compare_hardlinks:
rgroups = _hardlink_groups(rsync_root)
fgroups = _hardlink_groups(fs_root)
else:
rgroups = fgroups = {}
diffs: List[str] = []
for rel in sorted(set(rtree) | set(ftree)):
r = rtree.get(rel)
f = ftree.get(rel)
if r == f:
continue
if r is None:
diffs.append(f"+ fastsync-only: {rel!r} {f}")
elif f is None:
diffs.append(f"- rsync-only: {rel!r} {r}")
else:
diffs.append(f"~ differs: {rel!r} rsync={r} fastsync={f}")
if case.compare_hardlinks:
for rel in sorted(set(rgroups) | set(fgroups)):
if rgroups.get(rel) != fgroups.get(rel):
diffs.append(
f"~ hardlink group: {rel!r} rsync={rgroups.get(rel)} "
f"fastsync={fgroups.get(rel)}")
return diffs
# ---------------------------------------------------------------------------
# stdout normalization
# ---------------------------------------------------------------------------
def _parse_bytes(text: str) -> str:
m = re.match(r"([\d,]+)", text.strip())
return m.group(1).replace(",", "") if m else text.strip()
def normalize_stdout(text: str, mode: Optional[str]) -> object:
if mode == STDOUT_ITEMIZE:
lines = []
for line in (text or "").splitlines():
line = line.rstrip()
if not line:
continue
if line.startswith("*deleting"):
lines.append(line)
continue
if not _ITEMIZE_RE.match(line):
continue
# Directories are not transfer entries in FastSync's recursive
# scanner, so rsync's `cd+++++++++ name/` lines have no counterpart
# (documented recursive-empty-dir residual). Compare file/link
# itemization only.
if line.rsplit(" ", 1)[-1].endswith("/"):
continue
lines.append(line)
return sorted(lines)
if mode == STDOUT_OUTFMT:
lines = []
for line in (text or "").splitlines():
line = line.rstrip()
if not line:
continue
# Directory entries are emitted by rsync but not by FastSync's
# recursive scanner (documented residual). Tokens are either
# `%n %l` (path first) or `%i %n` (path last); drop a line when
# either end-token is a directory path.
first = line.split(" ", 1)[0]
last = line.rsplit(" ", 1)[-1]
if first.endswith("/") or last.endswith("/"):
continue
lines.append(line)
return sorted(lines)
if mode == STDOUT_STATS:
found = {}
for line in (text or "").splitlines():
for key in STATS_KEYS:
if line.startswith(key + ":"):
found[key] = _parse_bytes(line.split(":", 1)[1])
return found
if mode == STDOUT_PROGRESS:
# rsync prints the file-list entries in sorted depth-first order while
# FastSync's streaming scan emits them in readdir/BFS order; only the
# entry set and deterministic fields are compared. The transfer-root
# `./` line's trigger condition is a separate documented residual, and
# the per-frame rate/elapsed/xfr#/to-chk numerator are wall-clock- or
# order-dependent, so only the `to-chk` denominator and the name set are
# asserted.
names = []
totals = set()
max_xfr = 0
for line in (text or "").splitlines():
line = line.rstrip()
if not line:
continue
if "%" in line:
m = _PROGRESS_TOTAL_RE.search(line)
if m:
totals.add(int(m.group(1)))
mx = _PROGRESS_XFR_RE.search(line)
if mx:
max_xfr = max(max_xfr, int(mx.group(1)))
continue
if line == "sending incremental file list":
continue
if line.startswith("created directory "):
continue
if line == "./":
continue
names.append(line)
return {"names": sorted(names), "total": sorted(totals), "xfr": max_xfr}
# raw
return sorted(l.rstrip() for l in (text or "").splitlines() if l.strip())
def stdout_diff(rsync_out: str, fs_out: str, mode: Optional[str]) -> List[str]:
r = normalize_stdout(rsync_out, mode)
f = normalize_stdout(fs_out, mode)
if r == f:
return []
if mode == STDOUT_STATS:
return [f"stats rsync={r}", f"stats fastsync={f}"]
if mode == STDOUT_PROGRESS:
return [f"progress rsync={r}", f"progress fastsync={f}"]
return list(difflib.unified_diff(
[str(x) for x in r], [str(x) for x in f],
fromfile="rsync", tofile="fastsync", lineterm=""))
# ---------------------------------------------------------------------------
# Running one case
# ---------------------------------------------------------------------------
def run_rsync(src: str, rdst: str, flags: List[str]) -> subprocess.CompletedProcess:
args = [RSYNC] + list(flags) + [src + "/", rdst + "/"]
return subprocess.run(
args, capture_output=True, text=True,
env=dict(os.environ, LC_ALL="C"), timeout=180)
def run_fastsync(src: str, fdst: str, flags: List[str], port: int):
return run_client(src, fdst, flags=list(flags), port=port)
def run_differential( # noqa: PLR0913 (explicit scenario parameters)
src: str,
rdst: str,
fdst: str,
rs_flags: List[str],
fs_flags: List[str],
server,
layout: str = MIRROR,
seed: Optional[Callable] = None,
stdout: Optional[str] = STDOUT_NONE,
compare_modes: bool = False,
compare_hardlinks: bool = False,
ignore_paths: Tuple[str, ...] = (),
extra_check: Optional[Callable] = None,
files_from: Optional[Tuple[str, ...]] = None,
fs_src_suffix: str = "",
compare_tree: bool = True,
) -> Dict[str, object]:
"""Run one rsync/FastSync pair and return the diff aspects.
Returned dict keys: ``rsync_rc``, ``fastsync_rc``, ``rsync_stderr``,
``fastsync_stderr``, ``tree``, ``stdout``, ``extra``.
"""
clean_dir(rdst)
clean_dir(fdst)
abs_src = os.path.abspath(src)
rel = abs_src.lstrip(os.sep)
if layout == RELATIVE:
rroot, froot = rdst, fdst
elif layout == MIRROR_ABS:
rroot, froot = os.path.join(rdst, rel), get_dest_received_dir(fdst, src)
else:
rroot, froot = rdst, get_dest_received_dir(fdst, src)
# rsync's destination root always exists (clean_dir created it). FastSync's
# logical transfer root is the mirror path below the destination argument,
# so pre-create it too: `Number of created files` counts the root only when
# it is genuinely absent, and the two tools must start from the same state.
os.makedirs(froot, exist_ok=True)
if seed:
seed(src, rroot, froot)
rs_flags = list(rs_flags)
fs_flags = list(fs_flags)
if files_from is not None:
list_path = os.path.join(TEST_DATA_DIR, "parity_" +
os.path.basename(src) + ".list")
write_list(list_path, files_from)
rs_flags.append(f"--files-from={list_path}")
fs_flags.append(f"--files-from={list_path}")
rs = run_rsync(src, rdst, rs_flags)
fs_result, _ = run_fastsync(src + fs_src_suffix, fdst, fs_flags, server.port)
class _View:
"""Adapter so tree_diff/extra_check keep the Case-shaped interface."""
def __init__(self) -> None:
self.compare_modes = compare_modes
self.compare_hardlinks = compare_hardlinks
self.ignore_paths = ignore_paths
result = {
"rsync_rc": rs.returncode,
"fastsync_rc": fs_result.returncode,
"rsync_stderr": rs.stderr,
"fastsync_stderr": fs_result.stderr or fs_result.stdout,
"tree": tree_diff(rroot, froot, _View()) if compare_tree else [],
"stdout": [],
"extra": [],
}
if stdout is not None:
result["stdout"] = stdout_diff(rs.stdout, fs_result.stdout, stdout)
if extra_check:
result["extra"] = list(extra_check(src, rroot, froot, rs, fs_result) or [])
return result
def execute_case(case: Case, server) -> Dict[str, object]:
"""Run a table-driven case and return the diff aspects."""
tag = case.id
src = os.path.join(TEST_DATA_DIR, f"parity_{tag}_src")
rdst = os.path.join(TEST_DATA_DIR, f"parity_{tag}_rdst")
fdst = os.path.join(TEST_DATA_DIR, f"parity_{tag}_fdst")
CORPORA[case.corpus](src)
return run_differential(
src, rdst, fdst,
case.flags, case.fs_flags(), server,
layout=case.layout, seed=case.seed, stdout=case.stdout,
compare_modes=case.compare_modes, compare_hardlinks=case.compare_hardlinks,
ignore_paths=case.ignore_paths, extra_check=case.extra_check,
files_from=case.files_from, fs_src_suffix=case.fs_src_suffix,
compare_tree=case.compare_tree,
)
def write_list(path: str, entries) -> str:
os.makedirs(os.path.dirname(path), exist_ok=True)
with open(path, "w", encoding="utf-8") as fh:
for e in entries:
fh.write(e + "\n")
return path
+213
View File
@@ -0,0 +1,213 @@
"""Differential tests for the client CLI's codec defaults and env lists.
Track 3a of the rsync-parity plan pins two rsync 3.4.1 behaviors that are
resolved entirely on the client:
* the per-codec default ``--compress-level`` (zstd 3, zlib/zlibx 6, lz4
ignored) applied when the user omits ``--compress-level``/``--zl``, with an
explicit level clamped to the codec's range; and
* the ``RSYNC_COMPRESS_LIST`` / ``RSYNC_CHECKSUM_LIST`` preference lists that
rsync's ``auto`` consults before its compiled-in order (whitespace-separated,
unknown names skipped, first supported wins, all-unknown is exit 4).
The rsync side is observed through ``--debug=NSTR1``; FastSync publishes its
resolved codec/level through ``--debug=util``. The checksum side is confirmed
byte-for-byte through ``--out-format %C``. The rsync-based tests skip cleanly
when rsync is not installed.
"""
import os
import re
import shutil
import subprocess
import sys
import pytest
sys.path.insert(0, os.path.dirname(__file__))
from common import (
TEST_DATA_DIR,
run_client,
clean_dir,
get_dest_received_dir,
)
RSYNC = shutil.which("rsync")
requires_rsync = pytest.mark.skipif(RSYNC is None, reason="rsync 3.4.1 not installed")
CODEC_ROOT = os.path.join(TEST_DATA_DIR, "cli_differential")
_COMPRESS_RE = re.compile(r"compress(?:ion)?: (\w+) \(level (-?\d+)\)")
def _rsync(args):
env = dict(os.environ, LC_ALL="C")
return subprocess.run([RSYNC] + args, capture_output=True, text=True, env=env, timeout=120)
def _scratch(tag):
path = os.path.join(CODEC_ROOT, tag)
clean_dir(path)
os.makedirs(path, exist_ok=True)
return path
def _make_corpus(root):
clean_dir(root)
os.makedirs(root, exist_ok=True)
with open(os.path.join(root, "big.bin"), "wb") as fh:
fh.write(b"FastSync codec payload " * 4096)
with open(os.path.join(root, "small.txt"), "wb") as fh:
fh.write(b"hello codec world\n" * 32)
return root
def _rsync_compress_level(choice, level):
src = _make_corpus(_scratch(f"lvl_src_{choice}_{level}"))
dst = _scratch(f"lvl_rsync_{choice}_{level}")
args = ["-a", "-z", f"--zc={choice}"]
if level is not None:
args.append(f"--zl={level}")
args += ["--debug=NSTR1", src + "/", dst + "/"]
result = _rsync(args)
assert result.returncode == 0, result.stderr
match = _COMPRESS_RE.search(result.stdout + result.stderr)
assert match, (result.stdout, result.stderr)
return match.group(1), int(match.group(2))
def _fastsync_compress_level(choice, level, shared_server):
src = _make_corpus(_scratch(f"lvl_src_fs_{choice}_{level}"))
dst = _scratch(f"lvl_fs_{choice}_{level}")
args = ["-a", "-z", f"--zc={choice}"]
if level is not None:
args.append(f"--zl={level}")
args += ["-v", "--debug=util"]
result, _ = run_client(src, dst, flags=args, port=shared_server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
match = _COMPRESS_RE.search(result.stdout)
assert match, result.stdout[:500]
return match.group(1), int(match.group(2))
class TestPerCodecCompressionLevelDefaults:
"""``--compress-level`` defaults and clamping match rsync per codec."""
# FastSync uses a positive lz4 placeholder because its "level > 0" gate
# enables compression; lz4_compress ignores the value, so rsync's level 0
# and FastSync's level 1 produce the same bytes.
CASES = [
("zstd", None, 3),
("zlib", None, 6),
("zlibx", None, 6),
("lz4", None, 1),
("zstd", 10, 10),
("zlib", 15, 9),
("zlib", 3, 3),
("lz4", 15, 1),
]
@requires_rsync
@pytest.mark.ci
@pytest.mark.parametrize("choice,level,fs_level", CASES)
def test_level_matches_rsync(self, choice, level, fs_level, shared_server):
rsync_algo, rsync_level = _rsync_compress_level(choice, level)
fs_algo, fs_level_actual = _fastsync_compress_level(choice, level, shared_server)
assert rsync_algo == choice
assert fs_algo == choice
if choice == "lz4":
assert rsync_level == 0 and fs_level_actual > 0
else:
assert rsync_level == fs_level
assert fs_level_actual == fs_level
class TestEnvPreferenceLists:
"""``RSYNC_COMPRESS_LIST`` / ``RSYNC_CHECKSUM_LIST`` drive auto like rsync."""
# (env value, expected codec, rsync level, FastSync level)
COMPRESS_CASES = [
("zlib lz4", "zlib", 6, 6),
("lz4 zstd", "lz4", 0, 1),
("bogus zstd zlib", "zstd", 3, 3),
(" ", "zstd", 3, 3),
]
@requires_rsync
@pytest.mark.ci
@pytest.mark.parametrize("env,algo,rsync_level,fs_level", COMPRESS_CASES)
def test_compress_list_matches_rsync(self, env, algo, rsync_level, fs_level, shared_server,
monkeypatch):
monkeypatch.setenv("RSYNC_COMPRESS_LIST", env)
src = _make_corpus(_scratch(f"envc_src_{algo}"))
rdst = _scratch(f"envc_rsync_{algo}")
rs = _rsync(["-a", "-z", "--debug=NSTR1", src + "/", rdst + "/"])
assert rs.returncode == 0, rs.stderr
rm = _COMPRESS_RE.search(rs.stdout + rs.stderr)
assert rm, (rs.stdout, rs.stderr)
assert rm.group(1) == algo
assert int(rm.group(2)) == rsync_level
fdst = _scratch(f"envc_fs_{algo}")
result, _ = run_client(src, fdst, flags=["-a", "-z", "-v", "--debug=util"],
port=shared_server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
fm = _COMPRESS_RE.search(result.stdout)
assert fm, result.stdout[:500]
assert fm.group(1) == algo
assert int(fm.group(2)) == fs_level
received = get_dest_received_dir(fdst, src)
assert _tree_bytes(received) == _tree_bytes(src)
@requires_rsync
@pytest.mark.ci
@pytest.mark.parametrize("env,algo", [("md5", "md5"), ("sha1", "sha1"), ("xxh3 md5", "xxh3")])
def test_checksum_list_matches_rsync(self, env, algo, shared_server, monkeypatch):
monkeypatch.setenv("RSYNC_CHECKSUM_LIST", env)
src = _make_corpus(_scratch(f"envcc_src_{algo}"))
rdst = _scratch(f"envcc_rsync_{algo}")
rs = _rsync(["-a", "--checksum", "--out-format=%C %n", src + "/", rdst + "/"])
assert rs.returncode == 0, rs.stderr
rs_digests = _digests(rs.stdout)
fdst = _scratch(f"envcc_fs_{algo}")
result, _ = run_client(src, fdst, flags=["-a", "--checksum", "--out-format=%C %n"],
port=shared_server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
assert _digests(result.stdout) == rs_digests
@requires_rsync
@pytest.mark.ci
def test_all_unknown_lists_fail_like_rsync(self, shared_server, monkeypatch):
src = _make_corpus(_scratch("envbad_src"))
monkeypatch.setenv("RSYNC_COMPRESS_LIST", "bogus")
rs = _rsync(["-a", "-z", src + "/", _scratch("envbad_rsync_c") + "/"])
assert rs.returncode == 4, rs.stderr
result, _ = run_client(src, _scratch("envbad_fs_c"), flags=["-a", "-z"],
port=shared_server.port)
assert result.returncode == 4, (result.stderr or result.stdout)[:200]
monkeypatch.setenv("RSYNC_CHECKSUM_LIST", "bogus")
rs = _rsync(["-a", src + "/", _scratch("envbad_rsync_s") + "/"])
assert rs.returncode == 4, rs.stderr
result, _ = run_client(src, _scratch("envbad_fs_s"), flags=["-a"],
port=shared_server.port)
assert result.returncode == 4, (result.stderr or result.stdout)[:200]
def _tree_bytes(root):
out = {}
for dirpath, _dirs, files in os.walk(root):
for name in files:
path = os.path.join(dirpath, name)
with open(path, "rb") as fh:
out[os.path.relpath(path, root)] = fh.read()
return out
def _digests(output):
out = {}
for line in output.splitlines():
parts = line.split()
if len(parts) == 2 and parts[0]:
out[parts[1]] = parts[0]
return out
+149
View File
@@ -9,6 +9,7 @@ directory when the shared test server is launched), so every scratch tree lives
under ``TEST_DATA_DIR`` rather than pytest's ``tmp_path``.
"""
import os
import random
import shutil
import subprocess
import sys
@@ -68,6 +69,84 @@ def _tree_bytes(root):
return out
_CC_DELTA_T0 = 1_600_000_000
_CC_DELTA_T1 = 1_600_000_100
_DELTA_STATS_KEYS = (
"Number of created files",
"Number of regular files transferred",
"Total transferred file size",
"Literal data",
"Matched data",
)
def _pin_tree(root, mtime):
for dirpath, dirnames, filenames in os.walk(root):
for name in dirnames + filenames:
path = os.path.join(dirpath, name)
if not os.path.islink(path):
os.utime(path, (mtime, mtime))
os.utime(root, (mtime, mtime))
def _make_delta_basis(src, size=512 * 1024):
"""Build a source and a matching pre-modification basis tree.
The source's ``big.bin`` is then modified in a few disjoint places and given
a newer mtime so both tools take the delta path. Returns the basis dir.
"""
clean_dir(src)
original = random.Random(20240101).randbytes(size)
with open(os.path.join(src, "big.bin"), "wb") as fh:
fh.write(original)
with open(os.path.join(src, "small.txt"), "wb") as fh:
fh.write(b"hello world\n")
_pin_tree(src, _CC_DELTA_T0)
basis = src.rstrip("/") + "_basis"
clean_dir(basis)
shutil.copy2(os.path.join(src, "big.bin"), os.path.join(basis, "big.bin"))
shutil.copy2(os.path.join(src, "small.txt"), os.path.join(basis, "small.txt"))
_pin_tree(basis, _CC_DELTA_T0)
modified = bytearray(original)
for off in (0, size // 3, 2 * size // 3, size - 64):
for i in range(32):
modified[off + i] ^= 0x5A
with open(os.path.join(src, "big.bin"), "wb") as fh:
fh.write(bytes(modified))
os.utime(os.path.join(src, "big.bin"), (_CC_DELTA_T1, _CC_DELTA_T1))
return basis
def _seed_from_basis(basis, target):
clean_dir(target)
for name in os.listdir(basis):
shutil.copy2(os.path.join(basis, name), os.path.join(target, name))
def _delta_stats(text):
found = {}
for line in text.splitlines():
for key in _DELTA_STATS_KEYS:
if line.startswith(key + ":"):
found[key] = line.split(":", 1)[1].strip()
return found
def _big_bin_outfmt(text):
"""The ``(c, C)`` pair from the ``big.bin`` out-format line (`%c|%C %n`)."""
for line in text.splitlines():
stripped = line.strip()
if "|" not in stripped or not stripped.endswith("big.bin"):
continue
c_field, rest = stripped.split("|", 1)
fields = rest.split()
return c_field.strip(), (fields[0] if fields else "")
return None, None
class TestCodecChoiceMatrix:
"""The CLI accept/reject set and exit codes must match rsync 3.4.1."""
@@ -191,6 +270,76 @@ class TestCodecTransferDifferential:
assert _tree_bytes(received) == _tree_bytes(rsync_dst)
class TestChecksumChoiceDeltaSurface:
"""--checksum-choice does not move the delta-transfer parity surface.
FastSync's delta BLOCK strong checksum is a fixed xxHash32, so the
negotiated algorithm only selects the whole-file comparison digest (and the
``%C`` transfer digest). A pre-seeded delta transfer must therefore land
byte-identical bytes and report the same counters for every choice, while
``%C`` -- the one token that tracks the choice -- stays byte-identical to
rsync. This pins the Track-3b reclassification in RSYNC_COMPAT.md.
"""
CHOICES = ["xxh64", "xxh128", "xxh3", "md5", "md4", "sha1",
"xxh64,sha1", "sha1,xxh64"]
@requires_rsync
@pytest.mark.ci
def test_delta_surface_invariant_to_checksum_choice(self, shared_server):
src = _scratch("ccdelta_src")
basis = _make_delta_basis(src)
source_bytes = _tree_bytes(src)
rsync_c, fastsync_c, digests = {}, {}, {}
rsync_stats, fastsync_stats = {}, {}
for choice in self.CHOICES:
tag = choice.replace(",", "_")
rsync_dst = _scratch(f"ccdelta_rs_{tag}")
fs_dst = _scratch(f"ccdelta_fs_{tag}")
fs_root = get_dest_received_dir(fs_dst, src)
_seed_from_basis(basis, rsync_dst)
_seed_from_basis(basis, fs_root)
# Pin the block size on both ends so the literal/matched split is
# comparable (rsync's adaptive default would otherwise differ from
# FastSync's 8192-byte default).
rsync_result = _rsync(["-a", "--no-whole-file", "-B8192", "--stats",
"--out-format=%c|%C %n", f"--cc={choice}",
src + "/", rsync_dst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
result, _ = run_client(
src, fs_dst,
flags=["-a", "--incremental", "--delta", "-B8192", "--stats",
"--out-format=%c|%C %n", f"--cc={choice}"],
port=shared_server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
assert _tree_bytes(rsync_dst) == source_bytes, choice
assert _tree_bytes(fs_root) == source_bytes, choice
assert _delta_stats(rsync_result.stdout) == _delta_stats(result.stdout), choice
rs_c, rs_C = _big_bin_outfmt(rsync_result.stdout)
fs_c, fs_C = _big_bin_outfmt(result.stdout)
assert rs_C == fs_C, f"{choice}: %C rsync={rs_C!r} fastsync={fs_C!r}"
rsync_c[choice] = rs_c
fastsync_c[choice] = fs_c
digests[choice] = fs_C
rsync_stats[choice] = _delta_stats(rsync_result.stdout)
fastsync_stats[choice] = _delta_stats(result.stdout)
# The choice is only observable in %C, and it is effective (the digests
# are not all the same algorithm's output).
assert len(set(digests.values())) > 1, digests
# The compared --stats counters are invariant across choices in each tool
# (and were asserted equal cross-tool inside the loop).
assert len({tuple(sorted(s.items())) for s in rsync_stats.values()}) == 1, rsync_stats
assert len({tuple(sorted(s.items())) for s in fastsync_stats.values()}) == 1, fastsync_stats
# The block-checksum token (%c) is invariant across choices in each tool.
assert len(set(rsync_c.values())) == 1, rsync_c
assert len(set(fastsync_c.values())) == 1, fastsync_c
class TestCodecNegotiationFallback:
"""FastSync's auto negotiation and deterministic fallback order."""
+117 -9
View File
@@ -63,6 +63,10 @@ DETACH_MODULE = os.path.join(MODULE_ROOT, "detach")
DETACH_CONF = os.path.join(TEST_DATA_DIR, "fastsyncd_detach.conf")
DETACH_PORT = None
# A dedicated config for the umask test: the daemon must be launched in the real
# (double-fork) detach path, whose daemonize() applies umask(022).
UMASK_CONF = os.path.join(TEST_DATA_DIR, "fastsyncd_umask.conf")
# Passwords are never sent as plaintext and never logged; these literals are
# only hashed into the server credential file / client password file.
ALICE_PASS = "alice-s3cret"
@@ -137,6 +141,7 @@ class DaemonManager:
def __init__(self):
self._proc = None
self._port = None
self.log_path = None
def start(self, config_path, port_override=None, extra_args=None, log_path=None):
self.stop()
@@ -150,7 +155,11 @@ class DaemonManager:
if extra_args:
cmd += extra_args
if log_path is None:
log_path = os.path.join(TEST_DATA_DIR, "fastsyncd.log")
# A unique log per manager: several managers run in one xdist
# worker, and a shared log lets one daemon's truncate/write offset
# corrupt the other's appended lines (a flaky log assertion).
log_path = os.path.join(TEST_DATA_DIR, f"fastsyncd_{id(self):x}.log")
self.log_path = log_path
log = open(log_path, "w")
self._proc = subprocess.Popen(
cmd, stdout=log, stderr=log, stdin=subprocess.DEVNULL, start_new_session=True)
@@ -220,6 +229,7 @@ def daemon_env():
"\n"
"[files]\n"
"path = %s\n"
"read only = no\n"
"\n"
"[readonly]\n"
"path = %s\n"
@@ -227,18 +237,22 @@ def daemon_env():
"\n"
"[locked]\n"
"path = %s\n"
"read only = no\n"
"auth users = alice\n"
"\n"
"[team]\n"
"path = %s\n"
"read only = no\n"
"auth users = alice,bob\n"
"\n"
"[owner]\n"
"path = %s\n"
"read only = no\n"
"client owner = yes\n"
"\n"
"[denied]\n"
"path = %s\n"
"read only = no\n"
"hosts deny = 127.0.0.1\n"
% (config_port, FILES_MODULE, READONLY_MODULE, AUTH_MODULE, TEAM_MODULE, OWNER_MODULE,
DENIED_MODULE))
@@ -257,7 +271,8 @@ def daemon_env():
global DETACH_PORT
DETACH_PORT = _find_free_port()
with open(DETACH_CONF, "w") as f:
f.write("port = %d\n\n[detach]\npath = %s\n" % (DETACH_PORT, DETACH_MODULE))
f.write("port = %d\n\n[detach]\npath = %s\nread only = no\n"
% (DETACH_PORT, DETACH_MODULE))
yield
_kill_by_cmdline_marker(DETACH_CONF)
@@ -332,6 +347,97 @@ class TestDaemonModuleSelection:
assert not missing, f"missing: {missing[:5]}"
assert not mismatches, f"mismatch: {mismatches[:5]}"
def test_daemon_new_dirs_not_world_writable(self):
"""The daemon must not force umask 0: implied parent directories created
without -p are the source default (0755 under the daemon's 022 umask),
never world-writable 0777.
This drives the real double-fork detach path, where the umask(022) fix
lives (daemonize()); the --no-detach path never calls it. The launcher
is run with umask 0, so without the fix the daemon would inherit 0 and
create a 0777 directory; with the fix the assertion below fails only if
the fix regresses."""
port = _find_free_port()
with open(UMASK_CONF, "w") as f:
f.write("port = %d\n\n[files]\npath = %s\nread only = no\n"
% (port, FILES_MODULE))
sub = os.path.join(FILES_MODULE, "umask_check")
shutil.rmtree(sub, ignore_errors=True)
os.makedirs(sub, exist_ok=True)
log_path = os.path.join(TEST_DATA_DIR, "fastsyncd_umask.log")
log = open(log_path, "w")
cmd = SERVER_CMD + ["--daemon", "--config", UMASK_CONF, "--allow-unauthenticated"]
proc = subprocess.Popen(cmd, stdout=log, stderr=log, stdin=subprocess.DEVNULL,
preexec_fn=lambda: os.umask(0))
try:
_wait_for_port(port, timeout=15)
result = _push("127.0.0.1::files/umask_check", port)
assert result.returncode == 0, result.stderr or result.stdout
received = get_dest_received_dir(sub, SOURCE_DIR)
nested = os.path.join(received, "nested")
assert os.path.isdir(nested), f"nested dir missing under {received}"
mode = stat.S_IMODE(os.stat(nested).st_mode)
assert (mode & 0o022) == 0, f"implied directory is group/other writable: {oct(mode)}"
finally:
_kill_by_cmdline_marker(UMASK_CONF)
log.close()
try:
proc.wait(timeout=5)
except subprocess.TimeoutExpired:
proc.kill()
class TestRsyncConfigCompat:
"""A real rsyncd.conf can be pointed at FastSync: the common rsync GLOBAL
and MODULE keys are accepted, the ones with a FastSync equivalent (port,
path, read only, max connections) take effect, and the inert ones (pid
file, log file, comment, use chroot, uid, gid, exclude, timeout, ...) are
documented no-ops. --dparam accepts the same expanded key set."""
@pytest.mark.ci
def test_rsync_style_config_round_trip(self):
module = os.path.join(MODULE_ROOT, "rsync_style")
shutil.rmtree(module, ignore_errors=True)
os.makedirs(module, exist_ok=True)
port = _find_free_port()
conf = os.path.join(TEST_DATA_DIR, "fastsyncd_rsync_style.conf")
with open(conf, "w") as f:
f.write(
"# an rsync 3.4.1-style rsyncd.conf\n"
"pid file = /tmp/fastsyncd_rsync_style.pid\n"
"log file = /tmp/fastsyncd_rsync_style.log\n"
"socket options = TCP_NODELAY\n"
"use chroot = no\n"
"uid = nobody\n"
"gid = nogroup\n"
"timeout = 600\n"
"max verbosity = 2\n"
"transfer logging = yes\n"
"port = %d\n"
"\n"
"[rsync_style]\n"
"path = %s\n"
"comment = rsync-style module\n"
"use chroot = no\n"
"exclude = *.tmp\n"
"read only = no\n"
"max connections = 4\n"
% (port, module))
d = DaemonManager()
# --dparam borrows rsync's compact spelling; `pidfile` is inert but must
# not be rejected, proving dparam reuses the expanded global key set.
d.start(conf, extra_args=["--dparam", "pidfile=/tmp/rsync_style.pid"],
log_path=os.path.join(TEST_DATA_DIR, "fastsyncd_rsync_style.log"))
try:
result = _push("127.0.0.1::rsync_style", d.port)
assert result.returncode == 0, result.stderr or result.stdout
received = get_dest_received_dir(module, SOURCE_DIR)
mismatches, missing = verify_transfer(SOURCE_DIR, received)
assert not missing, f"missing: {missing[:5]}"
assert not mismatches, f"mismatch: {mismatches[:5]}"
finally:
d.stop()
class TestDaemonRejection:
def _tree_files(self):
@@ -435,7 +541,7 @@ class TestDaemonRejection:
before any data lands. `accept` lists the log phrases that count as the
refusal (a non-root daemon refuses --copy-as earlier, at the privilege
check, so the caller accepts that phrase too)."""
log_path = os.path.join(TEST_DATA_DIR, "fastsyncd.log")
log_path = daemon.log_path
before = os.path.getsize(log_path) if os.path.exists(log_path) else 0
before_files = self._tree_files()
result, _ = run_client(SOURCE_DIR, f"127.0.0.1::{module}", port=daemon.port, flags=flags)
@@ -472,14 +578,13 @@ class TestDaemonRejection:
the refusal into a silent accept."""
port = _find_free_port()
d = DaemonManager()
log_path = os.path.join(TEST_DATA_DIR, "fastsyncd.log")
try:
d.start(CONF_FILE, port_override=port,
extra_args=["--password-file", CRED_FILE, "--no-super"])
result, _ = run_client(SOURCE_DIR, "127.0.0.1::files", port=d.port,
flags=["--super", "--preserve"])
assert result.returncode != 0, "the --no-super daemon must refuse --super"
with open(log_path, "rb") as f:
with open(d.log_path, "rb") as f:
tail = f.read().decode("utf-8", "replace")
assert "client-chosen ownership" in tail, (
f"daemon did not log the --super refusal: {tail[-400:]!r}"
@@ -968,7 +1073,7 @@ class TestDaemonAuthentication:
def test_auth_log_does_not_leak_password(self, daemon):
"""The daemon log must never contain the password or the store verifier."""
log_path = os.path.join(TEST_DATA_DIR, "fastsyncd.log")
log_path = daemon.log_path
before = os.path.getsize(log_path) if os.path.exists(log_path) else 0
_push_with_creds("127.0.0.1::locked", daemon.port, "alice", WRONG_PASS)
_push_with_creds("127.0.0.1::locked", daemon.port, "alice", ALICE_PASS)
@@ -995,7 +1100,7 @@ class TestDaemonAuthentication:
_push_with_creds("127.0.0.1::locked", port, "alice", ALICE_PASS)
_push_with_creds("127.0.0.1::locked", port, "alice", WRONG_PASS)
time.sleep(0.3)
log_path = os.path.join(TEST_DATA_DIR, "fastsyncd.log")
log_path = d.log_path
with open(log_path, "rb") as f:
log = f.read().decode("utf-8", "replace")
finally:
@@ -1030,7 +1135,8 @@ class TestDaemonMotd:
motd_line = "motd file = %s\n" % motd_path if motd_path else ""
os.makedirs(self.MOTD_MODULE, exist_ok=True)
with open(self.MOTD_CONF, "w") as f:
f.write("port = %d\n%s\n[files]\npath = %s\n" % (port, motd_line, self.MOTD_MODULE))
f.write("port = %d\n%s\n[files]\npath = %s\nread only = no\n"
% (port, motd_line, self.MOTD_MODULE))
d = DaemonManager()
d.start(self.MOTD_CONF, port_override=port)
return d, port
@@ -1214,12 +1320,12 @@ class TestDaemonTLSAuth:
_write_client_password_file(client_creds, "alice", ALICE_PASS)
d = DaemonManager()
port = _find_free_port()
log_path = os.path.join(TEST_DATA_DIR, "fastsyncd.log")
try:
d.start(CONF_FILE, port_override=port, extra_args=[
"--tls", "--cert", certs["server_cert"], "--key", certs["server_key"],
"--ca", certs["ca"], "--client-cn", "fastsync-client",
"--password-file", CRED_FILE])
log_path = d.log_path
before_files = _tree_file_count(AUTH_MODULE)
log_before = os.path.getsize(log_path) if os.path.exists(log_path) else 0
tls_flags = ["--tls",
@@ -1267,6 +1373,7 @@ class TestDaemonConnectionLimits:
"\n"
"[locked]\n"
"path = %s\n"
"read only = no\n"
"auth users = alice\n"
% (port, AUTH_MODULE))
d = DaemonManager()
@@ -1304,6 +1411,7 @@ class TestDaemonConnectionLimits:
"\n"
"[files]\n"
"path = %s\n"
"read only = no\n"
"max connections = 2\n"
% (port, FILES_MODULE))
d = DaemonManager()
@@ -0,0 +1,291 @@
"""Differential coverage for the delete-timing ABORT BOUNDARY (A9/A10).
rsync's generator runs ahead of its throttled sender, so on a mid-transfer abort
it has already removed every extra it planned. FastSync now transmits the
COMPLETE per-directory plan set before the first data frame, so an abort has the
same effect. Before that change FastSync only removed the extras of the
directories its (slower) data stream had reached, and ``-d/--dirs`` used an
end-of-transfer commit that removed nothing on abort.
These tests abort both tools mid-transfer and assert the destination extras
removed match real ``rsync 3.4.1``. The rsync side is driven locally with
``--bwlimit`` and a small timing window (its generator's delete list is computed
long before the throttled payload finishes); the FastSync side uses the
byte-deterministic slicing proxy from ``test_delete_timing_parity``.
"""
import os
import shutil
import subprocess
import sys
import time
import pytest
sys.path.insert(0, os.path.dirname(__file__))
from common import ( # noqa: E402
TEST_DATA_DIR,
ServerManager,
clean_dir,
get_dest_received_dir,
run_client,
)
from test_delete_timing_parity import _SlicingProxy # noqa: E402
RSYNC = shutil.which("rsync")
requires_rsync = pytest.mark.skipif(RSYNC is None, reason="rsync 3.4.1 not installed")
# Exceeds the 10 MiB scanner chunk, so the next directory lands in a later chunk
# (still unreached when the proxy cuts the stream).
BIG_BYTES = 16 * 1024 * 1024
# Cut well past the (small) config + delete-plan frames and into the big payload,
# so the receiver has provably processed every plan before the abort.
MID_TRANSFER_BYTES = 256 * 1024
PROXY_THROTTLE = 0.001
# Throttle rsync's sender so the generator has deleted long before the payload
# finishes, then interrupt it mid-transfer.
RSYNC_BWLIMIT = 512 # KiB/s -> ~32 s for 16 MiB
RSYNC_ABORT_DELAY = 1.5
def _write(path, content):
os.makedirs(os.path.dirname(path), exist_ok=True)
with open(path, "wb") as fh:
fh.write(content)
def _rsync_aborted(args, delay=RSYNC_ABORT_DELAY):
"""Start rsync, let its generator run, then interrupt it mid-transfer."""
env = dict(os.environ, LC_ALL="C")
proc = subprocess.Popen([RSYNC] + args, stdout=subprocess.PIPE, stderr=subprocess.PIPE,
text=True, env=env)
time.sleep(delay)
proc.terminate()
try:
proc.wait(timeout=10)
except subprocess.TimeoutExpired:
proc.kill()
proc.wait(timeout=5)
return proc
class TestDeleteDuringAbortBoundary:
"""A9: on an abort, every planned removal has already been applied."""
def _seed_recursive(self, tag):
source = os.path.join(TEST_DATA_DIR, f"dab_{tag}_src")
clean_dir(source)
# ``a/keep.bin`` sorts first, so the client streams it (and the proxy
# cuts) before the data pass ever reaches ``z/deep``.
_write(os.path.join(source, "a", "keep.bin"), b"B" * BIG_BYTES)
_write(os.path.join(source, "z", "deep", "keep.txt"), b"keep\n")
return source
@requires_rsync
def test_recursive_abort_removes_all_planned_extras(self):
# ---- FastSync: abort mid ``a/keep.bin``; ``z/deep`` is never reached.
source = self._seed_recursive("rec_fs")
dest = os.path.join(TEST_DATA_DIR, "dab_rec_fs_dst")
clean_dir(dest)
received = get_dest_received_dir(dest, source)
os.makedirs(os.path.join(received, "a"), exist_ok=True)
_write(os.path.join(received, "a", "a_extra"), b"stale\n")
os.makedirs(os.path.join(received, "z", "deep"), exist_ok=True)
_write(os.path.join(received, "z", "deep", "old_extra"), b"stale\n")
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
proxy = _SlicingProxy(server.port, forward_limit=MID_TRANSFER_BYTES,
throttle=PROXY_THROTTLE)
result, _ = run_client(source, dest, flags=["--delete-during"], port=proxy.port)
proxy.finish()
assert result.returncode != 0, "truncated transfer reported success"
assert not os.path.exists(os.path.join(received, "a", "a_extra"))
assert not os.path.exists(os.path.join(received, "z", "deep", "old_extra")), (
"FastSync left an extra in a directory it never reached before the abort"
)
# ---- rsync 3.4.1: same tree, same abort, same delete outcome.
source = self._seed_recursive("rec_rs")
rsync_dst = os.path.join(TEST_DATA_DIR, "dab_rec_rs_dst")
clean_dir(rsync_dst)
os.makedirs(os.path.join(rsync_dst, "a"), exist_ok=True)
_write(os.path.join(rsync_dst, "a", "a_extra"), b"stale\n")
os.makedirs(os.path.join(rsync_dst, "z", "deep"), exist_ok=True)
_write(os.path.join(rsync_dst, "z", "deep", "old_extra"), b"stale\n")
proc = _rsync_aborted(["-a", "--delete-during", f"--bwlimit={RSYNC_BWLIMIT}",
source + "/", rsync_dst + "/"])
assert proc.returncode != 0, "rsync was not actually interrupted"
assert not os.path.exists(os.path.join(rsync_dst, "a", "a_extra"))
assert not os.path.exists(os.path.join(rsync_dst, "z", "deep", "old_extra")), (
"rsync's generator did not delete ahead of its sender"
)
class TestDirsDeleteAbortBoundary:
"""A10: ``-d/--dirs`` uses per-directory plans like rsync.
The listed directory's direct extras are removed by the up-front plan while
a kept but untraversed subdirectory (and its destination content) is
shielded.
"""
def _seed_dirs(self, tag):
source = os.path.join(TEST_DATA_DIR, f"ddb_{tag}_src")
clean_dir(source)
_write(os.path.join(source, "big.bin"), b"B" * BIG_BYTES)
_write(os.path.join(source, "subdir", "keep.txt"), b"inner\n")
return source
@pytest.mark.parametrize("fs_flag,rs_flag", [("--delete-during", "--delete-during"),
("--delete", "--delete")])
@requires_rsync
def test_dirs_abort_removes_direct_extras_only(self, fs_flag, rs_flag):
label = f"{fs_flag.lstrip('-')}_{rs_flag.lstrip('-')}"
# ---- FastSync: ``-d`` lists the immediate children; big.bin streams and
# the abort lands mid-payload.
source = self._seed_dirs(f"dirs_{label}_fs")
dest = os.path.join(TEST_DATA_DIR, f"ddb_{label}_fs_dst")
clean_dir(dest)
received = get_dest_received_dir(dest, source)
_write(os.path.join(received, "old_extra"), b"stale\n")
os.makedirs(os.path.join(received, "subdir"), exist_ok=True)
_write(os.path.join(received, "subdir", "stale.txt"), b"stale inner\n")
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
proxy = _SlicingProxy(server.port, forward_limit=MID_TRANSFER_BYTES,
throttle=PROXY_THROTTLE)
result, _ = run_client(source + "/", dest, flags=["-d", fs_flag], port=proxy.port)
proxy.finish()
assert result.returncode != 0, f"{fs_flag}: truncated transfer reported success"
assert not os.path.exists(os.path.join(received, "old_extra")), (
f"{fs_flag}: the listed directory's direct extra survived the abort"
)
assert os.path.exists(os.path.join(received, "subdir", "stale.txt")), (
f"{fs_flag}: descended into a kept, untraversed subdirectory"
)
# ---- rsync 3.4.1: same shape and same abort.
source = self._seed_dirs(f"dirs_{label}_rs")
rsync_dst = os.path.join(TEST_DATA_DIR, f"ddb_{label}_rs_dst")
clean_dir(rsync_dst)
_write(os.path.join(rsync_dst, "old_extra"), b"stale\n")
os.makedirs(os.path.join(rsync_dst, "subdir"), exist_ok=True)
_write(os.path.join(rsync_dst, "subdir", "stale.txt"), b"stale inner\n")
proc = _rsync_aborted(["-d", rs_flag, f"--bwlimit={RSYNC_BWLIMIT}",
source + "/", rsync_dst + "/"])
assert proc.returncode != 0, "rsync was not actually interrupted"
assert not os.path.exists(os.path.join(rsync_dst, "old_extra")), (
f"rsync {rs_flag}: the listed directory's direct extra survived the abort"
)
assert os.path.exists(os.path.join(rsync_dst, "subdir", "stale.txt")), (
f"rsync {rs_flag}: descended into a kept, untraversed subdirectory"
)
def _tree(root):
out = []
for dirpath, dirs, files in os.walk(root):
for name in dirs:
out.append(os.path.relpath(os.path.join(dirpath, name), root))
for name in files:
out.append(os.path.relpath(os.path.join(dirpath, name), root))
return sorted(out)
class TestDirsDeleteFinalStateParity:
"""A10 completed run: ``-d DIR/ --delete`` (during default) and
``--delete-during`` match rsync's final tree, including a kept but
untraversed subdirectory whose destination content survives."""
@pytest.mark.parametrize("flag", ["--delete", "--delete-during"])
@requires_rsync
def test_dirs_final_state_matches_rsync(self, flag):
source = os.path.join(TEST_DATA_DIR, f"ddf_{flag.lstrip('-')}_src")
clean_dir(source)
_write(os.path.join(source, "keep.txt"), b"new keep\n")
_write(os.path.join(source, "subdir", "inner.txt"), b"inner\n")
def seed_dest(root):
clean_dir(root)
_write(os.path.join(root, "keep.txt"), b"old keep\n")
_write(os.path.join(root, "extra.txt"), b"extra\n")
_write(os.path.join(root, "extrasub", "ex.txt"), b"extra sub\n")
_write(os.path.join(root, "subdir", "stale.txt"), b"stale inner\n")
rsync_dst = os.path.join(TEST_DATA_DIR, f"ddf_{flag.lstrip('-')}_rs_dst")
seed_dest(rsync_dst)
env = dict(os.environ, LC_ALL="C")
rsync_result = subprocess.run(
[RSYNC, "-d", flag, source + "/", rsync_dst + "/"],
capture_output=True, text=True, env=env, timeout=120)
assert rsync_result.returncode == 0, rsync_result.stderr
rsync_tree = _tree(rsync_dst)
dest = os.path.join(TEST_DATA_DIR, f"ddf_{flag.lstrip('-')}_fs_dst")
clean_dir(dest)
received = get_dest_received_dir(dest, source)
seed_dest(received)
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
result, _ = run_client(source + "/", dest, flags=["-d", flag], port=server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
fastsync_tree = _tree(received)
assert fastsync_tree == rsync_tree, (
f"-d {flag}: fastsync tree {fastsync_tree} != rsync tree {rsync_tree}")
class TestOneFileSystemDeleteParity:
"""A9 side effect: the per-directory plan is now emitted only for directories
whose children were enumerated, so a ``-x`` mount-point directory that is
emitted but never traversed is shielded -- its destination content survives,
exactly as rsync keeps a non-descended mount point under ``--delete``."""
@requires_rsync
def test_mountpoint_content_survives_delete(self):
local = os.stat(".")
shm = "/dev/shm"
if not os.path.isdir(shm) or os.stat(shm).st_dev == local.st_dev:
pytest.skip("no cross-device filesystem available")
probe = os.path.join(shm, f"fastsync_dofs_{os.getpid()}")
clean_dir(probe)
_write(os.path.join(probe, "inside.txt"), b"cross\n")
try:
source = os.path.join(TEST_DATA_DIR, "dofs_src")
clean_dir(source)
_write(os.path.join(source, "keep.txt"), b"keep\n")
os.symlink(probe, os.path.join(source, "nested_link"))
def seed_dest(root):
clean_dir(root)
_write(os.path.join(root, "keep.txt"), b"old\n")
_write(os.path.join(root, "nested_link", "stale.txt"), b"stale\n")
rsync_dst = os.path.join(TEST_DATA_DIR, "dofs_rs_dst")
seed_dest(rsync_dst)
env = dict(os.environ, LC_ALL="C")
rsync_result = subprocess.run(
[RSYNC, "-a", "--copy-links", "-x", "--delete-during",
source + "/", rsync_dst + "/"],
capture_output=True, text=True, env=env, timeout=120)
assert rsync_result.returncode == 0, rsync_result.stderr
assert os.path.exists(os.path.join(rsync_dst, "nested_link", "stale.txt")), (
"rsync unexpectedly descended into the mount point")
dest = os.path.join(TEST_DATA_DIR, "dofs_fs_dst")
clean_dir(dest)
received = get_dest_received_dir(dest, source)
seed_dest(received)
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
result, _ = run_client(source, dest,
flags=["-a", "--copy-links", "-x", "--delete-during"],
port=server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
assert os.path.exists(os.path.join(received, "nested_link", "stale.txt")), (
"FastSync descended into a non-traversed mount point under --delete")
finally:
clean_dir(probe)
@@ -0,0 +1,185 @@
"""Differential coverage for ``--delete-delay`` + ``--max-delete`` with a
refilled deferred directory.
FastSync snapshots a directory's extras at plan time (``defer_add``) but charges
``--max-delete`` only when a path is actually removed, and its deferred commit
re-scans a queued directory and removes content created after the plan -- the
same rules as rsync. These tests run both tools on the same fixture and assert
both sides remove the late content (recursively) and bound the deletion with
``--max-delete`` identically.
They are not part of the fast PR gate because the rsync side needs a wide
real-time injection window (a throttled transfer), while the FastSync side uses
the existing byte-deterministic slicing proxy.
The refilled directory sits at the transfer ROOT, whose delete plan is always
processed before any subdirectory's, so the budget is deterministically charged
to the refilled entry; the second extra lives under ``b`` and is skipped.
"""
import os
import shutil
import subprocess
import sys
import threading
import time
import pytest
sys.path.insert(0, os.path.dirname(__file__))
from common import ( # noqa: E402
TEST_DATA_DIR,
ServerManager,
clean_dir,
get_dest_received_dir,
run_client,
)
from test_delete_timing_parity import _SlicingProxy # noqa: E402
RSYNC = shutil.which("rsync")
requires_rsync = pytest.mark.skipif(RSYNC is None, reason="rsync 3.4.1 not installed")
BIG_BYTES = 8 * 1024 * 1024
MID_TRANSFER_BYTES = 256 * 1024
PROXY_THROTTLE = 0.001
# rsync is driven locally, so the refill is injected on a wall-clock delay while
# a throttled ~8 s transfer is in flight. 1.5 s is safely after rsync's plan
# scan (t=0) and well before the deferred commit at the end.
RSYNC_BWLIMIT = 1024 # 1 MiB/s
RSYNC_INJECT_DELAY = 1.5
def _write(path, content):
os.makedirs(os.path.dirname(path), exist_ok=True)
with open(path, "wb") as fh:
fh.write(content)
def _seed_source(tag):
source = os.path.join(TEST_DATA_DIR, f"ddb_{tag}_src")
clean_dir(source)
_write(os.path.join(source, "a", "keep.bin"), b"B" * BIG_BYTES)
_write(os.path.join(source, "b", "keep.txt"), b"keep\n")
return source
def _seed_fastsync(tag):
"""FastSync mirrors the absolute source path under its receive root, so the
extras live below ``received``."""
source = _seed_source(tag)
dest = os.path.join(TEST_DATA_DIR, f"ddb_{tag}_dst")
clean_dir(dest)
received = get_dest_received_dir(dest, source)
os.makedirs(os.path.join(received, "xdir"), exist_ok=True)
os.makedirs(os.path.join(received, "b", "ydir"), exist_ok=True)
return source, dest, received
def _seed_rsync(tag):
"""rsync mirrors the source contents directly into the destination, so the
extras are flat under ``rsync_dst``."""
source = _seed_source(tag)
rsync_dst = os.path.join(TEST_DATA_DIR, f"ddb_{tag}_dst")
clean_dir(rsync_dst)
os.makedirs(os.path.join(rsync_dst, "xdir"), exist_ok=True)
os.makedirs(os.path.join(rsync_dst, "b", "ydir"), exist_ok=True)
return source, rsync_dst
def _deleted_count(text):
for line in text.splitlines():
if line.startswith("Number of deleted files:"):
return int(line.split(":", 1)[1].split()[0])
return None
def _rsync(args, timeout=120):
env = dict(os.environ, LC_ALL="C")
return subprocess.run([RSYNC] + args, capture_output=True, text=True, env=env, timeout=timeout)
class TestDeleteDelayRefilledDirVsRsync:
"""Both tools charge --max-delete on actual removals and recurse."""
def _fastsync_refilled(self, tag, max_delete=None):
"""Run FastSync with the refill injected deterministically by the proxy
hook (fired once the receiver has processed the plan frames)."""
source, dest, received = _seed_fastsync(tag)
late = os.path.join(received, "xdir", "new.txt")
def hook():
_write(late, b"created mid-transfer\n")
flags = ["--delete-delay", "--incremental", "--ignore-times", "--stats"]
if max_delete is not None:
flags.append(f"--max-delete={max_delete}")
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
proxy = _SlicingProxy(server.port, hook=hook, hook_after=MID_TRANSFER_BYTES,
throttle=PROXY_THROTTLE, wait_for_reply=True)
result, _ = run_client(source, dest, flags=flags, port=proxy.port)
proxy.finish()
assert proxy.hook_called.is_set(), "refill hook never fired"
return result, received, late
@requires_rsync
def test_max_delete_budget_bound_matches(self):
# --- FastSync: the one actual removal is the late file; dirs survive ---
result, received, late = self._fastsync_refilled("budget_fs", max_delete=1)
assert result.returncode == 25, (result.stderr or result.stdout)[:300]
assert _deleted_count(result.stdout) == 1, result.stdout
assert not os.path.exists(late), "FastSync kept the late content of a queued dir"
assert os.path.isdir(os.path.join(received, "xdir"))
assert os.path.isdir(os.path.join(received, "b", "ydir")), (
"FastSync did not bound the deletion with --max-delete=1"
)
# --- rsync: same budget rule and recursive removal ---
source, rsync_dst = _seed_rsync("budget_rsync")
def inject():
time.sleep(RSYNC_INJECT_DELAY)
_write(os.path.join(rsync_dst, "xdir", "new.txt"), b"created mid-transfer\n")
t = threading.Thread(target=inject)
t.start()
rsync_result = _rsync(
["-a", "--delete-delay", "--max-delete=1", "--stats",
f"--bwlimit={RSYNC_BWLIMIT}", source + "/", rsync_dst + "/"]
)
t.join()
assert rsync_result.returncode == 25, rsync_result.stderr
assert _deleted_count(rsync_result.stdout) == _deleted_count(result.stdout)
assert not os.path.exists(os.path.join(rsync_dst, "xdir", "new.txt")), (
"rsync kept late content inside a queued directory"
)
assert os.path.isdir(os.path.join(rsync_dst, "b", "ydir")), (
"rsync did not bound the deletion with --max-delete=1"
)
assert os.path.isdir(os.path.join(received, "b", "ydir"))
@requires_rsync
def test_refilled_extra_dir_recursive_removal_matches(self):
"""Without --max-delete both tools remove the refilled extra directory
(and its late content)."""
result, received, late = self._fastsync_refilled("recur_fs")
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
assert not os.path.exists(late), "FastSync kept the refilled directory's late content"
assert not os.path.isdir(os.path.join(received, "xdir"))
source, rsync_dst = _seed_rsync("recur_rsync")
def inject():
time.sleep(RSYNC_INJECT_DELAY)
_write(os.path.join(rsync_dst, "xdir", "new.txt"), b"created mid-transfer\n")
t = threading.Thread(target=inject)
t.start()
rsync_result = _rsync(
["-a", "--delete-delay", "--stats", f"--bwlimit={RSYNC_BWLIMIT}",
source + "/", rsync_dst + "/"]
)
t.join()
assert rsync_result.returncode == 0, rsync_result.stderr
assert not os.path.exists(os.path.join(rsync_dst, "xdir")), (
"rsync did not recursively remove the refilled extra directory"
)
+291 -19
View File
@@ -111,13 +111,20 @@ class _SlicingProxy:
"""
def __init__(self, target_port, forward_limit=None, hook=None, hook_after=0,
throttle=0.0, wait_for_reply=False):
throttle=0.0, wait_for_reply=False, hook_after_config_ack=False):
self.target = ("127.0.0.1", target_port)
self.forward_limit = forward_limit
self.hook = hook
self.hook_after = hook_after
self.throttle = throttle
self.wait_for_reply = wait_for_reply
# When set, the hook fires on the FIRST client->server bytes that follow
# the config-frame ack, BEFORE they are forwarded. For --delete-before
# those bytes are the keep-set manifest, so this runs the hook after the
# client's source pre-scan but before the receiver's delete ack releases
# the client into its data pass -- a deterministic late-file window.
self.hook_after_config_ack = hook_after_config_ack
self.config_acked = False
self.server_replied = threading.Event()
self.hook_called = threading.Event()
self.listener = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
@@ -165,6 +172,13 @@ class _SlicingProxy:
socks = []
break
data = data[:room]
if (self.hook_after_config_ack and self.config_acked and self.hook is not None
and not self.hook_called.is_set()):
# The first client bytes after the config ack are the
# pre-scan keep-set manifest: run the injection before
# forwarding so it is causally after the source scan.
self.hook()
self.hook_called.set()
backend.sendall(data)
forwarded += len(data)
self._maybe_hook(forwarded)
@@ -177,6 +191,7 @@ class _SlicingProxy:
client.sendall(data)
# Any server reply proves the receiver consumed the
# frames that precede it, so the hook barrier is met.
self.config_acked = True
self.server_replied.set()
self._maybe_hook(forwarded)
except OSError:
@@ -198,8 +213,11 @@ class _SlicingProxy:
def _maybe_hook(self, forwarded):
"""Fire the one-shot hook once its barrier is satisfied: enough client
bytes have been forwarded and, when ``wait_for_reply`` is set, the
server has sent a reply proving it processed the preceding frames."""
if self.hook is None or self.hook_called.is_set():
server has sent a reply proving it processed the preceding frames.
``hook_after_config_ack`` uses its own barrier (see ``_serve``), so the
byte/reply heuristic is bypassed entirely."""
if self.hook is None or self.hook_called.is_set() or self.hook_after_config_ack:
return
if forwarded < self.hook_after:
return
@@ -217,7 +235,21 @@ class _SlicingProxy:
class TestDeleteTimingFinalStateParity:
"""On a successful transfer the per-directory timings match rsync's result."""
"""On a successful transfer the per-directory timings match rsync's result.
Plain ``--delete`` has no rsync-incompatible spelling: it defaults to
delete-during on both tools, so it is compared against rsync's own default.
``--delete-commit`` is FastSync-only and selects the late whole-tree commit,
which is rsync's ``--delete-after`` timing.
"""
# (fastsync flag, rsync flag)
PAIRS = [
("--delete", "--delete"),
("--delete-during", "--delete-during"),
("--delete-delay", "--delete-delay"),
("--delete-commit", "--delete-after"),
]
def _run_fastsync(self, tag, timing):
source, dest, received = _seed_pair(tag)
@@ -226,28 +258,32 @@ class TestDeleteTimingFinalStateParity:
result, _ = run_client(source, dest, flags=[timing], port=server.port)
return result, received
@pytest.mark.parametrize("timing", ["--delete-during", "--delete-delay"])
@pytest.mark.parametrize("fs_timing,rs_timing", PAIRS)
@requires_rsync
def test_success_final_state_matches_rsync(self, timing):
def test_success_final_state_matches_rsync(self, fs_timing, rs_timing):
# Worker-safe names: xdist may run the parametrizations concurrently, so
# the flags are part of every fixture path.
label = f"{fs_timing.lstrip('-')}_vs_{rs_timing.lstrip('-')}"
# Build the rsync fixture from the same seed so both sides start equal.
source, dest, received = _seed_pair("parity_rsync")
source, dest, received = _seed_pair(f"parity_rsync_{label}")
source2 = source
rsync_dst = os.path.join(TEST_DATA_DIR, "dtp_parity_rsync_dst")
rsync_dst = os.path.join(TEST_DATA_DIR, f"dtp_rsync_{label}_dst")
clean_dir(rsync_dst)
# rsync mirrors src/ into dst/; seed the same extra.
_write(os.path.join(rsync_dst, "d", "old_extra"), b"stale extra\n")
rsync_result = _rsync(["-a", timing, source2 + "/", rsync_dst + "/"])
rsync_result = _rsync(["-a", rs_timing, source2 + "/", rsync_dst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
rsync_tree = _tree(rsync_dst)
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
result, _ = run_client(source, dest, flags=[timing], port=server.port)
result, _ = run_client(source, dest, flags=[fs_timing], port=server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
fastsync_tree = _tree(received)
assert fastsync_tree == rsync_tree, (
f"{timing}: fastsync tree {fastsync_tree} != rsync tree {rsync_tree}"
f"{fs_timing} vs rsync {rs_timing}: fastsync tree {fastsync_tree} != "
f"rsync tree {rsync_tree}"
)
@@ -258,7 +294,8 @@ class TestDeleteTimingTypeConflictParity:
@pytest.mark.parametrize("timing", ["--delete-during", "--delete-delay"])
@requires_rsync
def test_type_conflicts_match_rsync(self, timing):
source = os.path.join(TEST_DATA_DIR, "dtc_src")
label = timing.lstrip("-")
source = os.path.join(TEST_DATA_DIR, f"dtc_{label}_src")
clean_dir(source)
_write(os.path.join(source, "foo"), b"now a file\n")
_write(os.path.join(source, "bar", "inner.txt"), b"now a dir\n")
@@ -268,13 +305,13 @@ class TestDeleteTimingTypeConflictParity:
_write(os.path.join(root, "foo", "inner.txt"), b"was a dir\n")
_write(os.path.join(root, "bar"), b"was a file\n")
rsync_dst = os.path.join(TEST_DATA_DIR, "dtc_rsync_dst")
rsync_dst = os.path.join(TEST_DATA_DIR, f"dtc_{label}_rsync_dst")
seed_dest(rsync_dst)
rsync_result = _rsync(["-a", timing, source + "/", rsync_dst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
rsync_tree = _tree(rsync_dst)
dest = os.path.join(TEST_DATA_DIR, "dtc_dst")
dest = os.path.join(TEST_DATA_DIR, f"dtc_{label}_dst")
clean_dir(dest)
received = get_dest_received_dir(dest, source)
seed_dest(received)
@@ -288,17 +325,28 @@ class TestDeleteTimingTypeConflictParity:
class TestDeleteTimingFailure:
"""A mid-transfer failure distinguishes during from delay."""
"""A mid-transfer failure distinguishes the during timings from the late
commit timings.
Plain ``--delete`` must behave like ``--delete-during`` (the rsync default),
removing the extras of the directories already reached; ``--delete-commit``
must behave like ``--delete-after`` and remove nothing until the transfer
has fully succeeded.
"""
@pytest.mark.parametrize("mt", [False, True])
def test_during_removes_delay_preserves_on_failure(self, mt):
source, dest, received = _seed_pair("failure", big=True)
source, dest, received = _seed_pair(f"failure_mt{int(mt)}", big=True)
extra = os.path.join(received, "d", "old_extra")
assert os.path.exists(extra)
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
for timing, expect_removed in (("--delete-during", True),
("--delete-delay", False)):
for timing, expect_removed in (
("--delete-during", True),
("--delete", True),
("--delete-delay", False),
("--delete-commit", False),
("--delete-after", False)):
# Re-seed the extra before each run.
_write(extra, b"stale extra\n")
proxy = _SlicingProxy(server.port, forward_limit=MID_TRANSFER_BYTES, throttle=PROXY_THROTTLE)
@@ -313,13 +361,99 @@ class TestDeleteTimingFailure:
)
class TestDeleteDelayDeletedCount:
"""The reported deleted count must reflect entries actually removed."""
def test_refilled_deferred_dir_is_recursively_removed_and_counted(self):
"""A directory snapshotted into a --delete-delay plan that is refilled
before the commit is re-scanned and removed recursively (rsync parity):
the late file and the directory are both counted as deleted."""
source = os.path.join(TEST_DATA_DIR, "ddc_src")
dest = os.path.join(TEST_DATA_DIR, "ddc_dst")
clean_dir(source)
clean_dir(dest)
_write(os.path.join(source, "d", "keep.txt"), b"kept payload\n")
_write(os.path.join(source, "d", "big.bin"), b"B" * BIG_BYTES)
received = get_dest_received_dir(dest, source)
extra_dir = os.path.join(received, "d", "extradir")
os.makedirs(extra_dir, exist_ok=True)
def hook():
# Runs while big.bin is in flight, after d's delete plan was processed.
_write(os.path.join(extra_dir, "new.txt"), b"created mid-transfer\n")
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
proxy = _SlicingProxy(server.port, hook=hook, hook_after=MID_TRANSFER_BYTES,
throttle=PROXY_THROTTLE, wait_for_reply=True)
flags = ["--delete-delay", "--incremental", "--ignore-times", "--stats"]
result, _ = run_client(source, dest, flags=flags, port=proxy.port)
proxy.finish()
assert result.returncode == 0, (result.stderr or result.stdout)[:400]
assert proxy.hook_called.is_set(), "hook never fired"
assert not os.path.exists(os.path.join(extra_dir, "new.txt")), "late file survived"
assert not os.path.isdir(extra_dir), "refilled extra dir survived"
deleted = None
for line in result.stdout.splitlines():
if line.startswith("Number of deleted files:"):
deleted = int(line.split(":", 1)[1].split()[0])
assert deleted == 2, (deleted, result.stdout)
class TestDeleteDelayMaxDeleteParity:
"""--max-delete with --delete-delay: a partial deletion still reports the
number of entries actually removed, matching rsync (the exact surviving set
can differ; only the count is compared)."""
@requires_rsync
def test_max_delete_count_matches_rsync(self):
source = os.path.join(TEST_DATA_DIR, "ddm_src")
rsync_dst = os.path.join(TEST_DATA_DIR, "ddm_rsync_dst")
clean_dir(source)
clean_dir(rsync_dst)
_write(os.path.join(source, "d", "keep.txt"), b"keep\n")
for i in range(1, 6):
_write(os.path.join(rsync_dst, "d", f"e{i}.txt"), f"extra{i}\n".encode())
rsync_result = _rsync(["-a", "--delete-delay", "--max-delete=2", "--stats",
source + "/", rsync_dst + "/"])
# rsync exits 25 ("the --max-delete limit stopped deletions").
assert rsync_result.returncode == 25, rsync_result.stderr
rsync_count = _deleted_count(rsync_result.stdout)
assert rsync_count == 2, rsync_result.stdout
dest = os.path.join(TEST_DATA_DIR, "ddm_dst")
clean_dir(dest)
received = get_dest_received_dir(dest, source)
for i in range(1, 6):
_write(os.path.join(received, "d", f"e{i}.txt"), f"extra{i}\n".encode())
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
result, _ = run_client(
source, dest,
flags=["--delete-delay", "--max-delete=2", "--stats"],
port=server.port,
)
# A capped --max-delete commit is a successful transfer that both tools
# report with exit 25.
assert result.returncode == 25, (result.stderr or result.stdout)[:300]
assert _deleted_count(result.stdout) == rsync_count, result.stdout
def _deleted_count(text):
for line in text.splitlines():
if line.startswith("Number of deleted files:"):
return int(line.split(":", 1)[1].split()[0])
return None
class TestDeleteDelayVsAfterSnapshot:
"""A destination entry created after its directory's scan survives under
--delete-delay but is removed by --delete-after's fresh end scan."""
@pytest.mark.parametrize("mt", [False, True])
def test_late_created_extra_survives_delay_not_after(self, mt):
source, dest, received = _seed_pair("latecreate", big=True)
source, dest, received = _seed_pair(f"latecreate_mt{int(mt)}", big=True)
old_extra = os.path.join(received, "d", "old_extra")
new_extra = os.path.join(received, "d", "new_extra")
with ServerManager() as server:
@@ -356,3 +490,141 @@ class TestDeleteDelayVsAfterSnapshot:
f"{timing} (mt={mt}): new_extra present="
f"{os.path.exists(new_extra)}, expected survives={new_survives}"
)
class TestDeleteAfterThreadsKeepSet:
"""Regression: -j/--threads must still transmit the delete keep-set in every
timing. PipelineContextSender.delete_suppressed was left uninitialized, so a
garbage true silently skipped the late keep-set manifest under --threads.
Plain --delete now uses the per-directory plans, while --delete-commit /
--delete-after keep exercising the late whole-tree manifest."""
@pytest.mark.parametrize("delete_flag", ["--delete", "--delete-commit", "--delete-after"])
def test_threads_delete_after_sends_keep_set(self, delete_flag):
source, dest, received = _seed_pair("mtkeep")
extra = os.path.join(received, "d", "old_extra")
assert os.path.exists(extra)
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
result, _ = run_client(source, dest, flags=["--threads", delete_flag],
port=server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
assert not os.path.exists(extra), (
f"{delete_flag} --threads did not remove an extra: delete keep-set was suppressed"
)
class TestDeleteDelayMaxDeleteRefilledDir:
"""--delete-delay charges the --max-delete budget on ACTUAL removals: the
refilled directory's late content is removed first (consuming the one slot),
so the directory itself and a later extra are skipped, matching rsync.
The refilled directory is at the destination ROOT (its plan is always sent
first) and the skipped extra is under a separate source directory, so the
ordering that decides the budget charge is deterministic -- not readdir
order. The refill is injected through the byte-barrier proxy so it is
causally after the plan frame."""
def test_budget_charged_on_actual_removal(self):
source = os.path.join(TEST_DATA_DIR, "ddmb_src")
dest = os.path.join(TEST_DATA_DIR, "ddmb_dst")
clean_dir(source)
clean_dir(dest)
_write(os.path.join(source, "a", "keep.bin"), b"B" * BIG_BYTES)
_write(os.path.join(source, "b", "keep.txt"), b"keep\n")
received = get_dest_received_dir(dest, source)
refilled_dir = os.path.join(received, "xdir")
os.makedirs(refilled_dir, exist_ok=True)
later_dir = os.path.join(received, "b", "ydir")
os.makedirs(later_dir, exist_ok=True)
def hook():
_write(os.path.join(refilled_dir, "new.txt"), b"created mid-transfer\n")
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
proxy = _SlicingProxy(server.port, hook=hook, hook_after=MID_TRANSFER_BYTES,
throttle=PROXY_THROTTLE, wait_for_reply=True)
flags = ["--delete-delay", "--max-delete=1", "--incremental", "--ignore-times", "--stats"]
result, _ = run_client(source, dest, flags=flags, port=proxy.port)
proxy.finish()
assert result.returncode == 25, (result.stderr or result.stdout)[:400]
assert proxy.hook_called.is_set(), "hook never fired"
# The late content consumes the single budget slot; the refilled
# directory itself and the later extra are skipped.
assert not os.path.exists(os.path.join(refilled_dir, "new.txt")), "late file survived"
assert os.path.isdir(later_dir), "later extra was not skipped by the budget"
# The one actual removal is reported.
assert _deleted_count(result.stdout) == 1, result.stdout
class TestDeleteBeforeLateFileParity:
"""rsync builds its file list once, so a source file created after that scan
is NOT transferred and its destination extra is deleted. FastSync used to
re-scan the source in its data pass (single-threaded) or pipeline a fresh
re-scan against the pre-scan keep-set (``--threads``) and would transfer the
late file (a safe superset); both paths now replay the pre-scan file list
instead, matching rsync.
The late file is injected through the config-ack barrier: the first client
bytes after the config ack are the pre-scan keep-set manifest, so the hook
runs causally after the source scan and before the receiver's delete ack
releases the client into its data pass.
For ``--threads`` the pipeline scanner runs concurrently with the sender, so
the injection must land while that re-scan is still in flight to be observed
by it. The source is therefore a tree of ``_N_DIRS`` directories: the
injection writes the late file into EVERY directory, so it is enough that
any one directory is still unscanned when the hook fires. The tree is sized
so the hook (a localhost round trip) lands long before a full scan finishes;
a re-scanning pipeline then transfers the late files for the directories it
has not yet reached, which the tree comparison catches.
"""
_N_DIRS = 2000
@requires_rsync
@pytest.mark.parametrize("mt", [False, True])
def test_late_source_file_not_transferred_and_extra_deleted(self, mt):
tag = f"dblate_mt{int(mt)}"
source = os.path.join(TEST_DATA_DIR, f"{tag}_src")
dest = os.path.join(TEST_DATA_DIR, f"{tag}_dst")
rsync_dst = os.path.join(TEST_DATA_DIR, f"{tag}_rsync_dst")
clean_dir(source)
clean_dir(dest)
clean_dir(rsync_dst)
for i in range(self._N_DIRS):
_write(os.path.join(source, f"dir{i:05d}", "keep.txt"), b"kept payload\n")
# Both destinations carry the would-be late file as an extra.
for root in (dest, rsync_dst):
received = get_dest_received_dir(root, source)
for i in range(self._N_DIRS):
_write(os.path.join(received, f"dir{i:05d}", "late.txt"), b"stale extra\n")
# rsync reference: the same source with no late file; the extras are
# removed and nothing is transferred for the (never-scanned) late paths.
rsync_result = _rsync(["-a", "--delete-before", source + "/", rsync_dst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
rsync_tree = _tree(rsync_dst)
assert "dir00000/late.txt" not in rsync_tree
received = get_dest_received_dir(dest, source)
def hook():
# Runs after the pre-scan and before the receiver's delete ack.
for i in range(self._N_DIRS):
_write(os.path.join(source, f"dir{i:05d}", "late.txt"),
b"created after the scan\n")
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
proxy = _SlicingProxy(server.port, hook=hook, hook_after_config_ack=True)
flags = ["--delete-before"] + (["--threads=4"] if mt else [])
result, _ = run_client(source, dest, flags=flags, port=proxy.port)
proxy.finish()
assert result.returncode == 0, (result.stderr or result.stdout)[:400]
assert proxy.hook_called.is_set(), "late-file hook never fired"
assert _tree(received) == rsync_tree, (
f"late source {'multithreaded' if mt else 'single-threaded'} data pass re-scanned: "
f"{sum(1 for p in _tree(received) if p.endswith('late.txt'))} late files were "
"transferred"
)
@@ -0,0 +1,719 @@
"""Differential rsync-parity gate.
Runs real ``rsync 3.4.1`` and FastSync over the same corpora and flags, then
compares the destination trees and the normalized output of the
output-oriented flags. This is the executable counterpart of
``RSYNC_COMPAT.md``: the fast subset (``-m parity_ci``) guards the ✅ surface on
every pull request, and the full set (``-m parity``) burns the documented
⚠️/❌ residuals down.
Known, documented differences live in ``parity_caveats.py``; anything else
fails with a readable tree/stdout diff. A stale allowlist entry is reported
loudly (and fails when ``FASTSYNC_PARITY_STRICT=1``).
Run locally::
python3 -m pytest tests/integration/test_differential_parity.py -n 4 --dist=load -m parity_ci
python3 -m pytest tests/integration/test_differential_parity.py -n 4 --dist=load -m parity
"""
import os
import shutil
import sys
import warnings
import pytest
sys.path.insert(0, os.path.dirname(__file__))
from common import ( # noqa: E402
ServerManager,
TEST_DATA_DIR,
clean_dir,
get_dest_received_dir,
)
from parity_caveats import ASPECTS, caveat_for # noqa: E402
import parity_harness as H # noqa: E402
RSYNC = shutil.which("rsync")
requires_rsync = pytest.mark.skipif(RSYNC is None, reason="rsync 3.4.1 not installed")
# `--allow-super` matches the rest of the integration suite; `--allow-delete`
# is needed only by the delete cases.
SUPER = ("--allow-super",)
DELETE = ("--allow-super", "--allow-delete")
_OLD_MTIME = 1_500_000_000
parity = pytest.mark.parity
parity_ci = pytest.mark.parity_ci
@pytest.fixture(scope="session")
def parity_server_factory():
"""Lazily start one server per distinct extra-argument set, per xdist worker."""
servers = {}
def get(extra):
key = tuple(extra)
if key not in servers:
s = ServerManager()
s.start(extra_args=list(extra))
servers[key] = s
return servers[key]
yield get
for s in servers.values():
s.stop()
def _pin(path, mtime):
os.utime(path, (mtime, mtime))
def _mk(path, data, mtime=None):
os.makedirs(os.path.dirname(path), exist_ok=True)
with open(path, "wb") as fh:
fh.write(data)
if mtime is not None:
_pin(path, mtime)
# --- destination seeds ------------------------------------------------------
def seed_extras(_src, rroot, froot):
for root in (rroot, froot):
_mk(os.path.join(root, "extra.txt"), b"extra\n")
_mk(os.path.join(root, "extradir", "z.txt"), b"z\n")
def seed_update(_src, rroot, froot):
for root in (rroot, froot):
p = os.path.join(root, "a.txt")
_mk(p, b"destination is newer and longer\n", 2_000_000_000)
def seed_ignore_existing(_src, rroot, froot):
for root in (rroot, froot):
_mk(os.path.join(root, "a.txt"), b"destination-kept\n", _OLD_MTIME)
def seed_append(_src, rroot, froot):
for root in (rroot, froot):
_mk(os.path.join(root, "a.txt"), b"hello ", _OLD_MTIME)
def seed_backup(_src, rroot, froot):
for root in (rroot, froot):
_mk(os.path.join(root, "a.txt"), b"OLD-CONTENT\n", _OLD_MTIME)
def seed_size_only(_src, rroot, froot):
for root in (rroot, froot):
_mk(os.path.join(root, "a.txt"), b"XXXXXXXXXXX\n", _OLD_MTIME)
def seed_delete_excluded(_src, rroot, froot):
for root in (rroot, froot):
_mk(os.path.join(root, "drop.log"), b"stale log\n", _OLD_MTIME)
_mk(os.path.join(root, "extra.txt"), b"extra\n", _OLD_MTIME)
_mk(os.path.join(root, "keep.txt"), b"keep\n", _OLD_MTIME)
def seed_filter_protect(_src, rroot, froot):
"""Destination-only entries, including nested ones, for the receiver-side
`protect` rule: the `.log` extras must survive --delete, the rest go."""
for root in (rroot, froot):
_mk(os.path.join(root, "extra.log"), b"dest-only log\n", _OLD_MTIME)
_mk(os.path.join(root, "other.txt"), b"dest-only other\n", _OLD_MTIME)
_mk(os.path.join(root, "sub", "extra2.log"), b"nested dest-only log\n", _OLD_MTIME)
_mk(os.path.join(root, "sub", "other2.txt"), b"nested dest-only other\n", _OLD_MTIME)
def seed_max_delete(_src, rroot, froot):
for root in (rroot, froot):
_mk(os.path.join(root, "extra1.txt"), b"e1\n", _OLD_MTIME)
_mk(os.path.join(root, "extra2.txt"), b"e2\n", _OLD_MTIME)
def fuzzy_basis_seed(_src, rroot, froot):
"""Seed a same-suffix sibling whose name is one edit from the source and
whose content matches it, with a DIFFERENT mtime so rsync's exact
size+mtime pass cannot fire: both tools must select it via the
name-distance pass. Where the two tools' basis choices coincide the
block-level results are identical when the block size is pinned."""
for root in (rroot, froot):
_mk(os.path.join(root, "report_v1.txt"), H.FUZZY_PAYLOAD, _OLD_MTIME)
def max_delete_count_check(_src, rroot, froot, _rs, _fs):
"""The exact survivor set is order-dependent; the count must still match."""
r = H.snapshot(rroot)
f = H.snapshot(froot)
if len(r) != len(f):
return [f"survivor count differs: rsync={len(r)} fastsync={len(f)}"]
return []
# --- case table -------------------------------------------------------------
_CASES = [
# --- core archive / recursion -----------------------------------------
H.Case("archive", "basic", ["-a"], ci=True, ref="-a/--archive"),
H.Case("recursive", "basic", ["-r"], ci=True, ref="-r/--recursive"),
H.Case("unicode_names", "unicode", ["-a"], ci=True, ref="-a unicode names"),
H.Case("links_archive", "links", ["-a"], ci=True, ref="-l/--links"),
H.Case("copy_links", "links", ["-aL"], ref="-L/--copy-links"),
H.Case("hardlinks", "hardlinks", ["-a", "-H"], compare_hardlinks=True,
ci=True, ref="-H/--hard-links"),
H.Case("hardlinks_without_H", "hardlinks", ["-a"], compare_hardlinks=True,
ref="hardlinks without -H"),
H.Case("sparse", "sparse", ["-a", "-S"], ref="-S/--sparse"),
# --- compression / checksums ------------------------------------------
H.Case("compress_zstd", "basic", ["-a", "-z"], ci=True, ref="-z/--compress"),
H.Case("checksum", "basic", ["-a", "-c"], ref="-c/--checksum"),
H.Case("checksum_choice_xxh64", "basic",
["-a", "-c", "--checksum-choice=xxh64"], ref="--checksum-choice"),
# --- selection --------------------------------------------------------
H.Case("exclude", "filters", ["-a", "--exclude=*.log"], ci=True,
ref="--exclude"),
H.Case("include_exclude", "filters",
["-a", "--include=*.txt", "--exclude=*"], ci=True,
ref="--include/--exclude ordering"),
H.Case("filter_rules", "filters",
["-a", "-f", "- *.log", "-f", "+ *.txt", "-f", "- *"],
ref="--filter/-f grammar"),
H.Case("max_size", "basic", ["-a", "--max-size=1000"], ref="--max-size"),
H.Case("min_size", "basic", ["-a", "--min-size=1000"], ref="--min-size"),
# --- output-oriented --------------------------------------------------
H.Case("stats", "basic", ["-a", "--stats"], stdout=H.STDOUT_STATS,
ci=True, ref="--stats"),
H.Case("itemize", "links", ["-a", "-i"], stdout=H.STDOUT_ITEMIZE,
ci=True, ref="-i/--itemize-changes"),
H.Case("out_format_n_l", "basic", ["-a", "--out-format=%n %l"],
stdout=H.STDOUT_OUTFMT, ref="--out-format %n %l"),
H.Case("out_format_i_n", "basic", ["-a", "--out-format=%i %n"],
stdout=H.STDOUT_OUTFMT, ref="--out-format %i %n"),
H.Case("progress", "multidir", ["-a", "--progress"], stdout=H.STDOUT_PROGRESS,
ci=True, ref="--progress multi-directory file list"),
H.Case("progress_threads", "multidir", ["-a", "--progress"],
fastsync_flags=["-a", "--progress", "--threads"],
stdout=H.STDOUT_PROGRESS, ci=True,
ref="--progress multi-directory file list (--threads)"),
# --- transfer modifications -------------------------------------------
H.Case("update", "basic", ["-a", "--update"], seed=seed_update,
ref="-u/--update"),
H.Case("ignore_existing", "basic", ["-a", "--ignore-existing"],
seed=seed_ignore_existing, ci=True, ref="--ignore-existing"),
H.Case("size_only", "basic",
["-a", "--size-only"], fastsync_flags=["-a", "--incremental", "--size-only"],
seed=seed_size_only, ref="--size-only"),
H.Case("append", "basic", ["-a", "--append"], seed=seed_append,
ref="--append"),
H.Case("append_verify", "basic", ["-a", "--append-verify"], seed=seed_append,
ref="--append-verify"),
H.Case("backup", "basic", ["-a", "--backup"], seed=seed_backup,
ref="--backup"),
H.Case("chmod", "basic", ["-a", "--chmod=Fu+rwx"], compare_modes=True,
ci=True, ref="--chmod"),
# --- delta / similar-file basis (--fuzzy) -----------------------------
# Basis choices coincide here (same-suffix sibling, name distance one edit,
# content identical); with the block size pinned both tools report the same
# Matched/Literal/transferred counters. The residual (FastSync's narrower
# delta size window) is covered by TestFuzzy in test_parity_quickwins.py.
H.Case("fuzzy_basis", "fuzzy",
["-a", "--no-whole-file", "--fuzzy", "--stats", "-B8192"],
fastsync_flags=["-a", "--incremental", "--delta", "--fuzzy",
"--stats", "--delta-block=8192"],
seed=fuzzy_basis_seed, stdout=H.STDOUT_STATS, ci=True,
ref="-y/--fuzzy similar-file basis"),
# --- deletion ---------------------------------------------------------
H.Case("delete", "basic", ["-a", "--delete"], seed=seed_extras,
server_args=DELETE, ci=True, ref="--delete"),
H.Case("delete_before", "basic", ["-a", "--delete-before"], seed=seed_extras,
server_args=DELETE, ref="--delete-before"),
H.Case("delete_during", "basic", ["-a", "--delete-during"], seed=seed_extras,
server_args=DELETE, ref="--delete-during"),
H.Case("delete_delay", "basic", ["-a", "--delete-delay"], seed=seed_extras,
server_args=DELETE, ref="--delete-delay"),
H.Case("delete_after", "basic", ["-a", "--delete-after"], seed=seed_extras,
server_args=DELETE, ref="--delete-after"),
H.Case("delete_commit", "basic", ["-a", "--delete-after"], seed=seed_extras,
fastsync_flags=["-a", "--delete-commit"], server_args=DELETE,
ref="FastSync-only --delete-commit == rsync --delete-after"),
H.Case("delete_excluded", "filters",
["-a", "--delete", "--delete-excluded", "--exclude=*.log"],
seed=seed_delete_excluded, server_args=DELETE, ref="--delete-excluded"),
H.Case("exclude_protect_dest_only", "filters",
["-a", "--delete", "--exclude=*.log"],
seed=seed_delete_excluded, server_args=DELETE, ci=True,
ref="--delete protects a destination-only excluded entry like rsync"),
H.Case("max_delete", "basic", ["-a", "--delete", "--max-delete=1"],
seed=seed_max_delete, server_args=DELETE,
extra_check=max_delete_count_check, compare_tree=False,
ref="--max-delete"),
H.Case("filter_protect", "filters",
["-a", "--delete", "--filter=P *.log"],
seed=seed_filter_protect, server_args=DELETE, ci=True,
ref="--filter P/--protect receiver-side delete protection (default during)"),
H.Case("filter_protect_during", "filters",
["-a", "--delete-during", "--filter=P *.log"],
seed=seed_filter_protect, server_args=DELETE, ci=True,
ref="--filter P/--protect under --delete-during"),
H.Case("filter_protect_delay", "filters",
["-a", "--delete-delay", "--filter=P *.log"],
seed=seed_filter_protect, server_args=DELETE, ci=True,
ref="--filter P/--protect under --delete-delay"),
H.Case("filter_protect_after", "filters",
["-a", "--delete-after", "--filter=P *.log"],
seed=seed_filter_protect, server_args=DELETE, ci=True,
ref="--filter P/--protect under the whole-tree --delete-after commit"),
# --- relative / dirs --------------------------------------------------
H.Case("relative_general", "basic", ["-a", "-R"], layout=H.MIRROR_ABS,
compare_modes=True, ref="-R/--relative"),
H.Case("relative_no_implied_dirs", "basic",
["-a", "-R", "--no-implied-dirs"], layout=H.MIRROR_ABS,
compare_modes=True, ref="--no-implied-dirs"),
H.Case("files_from", "relative", ["--dirs", "-R"],
files_from=("dir1", "sub/x.txt"), layout=H.RELATIVE, ci=True,
ref="-d/--dirs + --files-from"),
H.Case("dirs_plain", "basic", ["-d"], fs_src_suffix="/",
ref="-d/--dirs (plain)"),
H.Case("empty_dirs_recursive", "empty_dir", ["-a"],
ref="recursive empty-directory residual"),
H.Case("empty_dirs_files_from", "empty_dir", ["--dirs", "-R"],
files_from=("emptydir",), layout=H.RELATIVE, ci=True,
ref="-d/--dirs explicit empty directory"),
# --- codecs -----------------------------------------------------------
H.Case("iconv_identity", "basic", ["-a", "--iconv=UTF-8,UTF-8"],
ref="--iconv identity"),
H.Case("iconv_convert", "iconv",
["-a", "--iconv=ISO-8859-1,UTF-8"],
server_args=("--allow-super", "--iconv=UTF-8"),
ref="--iconv conversion (receiver declares its own charset)"),
# rsync's spec is LOCAL,REMOTE and the destination end's charset is REMOTE
# on a push, so a default server writes the wire (UTF-8) names verbatim.
H.Case("iconv_default_server", "iconv",
["-a", "--iconv=ISO-8859-1,UTF-8"],
ref="--iconv push direction (default receiver charset = REMOTE)"),
# --- partial ----------------------------------------------------------
H.Case("partial_complete", "basic", ["-a", "--partial"], ref="--partial"),
]
# Cases that must always be tolerated (documented ⚠️/❌ residuals) get an
# allowlist entry; the table below stays the exact ✅ surface.
ALL_CASES = _CASES
def _params():
out = []
for case in ALL_CASES:
marks = [parity]
if case.ci:
marks.append(parity_ci)
out.append(pytest.param(case, id=case.id, marks=marks))
return out
def _aspects_to_check(result):
return {
"tree": result["tree"],
"stdout": result["stdout"],
"extra": result["extra"],
}
def _assert_no_unexpected(case_id, mismatches, caveat, ref=""):
unexpected = {a: v for a, v in mismatches.items() if v and a not in caveat}
if unexpected:
lines = [f"differential parity mismatch for case {case_id!r}:"]
lines.append(f" ref: {ref or 'see RSYNC_COMPAT.md'}")
for aspect, detail in unexpected.items():
lines.append(f" --- {aspect} ---")
lines.extend(" " + str(d) for d in detail)
lines.append("If this is a documented residual, add it to "
"tests/integration/parity_caveats.py with a RSYNC_COMPAT.md "
"reference. Do not allowlist an undocumented divergence.")
pytest.fail("\n".join(lines))
stale = [a for a in ASPECTS
if a in caveat and a != "rc" and not mismatches.get(a)]
if stale:
msg = (f"stale parity allowlist entry for case {case_id!r}, aspect(s) "
f"{stale}: FastSync now matches rsync. Remove it from "
f"parity_caveats.py (and update RSYNC_COMPAT.md if the row moved).")
if os.environ.get("FASTSYNC_PARITY_STRICT") == "1":
pytest.fail(msg)
warnings.warn(msg, stacklevel=2)
@requires_rsync
@pytest.mark.parametrize("case", _params())
def test_differential_case(case, parity_server_factory):
server = parity_server_factory(case.server_args)
result = H.execute_case(case, server)
caveat = caveat_for(case.id)
mismatches = _aspects_to_check(result)
if result["rsync_rc"] != result["fastsync_rc"]:
mismatches["rc"] = [
f"rsync rc={result['rsync_rc']} fastsync rc={result['fastsync_rc']} "
f"(rsync stderr: {result['rsync_stderr'][:200]!r}, "
f"fastsync stderr: {result['fastsync_stderr'][:200]!r})"]
_assert_no_unexpected(case.id, mismatches, caveat, ref=case.ref)
# ---------------------------------------------------------------------------
# Multi-run and setup-heavy scenarios (kept as explicit tests)
# ---------------------------------------------------------------------------
def _result_aspects(result):
return _aspects_to_check(result)
_STANDALONE_REFS = {
"incremental_modified": "-i/--itemize-changes + incremental second run",
"compare_dest": "--compare-dest",
"copy_dest": "--copy-dest",
"link_dest": "--link-dest",
"link_dest_stats": "--link-dest + --stats",
"verify_basis": "--verify-basis (FastSync-only)",
"verify_basis_default": "--verify-basis (default quick-check vs rsync)",
"added_and_deleted": "--delete across two runs",
"added_and_deleted_seed": "--delete across two runs",
"one_file_system": "-x/--one-file-system",
}
def _run_and_check(case_id, result, ref=""):
mismatches = _result_aspects(result)
if result["rsync_rc"] != result["fastsync_rc"]:
mismatches["rc"] = [
f"rsync rc={result['rsync_rc']} fastsync rc={result['fastsync_rc']} "
f"(rsync stderr: {result['rsync_stderr'][:200]!r}, "
f"fastsync stderr: {result['fastsync_stderr'][:200]!r})"]
_assert_no_unexpected(case_id, mismatches, caveat_for(case_id),
ref=ref or _STANDALONE_REFS.get(case_id, ""))
@requires_rsync
@parity
def test_incremental_modified_file(parity_server_factory):
"""A second run sends only the modified file; destinations stay identical."""
case_id = "incremental_modified"
src = os.path.join(TEST_DATA_DIR, "parity_inc_src")
rdst = os.path.join(TEST_DATA_DIR, "parity_inc_rdst")
fdst = os.path.join(TEST_DATA_DIR, "parity_inc_fdst")
H.CORPORA["basic"](src)
server = parity_server_factory(SUPER)
# Seed both destinations with the initial content.
H.run_differential(src, rdst, fdst, ["-a"], ["-a"], server,
extra_check=lambda *a: [])
with open(os.path.join(src, "a.txt"), "wb") as fh:
fh.write(b"hello world, now modified and longer\n")
_pin(os.path.join(src, "a.txt"), 1_650_000_000)
result = H.run_differential(
src, rdst, fdst, ["-a", "-i"], ["-a", "-i", "--incremental"], server,
stdout=H.STDOUT_ITEMIZE)
_run_and_check(case_id, result)
def _seed_basis(rel_entries):
def seed(src, rroot, froot):
for root in (rroot, froot):
os.makedirs(root, exist_ok=True)
for rel, data in rel_entries.items():
_mk(os.path.join(root, rel), data)
return seed
@requires_rsync
@parity
def test_compare_dest_skips_basis(parity_server_factory):
"""--compare-dest: a file present in the basis is not copied."""
case_id = "compare_dest"
src = os.path.join(TEST_DATA_DIR, "parity_cmpd_src")
rdst = os.path.join(TEST_DATA_DIR, "parity_cmpd_rdst")
fdst = os.path.join(TEST_DATA_DIR, "parity_cmpd_fdst")
clean_dir(src)
_mk(os.path.join(src, "f.txt"), b"basis-content\n")
_pin(os.path.join(src, "f.txt"), _OLD_MTIME)
server = parity_server_factory(SUPER)
rel = os.path.abspath(src).lstrip(os.sep)
# rsync resolves --compare-dest relative to the destination dir; FastSync
# resolves it under the receive root and appends the mirrored source path.
# Both rely on rsync's size+mtime quick-check, so the basis mtime is pinned
# to the source's to keep the match deterministic across a second boundary.
def seed(_src, rroot, froot):
_mk(os.path.join(rroot, "basis", "f.txt"), b"basis-content\n", _OLD_MTIME)
_mk(os.path.join(fdst, "basis", rel, "f.txt"), b"basis-content\n", _OLD_MTIME)
def extra(_src, rroot, froot, _rs, _fs):
out = []
for label, root in (("rsync", rroot), ("fastsync", froot)):
if os.path.exists(os.path.join(root, "f.txt")):
out.append(f"{label} copied a file that is present in the "
f"compare basis")
return out
result = H.run_differential(
src, rdst, fdst,
["-a", "--compare-dest=basis"],
["-a", f"--compare-dest={os.path.join(fdst, 'basis')}", "--incremental"],
server, seed=seed, ignore_paths=("basis",), extra_check=extra)
_run_and_check(case_id, result)
@requires_rsync
@parity
def test_link_dest_hardlinks_basis(parity_server_factory):
"""--link-dest: an unchanged file is hard-linked to the basis, not copied."""
case_id = "link_dest"
src = os.path.join(TEST_DATA_DIR, "parity_linkd_src")
rdst = os.path.join(TEST_DATA_DIR, "parity_linkd_rdst")
fdst = os.path.join(TEST_DATA_DIR, "parity_linkd_fdst")
clean_dir(src)
_mk(os.path.join(src, "f.txt"), b"link-basis-content\n")
_pin(os.path.join(src, "f.txt"), _OLD_MTIME)
server = parity_server_factory(SUPER)
rel = os.path.abspath(src).lstrip(os.sep)
def seed(_src, rroot, froot):
_mk(os.path.join(rroot, "basis", "f.txt"), b"link-basis-content\n", _OLD_MTIME)
_mk(os.path.join(fdst, "basis", rel, "f.txt"), b"link-basis-content\n", _OLD_MTIME)
def extra(_src, rroot, froot, _rs, _fs):
r_basis = os.stat(os.path.join(rroot, "basis", "f.txt")).st_ino
f_basis = os.stat(os.path.join(fdst, "basis", rel, "f.txt")).st_ino
out = []
for label, root, basis in (("rsync", rroot, r_basis),
("fastsync", froot, f_basis)):
target = os.path.join(root, "f.txt")
if not os.path.exists(target):
out.append(f"{label}: f.txt missing")
elif os.stat(target).st_ino != basis:
out.append(f"{label}: f.txt is not hard-linked to the basis")
return out
result = H.run_differential(
src, rdst, fdst,
["-a", "--link-dest=basis"],
["-a", f"--link-dest={os.path.join(fdst, 'basis')}", "--incremental"],
server, seed=seed, ignore_paths=("basis",), extra_check=extra)
_run_and_check(case_id, result)
@requires_rsync
@parity
def test_link_dest_stats_matches_rsync(parity_server_factory):
"""A basis hit must not be counted as created or literal data: rsync reports
zero for both, so FastSync's receiver tallies must too (regression for the
basis materialization over-report)."""
case_id = "link_dest_stats"
src = os.path.join(TEST_DATA_DIR, "parity_linkds_src")
rdst = os.path.join(TEST_DATA_DIR, "parity_linkds_rdst")
fdst = os.path.join(TEST_DATA_DIR, "parity_linkds_fdst")
clean_dir(src)
_mk(os.path.join(src, "f.txt"), b"link-basis-content\n")
_pin(os.path.join(src, "f.txt"), _OLD_MTIME)
server = parity_server_factory(SUPER)
rel = os.path.abspath(src).lstrip(os.sep)
def seed(_src, rroot, froot):
_mk(os.path.join(rroot, "basis", "f.txt"), b"link-basis-content\n", _OLD_MTIME)
_mk(os.path.join(fdst, "basis", rel, "f.txt"), b"link-basis-content\n", _OLD_MTIME)
result = H.run_differential(
src, rdst, fdst,
["-a", "--link-dest=basis", "--stats"],
["-a", f"--link-dest={os.path.join(fdst, 'basis')}", "--incremental", "--stats"],
server, seed=seed, ignore_paths=("basis",), stdout=H.STDOUT_STATS)
_run_and_check(case_id, result, ref="--link-dest + --stats")
@requires_rsync
@parity
def test_copy_dest_copies_basis(parity_server_factory):
"""--copy-dest: a basis match is materialized as an independent copy with the
source's attributes, matching rsync (copy then fix attributes)."""
case_id = "copy_dest"
src = os.path.join(TEST_DATA_DIR, "parity_copyd_src")
rdst = os.path.join(TEST_DATA_DIR, "parity_copyd_rdst")
fdst = os.path.join(TEST_DATA_DIR, "parity_copyd_fdst")
clean_dir(src)
_mk(os.path.join(src, "f.txt"), b"copy-basis-content\n")
_pin(os.path.join(src, "f.txt"), 1_600_000_000)
os.chmod(os.path.join(src, "f.txt"), 0o755)
server = parity_server_factory(SUPER)
rel = os.path.abspath(src).lstrip(os.sep)
def seed(_src, rroot, froot):
# Basis content matches the source; give the basis a different mode so a
# wrong "keep basis attributes" implementation is visible.
_mk(os.path.join(rroot, "basis", "f.txt"), b"copy-basis-content\n",
1_600_000_000)
os.chmod(os.path.join(rroot, "basis", "f.txt"), 0o644)
_mk(os.path.join(fdst, "basis", rel, "f.txt"), b"copy-basis-content\n",
1_600_000_000)
os.chmod(os.path.join(fdst, "basis", rel, "f.txt"), 0o644)
def extra(_src, rroot, froot, _rs, _fs):
out = []
bases = {"rsync": os.path.join(rroot, "basis", "f.txt"),
"fastsync": os.path.join(fdst, "basis", rel, "f.txt")}
for label, root in (("rsync", rroot), ("fastsync", froot)):
target = os.path.join(root, "f.txt")
if not os.path.exists(target):
out.append(f"{label}: f.txt missing")
continue
if os.stat(target).st_ino == os.stat(bases[label]).st_ino:
out.append(f"{label}: f.txt is hard-linked, not copied")
if (os.stat(target).st_mode & 0o777) != 0o755:
out.append(f"{label}: f.txt mode "
f"{oct(os.stat(target).st_mode & 0o777)} != 0o755")
return out
result = H.run_differential(
src, rdst, fdst,
["-a", "--copy-dest=basis"],
["-a", f"--copy-dest={os.path.join(fdst, 'basis')}", "--incremental"],
server, seed=seed, ignore_paths=("basis",), extra_check=extra,
compare_modes=True)
_run_and_check(case_id, result)
@requires_rsync
@parity
def test_verify_basis_restores_strict_content(parity_server_factory):
"""Default matches rsync's metadata quick-check; FastSync-only
`--verify-basis` restores strict content equality and transfers the source
when a same-size/different-content basis would otherwise be trusted."""
case_id = "verify_basis"
src = os.path.join(TEST_DATA_DIR, "parity_vbasis_src")
rdst = os.path.join(TEST_DATA_DIR, "parity_vbasis_rdst")
fdst = os.path.join(TEST_DATA_DIR, "parity_vbasis_fdst")
clean_dir(src)
_mk(os.path.join(src, "f.txt"), b"AAAA\n")
_pin(os.path.join(src, "f.txt"), _OLD_MTIME)
server = parity_server_factory(SUPER)
rel = os.path.abspath(src).lstrip(os.sep)
def seed(_src, rroot, froot):
# Same size and mtime as the source, different bytes: a metadata
# quick-check trusts it; --verify-basis must not.
for root, basis_rel in ((rroot, os.path.join("basis", "f.txt")),
(fdst, os.path.join("basis", rel, "f.txt"))):
_mk(os.path.join(root, basis_rel), b"BBBB\n", _OLD_MTIME)
# Default: both tools trust the basis (rsync's quick check), so the
# destination carries the basis bytes and the trees match.
result = H.run_differential(
src, rdst, fdst,
["-a", "--link-dest=basis"],
["-a", f"--link-dest={os.path.join(fdst, 'basis')}", "--incremental"],
server, seed=seed, ignore_paths=("basis",))
_run_and_check(case_id + "_default", result)
# --verify-basis (FastSync only): the digest mismatch rejects the basis and
# the source is transferred, so the destination is the source bytes. rsync
# has no such flag; assert the FastSync outcome directly against the source.
fdst2 = os.path.join(TEST_DATA_DIR, "parity_vbasis_fdst2")
clean_dir(fdst2)
for root, basis_rel in ((fdst2, os.path.join("basis", rel, "f.txt")),):
_mk(os.path.join(root, basis_rel), b"BBBB\n", _OLD_MTIME)
result, _ = H.run_fastsync(src, fdst2,
["-a", f"--link-dest={os.path.join(fdst2, 'basis')}",
"--incremental", "--verify-basis"], server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
target = os.path.join(get_dest_received_dir(fdst2, src), "f.txt")
with open(target, "rb") as fh:
assert fh.read() == b"AAAA\n", \
"--verify-basis must reject the same-size/different-content basis"
@requires_rsync
@parity
def test_added_and_deleted_between_runs(parity_server_factory):
"""A source deletion and addition sync correctly under --delete."""
case_id = "added_and_deleted"
src = os.path.join(TEST_DATA_DIR, "parity_addel_src")
rdst = os.path.join(TEST_DATA_DIR, "parity_addel_rdst")
fdst = os.path.join(TEST_DATA_DIR, "parity_addel_fdst")
server = parity_server_factory(DELETE)
H.CORPORA["basic"](src)
seed = seed_extras
result = H.run_differential(
src, rdst, fdst, ["-a", "--delete"], ["-a", "--delete"], server,
seed=seed)
_run_and_check(case_id + "_seed", result)
os.remove(os.path.join(src, "a.txt"))
_mk(os.path.join(src, "added.txt"), b"added between runs\n")
result = H.run_differential(
src, rdst, fdst, ["-a", "--delete", "-i"],
["-a", "--delete", "-i", "--incremental"], server,
stdout=H.STDOUT_ITEMIZE)
_run_and_check(case_id, result)
@requires_rsync
@parity
def test_one_file_system(parity_server_factory):
"""-x emits the mount-point directory but not its contents."""
case_id = "one_file_system"
local = os.stat(".")
shm = "/dev/shm"
if not os.path.isdir(shm):
pytest.skip("/dev/shm not available")
if os.stat(shm).st_dev == local.st_dev:
pytest.skip("no cross-device filesystem available")
src = os.path.join(TEST_DATA_DIR, "parity_ofs_src")
rdst = os.path.join(TEST_DATA_DIR, "parity_ofs_rdst")
fdst = os.path.join(TEST_DATA_DIR, "parity_ofs_fdst")
clean_dir(src)
_mk(os.path.join(src, "keep.txt"), b"keep\n")
probe = os.path.join(shm, f"fastsync_ofs_{os.getpid()}")
shutil.rmtree(probe, ignore_errors=True)
os.makedirs(probe)
_mk(os.path.join(probe, "inside.txt"), b"cross\n")
try:
os.symlink(probe, os.path.join(src, "nested_link"))
server = parity_server_factory(SUPER)
result = H.run_differential(
src, rdst, fdst,
["-a", "--copy-links", "-x"],
["-a", "--copy-links", "-x"], server)
_run_and_check(case_id, result)
finally:
shutil.rmtree(probe, ignore_errors=True)
@requires_rsync
@parity
def test_parity_caveats_reference_known_cases():
"""Every allowlist entry must name a real case id and aspect."""
from parity_caveats import CAVEATS
known = {c.id for c in ALL_CASES} | {
"incremental_modified", "compare_dest", "link_dest",
"added_and_deleted", "added_and_deleted_seed", "one_file_system",
}
problems = []
for case_id, entry in CAVEATS.items():
if case_id not in known:
problems.append(f"unknown case id in parity_caveats.py: {case_id!r}")
for aspect in entry:
if aspect not in ASPECTS:
problems.append(
f"{case_id!r}: unknown aspect {aspect!r} (expected {ASPECTS})")
assert not problems, "\n".join(problems)
+1 -1
View File
@@ -36,7 +36,7 @@ from common import ( # noqa: E402
verify_transfer,
)
PROTOCOL_VERSION = b"2.26.0"
PROTOCOL_VERSION = b"2.29.0"
STATUS_MANIFEST = 5
STATUS_OK = 0
+557 -132
View File
@@ -183,10 +183,11 @@ class TestDeviceSpecial:
)
@pytest.mark.setpriv
def test_devices_nonroot_receiver_skips_safely(self):
"""A receiver without CAP_MKNOD must skip a device entry with a warning
and never abort. A root runner drops the receiver (server) to nobody
via setpriv; on a non-root runner (or without setpriv) the test skips."""
def test_devices_nonroot_receiver_errors_like_rsync(self):
"""A receiver without CAP_MKNOD must report the failed device mknod as a
transfer error (rsync parity, partial failure) instead of silently
succeeding. A root runner drops the receiver (server) to nobody via
setpriv; on a non-root runner (or without setpriv) the test skips."""
if os.geteuid() != 0 or shutil.which("setpriv") is None:
pytest.skip("requires root + setpriv to run the receiver unprivileged")
self._setup()
@@ -202,16 +203,16 @@ class TestDeviceSpecial:
flags=["--devices"], port=port)
finally:
out, err = _stop_captured_server(server)
assert result.returncode == 0, f"Exit {result.returncode}: {result.stderr[:300]}"
received = get_dest_received_dir(DEVICE_DEST, DEVICE_SOURCE)
with open(os.path.join(received, "plain.txt")) as f:
assert f.read() == "regular content\n"
assert not os.path.lexists(os.path.join(received, "chardev")), (
"a receiver without CAP_MKNOD must skip the device node, not create it"
assert result.returncode != 0, (
f"a failed device mknod must be a transfer error like rsync (got exit 0): "
f"{(out + err)[:300]}"
)
assert ("cannot create device node" in (out + err)
or "device-node creation is not permitted" in (out + err)), (
f"receiver did not log the documented device skip: out={out!r} err={err!r}"
received = get_dest_received_dir(DEVICE_DEST, DEVICE_SOURCE)
assert not os.path.lexists(os.path.join(received, "chardev")), (
"a receiver without CAP_MKNOD must not create the device node"
)
assert "cannot create device" in (out + err), (
f"receiver did not log the device creation error: out={out!r} err={err!r}"
)
@pytest.mark.skipif(os.geteuid() != 0, reason="requires root to create device nodes")
@@ -583,6 +584,57 @@ class TestRemoteDryRun:
assert os.path.exists(extra), f"{flags} deleted an extra in dry-run"
assert _snapshot_tree(received) == before, f"{flags} mutated the destination"
@pytest.mark.skipif(shutil.which("rsync") is None, reason="rsync not installed")
def test_dry_run_delete_lines_match_rsync(self):
"""`-n --delete` lists exactly the destination extras rsync would remove.
Covers the three cases that a real run protects: the file being updated
(in the keep set), a filter-excluded source entry (protected prefix), and
a --max-size-pruned source entry (always-protected prefix). Track 4a
adds a fourth: a destination-only entry matching the exclude rule is
re-derived on the receiver and also protected, so only the genuine
destination-only `extra.txt` appears.
"""
source = os.path.join(TEST_DATA_DIR, "dryrep_src")
rdst = os.path.join(TEST_DATA_DIR, "dryrep_rdst")
fdst = os.path.join(TEST_DATA_DIR, "dryrep_fdst")
clean_dir(source)
clean_dir(rdst)
clean_dir(fdst)
for name, data in (("a.txt", b"new content\n"), ("keep.log", b"log\n"),
("big.bin", b"B" * 2000)):
with open(os.path.join(source, name), "wb") as fh:
fh.write(data)
os.utime(os.path.join(source, "a.txt"), (1_700_000_000, 1_700_000_000))
received = get_dest_received_dir(fdst, source)
os.makedirs(received, exist_ok=True)
for root in (rdst, received):
for name, data in (("a.txt", b"old\n"), ("keep.log", b"log\n"),
("big.bin", b"B" * 2000), ("extra.txt", b"extra\n"),
("stray.log", b"dest only\n")):
with open(os.path.join(root, name), "wb") as fh:
fh.write(data)
os.utime(os.path.join(root, name), (1_500_000_000, 1_500_000_000))
flags = ["-a", "-n", "-i", "--delete", "--exclude=*.log", "--max-size=1000"]
r = subprocess.run(["rsync", "-an", "-i", "--delete", "--exclude=*.log",
"--max-size=1000", source + "/", rdst + "/"],
capture_output=True, text=True,
env=dict(os.environ, LC_ALL="C"))
assert r.returncode == 0, r.stderr
rsync_del = sorted(l for l in r.stdout.splitlines() if l.startswith("*deleting"))
assert rsync_del == ["*deleting extra.txt"], f"unexpected rsync set: {rsync_del}"
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
result, _ = run_client(source, fdst, flags=flags, port=server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
fs_del = sorted(l for l in (result.stdout or "").splitlines()
if l.startswith("*deleting"))
assert fs_del == rsync_del, f"rsync={rsync_del}\nfastsync={fs_del}"
assert os.path.exists(os.path.join(received, "stray.log")), \
"destination-only exclude match must be protected in the dry-run report"
@pytest.mark.ci
def test_remote_dry_run_quiet_is_silent(self, shared_server):
source = os.path.join(TEST_DATA_DIR, "remote_dry_quiet_src")
@@ -2224,6 +2276,36 @@ class TestPartialDir:
partial = os.path.join(dest, ".partial", os.path.relpath(source_file, os.path.sep))
assert not os.path.exists(partial)
def test_partial_dir_alone_implies_partial(self, shared_server):
"""--partial-dir=DIR with no --partial implies --partial, like rsync.
rsync 3.4.1 retains the staged partial when --partial-dir is given by
itself; before the implication was added FastSync discarded it. The
transfer is made to fail deterministically by placing a non-empty
directory at the destination path, so the final partial-dir ->
destination rename fails and whatever was staged under the partial dir
stays on disk."""
source = os.path.join(TEST_DATA_DIR, "partial_dir_implied_src")
dest = os.path.join(TEST_DATA_DIR, "partial_dir_implied_dst")
clean_dir(source)
clean_dir(dest)
source_file = os.path.join(source, "f.bin")
with open(source_file, "wb") as f:
f.write(b"partial payload")
received = get_dest_received_dir(dest, source)
os.makedirs(os.path.join(received, "f.bin"))
with open(os.path.join(received, "f.bin", "keep"), "wb") as f:
f.write(b"keep")
result, _ = run_client(source, dest, flags=["--partial-dir=.partial"],
port=shared_server.port)
assert result.returncode != 0, "expected the blocked install to fail"
partial = os.path.join(dest, ".partial", os.path.relpath(source_file, os.path.sep))
assert os.path.exists(partial), \
"--partial-dir alone must imply --partial and retain the partial file"
class TestLargeFile:
def test_transfer_100mb_file(self, shared_server):
@@ -2376,6 +2458,61 @@ class TestTempDir:
assert not mismatches, f"Mismatch: {mismatches}"
self._assert_clean_scratch(os.path.join(dest, "scratch"))
@pytest.mark.skipif(shutil.which("rsync") is None, reason="rsync not installed")
def test_relative_temp_dir_matches_rsync_absolute_rejected(self):
"""Differential: a relative --temp-dir is resolved under the destination
by both (rsync 3.4.1 and FastSync), producing identical trees. An
absolute --temp-dir is used verbatim by rsync standalone, but the
receiver deliberately confines it to the receive root (security
invariant), so FastSync rejects it without writing outside the root.
"""
source = self._make_source("tempdir_diff_src")
rdst = os.path.join(TEST_DATA_DIR, "tempdir_diff_rdst")
fdst = os.path.join(TEST_DATA_DIR, "tempdir_diff_fdst")
clean_dir(rdst)
clean_dir(fdst)
os.makedirs(os.path.join(rdst, "scratch"), exist_ok=True)
os.makedirs(os.path.join(fdst, "scratch"), exist_ok=True)
r = subprocess.run(["rsync", "-a", "--temp-dir=scratch", source + "/", rdst + "/"],
capture_output=True, text=True,
env=dict(os.environ, LC_ALL="C"))
assert r.returncode == 0, r.stderr
with ServerManager() as server:
server.start()
result, _ = run_client(source, fdst, flags=["--temp-dir=scratch"],
port=server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
# rsync lays the source contents directly in rdst; FastSync mirrors the
# absolute source path below fdst. Compare the mirrored content trees
# (the scratch dir lives at each destination root).
rtree = sorted(os.path.relpath(os.path.join(dp, n), rdst)
for dp, dn, fn in os.walk(rdst)
for n in dn + fn if os.path.join(dp, n) != os.path.join(rdst, "scratch"))
mirror = get_dest_received_dir(fdst, source)
ftree = sorted(os.path.relpath(os.path.join(dp, n), mirror)
for dp, dn, fn in os.walk(mirror) for n in dn + fn)
assert rtree == ftree, f"relative temp-dir tree mismatch: {rtree} != {ftree}"
assert _walk_tmp_files(os.path.join(rdst, "scratch")) == []
assert _walk_tmp_files(os.path.join(fdst, "scratch")) == []
# Absolute temp dir: rsync accepts it; FastSync rejects it safely.
abs_scratch = os.path.join(TEST_DATA_DIR, "tempdir_diff_abs")
clean_dir(abs_scratch)
rdst2 = os.path.join(TEST_DATA_DIR, "tempdir_diff_rdst2")
clean_dir(rdst2)
r2 = subprocess.run(["rsync", "-a", "--temp-dir=" + abs_scratch, source + "/", rdst2 + "/"],
capture_output=True, text=True,
env=dict(os.environ, LC_ALL="C"))
assert r2.returncode == 0, r2.stderr
fdst2 = os.path.join(TEST_DATA_DIR, "tempdir_diff_fdst2")
clean_dir(fdst2)
with ServerManager() as server:
server.start()
result2, _ = run_client(source, fdst2, flags=["--temp-dir", abs_scratch],
port=server.port)
assert result2.returncode != 0, "an absolute --temp-dir must be rejected (confined)"
assert os.listdir(abs_scratch) == [], "receiver wrote into an unconfined temp dir"
def test_default_behavior_has_no_scratch_dir(self, shared_server):
source = self._make_source("tempdir_default_src")
dest = os.path.join(TEST_DATA_DIR, "tempdir_default_dst")
@@ -2500,35 +2637,30 @@ class TestTimeoutAndAllocLimits:
mismatches, missing = verify_transfer(source, received)
assert not missing and not mismatches
def test_temp_dir_cross_filesystem_fallback(self, shared_server):
"""A confined relative --temp-dir that resolves (via a symlink under the
destination root) to another filesystem must fall back to a non-atomic
copy instead of aborting (rsync parity). Skipped when no second
filesystem is available."""
shm = "/dev/shm"
if not os.path.isdir(shm):
pytest.skip("/dev/shm not available")
if os.stat(shm).st_dev == os.stat(TEST_DATA_DIR).st_dev:
pytest.skip("/dev/shm is on the same filesystem as the test data")
scratch = os.path.join(shm, f"fastsync_tmp_{os.getpid()}")
shutil.rmtree(scratch, ignore_errors=True)
os.makedirs(scratch)
def test_temp_dir_symlink_escape_rejected(self, shared_server):
"""A symlink planted inside the destination root pointing outside it
must not redirect receiver scratch files: --temp-dir=<that link> is
refused and nothing is written at the link target. An in-root symlink
(e.g. to a mount point that stays inside the authorized root) is still
accepted, preserving the engine's EXDEV cross-filesystem fallback."""
source, dest = self._seed("tempdir_escape_src")
outside = "/tmp/fastsync_tempdir_escape_%d" % os.getpid()
shutil.rmtree(outside, ignore_errors=True)
os.makedirs(outside)
link = os.path.join(dest, "escape_scratch")
if os.path.lexists(link):
os.unlink(link)
os.symlink(outside, link)
try:
source, dest = self._seed("tempdir_xdev_src")
# The receiver resolves a relative temp dir under the destination
# root; a symlink there points the scratch at the second filesystem.
link = os.path.join(dest, "xdev_scratch")
os.symlink(scratch, link)
result, _ = run_client(source, dest, flags=["--temp-dir", "xdev_scratch"],
result, _ = run_client(source, dest, flags=["--temp-dir", "escape_scratch"],
port=shared_server.port)
assert result.returncode == 0, f"cross-fs temp-dir failed: {result.stderr[:300]}"
assert result.returncode != 0, "an escaping --temp-dir symlink must be refused"
received = get_dest_received_dir(dest, source)
mismatches, missing = verify_transfer(source, received)
assert not missing, f"Missing: {missing}"
assert not mismatches, f"Mismatch: {mismatches}"
assert os.listdir(scratch) == [], "temp files left behind in the cross-fs scratch"
assert not os.path.exists(os.path.join(received, "f.txt")), \
"the receiver must not fall back to writing the file"
assert os.listdir(outside) == [], "receiver wrote outside the authorized root"
finally:
shutil.rmtree(scratch, ignore_errors=True)
shutil.rmtree(outside, ignore_errors=True)
class TestRemoteOptionTransport:
@@ -2655,7 +2787,12 @@ class TestItemizeChanges:
flags=["--preserve", "-i", "--incremental"],
port=shared_server.port)
assert result.returncode == 0, f"incremental itemize failed: {result.stderr[:200]}"
itemized = [line for line in result.stdout.splitlines() if line and line[0] in ">.<c"]
# -i also emits the transfer-root and directory lines; only FILE entries
# matter here, so drop any line whose name has a trailing '/'.
itemized = [
line for line in result.stdout.splitlines()
if line and line[0] in ">.<c" and not line.rsplit(" ", 1)[-1].endswith("/")
]
assert itemized == [], f"unchanged files were itemized: {itemized[:5]}"
def test_multithreaded_emits_same_itemize_lines(self, shared_server):
@@ -2837,6 +2974,41 @@ class TestDelayUpdates:
assert not os.path.isdir(os.path.join(delay_dest, self.STAGING)), \
"staging directory left behind after a successful delayed transfer"
@pytest.mark.skipif(shutil.which("rsync") is None, reason="rsync not installed")
def test_delay_updates_staging_name_collision_residual(self):
"""Documented residual (RSYNC_COMPAT.md `--delay-updates` row): FastSync
uses a fixed `.fastsync-stage` staging name and wipes a pre-existing tree
of that name at the start of a delayed run (crash-leftover cleanup),
even without `--delete`; rsync leaves a genuine destination entry of that
name untouched. Pins the divergence that keeps the row Divergent."""
source = self._make_source("delay_collide_src")
rdst = os.path.join(TEST_DATA_DIR, "delay_collide_rdst")
fdst = os.path.join(TEST_DATA_DIR, "delay_collide_fdst")
clean_dir(rdst)
clean_dir(fdst)
for root in (rdst, fdst):
with open(os.path.join(root, "top.txt"), "wb") as fh:
fh.write(b"old\n")
stage = os.path.join(root, self.STAGING)
os.makedirs(stage, exist_ok=True)
with open(os.path.join(stage, "keepme.txt"), "wb") as fh:
fh.write(b"genuine user data\n")
r = subprocess.run(["rsync", "-a", "--delay-updates", source + "/", rdst + "/"],
capture_output=True, text=True,
env=dict(os.environ, LC_ALL="C"))
assert r.returncode == 0, r.stderr
assert os.path.exists(os.path.join(rdst, self.STAGING, "keepme.txt")), \
"rsync removed an unrelated destination entry named like the staging dir"
with ServerManager() as server:
server.start()
result, _ = run_client(source, fdst, flags=["--delay-updates"],
port=server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
assert not os.path.exists(os.path.join(fdst, self.STAGING)), \
"FastSync did not wipe the reserved staging name (residual changed)"
@pytest.mark.parametrize("mt", [False, True])
def test_delay_updates_incremental_rerun_no_leftovers(self, shared_server, mt):
source = self._make_source("delay_rerun_src")
@@ -3523,23 +3695,25 @@ class TestMissingArgs:
class TestNoImpliedDirs:
"""--no-implied-dirs (only meaningful with -R + --files-from) refuses to
place a listed file whose parent directory is not itself listed."""
"""--no-implied-dirs (meaningful with -R) omits the source metadata of a
listed path's implied parent directories but still creates those parents
with default attributes, matching rsync 3.4.1."""
def _make(self):
return _make_relative_source("noimplied_src")
@pytest.mark.parametrize("mt", [False, True])
def test_implied_dir_only_fails_entry(self, shared_server, mt):
def test_implied_dir_created_with_default_attrs(self, shared_server, mt):
source = self._make()
dest = os.path.join(TEST_DATA_DIR, "noimplied_dst")
clean_dir(dest)
lst = _write_rel_list(b"a/b.txt\n") # "a" itself is not listed
flags = ["--files-from", lst, "-R", "--no-implied-dirs"] + (["--threads"] if mt else [])
result, _ = run_client(source, dest, flags=flags, port=shared_server.port)
assert result.returncode != 0, "implied parent directory was not rejected"
assert "--no-implied-dirs" in (result.stderr or result.stdout)
assert not os.path.exists(os.path.join(dest, "a", "b.txt"))
assert result.returncode == 0, \
f"implied parent directory was not created: {result.stderr[:200]}"
assert os.path.isdir(os.path.join(dest, "a")), "implied parent 'a' was not created"
assert _read_file(os.path.join(dest, "a", "b.txt")) == b"nested\n"
@pytest.mark.parametrize("mt", [False, True])
def test_listed_dir_allows_file(self, shared_server, mt):
@@ -3666,9 +3840,10 @@ class TestDirs:
files.extend(os.path.relpath(os.path.join(root, n), mirror) for n in names)
assert files == [], f"--dirs descended into contents: {files}"
def test_dirs_listed_dir_colliding_with_file_fails(self, shared_server):
"""A listed directory that already exists as a regular file at the
destination fails the transfer cleanly instead of clobbering the file."""
def test_dirs_listed_dir_replaces_blocking_file(self, shared_server):
"""rsync parity: a listed directory replaces a regular file already at
its destination path (rsync removes the non-directory and creates the
directory)."""
source = self._make()
dest = os.path.join(TEST_DATA_DIR, "dirs_coll_dst")
clean_dir(dest)
@@ -3678,8 +3853,10 @@ class TestDirs:
lst = _write_rel_list(b"dir1\n")
result, _ = run_client(source, dest, flags=["--files-from", lst, "--dirs", "-R"],
port=shared_server.port)
assert result.returncode != 0, "dir entry over an existing file did not fail"
assert os.path.isfile(blocker), "blocking regular file was clobbered"
assert result.returncode == 0, \
f"dir entry over an existing file failed: {(result.stderr or result.stdout)[:300]}"
assert os.path.isdir(blocker) and not os.path.islink(blocker), \
"blocking regular file was not replaced by the incoming directory"
class TestMkpath:
@@ -3836,15 +4013,16 @@ class TestDeleteTiming:
assert _read_file(os.path.join(received, "sub", "deep.txt")) == b"deeply nested file\n", \
f"{flag}: nested file was not written after the early deletion"
@pytest.mark.parametrize("flag", ["--delete", "--delete-after"])
@pytest.mark.parametrize("flag", ["--delete-commit", "--delete-after"])
@pytest.mark.parametrize("mt", [False, True])
def test_late_flags_commit_only_after_success(self, flag, mt):
"""Plain --delete/--delete-after defer deletion until the whole transfer
"""--delete-commit/--delete-after defer deletion until the whole transfer
succeeds: a mid-transfer write failure must leave every extra in place
(commit-style safety). The -m receiver must also keep the extras: the
deferred keep-set is committed by the server only after the disk-writer
thread has finished, and a failing writer means the manifest is freed,
never applied."""
(commit-style safety). Plain --delete no longer defers (it defaults to
delete-during), so only the explicitly late timings are exercised here.
The --threads receiver must also keep the extras: the deferred keep-set is
committed by the server only after the disk-writer thread has finished,
and a failing writer means the manifest is freed, never applied."""
source = self._seed("late")
dest = os.path.join(TEST_DATA_DIR, "deltiming_late_dst")
clean_dir(dest)
@@ -4463,11 +4641,11 @@ class TestDeletePolicy:
@pytest.mark.parametrize("mt", [False, True])
@pytest.mark.setpriv
def test_ignore_errors_keeps_deletion_active_on_scan_error(self, mt):
"""A source I/O error (unreadable subdirectory) aborts the run so no
deletion happens by default; --ignore-errors continues, still transfers
the readable tree and still deletes, single-threaded and under -m. Run
as an unprivileged user so the mode-000 directory is genuinely
unreadable."""
"""rsync's --ignore-errors semantics: a source I/O error (unreadable
subdirectory) makes the run continue and transfer the readable tree, but
the default suppresses deletion ("IO error encountered -- skipping file
deletion"); --ignore-errors lets deletion proceed. Both exit 23. Run as
an unprivileged user so the mode-000 directory is genuinely unreadable."""
if os.geteuid() != 0 or shutil.which("setpriv") is None:
pytest.skip("requires root + setpriv to drop privileges for the client")
tag = f"ioerr_{os.getpid()}_{mt}"
@@ -4487,11 +4665,15 @@ class TestDeletePolicy:
try:
os.chmod(os.path.join(source, "locked"), 0)
# Default: scan error aborts the run; nothing is deleted.
# Default: the scan continues past the unreadable dir and the
# readable tree transfers, but deletion is skipped (exit 23).
self._write(os.path.join(received, "extra.txt"), b"extra\n")
flags = ["--delete"] + (["--threads"] if mt else [])
result = self._run_client_as_nobody(source, dest, server.port, flags)
assert result.returncode != 0, "unreadable source dir did not fail the run"
assert result.returncode == 23, \
f"unreadable source dir should exit 23 (got {result.returncode})"
assert os.path.exists(os.path.join(received, "top.txt")), \
"readable tree did not transfer past the I/O error"
assert os.path.exists(os.path.join(received, "extra.txt")), \
"default run deleted although the scan hit an I/O error"
@@ -4499,6 +4681,8 @@ class TestDeletePolicy:
self._write(os.path.join(received, "extra.txt"), b"extra\n")
flags = ["--delete", "--ignore-errors"] + (["--threads"] if mt else [])
result = self._run_client_as_nobody(source, dest, server.port, flags)
assert result.returncode == 23, \
f"--ignore-errors run should still exit 23 (got {result.returncode})"
assert not os.path.exists(os.path.join(received, "extra.txt")), \
f"--ignore-errors did not keep deletion active: {result.stderr[:300]}"
assert not os.path.exists(os.path.join(received, "locked")), \
@@ -4543,11 +4727,11 @@ class TestDeletePolicy:
finally:
os.chmod(source, 0o755)
def test_delete_excluded_protection_is_sender_derived(self):
def test_delete_protection_reapplied_on_receiver(self):
"""Plain --delete protects destination mirrors of files the SOURCE scan
excluded, but a destination-only file that merely matches an exclude
rule is still an extra and is removed (protection never re-applies rules
to the destination)."""
excluded, and (track 4a) also protects a destination-only file matching
an exclude rule because the compiled rule set is re-applied on the
receiver, matching rsync."""
source = os.path.join(TEST_DATA_DIR, "senderderived_src")
clean_dir(source)
self._write(os.path.join(source, "keep.txt"), b"kept\n")
@@ -4567,8 +4751,8 @@ class TestDeletePolicy:
f"delete sync failed: {(result.stderr or result.stdout)[:300]}"
assert os.path.exists(os.path.join(received, "secret.log")), \
"source-excluded mirror was deleted under plain --delete"
assert not os.path.exists(os.path.join(received, "stray.log")), \
"destination-only file matching the exclude rule was left (should be deleted)"
assert os.path.exists(os.path.join(received, "stray.log")), \
"destination-only file matching the exclude rule must be protected like rsync"
def _pin_mtime(path, ts):
@@ -4588,11 +4772,12 @@ class TestBasisDestDirs:
STAGING = ".fastsync-stage"
TS = 1577836800 # 2020-01-01 00:00:00 UTC, used to pin matching mtimes
# fixture files: source and basis share the mtime pin, so a basis "match"
# is decided purely by content (xxHash). unchanged.txt is byte-identical;
# changed.txt is byte-DIFFERENT but has the SAME SIZE as the source (and
# the same pinned mtime), which is what forces the content-hash gate;
# added.txt does not exist in the basis at all.
# fixture files: source and basis share the mtime pin, so the DEFAULT
# (rsync-parity) quick-check is a size+mtime match and trusts the basis even
# when the body differs. unchanged.txt is byte-identical; changed.txt is
# byte-DIFFERENT but has the SAME SIZE as the source (and the same pinned
# mtime), which is what the FastSync-only --verify-basis content gate
# rejects; added.txt does not exist in the basis at all.
UNCHANGED = "unchanged.txt"
CHANGED = "changed.txt"
ADDED = "added.txt"
@@ -4632,18 +4817,19 @@ class TestBasisDestDirs:
}
def _basis_tree(self, prefix):
# unchanged.txt is identical to the source; changed.txt has the SAME
# byte size and pinned mtime but a different body (equal size forces
# the xxHash gate); added.txt is missing from the basis.
# unchanged.txt is identical to the source; changed.txt has a DIFFERENT
# size (and body) so the size leg of the quick-check fails and it is
# transferred normally; added.txt is missing from the basis.
return {
self.UNCHANGED: b"stable content v1\n",
self.CHANGED: b"CHANGED CONTENT NOW\n",
self.CHANGED: b"CHANGED CONTENT NOW AND LONGER\n",
}
def test_same_size_different_content_is_not_a_basis_match(self, shared_server):
# Core safety property: equal size + pinned mtime but different content
# must NEVER be hard-linked or copied from the basis -- the xxHash gate
# rejects it and the sender's data is transferred instead.
def test_same_size_different_content_default_trusts_quick_check(self, shared_server):
# Default rsync-parity behavior: equal size + pinned mtime is a basis
# match, so the basis body is materialized/linked without reading it.
# This mirrors rsync 3.4.1's quick check (differential-tested in
# test_differential_parity.py::test_verify_basis_restores_strict_content).
for flag, basis_dir in (("--link-dest", "szlb"), ("--copy-dest", "szcp"),
("--compare-dest", "szcmp")):
source = self._make_source("basis_same_size_src",
@@ -4655,7 +4841,36 @@ class TestBasisDestDirs:
result, _ = run_client(source, dest, flags=[f"{flag}={basis_dir}"],
port=shared_server.port)
assert result.returncode == 0, \
f"{flag} same-size mismatch failed: {result.stderr[:300]}"
f"{flag} same-size quick-check failed: {result.stderr[:300]}"
received = get_dest_received_dir(dest, source)
dest_file = os.path.join(received, self.UNCHANGED)
if flag == "--compare-dest":
assert not os.path.exists(dest_file), \
f"{flag}: compare-dest must leave a matching file sparse"
else:
assert _read_file(dest_file) == b"SAME LENGTH BODY!", \
f"{flag}: default quick-check did not trust the basis body"
if flag == "--link-dest":
assert os.stat(dest_file).st_ino == os.stat(basis_file).st_ino, \
f"{flag}: basis was not hard-linked"
def test_verify_basis_rejects_same_size_different_content(self, shared_server):
# FastSync-only --verify-basis: the whole-file digest gate rejects the
# same-size/different-content basis, so the source data is transferred
# instead of the wrong basis bytes.
for flag, basis_dir in (("--link-dest", "vszlb"), ("--copy-dest", "vszcp"),
("--compare-dest", "vszcmp")):
source = self._make_source("basis_verify_src",
{self.UNCHANGED: b"same length body\n"})
dest = os.path.join(TEST_DATA_DIR, f"basis_verify_dst_{basis_dir}")
clean_dir(dest)
basis_file = self._seed_basis_file(dest, source, basis_dir, self.UNCHANGED,
b"SAME LENGTH BODY!")
result, _ = run_client(source, dest,
flags=[f"{flag}={basis_dir}", "--verify-basis"],
port=shared_server.port)
assert result.returncode == 0, \
f"{flag} --verify-basis failed: {result.stderr[:300]}"
received = get_dest_received_dir(dest, source)
dest_file = os.path.join(received, self.UNCHANGED)
assert _read_file(dest_file) == b"same length body\n", \
@@ -4687,11 +4902,13 @@ class TestBasisDestDirs:
self._source_tree("c")[self.ADDED], "added file not transferred"
@pytest.mark.ci
def test_dry_run_compare_dest_does_not_read_basis(self, shared_server):
# A dry-run --compare-dest must never read/hash the basis file: doing so
# is a 1-bit content oracle against the client-supplied digest. Even a
# byte-identical basis with a matching size+mtime is therefore reported
# as would-transfer, and nothing is created.
def test_dry_run_compare_dest_quick_check_does_not_read_basis(self, shared_server):
# A dry-run --compare-dest must never read/hash the basis file. Under
# the default metadata quick-check a matching basis is reported as a
# skip (matching rsync) without reading it; nothing is created. Under
# --verify-basis, which would require hashing, the dry-run cannot
# confirm the hit (that would be a 1-bit content oracle) and reports
# would-transfer instead.
source = self._make_source("basis_dry_src", {self.UNCHANGED: b"stable content v1\n"})
dest = os.path.join(TEST_DATA_DIR, "basis_dry_dst")
clean_dir(dest)
@@ -4702,11 +4919,30 @@ class TestBasisDestDirs:
port=shared_server.port)
assert result.returncode == 0, \
f"dry-run compare-dest failed: {result.stderr[:300]}"
assert self.UNCHANGED in result.stdout, (
"dry-run compare-dest silently skipped: receiver read the basis content"
assert self.UNCHANGED not in result.stdout, (
"dry-run compare-dest did not honor the metadata quick-check "
"(reported would-transfer for a matching basis)"
)
assert _snapshot_tree(dest) == before, "dry-run compare-dest mutated the destination"
# --verify-basis: the hit needs the basis content, which a dry-run must
# not read, so the file is reported as would-transfer.
dest2 = os.path.join(TEST_DATA_DIR, "basis_dry_verify_dst")
clean_dir(dest2)
self._seed_basis(dest2, source, "drybasis", {self.UNCHANGED: b"stable content v1\n"})
before2 = _snapshot_tree(dest2)
result, _ = run_client(source, dest2,
flags=["--compare-dest=drybasis", "--dry-run",
"--verify-basis"],
port=shared_server.port)
assert result.returncode == 0, \
f"dry-run --verify-basis compare-dest failed: {result.stderr[:300]}"
assert self.UNCHANGED in result.stdout, (
"dry-run --verify-basis must not read the basis to confirm a hit"
)
assert _snapshot_tree(dest2) == before2, \
"dry-run --verify-basis compare-dest mutated the destination"
def test_compare_dest_content_mismatch_forces_transfer(self, shared_server):
# The basis holds a file with a DIFFERENT body: even though it shares
# the mtime pin, the xxHash check fails and the data must be sent.
@@ -4945,27 +5181,36 @@ class TestBasisDestDirs:
assert os.stat(dest_file).st_ino != os.stat(basis_file).st_ino, \
"--ignore-times must not hard-link to a basis file"
def test_basis_refuses_file_above_whole_file_limit(self, shared_server):
# Every whole-file payload path in FastSync (basis dirs included) is
# bounded by MAX_RECEIVE_WHOLE_FILE_SIZE. rsync supports basis dirs for
# arbitrary sizes; FastSync refuses such a run up front with a clear
# diagnostic instead of letting the receiver abort the whole transfer
# mid-stream with no client-side explanation.
def test_basis_handles_file_above_whole_file_limit(self, shared_server):
# Track 5a: a basis hit streams the copy (and the --verify-basis digest
# streams the basis), so a source larger than the whole-file payload
# bound is supported for basis dirs exactly like rsync. A basis MISS
# still falls back to the normal transfer, which keeps its own bound.
source = self._make_source("basis_oversize_src", {"small.txt": b"ok\n"})
big = os.path.join(source, "huge.bin")
with open(big, "wb") as fh:
os.ftruncate(fh.fileno(), 256 * 1024 * 1024 + 4096)
dest = os.path.join(TEST_DATA_DIR, "basis_oversize_dst")
clean_dir(dest)
result, _ = run_client(source, dest, flags=["--link-dest=nope"],
port=shared_server.port)
assert result.returncode != 0, \
"basis run with an over-limit file unexpectedly succeeded"
assert "larger than" in result.stderr, \
f"no clear over-limit diagnostic: {result.stderr[:300]}"
received = get_dest_received_dir(dest, source)
assert not os.path.exists(received), \
"over-limit basis run transferred files before failing"
rel = os.path.relpath(received, dest)
basis_big = os.path.join(dest, "ob", rel, "huge.bin")
os.makedirs(os.path.dirname(basis_big), exist_ok=True)
shutil.copyfile(big, basis_big)
os.utime(basis_big, (self.TS, self.TS))
os.utime(big, (self.TS, self.TS))
result, _ = run_client(source, dest,
flags=["--link-dest=ob", "--incremental"],
port=shared_server.port)
assert result.returncode == 0, \
f"over-limit basis run failed: {result.stderr[:300]}"
dest_big = os.path.join(received, "huge.bin")
assert os.path.exists(dest_big), "over-limit basis hit was not materialized"
assert os.path.getsize(dest_big) == 256 * 1024 * 1024 + 4096
assert os.stat(dest_big).st_ino == os.stat(basis_big).st_ino, \
"over-limit --link-dest did not hard-link to the basis"
assert _read_file(os.path.join(received, "small.txt")) == b"ok\n"
def _random_payloads(size=2 * 1024 * 1024, changed=64 * 1024, seed=1234):
@@ -5568,6 +5813,35 @@ class TestStandaloneSuperDefault:
"standalone server accepted --copy-as without --allow-super"
)
@pytest.mark.skipif(
os.geteuid() != 0,
reason="root triggers the SUPER_MODE_OFF default and can create setuid sources",
)
def test_special_bits_masked_without_allow_super(self):
"""A root standalone server without --allow-super forces SUPER_MODE_OFF,
so client-supplied setuid/setgid/sticky bits must be stripped even under
-p (they are super-user activities just like device-node creation)."""
source = os.path.join(TEST_DATA_DIR, "super_default_mode_src")
dest = os.path.join(TEST_DATA_DIR, "super_default_mode_dst")
clean_dir(source)
clean_dir(dest)
src_file = os.path.join(source, "priv.sh")
with open(src_file, "wb") as f:
f.write(b"#!/bin/sh\necho hi\n")
os.chmod(src_file, 0o4755)
server = ServerManager()
server.start() # deliberately no --allow-super -> SUPER_MODE_OFF as root
try:
result, _ = run_client(source, dest, flags=["-p"], port=server.port)
finally:
server.stop()
assert result.returncode == 0, f"exit {result.returncode}: {(result.stderr or '')[:200]}"
received = get_dest_received_dir(dest, source)
mode = stat.S_IMODE(os.stat(os.path.join(received, "priv.sh")).st_mode)
assert (mode & (stat.S_ISUID | stat.S_ISGID | stat.S_ISVTX)) == 0, \
f"--no-super receiver kept a privileged bit: {oct(mode)}"
assert (mode & 0o777) == 0o755, f"ordinary permission bits lost: {oct(mode)}"
@pytest.mark.skipif(os.geteuid() != 0, reason="root can create the source device node")
def test_devices_skipped_without_allow_super(self):
"""Root standalone server without --allow-super must skip device-node
@@ -6326,7 +6600,7 @@ class TestExtendedAttributes:
os.getxattr(os.path.join(received, "data.txt"), "user.foo")
def test_reserved_fake_super_key_not_forwarded(self, shared_server):
"""A source file that already carries the reserved user.fastsync.stat
"""A source file that already carries the reserved user.rsync.%stat
record must NOT have it planted on the receiver during a plain -X run
(it is receiver-only, so it cannot be spoofed for a later privileged
restore)."""
@@ -6336,7 +6610,7 @@ class TestExtendedAttributes:
fh.write(b"reserved\n")
if not _xattr_supported(f):
pytest.skip("filesystem does not support user xattrs")
os.setxattr(f, "user.fastsync.stat", b"0:0:644:0:0")
os.setxattr(f, "user.rsync.%stat", b"100644 0,0 0:0")
# A normal user.* attr still travels alongside.
os.setxattr(f, "user.keep", b"yes")
@@ -6346,7 +6620,7 @@ class TestExtendedAttributes:
received = get_dest_received_dir(dest, source)
assert os.getxattr(os.path.join(received, "data.txt"), "user.keep") == b"yes"
with pytest.raises(OSError):
os.getxattr(os.path.join(received, "data.txt"), "user.fastsync.stat")
os.getxattr(os.path.join(received, "data.txt"), "user.rsync.%stat")
@pytest.mark.ci
def test_xattrs_multithreaded(self, shared_server):
@@ -6363,6 +6637,68 @@ class TestExtendedAttributes:
received = get_dest_received_dir(dest, source)
assert os.getxattr(os.path.join(received, "data.txt"), "user.k") == b"v"
@pytest.mark.ci
def test_symlink_own_xattrs_never_referent(self, shared_server):
"""Protocol 2.29.0: a symlink's STATUS_SYMLINK frame carries a trailing
xattr block captured with llistxattr/lgetxattr (no follow) and applied
with lsetxattr on the link itself. Linux's VFS refuses to associate
xattrs with a symlink at all, so the portable guarantee asserted here is
the no-follow one: a referent that carries user.* must NOT have those
attributes appear on the destination symlink entry (the old
path-following capture would have copied the referent's attrs onto the
link). On a platform/filesystem that does support symlink xattrs the
full round-trip of the link's own attribute is asserted too."""
source, dest = self._source_and_dest("symlink_xattr")
target = os.path.join(source, "target.txt")
with open(target, "wb") as fh:
fh.write(b"referent payload\n")
if not _xattr_supported(target):
pytest.skip("filesystem does not support user xattrs")
os.setxattr(target, "user.referent-only", b"referent-value")
link = os.path.join(source, "link")
os.symlink("target.txt", link)
link_xattr_supported = False
try:
os.setxattr(link, "user.link-own", b"link-value", follow_symlinks=False)
link_xattr_supported = os.getxattr(
link, "user.link-own", follow_symlinks=False
) == b"link-value"
except (OSError, AttributeError, NotImplementedError):
link_xattr_supported = False
result, _ = run_client(source, dest, flags=["-aX"], port=shared_server.port)
assert result.returncode == 0, \
f"-aX symlink sync failed: {(result.stderr or result.stdout)[:300]}"
received = get_dest_received_dir(dest, source)
dst_link = os.path.join(received, "link")
assert os.path.islink(dst_link), "destination link entry is not a symlink"
assert os.readlink(dst_link) == "target.txt"
# The no-follow guarantee. Checking only the link's own xattr list is
# vacuous on Linux (lsetxattr on a symlink always fails EPERM), so also
# prove the apply never followed the link: the destination REFERENT must
# keep its own user.* value untouched.
dst_target = os.path.join(received, "target.txt")
assert os.getxattr(dst_target, "user.referent-only") == b"referent-value", (
"the destination symlink apply followed the link and rewrote the "
"referent's xattr"
)
link_names = os.listxattr(dst_link, follow_symlinks=False)
assert "user.referent-only" not in link_names, (
"the destination symlink captured its REFERENT's xattr "
"(path-following capture bug)"
)
if sys.platform.startswith("linux"):
assert link_names == [], (
"Linux associates no xattrs with a symlink; the link entry must "
"carry none"
)
if link_xattr_supported:
assert os.getxattr(
dst_link, "user.link-own", follow_symlinks=False
) == b"link-value", "the symlink's own xattr did not round-trip"
@pytest.mark.ci
def test_acls_via_posix_acl_xattr(self, shared_server):
source, dest = self._source_and_dest("acl")
@@ -6435,10 +6771,17 @@ class TestExtendedAttributes:
assert result.returncode == 0, \
f"--fake-super sync failed: {(result.stderr or result.stdout)[:300]}"
received = get_dest_received_dir(dest, source)
record = os.getxattr(os.path.join(received, "data.txt"), "user.fastsync.stat").decode()
fields = record.split(":")
assert len(fields) == 5
assert fields[0] == str(uid), f"reserved uid field {fields[0]} != source uid {uid}"
record = os.getxattr(os.path.join(received, "data.txt"), "user.rsync.%stat").decode()
# rsync 3.4.1 grammar: "<octal st_mode> <rdev_major>,<rdev_minor> <uid>:<gid>".
fields = record.split()
assert len(fields) == 3, f"unexpected rsync fake-super record {record!r}"
mode_field, rdev_field, owner_field = fields
assert rdev_field == "0,0", f"regular file rdev must be 0,0, got {rdev_field!r}"
assert int(mode_field, 8) & 0o170000 == stat.S_IFREG, (
f"recorded mode {mode_field!r} must carry S_IFREG"
)
assert owner_field.split(":")[0] == str(uid), \
f"recorded uid {owner_field!r} != source uid {uid}"
@pytest.mark.ci
def test_fake_super_records_resolved_chown_without_real_chown(self, shared_server):
@@ -6457,12 +6800,71 @@ class TestExtendedAttributes:
assert result.returncode == 0, \
f"--fake-super --chown sync failed: {(result.stderr or result.stdout)[:300]}"
dst = os.path.join(get_dest_received_dir(dest, source), "data.txt")
record = os.getxattr(dst, "user.fastsync.stat").decode().split(":")
assert record[0] == "33333", f"recorded owner {record[0]} != resolved 33333"
assert record[1] == "44444", f"recorded group {record[1]} != resolved 44444"
record = os.getxattr(dst, "user.rsync.%stat").decode().split()
owner = record[2].split(":")
assert owner == ["33333", "44444"], (
f"recorded owner {record[2]!r} != resolved 33333:44444"
)
st = os.stat(dst)
assert st.st_uid != 33333, "--fake-super must not real-chown the recorded owner"
@pytest.mark.ci
def test_fake_super_rsync_interop(self, shared_server):
"""A fake-super tree written by FastSync is readable by rsync 3.4.1:
rsync reads the `user.rsync.%stat` record (mode/rdev/uid:gid) and, when
it re-emits a fake-super tree, reproduces the same record. This pins
the on-disk key and value grammar against the real tool."""
rsync = shutil.which("rsync")
if rsync is None:
pytest.skip("rsync not installed")
source, dest = self._source_and_dest("fakesuper_interop")
f = os.path.join(source, "data.txt")
with open(f, "wb") as fh:
fh.write(b"interop\n")
if not _xattr_supported(f):
pytest.skip("filesystem does not support user xattrs")
# rsync's fake-super receiver only writes a %stat% record when it has
# something to fake; a root-owned file with a matching root stat is
# a no-op. When privileged, record a non-root owner so the round-trip
# actually exercises the parser (non-root CI already has a non-zero uid).
if os.geteuid() == 0:
try:
os.chown(f, 12345, 12346)
except OSError:
pass
# A setuid bit exercises the full st_mode encoding; set it AFTER any
# chown (chown clears setuid/setgid), and note that neither tool installs
# it on the real destination file.
os.chmod(f, 0o4711)
result, _ = run_client(source, dest, flags=["--fake-super"],
port=shared_server.port)
assert result.returncode == 0, \
f"--fake-super sync failed: {(result.stderr or result.stdout)[:300]}"
received = get_dest_received_dir(dest, source)
rec = os.getxattr(os.path.join(received, "data.txt"), "user.rsync.%stat").decode()
rec_fields = rec.split()
assert len(rec_fields) == 3 and rec_fields[1] == "0,0", (
f"FastSync did not write rsync's stat grammar: {rec!r}"
)
assert int(rec_fields[0], 8) & 0o7777 == 0o4711, (
f"FastSync did not record the source mode in rsync's grammar: {rec!r}"
)
out = os.path.join(TEST_DATA_DIR, "fakesuper_interop_rsync")
clean_dir(out)
rs = subprocess.run([rsync, "-aX", "--fake-super",
received + "/", out + "/"],
capture_output=True, text=True, timeout=120)
assert rs.returncode == 0, (
f"rsync could not read FastSync's fake-super tree: {rs.stderr[:300]}"
)
out_rec = os.getxattr(os.path.join(out, "data.txt"), "user.rsync.%stat").decode()
assert out_rec == rec, (
"rsync re-emitted a different fake-super record; FastSync's grammar "
f"is not interoperable: ours={rec!r} rsync={out_rec!r}"
)
@pytest.mark.ci
def test_directory_xattrs_preserved(self, shared_server):
"""#286.3: -aX must preserve user.* xattrs on DIRECTORIES, not just files."""
@@ -6482,6 +6884,31 @@ class TestExtendedAttributes:
assert os.getxattr(received, "user.rootdir") == b"r"
assert os.getxattr(os.path.join(received, "sub"), "user.subdir") == b"s"
@pytest.mark.ci
@pytest.mark.parametrize("mt", [False, True])
def test_dirs_directory_xattr_applied(self, shared_server, mt):
"""#286.3: -d/-X must apply a transferred directory's user.* xattr at the
destination through the --dirs STATUS_MKDIR path (both the
single-threaded and -m/--threads receiver paths)."""
source, dest = self._source_and_dest("dirsxattr")
sub = os.path.join(source, "sub")
os.makedirs(sub)
if not _xattr_supported(sub):
pytest.skip("filesystem does not support user xattrs")
os.setxattr(sub, "user.dirsdir", b"dirs-value")
lst = os.path.join(TEST_DATA_DIR, "dirs_xattr_list.txt")
with open(lst, "wb") as fh:
fh.write(b"sub\n")
flags = ["--files-from", lst, "--dirs", "-R", "-X"] + (["--threads"] if mt else [])
result, _ = run_client(source, dest, flags=flags, port=shared_server.port)
assert result.returncode == 0, \
f"--dirs -X sync failed: {(result.stderr or result.stdout)[:300]}"
received = os.path.join(dest, "sub")
assert os.path.isdir(received), "--dirs directory entry was not created"
assert os.getxattr(received, "user.dirsdir") == b"dirs-value", \
"the --dirs directory's user.* xattr was not applied at the destination"
@pytest.mark.ci
def test_directory_default_acl_preserved(self, shared_server):
"""#286.3: -aA must preserve a directory's default POSIX ACL (the
@@ -6690,11 +7117,10 @@ class TestDirectoryAndSymlinkTimes:
@pytest.mark.ci
@pytest.mark.parametrize("mt", [False, True])
def test_preserve_does_not_create_empty_source_dir(self, shared_server, mt):
"""P7 Wave D #1: a captured-but-EMPTY source directory is never created
at the destination. The scanner records its time (it is transmitted via
STATUS_DIR_TIMES), but the receiver treats that entry as record-only, so
`-a` keeps the documented "empty dirs are never transferred" behavior."""
def test_preserve_creates_empty_source_dir(self, shared_server, mt):
"""rsync parity: a recursive `-a` transfer recreates an empty source
directory at the destination (the scanner emits it as an explicit
directory entry)."""
source = os.path.join(TEST_DATA_DIR, f"empty_dir_{'m' if mt else 's'}_src")
dest = os.path.join(TEST_DATA_DIR, f"empty_dir_{'m' if mt else 's'}_dst")
clean_dir(source)
@@ -6705,8 +7131,8 @@ class TestDirectoryAndSymlinkTimes:
flags = ["-a"] + (["--threads"] if mt else [])
received = self._run(source, dest, flags, shared_server)
assert os.path.isfile(os.path.join(received, "keep.txt")), "regular file missing"
assert not os.path.lexists(os.path.join(received, "empty_sub")), \
f"-a created an empty source directory at {received}/empty_sub"
assert os.path.isdir(os.path.join(received, "empty_sub")), \
f"-a did not recreate the empty source directory at {received}/empty_sub"
@pytest.mark.ci
@pytest.mark.parametrize("mt", [False, True])
@@ -6729,10 +7155,10 @@ class TestDirectoryAndSymlinkTimes:
@pytest.mark.ci
@pytest.mark.parametrize("mt", [False, True])
def test_collision_at_dir_time_path_does_not_abort(self, shared_server, mt):
"""P7 Wave D #1: a pre-existing regular file at a source-empty-dir's
mirror path must not abort the transfer (the old mkdir failed and failed
the run) and must not be clobbered."""
def test_collision_at_empty_dir_path_replaces_blocker(self, shared_server, mt):
"""rsync parity: a pre-existing regular file at a source empty-dir's
mirror path is replaced by the incoming directory (rsync removes the
non-directory and creates the directory); the run succeeds."""
source = os.path.join(TEST_DATA_DIR, f"dirtime_collide_{'m' if mt else 's'}_src")
dest = os.path.join(TEST_DATA_DIR, f"dirtime_collide_{'m' if mt else 's'}_dst")
clean_dir(source)
@@ -6751,10 +7177,8 @@ class TestDirectoryAndSymlinkTimes:
assert result.returncode == 0, \
f"-a aborted on a pre-existing file at an empty-dir path: " \
f"{(result.stderr or result.stdout)[:400]}"
assert os.path.isfile(blocker) and not os.path.islink(blocker), \
"the pre-existing blocker was replaced by a directory"
with open(blocker, "rb") as fh:
assert fh.read() == b"pre-existing blocker\n", "the blocker file was clobbered"
assert os.path.isdir(blocker) and not os.path.islink(blocker), \
"the pre-existing blocker was not replaced by the incoming directory"
assert os.path.isfile(os.path.join(received, "keep.txt")), "regular file missing"
@@ -6960,9 +7384,10 @@ class TestCopyAs:
)
received = get_dest_received_dir(dest, source)
dst = os.path.join(received, "mixed.txt")
record = os.getxattr(dst, "user.fastsync.stat").decode().split(":")
assert (record[0], record[1]) == ("65534", "65534"), (
f"fake-super must record the resolved copy-as ownership: {record[:2]}"
record = os.getxattr(dst, "user.rsync.%stat").decode().split()
owner = record[2].split(":")
assert owner == ["65534", "65534"], (
f"fake-super must record the resolved copy-as ownership: {owner}"
)
st = os.lstat(dst)
assert (st.st_uid, st.st_gid) != (12345, 12346), (
+34 -23
View File
@@ -1,10 +1,12 @@
"""--iconv=CONVERT_SPEC file-NAME charset conversion integration tests.
The client converts every source file name from LOCAL to REMOTE before it goes
on the wire, and the receiver converts it back from REMOTE to LOCAL, so a
source tree using one charset can be written into a destination tree using
another (rsync compatibility; content bytes are never touched).
rsync's spec is ``--iconv=LOCAL,REMOTE`` (the order is the same push or pull).
The sender converts each source name from LOCAL to REMOTE for the wire, and on
a PUSH the receiver's charset is the spec's REMOTE half, so it writes the wire
bytes verbatim (only a server with its own ``--iconv`` declares a different
destination charset and re-converts). Content bytes are never touched.
"""
import codecs
import os
import shutil
@@ -16,6 +18,11 @@ LATIN1_NAME = b"caf\xe9.txt"
UTF8_NAME = "caf\u00e9.txt".encode("utf-8")
def _to_utf8(name_bytes):
"""The UTF-8 encoding of a name that is stored as ISO-8859-1 bytes."""
return codecs.encode(codecs.decode(name_bytes, "iso-8859-1"), "utf-8")
def _make(tag):
source = os.path.join(TEST_DATA_DIR, f"iconv_{tag}_src")
dest = os.path.join(TEST_DATA_DIR, f"iconv_{tag}_dst")
@@ -41,10 +48,10 @@ def _dest_file(source, dest, name):
@pytest.mark.ci
def test_iconv_latin1_roundtrip(shared_server):
"""A source file whose name is ISO-8859-1 bytes is transferred with
--iconv=iso-8859-1,utf-8 and lands on the destination with the ORIGINAL
latin1 name (the wire carried it as UTF-8)."""
def test_iconv_latin1_to_utf8_dest(shared_server):
"""rsync push parity: --iconv=iso-8859-1,utf-8 converts a latin1 source name
to the spec's REMOTE (UTF-8) on the wire and the default receiver writes it
verbatim, so the destination name is UTF-8 (not the source's latin1)."""
source, dest = _make("latin1")
_place_bytes(source, LATIN1_NAME)
@@ -53,8 +60,10 @@ def test_iconv_latin1_roundtrip(shared_server):
)
assert result.returncode == 0, (result.stderr or result.stdout)[:400]
dst = _dest_file(source, dest, LATIN1_NAME)
assert os.path.exists(dst), f"dest latin1-named file not found under {dest}"
dst = _dest_file(source, dest, UTF8_NAME)
assert os.path.exists(dst), f"dest UTF-8-named file not found under {dest}"
assert not os.path.exists(_dest_file(source, dest, LATIN1_NAME)), \
"destination kept the latin1 name instead of the wire (UTF-8) charset"
@pytest.mark.ci
@@ -153,7 +162,7 @@ def test_iconv_expanding_name_growth(shared_server):
)
assert result.returncode == 0, (result.stderr or result.stdout)[:400]
assert os.path.exists(_dest_file(source, dest, name_bytes))
assert os.path.exists(_dest_file(source, dest, _to_utf8(name_bytes)))
def test_iconv_symlink_path_and_target(shared_server):
@@ -170,11 +179,13 @@ def test_iconv_symlink_path_and_target(shared_server):
)
assert result.returncode == 0, (result.stderr or result.stdout)[:400]
dst_target = _dest_file(source, dest, target)
dst_link = _dest_file(source, dest, b"link\xe9")
assert os.path.exists(dst_target), "dest latin1 target file missing"
assert os.path.islink(dst_link), "dest latin1 symlink missing"
assert os.readlink(dst_link) == target, "symlink target not preserved/decoded"
utf8_target = _to_utf8(target)
utf8_link = _to_utf8(b"link\xe9")
dst_target = _dest_file(source, dest, utf8_target)
dst_link = _dest_file(source, dest, utf8_link)
assert os.path.exists(dst_target), "dest UTF-8 target file missing"
assert os.path.islink(dst_link), "dest UTF-8 symlink missing"
assert os.readlink(dst_link) == utf8_target, "symlink target not wire-converted"
with open(dst_link, "rb") as fh:
assert fh.read() == b"t\n"
@@ -198,8 +209,8 @@ def test_iconv_hardlink_path_and_target(shared_server):
)
assert result.returncode == 0, (result.stderr or result.stdout)[:400]
dst_a = _dest_file(source, dest, a)
dst_b = _dest_file(source, dest, b)
dst_a = _dest_file(source, dest, _to_utf8(a))
dst_b = _dest_file(source, dest, _to_utf8(b))
assert os.path.exists(dst_a) and os.path.exists(dst_b)
assert os.stat(dst_a).st_ino == os.stat(dst_b).st_ino, \
"hard-link relationship not preserved across the transfer"
@@ -221,16 +232,16 @@ def test_iconv_delete_manifest_consistent(shared_server):
flags = ["--iconv=iso-8859-1,utf-8"]
result, _ = run_client(source, dest, flags=flags, port=server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:400]
assert os.path.exists(_dest_file(source, dest, keep))
assert os.path.exists(_dest_file(source, dest, gone))
assert os.path.exists(_dest_file(source, dest, _to_utf8(keep)))
assert os.path.exists(_dest_file(source, dest, _to_utf8(gone)))
os.remove(os.path.join(os.fsencode(source), gone))
result, _ = run_client(
source, dest, flags=flags + ["--delete"], port=server.port
)
assert result.returncode == 0, (result.stderr or result.stdout)[:400]
assert os.path.exists(_dest_file(source, dest, keep)), "kept file deleted"
assert not os.path.exists(_dest_file(source, dest, gone)), \
assert os.path.exists(_dest_file(source, dest, _to_utf8(keep))), "kept file deleted"
assert not os.path.exists(_dest_file(source, dest, _to_utf8(gone))), \
"missing file was not deleted"
@@ -247,4 +258,4 @@ def test_iconv_chunk_serialization_blob(shared_server):
)
assert result.returncode == 0, (result.stderr or result.stdout)[:400]
assert os.path.exists(_dest_file(source, dest, name))
assert os.path.exists(_dest_file(source, dest, _to_utf8(name)))
+538
View File
@@ -0,0 +1,538 @@
"""Differential parity tests for the option wave (bwlimit, --info=*, -M,
--ignore-errors, --filter protect).
Every differential here runs the SAME scenario with real ``rsync 3.4.1`` and
with fastsync and compares the observable result, so the modules are skipped
when rsync is unavailable. The privilege-dependent --ignore-errors differential
drops the client to an unprivileged uid so a mode-000 source directory is
genuinely unreadable; it is marked ``setpriv`` (run as root locally, excluded
from the root PR gate exactly like the other privilege tests).
"""
import os
import shutil
import subprocess
import sys
import time
import pytest
sys.path.insert(0, os.path.dirname(__file__))
from common import ( # noqa: E402
CLIENT_CMD,
TEST_DATA_DIR,
ServerManager,
clean_dir,
get_dest_received_dir,
run_client,
)
RSYNC = shutil.which("rsync")
requires_rsync = pytest.mark.skipif(RSYNC is None, reason="rsync 3.4.1 not installed")
def _rsync(args, timeout=120, as_nobody=False):
env = dict(os.environ, LC_ALL="C")
cmd = [RSYNC] + args
if as_nobody:
cmd = ["setpriv", "--reuid=65534", "--regid=65534", "--clear-groups"] + cmd
return subprocess.run(cmd, capture_output=True, text=True, env=env, timeout=timeout)
def _write(path, content):
os.makedirs(os.path.dirname(path), exist_ok=True)
with open(path, "wb") as fh:
fh.write(content)
class TestBwlimitParity:
"""--bwlimit must accept rsync 3.4.1's spellings and pace like it."""
ACCEPTED = ["100", "0", "1.5", "100K", "100KB", "100KiB", "1M", "1MB", "1m", "1G", "512"]
REJECTED = ["-1", "abc", "1x", "1 000"]
@requires_rsync
@pytest.mark.ci
def test_parse_acceptance_matches_rsync(self, shared_server):
source = os.path.join(TEST_DATA_DIR, "bwp_src")
clean_dir(source)
_write(os.path.join(source, "f.txt"), b"payload\n")
for value in self.ACCEPTED + self.REJECTED:
rdst = os.path.join(TEST_DATA_DIR, "bwp_rdst")
clean_dir(rdst)
rsync_result = _rsync(["-a", "--bwlimit=" + value, source + "/", rdst + "/"])
dest = os.path.join(TEST_DATA_DIR, "bwp_dst")
clean_dir(dest)
result, _ = run_client(source, dest, flags=["-a", "--bwlimit=" + value],
port=shared_server.port)
assert (result.returncode == 0) == (rsync_result.returncode == 0), (
f"--bwlimit={value}: fastsync rc={result.returncode} "
f"({(result.stderr or result.stdout)[:120]!r}) "
f"rsync rc={rsync_result.returncode} ({rsync_result.stderr[:120]!r})"
)
@requires_rsync
@pytest.mark.ci
def test_throttle_rate_matches_rsync(self, shared_server):
"""A 4 MiB transfer at --bwlimit=2048 (2 MiB/s) must take about the same
wall-clock time for both tools (~2 s with rsync's leaky bucket)."""
source = os.path.join(TEST_DATA_DIR, "bwt_src")
clean_dir(source)
_write(os.path.join(source, "big.bin"), os.urandom(4 * 1024 * 1024))
dest = os.path.join(TEST_DATA_DIR, "bwt_dst")
rdst = os.path.join(TEST_DATA_DIR, "bwt_rdst")
clean_dir(rdst)
start = time.monotonic()
rsync_result = _rsync(["-a", "--bwlimit=2048", source + "/", rdst + "/"])
rsync_secs = time.monotonic() - start
assert rsync_result.returncode == 0, rsync_result.stderr
clean_dir(dest)
result, fast_secs = run_client(source, dest, flags=["-a", "--bwlimit=2048"],
port=shared_server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
# 4 MiB at 2 MiB/s rendezvous near 2 s. Use a coarse band on each side
# (an unthrottled transfer finishes well under 1.5 s) plus a generous
# cross-tolerance so a loaded CI runner cannot flake the parity assert.
lo, hi = 1.5, 4.5
assert lo <= fast_secs <= hi, f"fastsync throttle out of band: {fast_secs:.2f}s"
assert lo <= rsync_secs <= hi, f"rsync throttle out of band: {rsync_secs:.2f}s"
assert abs(fast_secs - rsync_secs) < 2.0, (
f"fastsync {fast_secs:.2f}s vs rsync {rsync_secs:.2f}s"
)
def _output_tree(root):
clean_dir(root)
os.makedirs(os.path.join(root, "sub"))
_write(os.path.join(root, "a.txt"), b"top\n")
_write(os.path.join(root, "sub", "b.txt"), b"nested\n")
os.symlink("a.txt", os.path.join(root, "link"))
class TestInfoParity:
"""The --info categories that map to a FastSync event must print rsync's
line format."""
@requires_rsync
@pytest.mark.ci
def test_info_flist_matches_rsync(self, shared_server):
source = os.path.join(TEST_DATA_DIR, "inf_fl_src")
dest = os.path.join(TEST_DATA_DIR, "inf_fl_dst")
rdst = os.path.join(TEST_DATA_DIR, "inf_fl_rdst")
_output_tree(source)
clean_dir(dest)
clean_dir(rdst)
rsync_result = _rsync(["-a", "--info=flist", source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
result, _ = run_client(source, dest, flags=["-a", "--info=flist"],
port=shared_server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
assert "sending incremental file list" in result.stdout
assert "sending incremental file list" in rsync_result.stdout
@requires_rsync
@pytest.mark.ci
def test_info_name_matches_rsync(self, shared_server):
source = os.path.join(TEST_DATA_DIR, "inf_nm_src")
dest = os.path.join(TEST_DATA_DIR, "inf_nm_dst")
rdst = os.path.join(TEST_DATA_DIR, "inf_nm_rdst")
_output_tree(source)
clean_dir(dest)
clean_dir(rdst)
rsync_result = _rsync(["-a", "--info=name", source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
result, _ = run_client(source, dest, flags=["-a", "--info=name"],
port=shared_server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
def entries(text):
# Compare the transferred entries only: rsync also prints the
# transfer-root `./` and every directory (FastSync records dirs),
# which are a separate documented divergence.
out = []
for line in text.splitlines():
if not line or line.startswith("sending ") or line.startswith("created "):
continue
if line == "./" or line.endswith("/"):
continue
out.append(line)
return sorted(out)
assert entries(result.stdout) == entries(rsync_result.stdout), (
f"rsync={entries(rsync_result.stdout)} fastsync={entries(result.stdout)}"
)
@requires_rsync
@pytest.mark.ci
def test_info_name_root_line_matches_rsync(self, shared_server):
"""A fresh destination: rsync prints `created directory`, then the
transfer-root `./` name line before the entries; FastSync must emit the
same `./` line."""
source = os.path.join(TEST_DATA_DIR, "inf_root_src")
dest = os.path.join(TEST_DATA_DIR, "inf_root_dst")
rdst = os.path.join(TEST_DATA_DIR, "inf_root_rdst")
clean_dir(source)
_write(os.path.join(source, "f.bin"), b"payload\n")
clean_dir(dest)
shutil.rmtree(rdst, ignore_errors=True)
rsync_result = _rsync(["-a", "--info=name", source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
result, _ = run_client(source, dest, flags=["-a", "--info=name"],
port=shared_server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
def names(text):
return [l for l in text.splitlines()
if l and not l.startswith("created directory")
and not (l.endswith("/") and l != "./")]
assert names(rsync_result.stdout) == ["./", "f.bin"], names(rsync_result.stdout)
assert names(result.stdout) == ["./", "f.bin"], names(result.stdout)
@requires_rsync
@pytest.mark.ci
def test_info_name2_uptodate_matches_rsync(self, shared_server):
"""--info=name2 prints `NAME is uptodate` for entries the receiver
already has, matching rsync byte-for-byte."""
source = os.path.join(TEST_DATA_DIR, "inf_up_src")
dest = os.path.join(TEST_DATA_DIR, "inf_up_dst")
rdst = os.path.join(TEST_DATA_DIR, "inf_up_rdst")
clean_dir(source)
os.makedirs(os.path.join(source, "sub"))
_write(os.path.join(source, "a.txt"), b"a\n")
_write(os.path.join(source, "sub", "b.txt"), b"b\n")
clean_dir(rdst)
assert _rsync(["-a", source + "/", rdst + "/"]).returncode == 0
rsync_result = _rsync(["-a", "--info=name2", source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
clean_dir(dest)
seed, _ = run_client(source, dest, flags=["-a", "--incremental"],
port=shared_server.port)
assert seed.returncode == 0, (seed.stderr or seed.stdout)[:200]
result, _ = run_client(source, dest, flags=["-a", "--incremental", "--info=name2"],
port=shared_server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
rsync_lines = sorted(l for l in rsync_result.stdout.splitlines()
if l.endswith("is uptodate"))
fast_lines = sorted(l for l in result.stdout.splitlines()
if l.endswith("is uptodate"))
assert fast_lines == rsync_lines, (rsync_lines, fast_lines)
assert fast_lines == ["a.txt is uptodate", "sub/b.txt is uptodate"], fast_lines
@requires_rsync
@pytest.mark.ci
def test_info_nonreg_matches_rsync(self, shared_server):
source = os.path.join(TEST_DATA_DIR, "inf_nr_src")
dest = os.path.join(TEST_DATA_DIR, "inf_nr_dst")
rdst = os.path.join(TEST_DATA_DIR, "inf_nr_rdst")
clean_dir(source)
os.mkfifo(os.path.join(source, "fifo"))
_write(os.path.join(source, "a.txt"), b"a\n")
clean_dir(dest)
clean_dir(rdst)
rsync_result = _rsync(["-rlt", "--info=nonreg", source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
result, _ = run_client(source, dest, flags=["-rlt", "--info=nonreg"],
port=shared_server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
rsync_lines = sorted(l for l in rsync_result.stdout.splitlines()
if l.startswith("skipping non-regular"))
fast_lines = sorted(l for l in result.stdout.splitlines()
if l.startswith("skipping non-regular"))
assert fast_lines == rsync_lines, (rsync_lines, fast_lines)
assert fast_lines, "no non-regular skip line emitted"
@requires_rsync
@pytest.mark.ci
def test_info_del_real_matches_rsync(self, shared_server):
source = os.path.join(TEST_DATA_DIR, "inf_dl_src")
dest = os.path.join(TEST_DATA_DIR, "inf_dl_dst")
rdst = os.path.join(TEST_DATA_DIR, "inf_dl_rdst")
clean_dir(source)
_write(os.path.join(source, "keep.txt"), b"keep\n")
clean_dir(rdst)
_write(os.path.join(rdst, "extra.txt"), b"x\n")
_write(os.path.join(rdst, "extra2.txt"), b"y\n")
rsync_result = _rsync(["-a", "--delete", "--info=del", source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
rsync_lines = sorted(l for l in rsync_result.stdout.splitlines()
if l.startswith("deleting "))
clean_dir(dest)
received = get_dest_received_dir(dest, source)
_write(os.path.join(received, "extra.txt"), b"x\n")
_write(os.path.join(received, "extra2.txt"), b"y\n")
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
result, _ = run_client(source, dest, flags=["-a", "--delete", "--info=del"],
port=server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
fast_lines = sorted(l for l in result.stdout.splitlines()
if l.startswith("deleting "))
assert fast_lines == rsync_lines, (rsync_lines, fast_lines)
assert fast_lines, "no deletion lines emitted"
@requires_rsync
@pytest.mark.ci
def test_info_del_itemize_real_matches_rsync(self, shared_server):
source = os.path.join(TEST_DATA_DIR, "inf_di_src")
dest = os.path.join(TEST_DATA_DIR, "inf_di_dst")
rdst = os.path.join(TEST_DATA_DIR, "inf_di_rdst")
clean_dir(source)
_write(os.path.join(source, "keep.txt"), b"keep\n")
clean_dir(rdst)
_write(os.path.join(rdst, "extra.txt"), b"x\n")
rsync_result = _rsync(["-a", "-i", "--delete", source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
rsync_lines = sorted(l for l in rsync_result.stdout.splitlines()
if l.startswith("*deleting"))
clean_dir(dest)
received = get_dest_received_dir(dest, source)
_write(os.path.join(received, "extra.txt"), b"x\n")
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
result, _ = run_client(source, dest, flags=["-a", "-i", "--delete"],
port=server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
fast_lines = sorted(l for l in result.stdout.splitlines()
if l.startswith("*deleting"))
assert fast_lines == rsync_lines, (rsync_lines, fast_lines)
@requires_rsync
@pytest.mark.ci
def test_info_del_dry_run_matches_rsync(self, shared_server):
source = os.path.join(TEST_DATA_DIR, "inf_dd_src")
dest = os.path.join(TEST_DATA_DIR, "inf_dd_dst")
rdst = os.path.join(TEST_DATA_DIR, "inf_dd_rdst")
clean_dir(source)
_write(os.path.join(source, "keep.txt"), b"keep\n")
clean_dir(rdst)
_write(os.path.join(rdst, "extra.txt"), b"x\n")
rsync_result = _rsync(["-a", "-n", "--delete", "--info=del", source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
rsync_lines = sorted(l for l in rsync_result.stdout.splitlines()
if l.startswith("deleting "))
clean_dir(dest)
received = get_dest_received_dir(dest, source)
_write(os.path.join(received, "extra.txt"), b"x\n")
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
result, _ = run_client(source, dest, flags=["-a", "-n", "--delete", "--info=del"],
port=server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
fast_lines = sorted(l for l in result.stdout.splitlines()
if l.startswith("deleting "))
assert fast_lines == rsync_lines, (rsync_lines, fast_lines)
@requires_rsync
@pytest.mark.ci
def test_info_remove_matches_rsync(self, shared_server):
tag = "inf_rm"
rsync_src = os.path.join(TEST_DATA_DIR, f"{tag}_rsrc")
rsync_dst = os.path.join(TEST_DATA_DIR, f"{tag}_rdst")
fast_src = os.path.join(TEST_DATA_DIR, f"{tag}_fsrc")
fast_dst = os.path.join(TEST_DATA_DIR, f"{tag}_fdst")
for root in (rsync_src, rsync_dst, fast_src, fast_dst):
clean_dir(root)
_write(os.path.join(rsync_src, "a.txt"), b"a\n")
_write(os.path.join(rsync_src, "sub", "b.txt"), b"b\n")
_write(os.path.join(fast_src, "a.txt"), b"a\n")
_write(os.path.join(fast_src, "sub", "b.txt"), b"b\n")
rsync_result = _rsync(["-a", "--remove-source-files", "--info=remove",
rsync_src + "/", rsync_dst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
rsync_lines = sorted(l for l in rsync_result.stdout.splitlines()
if l.startswith("sender removed "))
result, _ = run_client(fast_src, fast_dst,
flags=["-a", "--remove-source-files", "--info=remove"],
port=shared_server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
fast_lines = sorted(l for l in result.stdout.splitlines()
if l.startswith("sender removed "))
assert fast_lines == rsync_lines, (rsync_lines, fast_lines)
assert fast_lines, "no source-removal lines emitted"
class TestIgnoreErrorsParity:
"""--ignore-errors: a source I/O error skips deletion by default; the flag
lets deletion proceed. Both exit 23. Run the client as an unprivileged user
so the mode-000 directory is genuinely unreadable."""
@pytest.mark.setpriv
def test_delete_after_io_error_matches_rsync(self):
if os.geteuid() != 0 or shutil.which("setpriv") is None:
pytest.skip("requires root + setpriv to drop privileges for the client")
tag = f"ie_{os.getpid()}"
source = os.path.join(TEST_DATA_DIR, f"{tag}_src")
rsync_dst = os.path.join(TEST_DATA_DIR, f"{tag}_rdst")
dest = os.path.join(TEST_DATA_DIR, f"{tag}_dst")
clean_dir(source)
clean_dir(rsync_dst)
clean_dir(dest)
_write(os.path.join(source, "top.txt"), b"top\n")
_write(os.path.join(source, "locked", "blocked.txt"), b"blocked\n")
os.chmod(os.path.join(source, "locked"), 0)
os.chmod(TEST_DATA_DIR, 0o777)
os.chmod(source, 0o755)
os.chmod(rsync_dst, 0o777)
os.chmod(dest, 0o777)
try:
for ignore in (False, True):
flags = ["-a", "--delete-after"] + (["--ignore-errors"] if ignore else [])
# rsync side
_write(os.path.join(rsync_dst, "extra.txt"), b"x\n")
os.chmod(os.path.join(rsync_dst, "extra.txt"), 0o666)
rres = _rsync(flags + [source + "/", rsync_dst + "/"], as_nobody=True)
rsync_extra = os.path.exists(os.path.join(rsync_dst, "extra.txt"))
# fastsync side
received = get_dest_received_dir(dest, source)
_write(os.path.join(received, "extra.txt"), b"x\n")
os.chmod(os.path.join(received, "extra.txt"), 0o666)
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
fflags = (["--delete", "--ignore-errors"] if ignore else ["--delete"])
cmd = CLIENT_CMD + ["--source-dir", source, "--dest-dir", dest,
"--save-to-disk", "--server-port", str(server.port)] + fflags
fres = subprocess.run(
["setpriv", "--reuid=65534", "--regid=65534", "--clear-groups"] + cmd,
text=True, capture_output=True)
fast_extra = os.path.exists(os.path.join(received, "extra.txt"))
assert rres.returncode == 23, (ignore, rres.returncode, rres.stderr[:200])
assert fres.returncode == 23, (ignore, fres.returncode, fres.stderr[:200])
assert rsync_extra == fast_extra, (
f"ignore_errors={ignore}: rsync extra={rsync_extra} fastsync extra={fast_extra}"
)
assert fast_extra is (not ignore), (ignore, fast_extra)
finally:
os.chmod(os.path.join(source, "locked"), 0o755)
class TestRemoteOptionDaemon:
"""rsync forwards -M/--remote-option to its remote process over a daemon
connection; FastSync's daemon has no per-connection argv channel and rejects
it. This pins the documented divergence with evidence."""
@requires_rsync
def test_rsync_forwards_M_over_daemon_and_fastsync_rejects(self, tmp_path):
import socket
with socket.socket() as probe:
probe.bind(("127.0.0.1", 0))
port = probe.getsockname()[1]
module_root = tmp_path / "mod"
module_root.mkdir()
os.chmod(module_root, 0o777)
source = tmp_path / "src"
source.mkdir()
(source / "a.txt").write_bytes(b"hello\n")
conf = tmp_path / "rsyncd.conf"
conf.write_text(
f"port = {port}\nuse chroot = no\n[m]\npath = {module_root}\nread only = no\n"
)
daemon = subprocess.Popen(
[RSYNC, "--daemon", "--no-detach", "--port", str(port), "--config", str(conf)],
stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
try:
deadline = time.monotonic() + 5
while time.monotonic() < deadline:
try:
with socket.create_connection(("127.0.0.1", port), timeout=0.3):
break
except OSError:
time.sleep(0.05)
else:
pytest.skip("rsync daemon did not start")
# A well-formed -M option is forwarded and accepted by the daemon...
ok = _rsync(["-a", "-M--safe-links", source.as_posix() + "/",
f"rsync://127.0.0.1:{port}/m/"])
# ...and a bogus one is rejected ON THE REMOTE with "unknown option",
# which proves the option reached the daemon's parser.
bogus = _rsync(["-a", "-M--totally-bogus", source.as_posix() + "/",
f"rsync://127.0.0.1:{port}/m/"])
assert bogus.returncode != 0
assert "unknown option" in (bogus.stderr + bogus.stdout), bogus.stderr
del ok
finally:
daemon.terminate()
try:
daemon.wait(timeout=5)
except subprocess.TimeoutExpired:
daemon.kill()
# FastSync rejects -M for a non-SSH transport up front.
dest = os.path.join(TEST_DATA_DIR, "ro_dst")
clean_dir(dest)
result, _ = run_client(source.as_posix(), dest, flags=["-a", "-M--safe-links"])
assert result.returncode != 0
assert "remote-option" in (result.stderr + result.stdout)
class TestFilterProtect:
"""Receiver-derived delete protection: a `protect`/`P` rule is compiled by
the sender and sent on the config frame, so the receiver shields a
destination-only entry that never appeared on the sender, matching rsync."""
@requires_rsync
@pytest.mark.ci
def test_protect_dest_only_matches_rsync(self, shared_server):
source = os.path.join(TEST_DATA_DIR, "fpd_src")
dest = os.path.join(TEST_DATA_DIR, "fpd_dst")
rdst = os.path.join(TEST_DATA_DIR, "fpd_rdst")
clean_dir(source)
_write(os.path.join(source, "keep.txt"), b"keep\n")
clean_dir(rdst)
_write(os.path.join(rdst, "extra.log"), b"extra\n")
_write(os.path.join(rdst, "other.txt"), b"other\n")
rsync_result = _rsync(["-a", "--delete", "--filter=P *.log", source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
assert os.path.exists(os.path.join(rdst, "extra.log")), "rsync did not protect extra.log"
assert not os.path.exists(os.path.join(rdst, "other.txt")), "rsync did not delete other.txt"
clean_dir(dest)
received = get_dest_received_dir(dest, source)
_write(os.path.join(received, "extra.log"), b"extra\n")
_write(os.path.join(received, "other.txt"), b"other\n")
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
result, _ = run_client(source, dest,
flags=["-a", "--delete", "--filter=P *.log"],
port=server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
assert os.path.exists(os.path.join(received, "extra.log")), (
"FastSync must protect a destination-only P match like rsync")
assert not os.path.exists(os.path.join(received, "other.txt"))
@pytest.mark.ci
def test_protect_dest_only_dry_run_enumeration(self, shared_server):
source = os.path.join(TEST_DATA_DIR, "fpd_nd_src")
dest = os.path.join(TEST_DATA_DIR, "fpd_nd_dst")
clean_dir(source)
_write(os.path.join(source, "keep.txt"), b"keep\n")
received = get_dest_received_dir(dest, source)
clean_dir(received)
_write(os.path.join(received, "keep.txt"), b"keep\n")
_write(os.path.join(received, "extra.log"), b"extra\n")
_write(os.path.join(received, "other.txt"), b"other\n")
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
result, _ = run_client(source, dest,
flags=["-a", "-n", "--delete", "--out-format=%n",
"--filter=P *.log"],
port=server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
assert "other.txt" in result.stdout, result.stdout
assert "extra.log" not in result.stdout, result.stdout
assert os.path.exists(os.path.join(received, "extra.log"))
assert os.path.exists(os.path.join(received, "other.txt"))
+455 -23
View File
@@ -26,6 +26,17 @@ def _rsync(args):
)
def _file_entry_line(text):
"""The file entry line for a single-file transfer.
-i/--out-format emit the transfer-root (and directory) lines too, so the
file entry is not necessarily the first line; for the one-file corpora used
by the wire-counter tests it is the last non-empty line.
"""
lines = [line for line in text.splitlines() if line.strip()]
return lines[-1] if lines else ""
def _make_selection_tree(root):
clean_dir(root)
os.makedirs(os.path.join(root, "sub"))
@@ -184,6 +195,120 @@ class TestItemizeParity:
)
assert fast_lines == rsync_lines, f"rsync={rsync_lines} fastsync={fast_lines}"
@requires_rsync
@pytest.mark.ci
def test_itemize_directory_lines_match_rsync(self, shared_server):
"""#292: -i/--out-format emit rsync's directory lines (including the
transfer root) in rsync's depth-first order."""
source = os.path.join(TEST_DATA_DIR, "out_itemdir_src")
dest = os.path.join(TEST_DATA_DIR, "out_itemdir_dst")
rdst = os.path.join(TEST_DATA_DIR, "out_itemdir_rdst")
clean_dir(source)
os.makedirs(os.path.join(source, "sub", "deep"))
os.makedirs(os.path.join(source, "emptydir"))
with open(os.path.join(source, "a.txt"), "wb") as fh:
fh.write(b"hello\n")
with open(os.path.join(source, "sub", "b.txt"), "wb") as fh:
fh.write(b"world\n")
with open(os.path.join(source, "sub", "deep", "d.txt"), "wb") as fh:
fh.write(b"deep\n")
clean_dir(dest)
clean_dir(rdst)
def dir_lines(text):
# Any line whose name ends with '/' is a directory entry.
return sorted(
line for line in text.splitlines()
if line.rsplit(" ", 1)[-1].endswith("/")
)
for fmt in (None, "%i %n%L"):
rsync_flags = ["-a", "-i"] if fmt is None else ["-a", "--out-format=" + fmt]
fast_flags = rsync_flags
clean_dir(rdst)
clean_dir(dest)
rsync_result = _rsync(rsync_flags + [source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
result, _ = run_client(source, dest, flags=fast_flags,
port=shared_server.port)
assert result.returncode == 0, result.stderr[:300]
expected = [l for l in dir_lines(rsync_result.stdout)
if not l.rsplit(" ", 1)[-1] == "./"]
fast = dir_lines(result.stdout)
assert [l for l in fast if not l.rsplit(" ", 1)[-1] == "./"] == expected, (
f"fmt={fmt} rsync={rsync_result.stdout!r} fastsync={result.stdout!r}"
)
assert ".d..t...... ./" in fast, f"missing root line: {result.stdout!r}"
@requires_rsync
@pytest.mark.ci
def test_itemize_info_flist_header_matches_rsync(self, shared_server):
"""`-i --info=flist` prints rsync's file-list header: the -i change
lines alone do not enable the flist category, but an explicit --info=flist
must not be suppressed when itemizing."""
source = os.path.join(TEST_DATA_DIR, "out_itemfl_src")
dest = os.path.join(TEST_DATA_DIR, "out_itemfl_dst")
rdst = os.path.join(TEST_DATA_DIR, "out_itemfl_rdst")
_make_output_tree(source)
clean_dir(dest)
clean_dir(rdst)
flags = ["-a", "-i", "--info=flist"]
rsync_result = _rsync(flags + [source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
result, _ = run_client(source, dest, flags=flags, port=shared_server.port)
assert result.returncode == 0, result.stderr[:300]
assert "sending incremental file list" in rsync_result.stdout
assert "sending incremental file list" in result.stdout, result.stdout
# -i alone (no explicit --info=flist) must stay silent like rsync.
clean_dir(dest)
clean_dir(rdst)
rsync_plain = _rsync(["-a", "-i", source + "/", rdst + "/"])
plain, _ = run_client(source, dest, flags=["-a", "-i"],
port=shared_server.port)
assert "sending incremental file list" not in rsync_plain.stdout
assert "sending incremental file list" not in plain.stdout, plain.stdout
@requires_rsync
@pytest.mark.ci
def test_itemize_files_from_dirs_root_and_ancestors(self, shared_server):
"""The -d/--files-from dirs generator emits the transfer-root line and
rsync's implied ancestor directory lines. The generator traverses no
directories, so those must be synthesized from the listed entries."""
source = os.path.join(TEST_DATA_DIR, "out_itemff_src")
dest = os.path.join(TEST_DATA_DIR, "out_itemff_dst")
rdst = os.path.join(TEST_DATA_DIR, "out_itemff_rdst")
clean_dir(source)
os.makedirs(os.path.join(source, "sub", "deep"))
with open(os.path.join(source, "sub", "deep", "d.txt"), "wb") as fh:
fh.write(b"deep\n")
clean_dir(dest)
clean_dir(rdst)
listing = os.path.join(TEST_DATA_DIR, "out_itemff.list")
with open(listing, "w") as fh:
fh.write("sub/deep/d.txt\n")
flags = ["-d", "-i", "--files-from=" + listing]
rsync_result = _rsync(flags + [source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
result, _ = run_client(source, dest, flags=flags, port=shared_server.port)
assert result.returncode == 0, result.stderr[:300]
def dir_lines(text):
return sorted(line for line in text.splitlines()
if line.rsplit(" ", 1)[-1].endswith("/"))
# rsync emits the implied parents (sub/, sub/deep/) but never the root
# here; FastSync emits the same set plus its unconditional root line.
expected = [line for line in dir_lines(rsync_result.stdout)
if not line.rsplit(" ", 1)[-1] == "./"]
fast = dir_lines(result.stdout)
assert [line for line in fast if not line.rsplit(" ", 1)[-1] == "./"] == expected, (
f"rsync={rsync_result.stdout!r} fastsync={result.stdout!r}"
)
assert "cd+++++++++ sub/" in fast, result.stdout
assert "cd+++++++++ sub/deep/" in fast, result.stdout
assert any(line.rsplit(" ", 1)[-1] == "./" for line in fast), result.stdout
@requires_rsync
@pytest.mark.ci
def test_itemize_modified_file_matches_rsync(self, shared_server):
@@ -262,6 +387,47 @@ class TestOutFormatParity:
if line:
assert pattern.match(line), f"bad %M format: {line!r}"
@requires_rsync
@pytest.mark.ci
def test_out_format_directory_metadata_with_delete_during(self):
"""--delete-during/--delete-delay reuse the per-directory plan pre-scan,
whose list carries no metadata. Directory %M/%B/%U/%G must still come
from the source, exactly as the plain recursive scan renders them."""
source = os.path.join(TEST_DATA_DIR, "out_fmtmeta_src")
dest = os.path.join(TEST_DATA_DIR, "out_fmtmeta_dst")
rdst = os.path.join(TEST_DATA_DIR, "out_fmtmeta_rdst")
clean_dir(source)
os.makedirs(os.path.join(source, "sub", "deep"))
with open(os.path.join(source, "a.txt"), "wb") as fh:
fh.write(b"hello\n")
with open(os.path.join(source, "sub", "b.txt"), "wb") as fh:
fh.write(b"world\n")
clean_dir(dest)
clean_dir(rdst)
def dir_lines(text):
# Directory names are the last whitespace-separated token.
return sorted(line for line in text.splitlines()
if line.rsplit(" ", 1)[-1].endswith("/")
and line.rsplit(" ", 1)[-1] != "./")
for timing in ("--delete-during", "--delete-delay"):
for fmt in ("%M %n", "%B %n", "%U %G %n"):
clean_dir(dest)
clean_dir(rdst)
flags = ["-a", "--out-format=" + fmt, timing]
rsync_result = _rsync(flags + [source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
result, _ = run_client(source, dest, flags=flags, port=server.port)
assert result.returncode == 0, result.stderr[:300]
assert dir_lines(result.stdout) == dir_lines(rsync_result.stdout), (
f"{timing} {fmt}: rsync={rsync_result.stdout!r} "
f"fastsync={result.stdout!r}"
)
assert "1970/" not in result.stdout, result.stdout
class TestListOnlyParity:
@requires_rsync
@@ -289,6 +455,49 @@ def _make_one_file(root, name="f.bin", size=100):
fh.write(bytes((i * 7 + 3) & 0xFF for i in range(size)))
def _make_multidir_tree(root):
"""Multi-directory corpus for the --progress file-list tests: nested files,
a directory-only branch, an empty directory and a symlink."""
clean_dir(root)
for rel, data in (("a.txt", b"alpha\n"), ("b.txt", b"bravo\n"),
("sub1/c.txt", b"charlie\n"), ("sub1/deep/d.txt", b"delta\n"),
("sub2/e.txt", b"echo\n")):
path = os.path.join(root, rel)
os.makedirs(os.path.dirname(path), exist_ok=True)
with open(path, "wb") as fh:
fh.write(data)
os.symlink("a.txt", os.path.join(root, "link1"))
os.makedirs(os.path.join(root, "emptydir"), exist_ok=True)
def _parse_progress(text):
"""Name lines and the `to-chk` denominators from a --progress run."""
names = []
totals = set()
for line in text.splitlines():
line = line.rstrip()
if not line or line == "sending incremental file list":
continue
if "%" in line:
match = re.search(r"to-chk=\d+/(\d+)", line)
if match:
totals.add(int(match.group(1)))
continue
if line == "./": # root-line trigger is a separate documented residual
continue
names.append(line)
return sorted(names), totals
def _pick_stats(text, keys):
out = {}
for line in text.splitlines():
for key in keys:
if line.startswith(key + ":"):
out[key] = line
return out
class TestWireStatsParity:
"""Wire-counter output parity: --out-format %b/%c/%C, --progress and
--stats versus real rsync 3.4.1."""
@@ -344,8 +553,8 @@ class TestWireStatsParity:
result, _ = run_client(source, dest, flags=["-a", "--out-format=" + fmt],
port=shared_server.port)
assert result.returncode == 0, result.stderr[:300]
rb, rl = (int(x) for x in rsync_result.stdout.split()[:2])
fb, fl = (int(x) for x in result.stdout.split()[:2])
rb, rl = (int(x) for x in _file_entry_line(rsync_result.stdout).split()[:2])
fb, fl = (int(x) for x in _file_entry_line(result.stdout).split()[:2])
assert rl == fl == 5000, (rsync_result.stdout, result.stdout)
assert rb > rl, f"rsync %b must include framing: {rsync_result.stdout!r}"
assert fb > fl, f"fastsync %b must include framing: {result.stdout!r}"
@@ -378,7 +587,9 @@ class TestWireStatsParity:
assert file_lines(result.stdout) == file_lines(rsync_result.stdout), (
f"rsync={rsync_result.stdout!r} fastsync={result.stdout!r}"
)
assert result.stdout.split()[0] == rsync_result.stdout.split()[0] == "16", (
fs_c = _file_entry_line(result.stdout).split()[0]
rs_c = _file_entry_line(rsync_result.stdout).split()[0]
assert fs_c == rs_c == "16", (
f"%c must be rsync's 16-byte sum header: {result.stdout!r}"
)
@@ -407,8 +618,8 @@ class TestWireStatsParity:
"--out-format=" + fmt],
port=shared_server.port)
assert result.returncode == 0, result.stderr[:300]
rs_c = int(rsync_result.stdout.split()[0])
fs_c = int(result.stdout.split()[0])
rs_c = int(_file_entry_line(rsync_result.stdout).split()[0])
fs_c = int(_file_entry_line(result.stdout).split()[0])
# No basis exists, so rsync still reports only its sum header.
assert rs_c == 16, rsync_result.stdout
# FastSync reports its own handshake bytes and is not aligned.
@@ -447,6 +658,99 @@ class TestWireStatsParity:
assert fast_frames[0] == rsync_frames[0], (rsync_frames[0], fast_frames[0])
assert "(xfr#1," in fast_frames[-1], fast_frames[-1]
@requires_rsync
@pytest.mark.ci
def test_progress_leading_root_line_and_to_chk_match_rsync(self, shared_server):
"""A single-file transfer: rsync emits the transfer-root `./` name line
and a `to-chk=0/2` denominator that counts that root entry. Both must
match FastSync byte-for-byte for the deterministic frames."""
source = os.path.join(TEST_DATA_DIR, "wire_pgroot_src")
dest = os.path.join(TEST_DATA_DIR, "wire_pgroot_dst")
rdst = os.path.join(TEST_DATA_DIR, "wire_pgroot_rdst")
_make_one_file(source, "f.bin", 100)
clean_dir(dest)
# rsync prints the `./` root line only when the transfer root itself is
# created, so make the rsync destination absent. The "created directory"
# line it then emits has no FastSync counterpart (different mirror
# layout), so only the name/frame lines are compared.
shutil.rmtree(rdst, ignore_errors=True)
rsync_result = _rsync(["-a", "--progress", source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
result, _ = run_client(source, dest, flags=["-a", "--progress"],
port=shared_server.port)
assert result.returncode == 0, result.stderr[:300]
# subprocess text mode normalizes \r to \n (universal newlines).
def lines_of(text):
return [ln for ln in text.splitlines() if ln and not ln.startswith("created directory")]
rsync_lines = lines_of(rsync_result.stdout)
fast_lines = lines_of(result.stdout)
rsync_names = [ln for ln in rsync_lines if "%" not in ln]
fast_names = [ln for ln in fast_lines if "%" not in ln]
assert rsync_names == ["sending incremental file list", "./", "f.bin"], rsync_names
assert fast_names == rsync_names, (rsync_names, fast_names)
# The final frame's to-chk denominator must include the source-root entry.
assert "to-chk=0/2" in fast_lines[-1], fast_lines[-1]
assert fast_lines[-1] == rsync_lines[-1], (rsync_lines[-1], fast_lines[-1])
@requires_rsync
@pytest.mark.ci
@pytest.mark.parametrize("mt", [False, True])
def test_progress_multidir_file_list_matches_rsync(self, shared_server, mt):
"""A multi-directory tree: the paths-only pre-count must reproduce
rsync's file-list set and `to-chk` denominator. Per-directory name
lines are emitted for directories, symlinks and the empty directory; the
name set and the denominator (every entry plus the transfer root) match
rsync, while the emitted *order* remains a documented residual (rsync
sorts depth-first, FastSync streams in readdir/BFS order)."""
source = os.path.join(TEST_DATA_DIR, "wire_pgmd_src")
dest = os.path.join(TEST_DATA_DIR, "wire_pgmd_dst")
rdst = os.path.join(TEST_DATA_DIR, "wire_pgmd_rdst")
_make_multidir_tree(source)
clean_dir(dest)
clean_dir(rdst)
rsync_result = _rsync(["-a", "--progress", source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
flags = ["-a", "--progress"] + (["--threads"] if mt else [])
result, _ = run_client(source, dest, flags=flags, port=shared_server.port)
assert result.returncode == 0, result.stderr[:300]
rsync_names, rsync_totals = _parse_progress(rsync_result.stdout)
fast_names, fast_totals = _parse_progress(result.stdout)
assert sorted(rsync_names) == [
"a.txt", "b.txt", "emptydir/", "link1 -> a.txt", "sub1/",
"sub1/c.txt", "sub1/deep/", "sub1/deep/d.txt", "sub2/", "sub2/e.txt",
], rsync_names
assert fast_names == rsync_names, (rsync_names, fast_names)
# 10 entries + the transfer-root "." counted by rsync's file list.
assert rsync_totals == {11}, rsync_totals
assert fast_totals == rsync_totals, (rsync_totals, fast_totals)
@pytest.mark.ci
def test_progress_delete_during_reuses_pre_scan(self):
"""--delete-during + --progress reuses the keep-set pre-scan instead of
walking the tree a second time: the file-list total and directory name
lines are identical to a plain --progress run."""
source = os.path.join(TEST_DATA_DIR, "wire_pgdel_src")
dest = os.path.join(TEST_DATA_DIR, "wire_pgdel_dst")
_make_multidir_tree(source)
clean_dir(dest)
server = ServerManager()
server.start(extra_args=["--allow-super", "--allow-delete"])
try:
result, _ = run_client(source, dest, flags=["-a", "--progress", "--delete-during"],
port=server.port)
finally:
server.stop()
assert result.returncode == 0, result.stderr[:300]
names, totals = _parse_progress(result.stdout)
assert totals == {11}, totals
assert "sub1/" in names and "sub1/deep/" in names and "emptydir/" in names, names
assert "link1 -> a.txt" in names, names
@requires_rsync
@pytest.mark.ci
@pytest.mark.parametrize("mt", [False, True])
@@ -459,6 +763,9 @@ class TestWireStatsParity:
_make_one_file(source, "f.bin", 6000)
clean_dir(dest)
clean_dir(rdst)
# Start both tools from the same state: rsync's destination root exists,
# so pre-create FastSync's mirrored logical root as well.
os.makedirs(get_dest_received_dir(dest, source), exist_ok=True)
rsync_result = _rsync(["-a", "--stats", source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
flags = ["-a", "--stats"] + (["--threads"] if mt else [])
@@ -488,24 +795,19 @@ class TestWireStatsParity:
@requires_rsync
@pytest.mark.ci
def test_stats_file_count_breakdown_residual(self, shared_server):
"""Residual (row #3): rsync prints the `Number of files` and
`Number of created files` lines with a per-type breakdown
(`(reg: X, dir: Y, link: Z)`).
FastSync cannot reproduce it from what the sender currently knows: the
scanner does not put directory entries in the transfer list (directories
are created implicitly), and without a per-entry destination-probe the
sender cannot tell which entries the receiver newly created. So FastSync
prints the bare transferred-entry count. This test pins the divergence
explicitly -- the row must not be marked ✅.
"""
def test_stats_file_count_breakdown_matches_rsync(self, shared_server):
"""`Number of files` and `Number of created files` both carry rsync's
per-type breakdown (protocol 2.28.0 reports the receiver-created
reg/dir/link/special split over STATUS_STATS)."""
source = os.path.join(TEST_DATA_DIR, "wire_stc_src")
dest = os.path.join(TEST_DATA_DIR, "wire_stc_dst")
rdst = os.path.join(TEST_DATA_DIR, "wire_stc_rdst")
_make_one_file(source, "f.bin", 6000)
clean_dir(dest)
clean_dir(rdst)
# Start both tools from the same state: rsync's destination root exists,
# so pre-create FastSync's mirrored logical root as well.
os.makedirs(get_dest_received_dir(dest, source), exist_ok=True)
rsync_result = _rsync(["-a", "--stats", source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
result, _ = run_client(source, dest, flags=["-a", "--stats"],
@@ -523,14 +825,144 @@ class TestWireStatsParity:
f_files = stats_line(result.stdout, "Number of files")
f_created = stats_line(result.stdout, "Number of created files")
# rsync always carries the type breakdown (the source root counts as a
# directory; the single regular file as reg).
assert re.match(r"Number of files: 2 \(reg: 1, dir: 1\)$", r_files), r_files
assert r_files == f_files, (r_files, f_files)
assert re.match(r"Number of created files: 1 \(reg: 1\)$", r_created), r_created
# FastSync prints only the bare count: no directory accounting and no
# per-entry "created" knowledge.
assert re.fullmatch(r"Number of files: 1", f_files), f_files
assert re.fullmatch(r"Number of created files: 1", f_created), f_created
assert f_created == r_created, (r_created, f_created)
@requires_rsync
@pytest.mark.ci
@pytest.mark.parametrize("mt", [False, True])
def test_stats_r_directory_breakdown_matches_rsync(self, shared_server, mt):
"""A recursive `-r` scan (no -t/-p) exposes no directory metadata, but
rsync still counts every directory in `Number of files`; the sender's
lightweight directory counter must reproduce the `dir: N` category."""
source = os.path.join(TEST_DATA_DIR, "wire_stdir_src")
dest = os.path.join(TEST_DATA_DIR, "wire_stdir_dst")
rdst = os.path.join(TEST_DATA_DIR, "wire_stdir_rdst")
clean_dir(source)
clean_dir(dest)
clean_dir(rdst)
os.makedirs(os.path.join(source, "sub", "deep"))
os.makedirs(os.path.join(source, "empty"))
for rel in ("a.txt", os.path.join("sub", "b.txt"), os.path.join("sub", "deep", "c.txt")):
with open(os.path.join(source, rel), "wb") as fh:
fh.write(b"x\n")
os.makedirs(get_dest_received_dir(dest, source), exist_ok=True)
rsync_result = _rsync(["-r", "--stats", source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
flags = ["-r", "--stats"] + (["--threads"] if mt else [])
result, _ = run_client(source, dest, flags=flags, port=shared_server.port)
assert result.returncode == 0, result.stderr[:300]
def stats_line(text, key):
for line in text.splitlines():
if line.startswith(key + ":"):
return line
return None
r_files = stats_line(rsync_result.stdout, "Number of files")
f_files = stats_line(result.stdout, "Number of files")
# 3 regular files, 4 directories (root, sub, sub/deep, empty).
assert re.match(r"Number of files: 7 \(reg: 3, dir: 4\)$", r_files), r_files
assert f_files == r_files, (r_files, f_files)
assert (stats_line(result.stdout, "Number of regular files transferred") ==
stats_line(rsync_result.stdout, "Number of regular files transferred"))
@requires_rsync
@pytest.mark.ci
@pytest.mark.parametrize("mt", [False, True])
def test_stats_created_and_literal_fresh_update_delta(self, shared_server, mt):
"""The receiver-observed counters must match rsync for the three
transfer shapes: a fresh create (created breakdown + whole-file literal),
an update (created == 0, whole-file literal), and a delta update (only
the literal delta fragments are counted, not the whole file)."""
source = os.path.join(TEST_DATA_DIR, "wire_stcd_src")
dest = os.path.join(TEST_DATA_DIR, "wire_stcd_dst")
rdst = os.path.join(TEST_DATA_DIR, "wire_stcd_rdst")
clean_dir(source)
clean_dir(dest)
clean_dir(rdst)
os.makedirs(source, exist_ok=True)
os.makedirs(get_dest_received_dir(dest, source), exist_ok=True)
with open(os.path.join(source, "big.bin"), "wb") as fh:
fh.write(bytes(range(256)) * 4096) # 1 MiB
mt_flag = ["--threads"] if mt else []
def compare(tag):
# Pin the delta block size on both ends: rsync's adaptive block size
# would otherwise make the literal/matched split non-comparable.
rsync_result = _rsync(["-a", "--stats", "--no-whole-file", "-B8192",
source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
result, _ = run_client(
source, dest,
flags=["-a", "--stats", "--incremental", "--delta", "-B8192"] + mt_flag,
port=shared_server.port)
assert result.returncode == 0, result.stderr[:300]
keys = ("Number of created files", "Literal data", "Matched data",
"Total transferred file size")
r = _pick_stats(rsync_result.stdout, keys)
f = _pick_stats(result.stdout, keys)
assert r == f, f"{tag}: rsync={r} fastsync={f}"
return r
fresh = compare("fresh")
assert re.match(r"Number of created files: 1 \(reg: 1\)$",
fresh["Number of created files"]), fresh
# Update the source and re-run: the destination already exists.
sleep_mtime = os.path.getmtime(os.path.join(source, "big.bin")) + 2
with open(os.path.join(source, "big.bin"), "r+b") as fh:
fh.seek(100)
fh.write(b"XXXXXXXXXX")
os.utime(os.path.join(source, "big.bin"), (sleep_mtime, sleep_mtime))
update = compare("update")
assert update["Number of created files"] == "Number of created files: 0", update
# Second delta update: change bytes far apart, so rsync ships only the
# literal fragments and FastSync must report the same Literal data.
sleep_mtime = os.path.getmtime(os.path.join(source, "big.bin")) + 2
with open(os.path.join(source, "big.bin"), "r+b") as fh:
fh.seek(500000)
fh.write(b"YYYYYYYYYY")
os.utime(os.path.join(source, "big.bin"), (sleep_mtime, sleep_mtime))
delta = compare("delta")
assert delta["Number of created files"] == "Number of created files: 0", delta
lit = int(delta["Literal data"].split(":", 1)[1].strip().split()[0].replace(",", ""))
assert 0 < lit < 1024 * 1024, delta
@requires_rsync
@pytest.mark.ci
@pytest.mark.parametrize("choice", ["xxh128", "xxh64", "xxh3", "md5", "md4", "sha1", "none"])
def test_out_format_C_selected_algorithm_matches_rsync(self, shared_server, choice):
"""`%C` must use the algorithm selected by --checksum-choice, not always
xxh128, and render it exactly like rsync (big-endian for the 64-bit
hashes, high-then-low for xxh128, standard hex for md5/md4/sha1)."""
source = os.path.join(TEST_DATA_DIR, f"wire_cc_{choice}_src")
dest = os.path.join(TEST_DATA_DIR, f"wire_cc_{choice}_dst")
rdst = os.path.join(TEST_DATA_DIR, f"wire_cc_{choice}_rdst")
_make_one_file(source, "f.bin", 200000)
clean_dir(dest)
clean_dir(rdst)
fmt = "%C %l %n"
rsync_result = _rsync(["-a", "--checksum-choice=" + choice,
"--out-format=" + fmt, source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
result, _ = run_client(source, dest,
flags=["-a", "--checksum-choice=" + choice,
"--out-format=" + fmt],
port=shared_server.port)
assert result.returncode == 0, result.stderr[:300]
def file_lines(text):
return [line for line in text.splitlines()
if line and not line.rsplit(" ", 1)[-1].endswith("/")]
assert file_lines(result.stdout) == file_lines(rsync_result.stdout), (
f"choice={choice}: rsync={rsync_result.stdout!r} fastsync={result.stdout!r}"
)
@requires_rsync
@pytest.mark.ci
@@ -0,0 +1,198 @@
"""Differential rsync-parity coverage for two residuals closed on this branch.
* A4 -- ``--compare-dest``/``--copy-dest``/``--link-dest`` relative-DIR
resolution: rsync resolves a relative DIR against the destination directory
and appends the file's TRANSFER-RELATIVE name. FastSync's default transfer
mirrors the absolute source path below its receive root, so a naive relative
DIR used to probe a different tree. These tests seed the basis at rsync's
spelling and assert FastSync finds it (byte-exact / hard-linked / sparse),
matching real rsync 3.4.1.
* A5 -- ``-y``/``--fuzzy`` candidate eligibility: rsync's ``find_fuzzy`` has no
delta-size gate, so it reuses an oversized (>10x) or sub-16-KiB sibling;
FastSync used to decline both. These tests assert FastSync now uses the same
sibling as rsync (observable as ``Matched data``) with a byte-exact result.
Every test skips cleanly when rsync is absent.
"""
import os
import shutil
import subprocess
import sys
import pytest
sys.path.insert(0, os.path.dirname(__file__))
from common import ( # noqa: E402
TEST_DATA_DIR,
clean_dir,
get_dest_received_dir,
run_client,
)
RSYNC = shutil.which("rsync")
requires_rsync = pytest.mark.skipif(RSYNC is None, reason="rsync 3.4.1 not installed")
OLD_MTIME = 1_500_000_000
def _write(path, content, mtime=None):
os.makedirs(os.path.dirname(path), exist_ok=True)
with open(path, "wb") as fh:
fh.write(content)
if mtime is not None:
os.utime(path, (mtime, mtime))
def _read(path):
with open(path, "rb") as fh:
return fh.read()
def _rsync(args):
env = dict(os.environ, LC_ALL="C")
return subprocess.run([RSYNC] + args, capture_output=True, text=True, env=env, timeout=120)
def _stat_bytes(text, label):
for line in text.splitlines():
if line.startswith(label + ":"):
return int(line.split(":", 1)[1].strip().split()[0].replace(",", ""))
return None
class TestRelativeBasisDirResolution:
"""A4: a relative basis DIR must resolve to the same tree as rsync's."""
_FILES = {
"root.txt": b"root-basis-content\n",
"sub/nested.txt": b"nested-basis-content\n",
}
def _seed_source(self, source):
clean_dir(source)
for rel, data in self._FILES.items():
_write(os.path.join(source, rel), data, OLD_MTIME)
return self._FILES
@requires_rsync
@pytest.mark.parametrize("flag", ["--compare-dest", "--link-dest"])
def test_relative_dir_resolves_like_rsync(self, shared_server, flag):
tag = flag.lstrip("-")
source = os.path.join(TEST_DATA_DIR, f"relbasis_{tag}_src")
rdst = os.path.join(TEST_DATA_DIR, f"relbasis_{tag}_rdst")
fdst = os.path.join(TEST_DATA_DIR, f"relbasis_{tag}_fdst")
self._seed_source(source)
# rsync: relative DIR -> dest/basis/<transfer-relative name>.
clean_dir(rdst)
for rel, data in self._FILES.items():
_write(os.path.join(rdst, "basis", rel), data, OLD_MTIME)
rs = _rsync(["-a", f"{flag}=basis", source + "/", rdst + "/"])
assert rs.returncode == 0, rs.stderr
# FastSync: the SAME relative spelling seeded at the SAME
# transfer-relative location under its destination root.
clean_dir(fdst)
for rel, data in self._FILES.items():
_write(os.path.join(fdst, "basis", rel), data, OLD_MTIME)
result, _ = run_client(source, fdst,
flags=["-a", f"{flag}=basis", "--incremental"],
port=shared_server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
received = get_dest_received_dir(fdst, source)
for rel, data in self._FILES.items():
rfile = os.path.join(rdst, rel)
ffile = os.path.join(received, rel)
basis = os.path.join(fdst, "basis", rel)
if flag == "--compare-dest":
# compare-dest never copies: both destinations stay sparse.
assert not os.path.exists(rfile), f"rsync copied {rel}"
assert not os.path.exists(ffile), (
f"FastSync did not resolve the relative basis DIR at {basis!r} "
f"(expected {rel!r} to stay sparse like rsync)")
else:
# link-dest hard-links; a basis miss would transfer a new file.
assert os.path.exists(ffile), f"FastSync lost {rel}"
assert _read(ffile) == data
assert os.stat(ffile).st_ino == os.stat(basis).st_ino, (
f"FastSync did not hard-link {rel!r} to the relative basis at "
f"{basis!r} (basis not resolved like rsync)")
@requires_rsync
def test_relative_dir_copy_dest_content(self, shared_server):
source = os.path.join(TEST_DATA_DIR, "relbasis_copy_src")
fdst = os.path.join(TEST_DATA_DIR, "relbasis_copy_fdst")
self._seed_source(source)
clean_dir(fdst)
for rel, data in self._FILES.items():
_write(os.path.join(fdst, "basis", rel), data, OLD_MTIME)
result, _ = run_client(source, fdst,
flags=["-a", "--copy-dest=basis", "--incremental"],
port=shared_server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
received = get_dest_received_dir(fdst, source)
for rel, data in self._FILES.items():
ffile = os.path.join(received, rel)
assert os.path.exists(ffile), f"copy-dest did not materialize {rel}"
assert _read(ffile) == data
assert os.stat(ffile).st_ino != os.stat(os.path.join(fdst, "basis", rel)).st_ino
class TestFuzzyEligibilityWindow:
"""A5: --fuzzy candidate eligibility must match rsync's uncapped window."""
BASE = b"the quick brown fox jumps over the lazy dog\n" * 4000
def _run_pair(self, shared_server, tag, payload, sibling):
source = os.path.join(TEST_DATA_DIR, f"fzw_{tag}_src")
dest = os.path.join(TEST_DATA_DIR, f"fzw_{tag}_dst")
rdst = os.path.join(TEST_DATA_DIR, f"fzw_{tag}_rdst")
clean_dir(source)
clean_dir(dest)
clean_dir(rdst)
_write(os.path.join(source, "report_v2.txt"), payload)
for root in (rdst, get_dest_received_dir(dest, source)):
_write(os.path.join(root, "report_v1.txt"), sibling)
rs = _rsync(["-a", "--no-whole-file", "--fuzzy", "--stats",
source + "/", rdst + "/"])
assert rs.returncode == 0, rs.stderr
result, _ = run_client(
source, dest,
flags=["-a", "--incremental", "--delta", "--fuzzy", "--stats"],
port=shared_server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
# The reconstructed file is byte-exact in every case.
assert _read(os.path.join(get_dest_received_dir(dest, source),
"report_v2.txt")) == payload
return rs, result
@requires_rsync
def test_oversized_sibling_eligible_like_rsync(self, shared_server):
"""A sibling 20x the source is used by rsync; FastSync must too (its old
10x delta-size gate declined it)."""
n = 65536
payload = (self.BASE * ((n // len(self.BASE)) + 1))[:n]
sibling = (self.BASE * 200)[: n * 20]
rs, result = self._run_pair(shared_server, "big", payload, sibling)
assert _stat_bytes(rs.stdout, "Matched data") > 0, \
"rsync should use a >10x fuzzy basis"
assert _stat_bytes(result.stdout, "Matched data") > 0, (
"FastSync's fuzzy eligibility must accept a >10x sibling like rsync "
f"(Matched data={_stat_bytes(result.stdout, 'Matched data')})")
@requires_rsync
def test_small_source_sibling_eligible_like_rsync(self, shared_server):
"""A sub-16-KiB source with an identical sibling is used by rsync;
FastSync's old 16 KiB delta minimum declined it."""
n = 8192
payload = (self.BASE * ((n // len(self.BASE)) + 1))[:n]
rs, result = self._run_pair(shared_server, "small", payload, payload)
assert _stat_bytes(rs.stdout, "Matched data") > 0, \
"rsync applies --fuzzy below 16 KiB"
assert _stat_bytes(result.stdout, "Matched data") > 0, (
"FastSync's fuzzy eligibility must accept a sub-16-KiB source like "
f"rsync (Matched data={_stat_bytes(result.stdout, 'Matched data')})")

Some files were not shown because too many files have changed in this diff Show More