10 Commits
Author SHA1 Message Date
TapTap cd7b96d0bb chore: remove manual-test scratch trees from branch
CI / lint (pull_request) Successful in 1m59s
CI / parity-full (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / parity-fast (pull_request) Successful in 17s
CI / build-and-test (pull_request) Successful in 49s
2026-09-19 17:10:12 +02:00
TapTap f6f49d536e test(delete): cover apply=false config-only carrier frame
Adds a unit test proving the config-only STATUS_DELETE_PLAN frame applies
--delete-missing-args exact deletions while walking no directory (the
--files-from-with-no-synced-dir fix).
2026-09-19 17:07:56 +02:00
TapTap 38d304103c fix(parity): review-wave fixes (basis stats over-report, filter-rule bounds)
- receiver stats: exclude basis-dir materializations (--link-dest/--copy-dest)
  from created/literal tallies; rsync reports 0 for a basis hit, so a fresh
  --link-dest --stats run now matches (differential test_link_dest_stats_matches_rsync)
- filter wire block: reject a pattern above the glob evaluation bound and lower
  MAX_FILTER_RULES to 1024, so a crafted rule list cannot amplify delete-walk
  glob work or install a rule that silently never protects
- negative tests for over-cap count and over-long pattern
- README: correct --delete default, --stats/--progress description, add
  --delete-commit; RSYNC_COMPAT stale version labels/overclaim fixed
2026-09-19 17:04:12 +02:00
TapTap 4b09213b88 chore: remove accidental scratch tree from branch 2026-09-19 16:30:21 +02:00
TapTap 36fd0774e8 feat(delete): default --delete to rsync delete-during; add --delete-commit
- plain --delete with no timing flag now selects delete-during (progressive
  deletion, matching rsync and avoiding the full old+new tree peak)
- new long-only FastSync --delete-commit restores the old atomic behavior
  (delete only after the whole transfer succeeds); timing-identical to
  --delete-after, implemented via the same wire bool
- timing flags are mutually exclusive; --delete-commit conflicts with other
  timings; --delete-before/--delete-during rows reworded per Phase-0 probes
- CHANGELOG migration note; tally unchanged 116/14/27
2026-09-19 16:29:48 +02:00
TapTap 711b7e50b3 docs(parity): --fuzzy name heuristic matches rsync; reclassify to caveat
Probe shows the candidate choice is observable only as --stats bandwidth
counters (tree/exit always identical). FastSync already uses rsync's
fuzzy_distance/find_filename_suffix name heuristic; the residual is the
narrower delta eligibility window (>=16 KiB, <=10x) vs rsync's wider one.
Row -> caveat; tally 116/14/27.
2026-09-19 16:06:24 +02:00
TapTap 67076bf218 feat(basis): rsync quick-check default + FastSync-only --verify-basis
- default basis match is rsync's metadata quick-check (size + mtime; size-only
  drops mtime; -I disables), no mandatory content digest
- new long-only --verify-basis (wire bool, protocol stays 2.28.0) restores the
  strict whole-file content equality
- --copy-dest re-applies source attributes; basis-hit 256 MiB cap removed by
  streaming the copy/hash; basis miss keeps the normal payload bound
- compare/copy/link-dest rows -> caveat; tally 116/13/28
2026-09-19 15:55:51 +02:00
TapTap 8ec8cb7203 test(parity): dest-only excluded entry protected under default --delete
Adds the exclude_protect_dest_only differential and flips --delete-excluded
to parity now that receiver-side rules protect a destination-only excluded
entry like rsync; tally 116/10/31.
2026-09-19 13:53:50 +02:00
TapTap 82959395fb feat(filter): receiver-side protect/risk engine for dest-only entries
- new bounded config-wire block (BLOCK_PROTECT_RULES) serializes the sender's
  compiled filter rules to the receiver (bounded count + 256 KiB patterns;
  strict action/sides validation)
- receiver evaluates protect/risk in the whole-tree extras walk and the
  per-directory delete plans, so a dest-only entry matching 'P' is kept like
  rsync; dry-run would-delete enumeration also honours it
- --filter flips to parity (115/11/31); per-dir merge receiver re-derivation
  remains the documented residual
2026-09-19 13:52:48 +02:00
TapTap c9f94ea46e docs(parity): --checksum-choice observable-equivalent to rsync; reclassify
Probe shows the block-checksum choice is not observable in the parity surface:
%c, Matched/Literal data and the destination tree are invariant across
xxh64/xxh128/xxh3/md5/md4/sha1 and the transfer,pre-transfer form; only %C
changes, and it is byte-identical to rsync. FastSync's fixed xxHash32 block
strong sum is collision-safe within the payload cap. Row -> parity (114/11/32).
2026-09-19 13:23:02 +02:00
40 changed files with 2125 additions and 395 deletions
+31
View File
@@ -4,6 +4,37 @@ All notable changes to FastSync are documented here. Versions match
`PROTOCOL_VERSION` (printed by `fastsync --version`); the client and server must
run the same version because the handshake is strict.
## [Unreleased]
### Changed
- **`--delete` now defaults to delete-during (rsync `--del`) timing.** With no
explicit timing flag, a plain `--delete` removes each directory's extras as
that directory is processed instead of committing one whole-tree deletion only
after the entire transfer succeeds. This matches rsync, frees destination
space progressively, and avoids the whole-old+new-tree peak that could
`ENOSPC` a tight destination. The client maps the default onto the existing
`delete_during` wire boolean, so `PROTOCOL_VERSION` stays `2.28.0`.
- Added the FastSync-only long option **`--delete-commit`** (implies
`--delete`): it selects the old late whole-tree commit and is timing-identical
to `--delete-after` (the same `delete_after` wire boolean). Explicit timing
flags always win over the default, at most one timing flag may be given, and a
timing flag combined with `--no-delete` is still rejected.
- The per-directory `STATUS_DELETE_PLAN` frame gained a one-int `apply` flag:
the one-shot per-run config block (protected prefixes, size-pruned mirrors,
`--delete-missing-args` exact paths) is now always transmitted first on a
config-only carrier (`apply=false`), fixing a latent bug where a
`--delete-missing-args` run whose `--files-from` list synchronized no directory
never sent its exact deletions.
### Migration
- Scripts that relied on plain `--delete` deleting nothing until the transfer
fully succeeded must pass **`--delete-commit`** (or `--delete-after`) to keep
that behavior. Plain `--delete` now removes reached directories' extras during
the transfer, exactly like rsync's default; on a completed run the final tree
is unchanged.
## [2.26.0] - 2026-09-17
### Added
+87 -5
View File
@@ -52,9 +52,10 @@
`tests/test_config.c`). Two residuals were reclassified **divergent**: `-M`
over daemon/TCP (no argv channel in FastSync's binary config handshake;
rsync-daemon differential pins the rsync behavior) and receiver-side
`protect`/`risk` re-derivation for destination-only entries (would need a
receiver filter engine; differential pins the divergence). The options pass
stands at **110 ✅ / 21 ⚠️ / 26 ❌**. New `tests/integration/test_option_parity.py`
`protect`/`risk` re-derivation for destination-only entries (would need a
receiver filter engine; differential pins the divergence — **reversed by
track 4a below**, which adds that engine). The options pass
stands at **110 ✅ / 21 ⚠️ / 26 ❌**. New `tests/integration/test_option_parity.py`
holds the rsync differentials (bwlimit parse+rate, info lines, real-setpriv
`--ignore-errors`, rsync-daemon `-M`, filter-protect pin).
@@ -74,7 +75,8 @@
heuristic with a 10× size window, not rsync's matcher), but its residual is the
candidate-selection heuristic itself: the final tree is byte-exact by design, so
it is pinned by the `TestFuzzy` threshold suite rather than a byte-level rsync
differential. The parity-review pass then moved `--delete-delay` to ⚠️ (the
differential. (Track 5b later found the name heuristic is rsync's own and moved
the row ❌ → ⚠️, leaving only the narrower delta size window; see entry 15.) The parity-review pass then moved `--delete-delay` to ⚠️ (the
plan-time `--max-delete` charge and non-recursive deferred removal differ from
rsync when a snapshotted entry fails removal). Differential-gate allowlist
entries `min_size`/`empty_dirs_recursive`/`dirs_plain` were removed. The
@@ -85,7 +87,8 @@
`-n --delete` now sends the same filter-excluded + size-pruned protected
prefixes and synchronized-directory scope as a real run (dry-run would-delete
matches rsync for source-derived protections; the destination-only exclude
residual and readdir ordering remain); `--delete-delay` now charges
residual was later closed by track 4a, readdir ordering remains);
`--delete-delay` now charges
`--max-delete` on actual removals and re-scans a queued directory at commit
to remove content created after the plan, with an independent deferred-list
cap (only partial-delete ordering remains); and `--info=name2` emits `NAME is
@@ -110,6 +113,85 @@
caveats, so the row stays ⚠️ and the matrix is unchanged at
**111 ✅ / 14 ⚠️ / 32 ❌ = 157**.
13. **Wire parity track 4a** on `feat/parity-2.28` (`PROTOCOL_VERSION` stays
`2.28.0`): the receiver now has a delete-time filter engine. The sender
compiles its root-level selection rules exactly as the scanner does
(`filter_base_build`) and streams them as one bounded, self-describing
config-frame block (action, sides, anchored, dir-only, negate, owner,
pattern; bounded rule count and pattern bytes, unknown action/sides is a
protocol error). The receiver reconstructs `protect_rules` and applies them
first-match-wins to each extraneous destination path in every delete timing
(the whole-tree commit walker, the `--delete-during`/`--delete-delay`
per-directory plans, and the `-n` would-delete enumeration), so a
`P *.log` rule protects a destination-only `extra.log` like rsync (with
`risk` cancelling); the sender-derived protected-prefix behavior is
preserved when no rules are sent and `--delete-excluded` semantics are
unchanged. Per-directory merge (`:`/`.`) receiver re-derivation remains the
residual. `TestFilterProtect` (real + dry-run) plus differential cases
`filter_protect`, `filter_protect_during`, `filter_protect_delay` added and
the `--filter=RULE` row moves ❌ → ✅: matrix now
**115 ✅ / 11 ⚠️ / 31 ❌ = 157**; unit tests, the three named integration
files, clang-format and cppcheck clean.
14. **Wire parity track 5a** on `feat/parity-2.28` (`PROTOCOL_VERSION` stays
`2.28.0` by project decision): the three basis-dir options now default to
rsync's metadata quick-check (equal size + equal mtime, or size alone under
`--size-only`; `-I` disables matching) instead of FastSync's historical
xxHash64 content equality, so a same-size/different-content basis is trusted
exactly as rsync trusts it. A new FastSync-only, long-only `--verify-basis`
flag restores the strict whole-file content equality; its bool is appended to
the basis block of the config frame (golden wire frame 882 → 886 bytes).
`--verify-basis` streams the confined basis descriptor to hash it, and a
basis hit is no longer capped at the 256 MiB whole-file payload bound:
`--copy-dest` streams the basis through a bounded buffer and `--link-dest`'s
copy fallback streams from the basis, so an over-limit hit materializes (a
basis MISS still falls back to the normal transfer and keeps its own bound).
A `--copy-dest` hit re-applies the SOURCE attributes (the sender transmits
the source metadata with the basis check frame), matching rsync's
"copy then fix attributes"; a `--link-dest` success keeps the shared inode's
attributes (writing through it would mutate the basis). Differential cases
`copy_dest` and `verify_basis` added; `test_basis_dir_size_only_content_residual`
converted to a passing parity assertion; `TestBasisDestDirs` updated for the
new default + `--verify-basis`; unit tests cover the quick-check/verify
decision and the same-size/different-content handshake. The
`--compare-dest`/`--copy-dest`/`--link-dest` rows move ❌ → ⚠️ (relative-DIR
resolution base and over-limit MISS refusal): matrix now
**116 ✅ / 13 ⚠️ / 28 ❌ = 157**.
15. **No-wire parity track 5b** on `feat/parity-2.28` (`PROTOCOL_VERSION` stays
`2.28.0` by project decision): `-y`/`--fuzzy` reclassified ❌ → ⚠️. A probe
against real rsync 3.4.1 (pinned `-B8192`, repeated-content 64 KiB corpus)
showed the name heuristic is already rsync's (`util1.c fuzzy_distance` /
`find_filename_suffix` + the exact size+mtime pass) and the output is always
byte-exact; the only residual is candidate ELIGIBILITY, because FastSync's
`delta_should_attempt` gate caps the size ratio at 10× and requires both
files ≥ 16 KiB while rsync will reuse a basis from 0.25× to 10000× and below
16 KiB. The choice is observable only as `--stats` bandwidth counters. Added
differential case `fuzzy_basis` (same-suffix sibling, one name edit,
identical content, block size pinned) asserting tree **and** normalized
`--stats` parity where the choices coincide, plus `TestFuzzy` pinning the
window boundary on both sides (>10× and <16 KiB siblings declined by
FastSync while rsync uses them, both trees byte-identical). Matrix now
**116 ✅ / 14 ⚠️ / 27 ❌ = 157**.
16. **Lockstep delete-default track 6** on `feat/parity-2.28` (`PROTOCOL_VERSION`
stays `2.28.0`): plain `--delete` now defaults to rsync's delete-during
(`--del`) timing, normalized on the client onto the existing `delete_during`
wire bool. The old late whole-tree commit is opt-in via `--delete-after` or
the FastSync-only long `--delete-commit` (identical `delete_after` timing).
`-d/--dirs` still falls back to the end commit, `--delay-updates` still
deletes before publication, and `--files-from`/`-R` scope is unchanged. The
`STATUS_DELETE_PLAN` frame gained a one-int `apply` flag so the per-run
config block (including `--delete-missing-args` exact paths) is always
transmitted, on a config-only carrier when the scope allows no directory
plan — fixing a latent bug with a file-only `--files-from` list. Differential
cases `delete`/`delete_commit`/`filter_protect_after` plus the extended
`test_delete_timing_parity.py` (plain `--delete` mid-abort removes reached
extras, `--delete-commit` defers) pass; full `-m "not setpriv"` suite,
clang-format and cppcheck clean. Matrix unchanged at
**116 ✅ / 14 ⚠️ / 27 ❌ = 157** (the `--delete`/`--delete-during` rows stay
⚠️ for the abort boundary; `--delete-after` stays ✅).
## Next steps
1. **Merge PR #284** (`dev` -> `main`) once reviewed (protected branch).
2. **Deferred security items** (documented, not implemented):
+10 -4
View File
@@ -113,9 +113,11 @@ matrix is classified as parity, caveat, or divergent in
(`-B1000`, `-essh`, `-MOPT`, `--opt=value`) are accepted, matching rsync.
- `-r`, `-b`, `-L`, and `-B` are parsed with the rsync short names.
- `--stats` prints the counters FastSync can observe plus the receiver-only
counters (`Matched data`, deleted files) reported over the wire; rsync's
per-type `Number of files` breakdown is not reproduced. `--progress` prints
rsync-style per-file blocks (without rsync's leading `./` line).
counters reported over the wire (`Matched data`, deleted files, and the
created/literal counters); `Number of files` and `Number of created files`
carry rsync's per-type breakdown. `--progress` prints rsync-style per-file
blocks including the leading `./` line, and (when progress is requested) a
paths-only pre-count supplies rsync's `to-chk` denominator.
- Codecs match rsync 3.4.1: `zstd`/`lz4`/`zlib`/`zlibx` compression and
`xxh128`/`xxh3`/`xxh64`/`md5`/`md4`/`sha1`/`none` checksums. `auto` honors
`RSYNC_COMPRESS_LIST`/`RSYNC_CHECKSUM_LIST` and otherwise follows rsync's
@@ -205,11 +207,13 @@ This produces `./build/client` and `./build/server`. `compile_commands.json` is
| `--compare-dest <dir>` | Extra comparison basis: unchanged files are not transferred (requires/implies `--incremental`) |
| `--copy-dest <dir>` | Like `--compare-dest`, but copies the unchanged file from DIR into the destination |
| `--link-dest <dir>` | Like `--copy-dest`, but hard-links the unchanged file from DIR (repeatable; earlier DIRs win) |
| `--delete` | Delete files on receiver not present in source (default timing: delete-after, i.e. only after the whole transfer succeeded). Scoped to the synchronized directories, so `--files-from` subsets are safe |
| `--verify-basis` | FastSync-only: require a basis hit (`--compare-dest`/`--copy-dest`/`--link-dest`) to match the source by whole-file digest instead of trusting the size+mtime quick-check (default matches rsync) |
| `--delete` | Delete files on receiver not present in source (default timing: delete-during, matching rsync, so destination space is freed progressively). Scoped to the synchronized directories, so `--files-from` subsets are safe |
| `--delete-before` | Delete extras before the transfer starts (implies `--delete`) |
| `--delete-during`, `--del` | Delete extras once the keep-set is known, before data is applied (implies `--delete`) |
| `--delete-delay` | Delete extras only after a successful transfer (implies `--delete`) |
| `--delete-after` | Explicit delete-after timing (implies `--delete`) |
| `--delete-commit` | FastSync-only: keep the pre-2.28 atomic timing — delete only after the whole transfer succeeded (identical timing to `--delete-after`) |
| `--delete-excluded` | Also delete filter-excluded destination mirrors (size-pruned mirrors stay protected) |
| `--max-delete <n>` | Delete at most n destination entries; the rest are skipped and the run exits 25 (partial), matching rsync |
| `--delay-updates` | Put updated files into place only at the end of the transfer (`--force` is honored at publication) |
@@ -567,6 +571,7 @@ remote SSH argv is already built injection-safe.
| `--compare-dest <dir>` | Extra comparison basis: unchanged files are not transferred (requires/implies `--incremental`). |
| `--copy-dest <dir>` | Like `--compare-dest`, but copies the unchanged file from DIR into the destination. |
| `--link-dest <dir>` | Like `--copy-dest`, but hard-links the unchanged file from DIR (repeatable; earlier DIRs win). |
| `--verify-basis` | FastSync-only: require a basis hit to match the source by whole-file digest instead of trusting the size+mtime quick-check (default matches rsync). |
| `--preallocate` | Allocate destination file space up front (fail-fast on a full disk). |
| `--append` | Resume a shorter destination by appending only its tail (prefix not verified; requires `--incremental`). |
| `--append-verify` | Like `--append`, but verifies the retained prefix checksum first (falls back to a full transfer on mismatch). |
@@ -574,6 +579,7 @@ remote SSH argv is already built injection-safe.
| `--delete-before` | Delete extras before the transfer starts (implies `--delete`). |
| `--delete-during`, `--del` | Delete extras once the keep-set manifest is known, before data is applied (implies `--delete`; early mode, same engine behaviour as `--delete-before`). |
| `--delete-delay` | Delete extras only after a successful transfer (implies `--delete`; commit mode, same behaviour as `--delete-after`). |
| `--delete-commit` | FastSync-only: atomic delete-after timing (only after the whole transfer succeeded). |
| `--delete-after` | Explicit delete-after timing: delete only after the transfer succeeded (implies `--delete`). |
| `--delete-excluded` | Also delete filter-excluded destination mirrors (size-pruned mirrors stay protected). |
| `--max-delete <n>` | Delete at most n destination entries; the rest are skipped and the run exits 25 (partial), matching rsync. |
+64 -35
View File
File diff suppressed because one or more lines are too long
+27 -1
View File
@@ -997,6 +997,12 @@ static const OptionEntry OPTION_TABLE[] = {
{"--delete-during", "--del", OPT_FLAG, offsetof(Config, delete_during)},
{"--delete-delay", NULL, OPT_FLAG, offsetof(Config, delete_delay)},
{"--delete-after", NULL, OPT_FLAG, offsetof(Config, delete_after)},
/* FastSync-only long spelling of the late whole-tree commit, which selects
the same timing as rsync's --delete-after in FastSync (the whole-tree
keep-set manifest is committed only after the entire transfer succeeded).
Plain --delete now defaults to delete-during, so this restores the old
FastSync behavior; it maps onto the same delete_after wire field. */
{"--delete-commit", NULL, OPT_FLAG, offsetof(Config, delete_after)},
{"--delete-excluded", NULL, OPT_FLAG, offsetof(Config, delete_excluded)},
{"--max-delete", NULL, OPT_SIGNED_INT, offsetof(Config, max_delete)},
{"--ignore-errors", NULL, OPT_FLAG, offsetof(Config, ignore_errors)},
@@ -1054,6 +1060,11 @@ static const OptionEntry OPTION_TABLE[] = {
* --remote-option is parsed. --trust-sender is a local receiver policy and
* never travels to the remote peer. */
{"--trust-sender", NULL, OPT_FLAG, offsetof(Config, trust_sender)},
/* FastSync-only (not an rsync option): require a basis-hit's content to
* match the source by whole-file digest instead of trusting rsync's
* size+mtime quick-check. Long-only; crosses the wire so the receiver
* performs the extra read/hash. */
{"--verify-basis", NULL, OPT_FLAG, offsetof(Config, verify_basis)},
};
/* Only boolean options with no required argument are safe to negate. */
@@ -1096,6 +1107,7 @@ static const NegatableOption NEGATABLE_OPTIONS[] = {
{"xattrs", "X", offsetof(Config, preserve_xattrs)},
{"acls", "A", offsetof(Config, preserve_acls)},
{"fake-super", NULL, offsetof(Config, fake_super)},
{"verify-basis", NULL, offsetof(Config, verify_basis)},
};
static bool opt_is(const char* arg, const char* name, const char* alias) {
@@ -1507,7 +1519,9 @@ static bool cli_handle_table_option(CliParseCtx* ctx) {
if (entry->offset == offsetof(Config, per_dir_filter) && config->per_dir_filter_count < INT_MAX)
config->per_dir_filter_count++;
/* A delete-timing flag selects when --delete removes extras, so it
implies --delete exactly like the rsync options do. */
implies --delete exactly like the rsync options do. --delete-commit (the
FastSync-only late-commit spelling) is mapped onto delete_after and so is
covered here too. */
if (entry->offset == offsetof(Config, delete_before) ||
entry->offset == offsetof(Config, delete_during) ||
entry->offset == offsetof(Config, delete_delay) ||
@@ -2561,6 +2575,18 @@ static bool cli_handle_outbuf_option(CliParseCtx* ctx) {
* -1 on error. */
static int cli_finalize_config(Config* config, bool verbose, bool no_delta, bool no_incremental) {
set_log_level(config->quiet ? LOG_LEVEL_ERROR : (verbose ? LOG_LEVEL_DEBUG : LOG_LEVEL_WARNING));
/* rsync's plain --delete defaults to delete-during (--del): each directory's
extras are removed as that directory is processed, so space is freed
progressively and a tight destination never has to hold the whole old+new
tree at once. The late whole-tree commit FastSync historically used is
still selected explicitly by --delete-after or by the FastSync-only long
spelling --delete-commit (an exact alias for --delete-after). Resolve the
default on the client, before validation and before the config crosses the
wire, so exactly one timing flag is ever set; an explicit timing (including
--delete-commit) always wins. */
if (config->use_delete && !config->delete_before && !config->delete_during &&
!config->delete_delay && !config->delete_after)
config->delete_during = true;
if (config->compress_choice) {
int algo = compression_algo_from_name(config->compress_choice);
if (algo >= 0) {
+18 -69
View File
@@ -1047,53 +1047,6 @@ static bool files_from_list_check(const Config* config, ArrayList* missing_dest,
return true;
}
/* Basis directories are honored by the receiver's per-file incremental check,
which (like every whole-file payload path in FastSync) is bounded by
MAX_RECEIVE_WHOLE_FILE_SIZE. rsync would apply basis dirs to files of any
size; FastSync cannot, so when basis dirs are requested this preflight scan
refuses the run up front with a clear diagnostic instead of letting the
receiver abort the whole transfer mid-stream with no client explanation.
Returns true when the tree can be transferred. */
static bool basis_oversize_preflight(const Config* config) {
PreparedScanner prepared;
if (!prepare_scanner(config, 0, &prepared))
return false;
DirectoryScanner* scanner =
directory_scanner_create_with_options(config->send_directory, &prepared.options);
if (!scanner) {
prepared_scanner_destroy(&prepared);
return false;
}
bool ok = true;
Chunk* chunk;
while ((chunk = directory_scanner_next(scanner)) != NULL) {
for (int i = 0; i < chunk->element_count; i++) {
File* f = chunk->items[i];
if (f == NULL || f->is_dir || f->data == NULL || f->data->size <= MAX_RECEIVE_WHOLE_FILE_SIZE)
continue;
char* escaped = output_escape(file_wire_path(f), config->eight_bit_output);
log_message(LOG_LEVEL_ERROR,
"%s is %llu bytes, larger than the %llu-byte whole-file transfer limit; "
"--compare-dest/--copy-dest/--link-dest cannot sync files above this limit",
escaped ? escaped : "<allocation failed>", (unsigned long long)f->data->size,
(unsigned long long)MAX_RECEIVE_WHOLE_FILE_SIZE);
free(escaped);
ok = false;
break;
}
chunk_destroy(chunk);
if (!ok)
break;
}
if (directory_scanner_failed(scanner) || directory_scanner_had_io_error(scanner))
ok = false;
/* The scanner borrows prepared.options' base_filters/hardlinks pointers, so
prepared must outlive the scanner. */
directory_scanner_destroy(scanner);
prepared_scanner_destroy(&prepared);
return ok;
}
/* Read the daemon's MOTD frame and, unless --no-motd, display it on stdout.
*
* The daemon sends the MOTD as the first thing after the config-frame STATUS_OK
@@ -1939,11 +1892,13 @@ static int incremental_check(Client* client, File* file, const Config* config,
return -1;
if (!send_n_data(client->file_descriptor, &mtime_nsec, sizeof(mtime_nsec)))
return -1;
/* With alternate basis directories the receiver must be able to verify the
* content of every candidate basis file, so the sender supplies its whole-file
* digest (computed with the negotiated --checksum-choice algorithm and
* --checksum-seed) for every file even when --checksum was not requested. */
if (config->checksum || config_has_basis(config)) {
/* The whole-file digest (negotiated --checksum-choice algorithm and
* --checksum-seed) is only needed when it drives a decision: --checksum's
* per-file quick check, or a --verify-basis content equality. Under the
* default metadata quick-check the receiver never reads it, so the sender
* skips the full-file read/hash exactly as rsync does for a plain
* --link-dest run. */
if (config->checksum || config->verify_basis) {
uint8_t digest[CHECKSUM_MAX_DIGEST_LEN];
size_t digest_len = 0;
if (!file_checksum(file, (ChecksumAlgo)config->checksum_algo, config->checksum_seed, digest,
@@ -1954,6 +1909,15 @@ static int incremental_check(Client* client, File* file, const Config* config,
!send_n_data(client->file_descriptor, digest, wire_len))
return -1;
}
/* Basis directories: the receiver materializes a hit from the basis without a
* data frame, so it would otherwise only have the basis inode's metadata.
* Transmit the SOURCE metadata with the check (rsync's copy-then-fix) so a
* --copy-dest hit / --link-dest copy fallback applies the source's
* attributes. Symmetric with incremental_check_receive_request. */
if (config_has_basis(config) && config->use_metadata) {
if (!metadata_send(client->file_descriptor, file->metadata))
return -1;
}
Status s;
if (!receive_status(client->file_descriptor, &s))
return -1;
@@ -2193,12 +2157,6 @@ static int send_dry_run_remote(Config* config) {
}
if (missing_args)
array_list_delete(missing_args);
/* Alternate basis dirs force the whole-file per-file check on the real
receiver; refuse an oversize source up front exactly as send_files does so
dry-run reports the same clear diagnostic instead of aborting mid-stream. */
if (config_has_basis(config) && !basis_oversize_preflight(config))
return 1;
/* A live session may follow, so arm graceful abort handling. */
client_set_abort_armed(true);
Client* client = connect_transfer_client(config);
@@ -3240,11 +3198,6 @@ int send_files(Config* config) {
array_list_delete(missing_args);
return 1;
}
if (config_has_basis(config) && !basis_oversize_preflight(config)) {
if (missing_args)
array_list_delete(missing_args);
return 1;
}
/* From here on a server session may be live, so Ctrl-C/SIGTERM should set the
abort flag (and be forwarded as STATUS_ABORT) instead of terminating. */
@@ -3334,7 +3287,8 @@ int send_files(Config* config) {
prepared.options.synced_dirs = synced_dirs;
}
}
/* The late-timing modes (plain --delete / --delete-after) build the manifest
/* The late-timing modes (--delete-after/--delete-commit and a plain --delete
that fell back from per-dir mode because of -d/--dirs) build the manifest
while streaming and send it after the last data frame. --delete-before
sends a whole-tree keep-set up front; --delete-during/--delete-delay build a
per-directory plan set up front (paths only) and stream the plans alongside
@@ -3682,11 +3636,6 @@ int send_files_multithreaded(Config** config_ptr) {
array_list_delete(missing_args);
return 1;
}
if (config_has_basis(config) && !basis_oversize_preflight(config)) {
if (missing_args)
array_list_delete(missing_args);
return 1;
}
/* Armed only once a session may go live (see send_files). */
client_set_abort_armed(true);
+9 -4
View File
@@ -62,8 +62,7 @@ void print_usage(void) {
printf(" NOTE: the FastSync batch format is NOT interoperable with rsync's batch\n");
printf(" files (different container format); do not mix the two tools.\n");
printf(" --delete Delete files on receiver not in source\n");
printf(" (default timing: delete only after the whole\n");
printf(" transfer has succeeded)\n");
printf(" (default timing: delete-during, like rsync --del)\n");
printf(" --delete-before Delete extras before the transfer starts\n");
printf(" (implies --delete)\n");
printf(" --delete-during Delete a directory's extras as that directory is\n");
@@ -72,7 +71,9 @@ void print_usage(void) {
printf(" --delete-delay Record the extras during the scan but remove them\n");
printf(" only after a successful transfer (implies --delete)\n");
printf(" --delete-after Delete only after the whole transfer succeeded\n");
printf(" (the default --delete timing; implies --delete)\n");
printf(" (implies --delete)\n");
printf(" --delete-commit FastSync-only: restore the late whole-tree commit\n");
printf(" (identical to --delete-after; implies --delete)\n");
printf(" --delete-excluded Also delete destination files that were excluded on\n");
printf(" the source (default protects them, matching rsync)\n");
printf(" --max-delete=NUM Delete at most NUM destination entries per run; if the\n");
@@ -93,7 +94,8 @@ void print_usage(void) {
printf(" -m, --prune-empty-dirs Do not create empty directories (a recursive transfer\n");
printf(" otherwise recreates them, like rsync)\n");
printf(" Note: each timing flag implies --delete. Combining a timing flag with\n");
printf(" --no-delete (in either order) is rejected as a config error.\n");
printf(" --no-delete (in either order) is rejected as a config error, as is more\n");
printf(" than one timing flag.\n");
printf(" --ignore-existing Skip files that already exist on receiver\n");
printf(" --delay-updates Put updated files into place only at the end of transfer\n");
printf(" --dirs, -d, --old-dirs, --old-d Transfer the named directory entries without\n");
@@ -138,6 +140,9 @@ void print_usage(void) {
printf(" into the destination instead of transferring its data\n");
printf(" --link-dest <dir> Like --copy-dest, but hard-links the unchanged file from DIR\n");
printf(" into the destination (repeatable; earlier DIRs win)\n");
printf(" --verify-basis FastSync-only: require a basis hit's content to match the\n");
printf(" source by whole-file digest instead of trusting rsync's\n");
printf(" size+mtime (or --size-only) quick-check\n");
printf(" --checksum-choice, --cc <alg> Whole-file checksum algorithm for --incremental/\n");
printf(" --checksum compares. Accepted: xxh128 (default), xxh3, xxh64\n");
printf(" (aka xxhash), md5, md4, sha1, or none. A two-name\n");
+6 -4
View File
@@ -305,14 +305,16 @@ int receiver_process(Config* config, int file_descriptor, const ReceiverSink* si
/* Runs the whole receive loop. The delete manifest may legitimately arrive
either FIRST (--delete-before / --delete-during: the sender transmits the
validated keep-set before any file data) or LAST (plain --delete /
--delete-after / --delete-delay: the manifest closes the data stream). In
validated keep-set before any file data) or LAST (--delete-after /
--delete-commit / --delete-delay: the manifest closes the data stream). In
the early modes the receiver deletes as soon as the manifest has been read
and acknowledges with STATUS_OK so the sender only starts streaming once the
deletion has committed (or failed); in the late modes the manifest is held
and the deletion is committed only after the terminal STATUS_FINISHED proves
the whole transfer succeeded. See receiver_process_pending() for how the -m
receiver defers that commit until its disk writer has drained. */
the whole transfer succeeded. A plain --delete defaults to the per-directory
delete-during plan mode (no manifest at all). See
receiver_process_pending() for how the -m receiver defers that commit until
its disk writer has drained. */
int receiver_process_pending(Config* config, int file_descriptor, const ReceiverSink* sink,
DeleteManifest** pending_manifest, DeletePlanSession** pending_plans) {
Status status;
+13 -11
View File
@@ -224,23 +224,31 @@ bool checksum_digest_file(ChecksumAlgo algo, uint64_t seed, const char* path, ui
if (fd < 0)
return false;
bool ok = checksum_digest_fd(algo, seed, fd, out, out_capacity, out_len);
close(fd);
return ok;
}
bool checksum_digest_fd(ChecksumAlgo algo, uint64_t seed, int fd, uint8_t* out, size_t out_capacity,
size_t* out_len) {
if (fd < 0 || !out || !out_len || out_capacity < CHECKSUM_MAX_DIGEST_LEN)
return false;
if (algo == CHECKSUM_ALGO_NONE) {
/* No checksum requested: nothing to read; an empty digest succeeds. */
close(fd);
*out_len = 0;
return true;
}
uint8_t buffer[64 * 1024];
bool ok = false;
lseek(fd, 0, SEEK_SET);
if (algo == CHECKSUM_ALGO_MD5 || algo == CHECKSUM_ALGO_SHA1) {
const EVP_MD* md = algo == CHECKSUM_ALGO_MD5 ? EVP_md5() : EVP_sha1();
EVP_MD_CTX* ctx = EVP_MD_CTX_new();
if (!ctx) {
close(fd);
if (!ctx)
return false;
}
unsigned int digest_len = 0;
if (EVP_DigestInit_ex(ctx, md, NULL) == 1) {
ok = true;
@@ -259,7 +267,6 @@ bool checksum_digest_file(ChecksumAlgo algo, uint64_t seed, const char* path, ui
ok = false;
}
EVP_MD_CTX_free(ctx);
close(fd);
return ok;
}
@@ -276,7 +283,6 @@ bool checksum_digest_file(ChecksumAlgo algo, uint64_t seed, const char* path, ui
md4_final(&ctx, out);
*out_len = 16;
}
close(fd);
return ok;
}
@@ -286,16 +292,13 @@ bool checksum_digest_file(ChecksumAlgo algo, uint64_t seed, const char* path, ui
XXH64_reset(&xxh64, seed);
} else if (algo == CHECKSUM_ALGO_XXH3 || algo == CHECKSUM_ALGO_XXH128) {
xxh3 = XXH3_createState();
if (!xxh3) {
close(fd);
if (!xxh3)
return false;
}
if (algo == CHECKSUM_ALGO_XXH3)
XXH3_64bits_reset_withSeed(xxh3, seed);
else
XXH3_128bits_reset_withSeed(xxh3, seed);
} else {
close(fd);
return false;
}
@@ -329,7 +332,6 @@ bool checksum_digest_file(ChecksumAlgo algo, uint64_t seed, const char* path, ui
}
if (xxh3)
XXH3_freeState(xxh3);
close(fd);
return ok;
}
+7
View File
@@ -51,6 +51,13 @@ bool checksum_digest(ChecksumAlgo algo, uint64_t seed, const void* data, size_t
bool checksum_digest_file(ChecksumAlgo algo, uint64_t seed, const char* path, uint8_t* out,
size_t out_capacity, size_t* out_len);
/* Descriptor form of the streaming digest: rewinds `fd` to the start and hashes
* to EOF without closing it. Used by the --verify-basis path to hash an
* already-open, root-confined basis descriptor. Same contract as
* checksum_digest_file. */
bool checksum_digest_fd(ChecksumAlgo algo, uint64_t seed, int fd, uint8_t* out, size_t out_capacity,
size_t* out_len);
/* Resolve a --checksum-choice string (case-insensitive) to an algorithm id.
* Accepts "xxh64"/"xxhash", "xxh3", "xxh128", "md5", "md4", "sha1", "none".
* "auto" is not an algorithm here; the caller resolves it to the negotiated
+129 -2
View File
@@ -791,6 +791,8 @@ void config_delete(Config* config) {
if (config->filters) {
array_list_delete(config->filters);
}
filter_rule_list_free(config->protect_rules);
config->protect_rules = NULL;
/* A --delay-updates staging tree is transient receiver state: remove any
leftovers on every exit path (success already emptied it). */
if (config->delay_context)
@@ -1026,6 +1028,124 @@ static bool receive_basis_entries(int fd, Config* c, ConfigStringBudget* budget)
return true;
}
/* Receiver-side delete-protection rules (protocol 2.28.0). The sender compiles
* its command-line selection rules exactly as the scanner does and streams the
* result as one bounded, self-describing block (count + per-rule records); the
* receiver reconstructs a FilterRuleList for the --delete extras walk. owner
* and pattern are charged through the shared ConfigStringBudget and the block
* additionally enforces MAX_FILTER_RULES / MAX_FILTER_BYTES. */
static bool send_protect_entries(int fd, const Config* c) {
int count = c->filters ? c->filters->size : 0;
const char** texts = NULL;
if (count > 0) {
texts = malloc((size_t)count * sizeof(char*));
if (!texts)
return false;
for (int i = 0; i < count; i++)
texts[i] = (const char*)c->filters->items[i];
}
char err[160];
FilterRuleList* rules =
filter_base_build(texts, count, c->cvs_exclude, c->delete_excluded, err, sizeof(err));
free(texts);
if (!rules) {
log_message(LOG_LEVEL_ERROR, "invalid filter rule: %s", err);
return false;
}
bool ok = send_int(fd, rules->count);
for (int i = 0; ok && i < rules->count; i++) {
const FilterRule* r = rules->items[i];
/* Mirror the receiver's limit so the peer never receives a rule it will
reject as a protocol error. */
if (r->pattern && strlen(r->pattern) > MAX_PROTECT_PATTERN_LEN) {
log_message(LOG_LEVEL_ERROR, "filter pattern exceeds %d bytes", MAX_PROTECT_PATTERN_LEN);
filter_rule_list_free(rules);
return false;
}
ok = send_int(fd, (int)r->action) && send_int(fd, (int)r->sides) &&
send_int(fd, r->anchored ? 1 : 0) && send_int(fd, r->dir_only ? 1 : 0) &&
send_int(fd, r->negate ? 1 : 0) && send_str(fd, r->owner ? r->owner : "") &&
send_str(fd, r->pattern ? r->pattern : "");
}
filter_rule_list_free(rules);
return ok;
}
static bool receive_protect_entries(int fd, Config* c, ConfigStringBudget* budget) {
int count;
if (!receive_int(fd, &count))
return false;
if (count < 0 || count > MAX_FILTER_RULES)
return false;
if (count == 0)
return true;
FilterRuleList* list = filter_rule_list_create();
if (!list)
return false;
size_t pattern_bytes = 0;
for (int i = 0; i < count; i++) {
int action;
int sides;
bool anchored;
bool dir_only;
bool negate;
if (!receive_int(fd, &action) ||
(action != FILTER_ACTION_EXCLUDE && action != FILTER_ACTION_INCLUDE) ||
!receive_int(fd, &sides) || sides < (int)FILTER_SIDE_SENDER ||
sides > (int)(FILTER_SIDE_SENDER | FILTER_SIDE_RECEIVER) ||
!receive_wire_bool(fd, &anchored) || !receive_wire_bool(fd, &dir_only) ||
!receive_wire_bool(fd, &negate))
goto fail;
char* owner = config_receive_str(fd, budget);
if (!owner)
goto fail;
char* pattern = config_receive_str(fd, budget);
if (!pattern || pattern[0] == '\0') {
free(owner);
free(pattern);
goto fail;
}
/* A pattern too long to be evaluated by glob_match against a PATH_MAX path
would silently fail to match and leave a protect rule inert (fail-open:
the entry is then deleted). Reject it up front as a protocol error
rather than accept a rule that can never shield anything. */
if (strlen(pattern) > MAX_PROTECT_PATTERN_LEN) {
free(owner);
free(pattern);
goto fail;
}
size_t bytes = strlen(owner) + strlen(pattern);
if (bytes > MAX_FILTER_BYTES - pattern_bytes) {
free(owner);
free(pattern);
goto fail;
}
pattern_bytes += bytes;
FilterRule* rule = calloc(1, sizeof(FilterRule));
if (!rule) {
free(owner);
free(pattern);
goto fail;
}
rule->action = (FilterAction)action;
rule->sides = (unsigned)sides;
rule->anchored = anchored;
rule->dir_only = dir_only;
rule->negate = negate;
rule->owner = owner;
rule->pattern = pattern;
if (!filter_rule_list_add(list, rule)) {
filter_rule_free(rule);
goto fail;
}
}
c->protect_rules = list;
return true;
fail:
filter_rule_list_free(list);
return false;
}
static bool send_identity_entries(int fd, const IdentityMap* map, int count) {
for (int i = 0; i < count; i++) {
if (!send_int(fd, map[i].from) || !send_int(fd, map[i].from_hi) || !send_int(fd, map[i].to) ||
@@ -1153,6 +1273,9 @@ fail:
#define CONFIG_RECV_BLOCK_IDMAP(name) \
receive_identity_entries(fd, budget, c->name##_count, &c->name)
#define CONFIG_SEND_BLOCK_PROTECT_RULES(name) send_protect_entries(fd, c)
#define CONFIG_RECV_BLOCK_PROTECT_RULES(name) receive_protect_entries(fd, c, budget)
/* One table entry, applied in sequence. XSEND/XRECV are statement macros so
* consecutive entries read as a plain sequence of assignments. */
#define XSEND(name, ctype, def, kind) ok = ok && (CONFIG_SEND_##kind(name));
@@ -1190,6 +1313,7 @@ CONFIG_DEFINE_SEND(send_privilege_options, CONFIG_WIRE_PRIVILEGE_FIELDS)
CONFIG_DEFINE_SEND(send_copy_as_options, CONFIG_WIRE_COPY_AS_FIELDS)
CONFIG_DEFINE_SEND(send_output_options, CONFIG_WIRE_OUTPUT_FIELDS)
CONFIG_DEFINE_SEND(send_codec_options, CONFIG_WIRE_CODEC_FIELDS)
CONFIG_DEFINE_SEND(send_protect_options, CONFIG_WIRE_PROTECT_FIELDS)
CONFIG_DEFINE_RECV(receive_core_fields, CONFIG_WIRE_CORE_FIELDS)
CONFIG_DEFINE_RECV(receive_delta_fields, CONFIG_WIRE_DELTA_FIELDS)
@@ -1210,6 +1334,7 @@ CONFIG_DEFINE_RECV(receive_privilege_options, CONFIG_WIRE_PRIVILEGE_FIELDS)
CONFIG_DEFINE_RECV(receive_copy_as_options, CONFIG_WIRE_COPY_AS_FIELDS)
CONFIG_DEFINE_RECV(receive_output_options, CONFIG_WIRE_OUTPUT_FIELDS)
CONFIG_DEFINE_RECV(receive_codec_options, CONFIG_WIRE_CODEC_FIELDS)
CONFIG_DEFINE_RECV(receive_protect_options, CONFIG_WIRE_PROTECT_FIELDS)
#undef XSEND
#undef XRECV
@@ -1328,7 +1453,8 @@ bool config_send_wire_block(int file_descriptor, const Config* config) {
send_privilege_options(file_descriptor, config) &&
send_copy_as_options(file_descriptor, config) &&
send_output_options(file_descriptor, config) &&
send_codec_options(file_descriptor, config);
send_codec_options(file_descriptor, config) &&
send_protect_options(file_descriptor, config);
}
bool config_send(int file_descriptor, const Config* config) {
@@ -1400,7 +1526,8 @@ Config* config_receive_with_validate(int file_descriptor, ConfigValidateFunc val
!receive_privilege_options(file_descriptor, config, &budget) ||
!receive_copy_as_options(file_descriptor, config, &budget) ||
!receive_output_options(file_descriptor, config, &budget) ||
!receive_codec_options(file_descriptor, config, &budget))
!receive_codec_options(file_descriptor, config, &budget) ||
!receive_protect_options(file_descriptor, config, &budget))
goto error;
/* Validate/normalize the negotiated codec. compress_choice is the human
* spelling (NULL or "" when -z was not given); compression_algo is the
+60 -8
View File
@@ -4,6 +4,7 @@
#include "array_list.h"
#include "checksum.h"
#include "compression.h"
#include "filter.h"
#include <stdbool.h>
#include <stdint.h>
#include <stdio.h>
@@ -197,9 +198,18 @@ typedef enum SuperMode { SUPER_MODE_AUTO = 0, SUPER_MODE_ON = 1, SUPER_MODE_OFF
X(skip_compress_count, int, 0, INT_SKIPCOUNT) \
X(skip_compress_suffixes, char**, NULL, BLOCK_SKIP_SUFFIXES)
/* FastSync-only --verify-basis (protocol 2.28.0, no version bump by project
* decision): restores the stricter content equality on a basis hit. By
* default a basis hit is accepted on rsync's metadata quick-check alone (equal
* size plus equal mtime, or size alone under --size-only); with this flag the
* receiver ALSO requires the basis bytes' whole-file digest (the negotiated
* --checksum-choice algorithm) to equal the sender's, exactly FastSync's
* historical behavior. It is a receiver policy and crosses the wire so the
* receiver knows whether to read and hash the basis content. */
#define CONFIG_WIRE_BASIS_FIELDS(X) \
X(basis_count, int, 0, INT_BASISCOUNT) \
X(basis_dirs, BasisDest*, NULL, BLOCK_BASIS)
X(basis_dirs, BasisDest*, NULL, BLOCK_BASIS) \
X(verify_basis, bool, false, BOOL)
#define CONFIG_WIRE_FUZZY_FIELDS(X) X(fuzzy, bool, false, BOOL)
@@ -293,6 +303,20 @@ typedef enum SuperMode { SUPER_MODE_AUTO = 0, SUPER_MODE_ON = 1, SUPER_MODE_OFF
#define CONFIG_WIRE_CODEC_FIELDS(X) \
X(compression_algo, int, COMPRESSION_ALGO_ZSTD, INT_COMPRESSION_ALGO)
/* Receiver-side delete-protection filter rules (protocol 2.28.0). The sender
* compiles its root-level selection rules exactly as the scanner does
* (filter_base_build over --filter/-f/--exclude/--include/-C) and streams them
* as one self-describing, bounded block (count followed by per-rule records).
* The receiver reconstructs `protect_rules` and evaluates them against
* DESTINATION-ONLY entries during the --delete extras walk, so a
* `protect`/`P` rule protects an extra that never appeared on the sender
* (rsync re-derives deletion protection from the filter list; FastSync
* historically derived it only from the source scan). `protect_rules` is NULL
* on the sender and is owned/freed by the receiver Config. Bounded by
* MAX_FILTER_RULES and MAX_FILTER_BYTES; an unknown action/sides is a protocol
* error. */
#define CONFIG_WIRE_PROTECT_FIELDS(X) X(protect_rules, FilterRuleList*, NULL, BLOCK_PROTECT_RULES)
/* All serialized fields, in exact wire order. Concatenating the per-segment
* lists here is what keeps the declaration order = the wire order. */
#define CONFIG_WIRE_FIELDS(X) \
@@ -315,7 +339,8 @@ typedef enum SuperMode { SUPER_MODE_AUTO = 0, SUPER_MODE_ON = 1, SUPER_MODE_OFF
CONFIG_WIRE_PRIVILEGE_FIELDS(X) \
CONFIG_WIRE_COPY_AS_FIELDS(X) \
CONFIG_WIRE_OUTPUT_FIELDS(X) \
CONFIG_WIRE_CODEC_FIELDS(X)
CONFIG_WIRE_CODEC_FIELDS(X) \
CONFIG_WIRE_PROTECT_FIELDS(X)
typedef struct Config {
/* -j/--threads=N: number of parallel scanner worker threads for the -m
@@ -617,10 +642,14 @@ typedef struct Config {
source directory is streamed in directory order, and the receiver removes
each directory's extras when its plan arrives (during) or snapshots them
and removes them only after a successful transfer (delay). delete_after
(and plain --delete) keep the whole-tree commit mode: extras are removed
from a fresh end-of-transfer destination scan only after the whole transfer
succeeded. See config_delete_timing_early()/config_delete_timing_per_dir()
below. */
keeps the whole-tree commit mode: extras are removed from a fresh
end-of-transfer destination scan only after the whole transfer succeeded.
A plain --delete with no explicit timing flag defaults to delete_during on
the client (cli_finalize_config), matching rsync's --del default; the old
late-commit behavior is selected explicitly by --delete-after or the
FastSync-only long spelling --delete-commit (an exact alias for
--delete-after, mapped onto the same wire field). See
config_delete_timing_early()/config_delete_timing_per_dir() below. */
/* partial_dir */
// PR #174: Partial transfer resumption
/* suffix */
@@ -1025,11 +1054,31 @@ typedef struct Config {
* same-version handshake (config_receive rejects a mismatched version before
* parsing anything else) keeps mixed deployments from ever reaching that
* state. */
/* (9) Receiver-side delete protection (still protocol 2.28.0): the config frame
* gains one trailing self-describing block carrying the sender's compiled base
* filter rules so the receiver can protect DESTINATION-ONLY entries from
* --delete with `protect`/`risk` rules (rsync parity). The block appends after
* compression_algo; see CONFIG_WIRE_PROTECT_FIELDS. */
#define PROTOCOL_VERSION "2.28.0"
#define DEFAULT_CHUNK_SIZE (10 * 1024 * 1024)
/* Upper bound on total basis-dir entries (rsync caps --link-dest at 20). */
#define MAX_BASIS_DIRS 64
/* Bounds on the received receiver-side delete-protection rule block. The rule
* count and the aggregate pattern+owner bytes are each capped so a hostile
* peer cannot pin unbounded pre-auth memory; both are validated strictly on
* receive (alongside the per-string ConfigStringBudget). */
/* A peer may supply protect rules; cap the list so a crafted config cannot make
* the receiver's delete walk evaluate an unbounded number of glob patterns per
* destination entry (glob_match is O(pattern x path)). 1024 is far above any
* legitimate selection. */
#define MAX_FILTER_RULES 1024
#define MAX_FILTER_BYTES (256 * 1024)
/* glob_match's DP is capped at 64 Mi work units; a pattern longer than this
* could exceed the cap against a PATH_MAX path and silently stop matching,
* leaving a protect rule inert. Reject such a rule at receive time. */
#define MAX_PROTECT_PATTERN_LEN 8192
/* Upper bound on the number of --skip-compress suffixes accepted from the wire.
* Each suffix is an independent wire string (up to MAX_STRING_SIZE = 64 KiB), so
* without this a hostile pre-auth client could otherwise retain
@@ -1131,8 +1180,11 @@ bool config_delete_timing_early(const Config* config);
* commits them only after a fully-successful transfer (delay). */
bool config_delete_timing_per_dir(const Config* config);
/* Delete-timing sanity: with deletion enabled at most one timing flag may be
* set (none = the default delete-after commit timing); without deletion no
* timing flag may be set (each timing flag implies --delete). */
* set; without deletion no timing flag may be set (each timing flag implies
* --delete). A plain --delete is normalized to delete_during by
* cli_finalize_config on the client, so a transmitted use_delete config always
* carries exactly one timing; the zero-timing case remains valid only for a
* config that has not been through the CLI. */
bool config_has_valid_delete_timing(const Config* config);
/* Single source of truth for the cross-field ("combination") invariants a
+59 -5
View File
@@ -331,6 +331,8 @@ static int send_plan_node(int fd, DeletePlanSender* sender, PlanNode* node) {
return -1;
sender->config_sent = true;
}
if (!send_int(fd, 1)) /* apply = true */
return -1;
if (!send_wire_str(fd, node->dir))
return -1;
if (send_str_section(fd, node->dirs) != 0 || send_str_section(fd, node->files) != 0)
@@ -339,6 +341,29 @@ static int send_plan_node(int fd, DeletePlanSender* sender, PlanNode* node) {
return 0;
}
/* Transmit the one-shot per-run config block (protected prefixes, size-pruned
* mirrors, --delete-missing-args exact paths) on its own carrier frame, with
* apply=false so the receiver consumes the config but walks nothing. This is
* how the config still reaches the receiver when the scope allows no directory
* plan at all (a --files-from list of bare files synchronizes no directory):
* without it, the missing-args exact deletions would be lost. Idempotent. */
static int send_config_only(int fd, DeletePlanSender* sender) {
if (!sender || sender->config_sent)
return 0;
if (!send_status(fd, STATUS_DELETE_PLAN) || !send_int(fd, 1))
return -1;
if (send_str_section(fd, sender->protected_prefixes) != 0 ||
send_str_section(fd, sender->size_skipped) != 0 ||
send_str_section(fd, sender->missing_args) != 0)
return -1;
sender->config_sent = true;
if (!send_int(fd, 0)) /* apply = false */
return -1;
if (!send_wire_str(fd, ".") || !send_int(fd, 0) || !send_int(fd, 0))
return -1;
return 0;
}
static int send_prefix_plan(int fd, DeletePlanSender* sender, const char* dir) {
PlanNode* node = plan_find(sender, dir);
if (!node || node->sent)
@@ -354,6 +379,10 @@ int delete_plan_send_root(int fd, DeletePlanSender* sender) {
const char* root = sender->walk_root ? sender->walk_root : ".";
if (!plan_ensure(sender, root))
return -1;
/* Put the config block on the wire first, on its own carrier frame, so the
receiver always sees it even when the scope permits no directory plan. */
if (send_config_only(fd, sender) != 0)
return -1;
return send_prefix_plan(fd, sender, root);
}
@@ -546,12 +575,18 @@ static int open_plan_dir(const Config* config, const char* dir) {
typedef struct PlanSkips {
DeleteSkipEntry* entries;
int count;
/* Receiver-side delete-protection rules received on the config frame (NULL
when the sender sent none). Evaluated per extra so a protect/risk rule is
honored under --delete-during/--delete-delay exactly like the whole-tree
commit walker. */
const FilterRuleList* protect_rules;
} PlanSkips;
static bool build_plan_skips(const Config* config, const DeletePlanSession* session,
PlanSkips* out) {
out->entries = NULL;
out->count = 0;
out->protect_rules = config->protect_rules;
int count = (config->delay_updates ? 1 : 0) + config->basis_count +
session->protected_prefixes->size + session->size_skipped->size;
if (count == 0)
@@ -733,6 +768,10 @@ static bool process_children(int dirfd, const char* dir_rel, const ArrayList* ke
bool is_dir = S_ISDIR(st.st_mode);
bool in_keep_dirs = is_dir && list_contains_str(keep_dirs, entry->d_name);
bool in_keep_files = !is_dir && list_contains_str(keep_files, entry->d_name);
bool rule_protected =
skips->protect_rules &&
filter_rules_apply_side(skips->protect_rules, child_rel, entry->d_name, is_dir,
FILTER_SIDE_RECEIVER) == FILTER_ACTION_PROTECT;
if (in_keep_dirs) {
local_survives = true;
} else if (keep_dirs && !is_dir && list_contains_str(keep_dirs, entry->d_name)) {
@@ -750,11 +789,18 @@ static bool process_children(int dirfd, const char* dir_rel, const ArrayList* ke
else if (!removed)
local_survives = true;
} else if (is_dir) {
bool removed = false;
if (!process_extra_dir(dirfd, entry->d_name, child_rel, force_now, skips, session, &removed))
operation_ok = false;
else if (!removed)
if (rule_protected) {
local_survives = true;
} else {
bool removed = false;
if (!process_extra_dir(dirfd, entry->d_name, child_rel, force_now, skips, session,
&removed))
operation_ok = false;
else if (!removed)
local_survives = true;
}
} else if (rule_protected) {
local_survives = true;
} else {
if (!process_extra_file(dirfd, entry->d_name, child_rel, force_now, session))
operation_ok = false;
@@ -833,6 +879,14 @@ int delete_plan_session_receive(DeletePlanSession* session, const Config* config
}
session->config_seen = true;
}
/* apply=false is the config-only carrier frame: the receiver consumes the
config (and the missing-args exact deletions) but must not walk any
directory. Every real plan carries apply=true. */
int apply;
if (!receive_int(fd, &apply) || (apply != 0 && apply != 1)) {
send_status(fd, STATUS_ERROR);
return -1;
}
char* dir = receive_wire_str(fd);
ArrayList* dirs = array_list_create(free);
ArrayList* files = array_list_create(free);
@@ -850,7 +904,7 @@ int delete_plan_session_receive(DeletePlanSession* session, const Config* config
if (!session->dry_run && enabled) {
if (!session->defer && !apply_missing(session, config))
ok = false;
if (ok && !apply_plan_dir(session, config, dir, dirs, files))
if (ok && apply && !apply_plan_dir(session, config, dir, dirs, files))
ok = false;
}
free(dir);
+5 -2
View File
@@ -48,11 +48,14 @@ void delete_plan_sender_finalize(DeletePlanSender* sender, const ArrayList* sync
Directory keep entries do not count, so an I/O error that hid every file
still refuses to delete. */
bool delete_plan_sender_empty(const DeletePlanSender* sender);
/* Attach the global config sections advertised on the first plan frame. */
/* Attach the global config sections advertised on the first plan frame. The
* block is always transmitted by delete_plan_send_root(), on a config-only
* carrier frame when the scope allows no directory plan. */
void delete_plan_sender_set_config(DeletePlanSender* sender, const ArrayList* protected_prefixes,
const ArrayList* size_skipped, const ArrayList* missing_args);
/* Send the root plan (even before any data, so root extras are handled like
* rsync's first generator directory). Returns -1 on I/O error. */
* rsync's first generator directory), after transmitting the per-run config
* block on its own carrier frame. Returns -1 on I/O error. */
int delete_plan_send_root(int fd, DeletePlanSender* sender);
/* Send the plans for every ancestor of `path` (root-first) and, when is_dir,
* for `path` itself; already-sent plans are skipped. */
+209 -1
View File
@@ -15,6 +15,7 @@
#include <unistd.h>
#include "data.h"
#include "checksum.h"
#include "delta.h"
#include "file.h"
#include "file_store.h"
@@ -24,6 +25,13 @@
#include "utils.h"
#include "protocol.h"
#include "xattr.h"
#include <fcntl.h>
#include <unistd.h>
/* Files larger than this are not loaded whole for transfer (the sender streams
* them); a whole-file digest is computed from the path instead. Kept in sync
* with the sender's streaming threshold. */
#define STREAM_THRESHOLD (64ULL * 1024 * 1024)
static bool write_all(int fd, const void* data, unsigned long long size) {
const unsigned char* p = data;
@@ -39,6 +47,31 @@ static bool write_all(int fd, const void* data, unsigned long long size) {
return true;
}
/* Streaming copy of an open source descriptor into the just-created destination
`fd` (already at offset 0). Used by the --copy-dest basis install so a basis
larger than any in-memory whole-file bound still materializes without
buffering the entire file. `expected_size` is the caller-verified basis
size; the copy must produce exactly that many bytes (a short source is a hard
error, never a silently truncated destination). The final ftruncate drops
any residual tail a raced-in longer source might have left. */
static bool copy_fd_all(int dst_fd, int src_fd, unsigned long long expected_size) {
unsigned char buf[1 << 20];
unsigned long long done = 0;
while (done < expected_size) {
unsigned long long remaining = expected_size - done;
size_t want = remaining < sizeof(buf) ? (size_t)remaining : sizeof(buf);
ssize_t n = read(src_fd, buf, want);
if (n < 0 && errno == EINTR)
continue;
if (n <= 0)
return false;
if (!write_all(dst_fd, buf, (unsigned long long)n))
return false;
done += (unsigned long long)n;
}
return ftruncate(dst_fd, (off_t)expected_size) == 0;
}
/* Preallocate `size` bytes on `fd` before any data is written (--preallocate).
* fallocate(2) reserves real disk blocks, so an out-of-space condition
* (ENOSPC/EDQUOT) surfaces up front instead of partway through a transfer;
@@ -140,6 +173,12 @@ bool file_checksum(File* file, ChecksumAlgo algo, uint64_t seed, uint8_t* out, s
if (file->data->size == 0) {
return checksum_digest(algo, seed, "", 0, out, out_capacity, out_len);
}
/* A streamed source (data not loaded) may exceed any in-memory whole-file
bound; hash it from the file path in bounded buffers instead of forcing a
full load. This is the same digest the receiver recomputes on the basis. */
if (!file->data->data && file->path && file->data->size > STREAM_THRESHOLD &&
checksum_digest_file(algo, seed, file->path, out, out_capacity, out_len))
return true;
if (!file->data->data && !file_load_data(file))
return false;
return checksum_digest(algo, seed, file->data->data, file->data->size, out, out_capacity,
@@ -176,6 +215,7 @@ File* file_create(const char* path) {
file->is_dir = false;
file->dir_time_only = false;
file->basis_link = NULL;
file->basis_copy = NULL;
file->link_group = 0;
file->link_first = false;
file->hardlink_target = NULL;
@@ -205,6 +245,8 @@ void file_destroy(void* item) {
file->send_path = NULL;
free(file->basis_link);
file->basis_link = NULL;
free(file->basis_copy);
file->basis_copy = NULL;
free(file->hardlink_target);
file->hardlink_target = NULL;
free(file->symlink_target);
@@ -1512,6 +1554,161 @@ bool file_to_disk_secure_attrs_counted(const char* path, const void* data,
* basis). Likewise `xattrs`/`fake_super` are applied only on the copy
* fallback, so a fallback copy preserves the per-file attributes instead of
* silently dropping them. */
/* Streaming --copy-dest basis install: atomically materialize `path` from the
* bytes of `basis_path` without holding the file in memory, so a basis larger
* than any whole-file bound still works. Mirrors the ordinary secure store
* path (confined parent walk, temp + rename, --update/--ignore-existing/
* --preallocate/--temp-dir) but sources the data from the basis descriptor
* rather than a caller buffer, and applies the SOURCE metadata (rsync copies
* then fixes attributes). A hard-link install that falls back to a byte copy
* also routes through here when the caller supplies the basis path. */
static bool file_copy_basis_stream_impl(const char* path, const char* basis_path,
unsigned long long expected_size, bool preallocate,
const FileMetadata* metadata, FileAttrPolicy policy,
bool update, bool no_replace, bool use_fsync,
const FileXattrList* xattrs, bool fake_super,
const char* temp_dir, unsigned* dirs_created,
const char* count_floor) {
if (!path || !basis_path)
return false;
char* leaf = NULL;
int dirfd = file_open_secure_parent_counted(path, &leaf, true, dirs_created, count_floor);
if (dirfd < 0)
return false;
char* basis_leaf = NULL;
int basis_dirfd = file_open_secure_parent(basis_path, &basis_leaf, false);
int src_fd = -1;
if (basis_dirfd >= 0 && basis_leaf != NULL) {
/* O_NONBLOCK rejects a raced-in FIFO without blocking; the S_ISREG gate
below is the real type check. */
src_fd = openat(basis_dirfd, basis_leaf, O_RDONLY | O_CLOEXEC | O_NOFOLLOW | O_NONBLOCK);
struct stat src_st;
if (src_fd >= 0 && (fstat(src_fd, &src_st) != 0 || !S_ISREG(src_st.st_mode))) {
close(src_fd);
src_fd = -1;
}
}
if (basis_dirfd >= 0)
close(basis_dirfd);
free(basis_leaf);
if (src_fd < 0) {
close(dirfd);
free(leaf);
return false;
}
struct stat destination_stat;
bool destination_is_regular = fstatat(dirfd, leaf, &destination_stat, AT_SYMLINK_NOFOLLOW) == 0 &&
S_ISREG(destination_stat.st_mode);
if (update && metadata && destination_is_regular && stat_is_newer(&destination_stat, metadata)) {
close(src_fd);
close(dirfd);
free(leaf);
return true;
}
if (no_replace && file_path_exists_secure(path)) {
close(src_fd);
close(dirfd);
free(leaf);
return true;
}
int scratch_dirfd = -1;
if (temp_dir) {
scratch_dirfd = file_open_temp_dir(temp_dir);
if (scratch_dirfd < 0) {
int saved_errno = errno;
log_message(LOG_LEVEL_ERROR,
"--temp-dir '%s' could not be opened (rsync requires it to already exist): %s",
temp_dir, strerror(saved_errno));
close(src_fd);
close(dirfd);
free(leaf);
return false;
}
}
int target_dirfd = scratch_dirfd >= 0 ? scratch_dirfd : dirfd;
int tmp_size = snprintf(NULL, 0, ".%s.tmp.%ld.%llu", leaf, (long)getpid(), ~0ULL);
char* tmp = NULL;
bool ok = false;
if (tmp_size >= 0)
tmp = malloc((size_t)tmp_size + 1);
if (tmp) {
for (unsigned int i = 0; i < 100 && !ok; ++i) {
if (scratch_dirfd >= 0)
snprintf(tmp, (size_t)tmp_size + 1, ".%s.tmp.%ld.%llu", leaf, (long)getpid(),
next_temp_sequence());
else
snprintf(tmp, (size_t)tmp_size + 1, ".%s.tmp.%ld.%u", leaf, (long)getpid(), i);
int fd =
openat(target_dirfd, tmp, O_WRONLY | O_CREAT | O_EXCL | O_CLOEXEC | O_NOFOLLOW, 0600);
if (fd < 0) {
if (errno != EEXIST)
break;
continue;
}
bool wrote = true;
if (preallocate && expected_size > 0 && preallocate_fd(fd, expected_size) != 0)
wrote = false;
if (wrote)
wrote = copy_fd_all(fd, src_fd, expected_size);
if (wrote && metadata) {
if (!policy.perms &&
fchmod(fd, file_mode_base(metadata, destination_is_regular,
destination_is_regular ? destination_stat.st_mode & 0777
: 0)) != 0)
wrote = false;
if (wrote)
wrote = file_restore_metadata_fd(fd, metadata, policy);
} else if (wrote && fchmod(fd, S_IRUSR | S_IWUSR | S_IRGRP | S_IROTH) != 0) {
wrote = false;
}
if (wrote)
restore_extra_fd(fd, metadata, xattrs, fake_super, policy);
if (wrote && use_fsync)
wrote = fsync(fd) == 0;
if (close(fd) != 0)
wrote = false;
if (wrote && renameat(target_dirfd, tmp, dirfd, leaf) != 0)
wrote = false;
if (!wrote)
unlinkat(target_dirfd, tmp, 0);
ok = wrote;
}
free(tmp);
}
if (!ok && scratch_dirfd >= 0) {
/* Retry once with no scratch dir (rsync's EXDEV fallback). */
close(scratch_dirfd);
close(src_fd);
close(dirfd);
free(leaf);
return file_copy_basis_stream_impl(path, basis_path, expected_size, preallocate, metadata,
policy, update, no_replace, use_fsync, xattrs, fake_super,
NULL, dirs_created, count_floor);
}
if (scratch_dirfd >= 0)
close(scratch_dirfd);
close(src_fd);
close(dirfd);
free(leaf);
return ok;
}
/* --copy-dest basis install (streaming). Applies the source metadata and the
per-file xattrs / --fake-super record. */
bool file_copy_basis_stream_attrs(const char* path, const char* basis_path,
unsigned long long expected_size, bool preallocate,
const FileMetadata* metadata, FileAttrPolicy policy, bool update,
bool use_fsync, const FileXattrList* xattrs, bool fake_super,
const char* temp_dir) {
return file_copy_basis_stream_impl(path, basis_path, expected_size, preallocate, metadata, policy,
update, false, use_fsync, xattrs, fake_super, temp_dir, NULL,
NULL);
}
static bool file_to_disk_secure_link_impl(const char* path, const char* basis_path,
const void* data, unsigned long long data_size,
bool preallocate, const FileMetadata* metadata,
@@ -1521,6 +1718,10 @@ static bool file_to_disk_secure_link_impl(const char* path, const char* basis_pa
const char* count_floor) {
if (!path || !basis_path)
return false;
/* The caller-supplied buffer is no longer used: the copy fallback streams
from the basis path (which may hold an over-limit file). Kept in the
signature for the existing API. */
(void)data;
char* leaf = NULL;
int dirfd = file_open_secure_parent_counted(path, &leaf, true, dirs_created, count_floor);
if (dirfd < 0)
@@ -1604,7 +1805,14 @@ static bool file_to_disk_secure_link_impl(const char* path, const char* basis_pa
close(dirfd);
free(leaf);
/* The basis file could not be linked in (missing, cross-device, refused
by the filesystem). Write a byte-identical local copy instead. */
by the filesystem). Stream a byte-identical local copy from the basis
itself (never the possibly-absent caller buffer) so an over-limit basis
still materializes. When the basis path is not a readable regular file
(e.g. a directory raced in), fall back to the caller-supplied bytes. */
if (file_copy_basis_stream_impl(path, basis_path, data_size, preallocate, metadata, policy,
false, false, use_fsync, xattrs, fake_super, temp_dir,
dirs_created, count_floor))
return true;
return file_to_disk_secure_attrs_counted(
path, data, data_size, false, false, preallocate, metadata, policy, false, false, use_fsync,
xattrs, fake_super, false, temp_dir, dirs_created, count_floor);
+10
View File
@@ -163,6 +163,16 @@ bool file_to_disk_secure_link_attrs(const char* path, const char* basis_path, co
const FileMetadata* metadata, FileAttrPolicy policy,
bool use_fsync, const FileXattrList* xattrs, bool fake_super,
const char* temp_dir);
/* Streaming --copy-dest install: atomically materialize `path` by copying the
* bytes of `basis_path` through a bounded buffer (no whole-file buffering, so
* an arbitrarily large basis works), applying the SOURCE metadata and the
* per-file xattrs / --fake-super record. `update` honors a newer destination;
* a --temp-dir scratch location falls back to a direct write on EXDEV. */
bool file_copy_basis_stream_attrs(const char* path, const char* basis_path,
unsigned long long expected_size, bool preallocate,
const FileMetadata* metadata, FileAttrPolicy policy, bool update,
bool use_fsync, const FileXattrList* xattrs, bool fake_super,
const char* temp_dir);
/* Protocol 2.28.0 receiver-stat variants: like the two above but additionally
* report through `dirs_created` (when non-NULL) how many parent directories the
* confined secure walk had to create that lie strictly below `count_floor` (a
+156 -101
View File
@@ -99,6 +99,12 @@ static FileSaveResult file_stage_delayed_update(const char* root_directory,
if (file->basis_link) {
ok = file_to_disk_secure_link(staged_path, file->basis_link, file->data->data, file->data->size,
config->preallocate, metadata, policy, config->use_fsync, NULL);
} else if (file->basis_copy) {
/* --copy-dest basis hit: stream the basis into the staging tree (bounded
buffers, so an over-limit basis still stages). */
ok = file_copy_basis_stream_attrs(staged_path, file->basis_copy, file->data->size,
config->preallocate, metadata, policy, config->update,
config->use_fsync, file->xattrs, config->fake_super, NULL);
} else {
ok =
file_to_disk_secure_attrs(staged_path, file->data->data, file->data->size, false, sparse,
@@ -652,7 +658,8 @@ FileSaveResult file_save_to_disk_full_ex(const char* root_directory, const File*
char* destination_path = NULL;
char *backup_path = NULL, *parent_copy = NULL;
if (!file || !file->path || !file->data || (file->data->size != 0 && !file->data->data) ||
if (!file || !file->path || !file->data ||
(file->data->size != 0 && !file->data->data && !file->basis_link && !file->basis_copy) ||
(!file_get_trust_sender() && has_path_traversal(file->path)) ||
(backup_enabled &&
(!backup_suffix || backup_suffix[0] == '\0' || strchr(backup_suffix, '/') != NULL ||
@@ -975,6 +982,13 @@ FileSaveResult file_save_to_disk_full_ex(const char* root_directory, const File*
disk_path, file->basis_link, file->data->data, file->data->size, config->preallocate,
metadata, policy, config->use_fsync, file->xattrs, config->fake_super, confined_temp,
created_dirs, count_floor);
} else if (config && file->basis_copy) {
/* --copy-dest: stream the basis bytes through a bounded buffer so a basis
larger than any whole-file bound still materializes. The source
metadata was transmitted with the check frame. */
ok = file_copy_basis_stream_attrs(
disk_path, file->basis_copy, file->data->size, config->preallocate, metadata, policy,
config->update, config->use_fsync, file->xattrs, config->fake_super, confined_temp);
} else {
/* The plain no-replace / update / with-fsync engines, plus per-file xattr
(-X/-A) and --fake-super application on the written fd. */
@@ -1024,6 +1038,13 @@ void receiver_stats_note_saved(ReceiverStats* stats, const File* file, bool crea
unsigned created_dirs) {
if (!stats || !file)
return;
/* A basis-dir hit (--link-dest/--copy-dest) materializes bytes the sender
* never transferred. rsync reports no literal data and no created entry for
* such a file, and does not count the parent directories it creates only to
* hold it, so exclude the whole entry from the receiver tallies. */
bool basis_sourced = file->basis_link != NULL || file->basis_copy != NULL;
if (basis_sourced)
return;
bool is_sibling = file->link_group != 0 && !file->link_first;
if (!file->is_dir && !file->is_symlink && !file->is_special && !is_sibling) {
unsigned long long literal = file->literal_bytes;
@@ -1298,16 +1319,17 @@ static File* receive_delta_file(int fd, const Config* config, const char* check_
/* ---- Alternate basis directories (--compare-dest / --copy-dest / --link-dest) ----
* The receiver consults the ordered basis-dir list only when the destination
* entry is NOT already up to date. An "exact match" requires an equal size,
* an equal mtime (unless --size-only), and an equal content xxHash64, so a
* hard link / local copy is only ever made from byte-identical content. */
* entry is NOT already up to date. By default an "exact match" is rsync's
* metadata quick-check: an equal size and an equal mtime (unless --size-only).
* The FastSync-only --verify-basis additionally requires an equal whole-file
* content digest, so a hard link / local copy is only then made from
* byte-verified content. */
typedef struct BasisMatch {
bool hit;
BasisDestType type;
char* basis_path; /* owned absolute path of the matched basis file */
struct stat st; /* fstat() of the matched basis file */
Data* content; /* owned basis bytes (or empty Data), NULL when not loaded */
} BasisMatch;
static void basis_match_free(BasisMatch* match) {
@@ -1315,8 +1337,6 @@ static void basis_match_free(BasisMatch* match) {
return;
free(match->basis_path);
match->basis_path = NULL;
data_destroy(match->content);
match->content = NULL;
match->hit = false;
match->type = BASIS_DEST_NONE;
}
@@ -1348,32 +1368,10 @@ static bool basis_open_regular(const char* path, unsigned long long expected_siz
return true;
}
/* Read the whole remaining content of an open descriptor. A zero-length file
yields an empty Data (data pointer NULL). */
static Data* basis_read_content(int fd, unsigned long long size) {
if (size == 0)
return data_create_reserve(0);
if (size > MAX_RECEIVE_WHOLE_FILE_SIZE || size > SIZE_MAX)
return NULL;
void* buf = protocol_alloc((size_t)size);
if (!buf)
return NULL;
size_t got = 0;
while (got < (size_t)size) {
ssize_t n = read(fd, (char*)buf + got, (size_t)size - got);
if (n <= 0) {
free(buf);
return NULL;
}
got += (size_t)n;
}
return data_create(buf, (size_t)size);
}
/* --ignore-times forces every file to be updated, so no basis hit is ever
declared (matching rsync, where -I prevents link-dest from linking). */
static bool basis_quick_matches(const Config* config, const struct stat* st, time_t check_mtime,
long check_mtime_nsec) {
bool file_basis_quick_match(const Config* config, const struct stat* st, time_t check_mtime,
long check_mtime_nsec) {
if (config->size_only)
return true;
long mtime_nsec = 0;
@@ -1384,28 +1382,39 @@ static bool basis_quick_matches(const Config* config, const struct stat* st, tim
config->modify_window);
}
/* Search the basis-dir list in command-line order and return the first exact
match. When load_content is true the matched bytes are kept in out->content
so the caller can materialize the file without re-reading it.
/* True when a basis hit must be confirmed by a whole-file content digest
(--verify-basis). False is the rsync-parity default: the metadata
quick-check alone decides a hit. */
bool file_basis_content_required(const Config* config) {
return config != NULL && config->verify_basis;
}
An exact match ALSO requires the basis bytes' digest to equal the source's,
so `hash_content` gates the content read/hash itself. A server-contacting
--dry-run passes hash_content=false: no basis file may be read or hashed
(that would be a 1-bit content oracle against a client-supplied digest), so a
metadata-only pass can never confirm a hit and declines it. The real path
always passes hash_content=true, keeping its behavior byte-for-byte. */
/* Search the basis-dir list in command-line order and return the first match.
By default (no --verify-basis) rsync's metadata quick-check is sufficient:
basis_open_regular has already required an equal size, and
file_basis_quick_match applies rsync's mtime (or --size-only) rule.
--verify-basis additionally requires the basis bytes' whole-file digest to
equal the sender's, restoring FastSync's historical content equality; that
digest is computed by streaming the open basis descriptor, so an arbitrarily
large basis is verified without buffering it. A copy/link install re-reads
the basis from its path in bounded buffers, so no content buffer is kept.
`hash_content` gates content READS under --verify-basis: a server-contacting
--dry-run passes false because hashing a basis against a client-supplied
digest would be a 1-bit content oracle. Without --verify-basis a dry-run can
still confirm the metadata-only hit without reading any basis bytes, matching
rsync's read-only quick-check. */
static bool basis_match_find(const Config* config, const char* check_path,
unsigned long long check_size, time_t check_mtime,
long check_mtime_nsec, const uint8_t* check_digest,
size_t check_digest_len, bool load_content, bool hash_content,
BasisMatch* out) {
size_t check_digest_len, bool hash_content, BasisMatch* out) {
memset(out, 0, sizeof(*out));
if (!config || !config_has_basis(config) || config->ignore_times)
return false;
/* Dry-run: never read/hash basis content. A hit cannot be decided from
metadata alone, so report no match (the caller treats it as would-transfer)
without touching the file's contents. */
if (!hash_content)
/* --verify-basis needs the basis content; a content-blind (dry-run) pass can
never confirm it and must not read the file, so decline without touching
the basis bytes. */
if (file_basis_content_required(config) && !hash_content)
return false;
for (int i = 0; i < config->basis_count; i++) {
const BasisDest* entry = &config->basis_dirs[i];
@@ -1424,29 +1433,26 @@ static bool basis_match_find(const Config* config, const char* check_path,
int fd;
struct stat st;
if (basis_open_regular(candidate, check_size, &fd, &st)) {
if (basis_quick_matches(config, &st, check_mtime, check_mtime_nsec)) {
Data* content = basis_read_content(fd, check_size);
if (content) {
if (file_basis_quick_match(config, &st, check_mtime, check_mtime_nsec)) {
bool hit = true;
if (file_basis_content_required(config)) {
uint8_t basis_digest[CHECKSUM_MAX_DIGEST_LEN];
size_t basis_len = 0;
bool hashed = checksum_digest((ChecksumAlgo)config->checksum_algo, config->checksum_seed,
content->data, content->size, basis_digest,
sizeof(basis_digest), &basis_len);
if (hashed && basis_len == check_digest_len && check_digest_len > 0 &&
memcmp(basis_digest, check_digest, check_digest_len) == 0) {
out->hit = true;
out->type = entry->type;
out->basis_path = candidate;
candidate = NULL; /* ownership transferred to out */
out->st = st;
out->content = load_content ? content : NULL;
if (!load_content)
data_destroy(content);
close(fd);
return true;
}
bool hashed =
checksum_digest_fd((ChecksumAlgo)config->checksum_algo, config->checksum_seed, fd,
basis_digest, sizeof(basis_digest), &basis_len);
hit = hashed && basis_len == check_digest_len && check_digest_len > 0 &&
memcmp(basis_digest, check_digest, check_digest_len) == 0;
}
if (hit) {
out->hit = true;
out->type = entry->type;
out->basis_path = candidate;
candidate = NULL; /* ownership transferred to out */
out->st = st;
close(fd);
return true;
}
data_destroy(content);
}
close(fd);
}
@@ -1871,6 +1877,12 @@ typedef struct {
long long check_mtime_nsec;
uint8_t check_digest[CHECKSUM_MAX_DIGEST_LEN];
size_t check_digest_len;
/* Source metadata carried alongside the check frame whenever a basis dir is
configured (rsync keeps the whole file list; FastSync's sender-driven
incremental path otherwise never transmits metadata for a SKIPPED file).
A basis materialization applies these SOURCE attributes instead of the
basis inode's, matching rsync's "copy then fix attributes". */
FileMetadata* source_metadata;
bool dest_exists; /* any destination entry exists (lstat succeeded) */
bool has_old_file;
int old_fd;
@@ -1903,6 +1915,8 @@ static void incremental_check_state_cleanup(IncrementalCheckState* state) {
if (state->old_fd >= 0)
close(state->old_fd);
state->old_fd = -1;
file_metadata_destroy(state->source_metadata);
state->source_metadata = NULL;
free(state->full_path);
state->full_path = NULL;
free(state->check_path);
@@ -1927,7 +1941,7 @@ static IncrementalCheckOutcome incremental_check_receive_request(IncrementalChec
send_error_detail(fd, "invalid check mtime nanoseconds");
return INCREMENTAL_ERROR;
}
if ((config->checksum || config_has_basis(config))) {
if ((config->checksum || config->verify_basis)) {
uint8_t wire_len;
if (!receive_n_data(fd, &wire_len, sizeof(wire_len)) || wire_len == 0 ||
wire_len > CHECKSUM_MAX_DIGEST_LEN ||
@@ -1939,8 +1953,24 @@ static IncrementalCheckOutcome incremental_check_receive_request(IncrementalChec
if (!receive_n_data(fd, state->check_digest, state->check_digest_len))
return INCREMENTAL_ERROR;
}
/* The sender transmits the source metadata with every basis-configured check
so a basis hit can be materialized with the SOURCE's attributes (rsync
copies/copies-then-fixes; the receiver would otherwise only have the basis
inode's stat). The block is symmetric and consumed unconditionally here,
whether or not this file ends up as a basis hit. */
if (config_has_basis(config) && config->use_metadata) {
int meta_ok = 1;
state->source_metadata = metadata_receive(fd, &meta_ok);
if (!meta_ok)
return INCREMENTAL_ERROR;
}
if (state->check_size > MAX_RECEIVE_WHOLE_FILE_SIZE) {
/* A basis-configured run may materialize a file larger than the whole-file
payload bound: a basis hit is streamed from the basis path (bounded
buffers), so the check size is not itself an allocation. Every other
path (delta/append/full) still applies MAX_RECEIVE_WHOLE_FILE_SIZE, and a
miss simply falls through to the normal transfer with its own bound. */
if (!config_has_basis(config) && state->check_size > MAX_RECEIVE_WHOLE_FILE_SIZE) {
send_error_detail(fd, "check size exceeds receiver limit");
return INCREMENTAL_ERROR;
}
@@ -2030,6 +2060,20 @@ incremental_check_ignore_existing(const IncrementalCheckState* state) {
return INCREMENTAL_SKIP;
}
/* Metadata for a materialized basis hit: prefer the SOURCE metadata the sender
transmitted with the check frame (rsync copies then fixes the destination to
the source's attributes); fall back to the basis inode's own stat when
metadata was not negotiated. Consumes state->source_metadata on success. */
static FileMetadata* basis_take_metadata(IncrementalCheckState* state,
const struct stat* basis_st) {
if (state->source_metadata) {
FileMetadata* meta = state->source_metadata;
state->source_metadata = NULL;
return meta;
}
return file_metadata_create(NULL, basis_st, false, false);
}
/* --link-dest relink of an already up-to-date destination. rsync hard-links a
destination entry to a matching basis even when the entry is already correct,
so a run over an existing tree still maximizes sharing with the basis. Only a
@@ -2047,7 +2091,7 @@ static IncrementalCheckOutcome incremental_check_link_dest_relink(IncrementalChe
BasisMatch basis;
basis_match_find(config, state->check_path, state->check_size, (time_t)state->check_mtime,
(long)state->check_mtime_nsec, state->check_digest, state->check_digest_len,
true, true, &basis);
true, &basis);
/* Only a link-dest hit relinks; a copy-dest/compare-dest hit (or a miss) lets
the up-to-date check below keep the existing destination. */
if (!basis.hit || basis.type != BASIS_DEST_LINK) {
@@ -2060,11 +2104,16 @@ static IncrementalCheckOutcome incremental_check_link_dest_relink(IncrementalChe
return INCREMENTAL_CONTINUE;
}
File* materialized = file_create(state->check_path);
if (materialized && basis.content) {
if (materialized) {
data_destroy(materialized->data);
materialized->data = basis.content;
basis.content = NULL;
materialized->metadata = file_metadata_create(NULL, &basis.st, false, false);
materialized->data = data_create_reserve((size_t)state->check_size);
if (!materialized->data) {
file_destroy(materialized);
materialized = NULL;
}
}
if (materialized) {
materialized->metadata = basis_take_metadata(state, &basis.st);
materialized->skip = true;
materialized->basis_link = basis.basis_path;
basis.basis_path = NULL;
@@ -2072,9 +2121,6 @@ static IncrementalCheckOutcome incremental_check_link_dest_relink(IncrementalChe
file_destroy(materialized);
materialized = NULL;
}
} else {
file_destroy(materialized);
materialized = NULL;
}
if (materialized) {
if (!send_status(state->fd, STATUS_OK)) {
@@ -2175,12 +2221,13 @@ static IncrementalCheckOutcome incremental_check_quick_skip(IncrementalCheckStat
materialize nothing (no basis link/copy, no append/delta/full transfer) and
the sender must send no data, so answer STATUS_DRY_RUN_TRANSFER and stop.
The basis lookup is deliberately content-blind: a real run would only accept
a --compare-dest exact hit after hashing the basis file and comparing it with
the client-supplied digest, which in a dry-run is a 1-bit content oracle.
Under dry_run no basis bytes may be read, so an otherwise-matching entry is
treated as would-transfer instead of a skip. Everything read here (the
destination file's metadata, basis candidates' metadata) is read-only. */
The basis lookup is content-blind: under the default metadata quick-check a
hit needs no basis bytes and is honored here just as in a real run; under
--verify-basis a real run hashes the basis against the client-supplied digest,
which in a dry-run is a 1-bit content oracle, so no basis bytes may be read
and an otherwise-matching entry is reported as would-transfer. Everything
read here (the destination file's metadata, basis candidates' metadata) is
read-only. */
static IncrementalCheckOutcome incremental_check_dry_run_shortcut(IncrementalCheckState* state,
bool* skipped,
bool* would_transfer) {
@@ -2191,12 +2238,14 @@ static IncrementalCheckOutcome incremental_check_dry_run_shortcut(IncrementalChe
bool skip_via_compare = false;
if (config_has_basis(config) && !config->ignore_times) {
BasisMatch basis;
/* hash_content=false: a dry-run must not read or hash the basis file. No
content comparison is possible, so no compare-dest hit can be confirmed
and an otherwise-matching file is reported as would-transfer. */
/* hash_content=false: a dry-run must not read or hash the basis file, so
under --verify-basis no compare-dest hit can be confirmed and an
otherwise-matching file is reported as would-transfer. Without
--verify-basis the metadata quick-check confirms it without touching any
basis bytes. */
basis_match_find(config, state->check_path, state->check_size, (time_t)state->check_mtime,
(long)state->check_mtime_nsec, state->check_digest, state->check_digest_len,
false, false, &basis);
false, &basis);
if (basis.hit && basis.type == BASIS_DEST_COMPARE && !state->has_old_file)
skip_via_compare = true;
basis_match_free(&basis);
@@ -2224,7 +2273,7 @@ static IncrementalCheckOutcome incremental_check_try_basis(IncrementalCheckState
BasisMatch basis;
basis_match_find(config, state->check_path, state->check_size, (time_t)state->check_mtime,
(long)state->check_mtime_nsec, state->check_digest, state->check_digest_len,
true, true, &basis);
true, &basis);
if (basis.hit) {
if (basis.type == BASIS_DEST_COMPARE) {
basis_match_free(&basis);
@@ -2234,24 +2283,30 @@ static IncrementalCheckOutcome incremental_check_try_basis(IncrementalCheckState
return INCREMENTAL_SKIP;
}
} else {
/* Copy/link installs source their bytes from the basis PATH at install
time (bounded buffers), so no whole-file content buffer is needed here
even for an over-limit basis. */
File* materialized = file_create(state->check_path);
if (materialized && basis.content) {
if (materialized) {
data_destroy(materialized->data);
materialized->data = basis.content;
basis.content = NULL;
materialized->metadata = file_metadata_create(NULL, &basis.st, false, false);
materialized->skip = true; /* receiver must not ack this as a data file */
if (basis.type == BASIS_DEST_LINK) {
materialized->basis_link = basis.basis_path;
basis.basis_path = NULL;
materialized->data = data_create_reserve((size_t)state->check_size);
if (!materialized->data) {
file_destroy(materialized);
materialized = NULL;
}
}
if (materialized) {
materialized->metadata = basis_take_metadata(state, &basis.st);
materialized->skip = true; /* receiver must not ack this as a data file */
if (basis.type == BASIS_DEST_LINK)
materialized->basis_link = basis.basis_path;
else
materialized->basis_copy = basis.basis_path;
basis.basis_path = NULL;
if (!materialized->metadata) {
file_destroy(materialized);
materialized = NULL;
}
} else {
file_destroy(materialized);
materialized = NULL;
}
if (materialized) {
if (!send_status(fd, STATUS_OK)) {
@@ -3328,7 +3383,7 @@ static bool delete_extras_budgeted_observed(const Config* config, DeleteManifest
size_t skipped = 0;
DeleteWalkResult result = delete_extras_limited_observed(
config->receive_root_directory, manifest->keeps, manifest->dirs, remaining, skips, used,
&deleted, &skipped, observer, observer_context);
config->protect_rules, &deleted, &skipped, observer, observer_context);
if (owned_prefixes) {
for (int i = 0; i < config->basis_count; i++)
free(owned_prefixes[i]);
@@ -3509,7 +3564,7 @@ static bool delete_missing_args_budgeted_observed(const Config* config, DeleteMa
PrefixedDeleteObserver nested = {observer, observer_context, rel};
DeleteWalkResult walk =
no_keeps ? delete_extras_limited_observed(full, no_keeps, NULL, remaining, NULL, 0,
&contents_deleted, &contents_skipped,
NULL, &contents_deleted, &contents_skipped,
observer ? prefixed_delete_observer : NULL,
observer ? &nested : NULL)
: DELETE_WALK_ERROR;
@@ -3620,7 +3675,7 @@ bool manifest_would_delete_list(const Config* config, DeleteManifest* manifest,
used = idx;
}
bool ok = delete_extras_list(config->receive_root_directory, manifest->keeps, manifest->dirs,
skips, used, out, count_out);
skips, used, config->protect_rules, out, count_out);
if (owned_prefixes) {
for (int i = 0; i < config->basis_count; i++)
free(owned_prefixes[i]);
+10
View File
@@ -24,6 +24,16 @@ File* file_receive_hardlink(int file_descriptor);
File* file_receive_symlink(int file_descriptor, const Config* config);
File* file_receive_special(int file_descriptor);
bool file_special_rdev_valid(int32_t major, int32_t minor, mode_t mode);
/* Testable basis quick-check / verification policy. file_basis_quick_match is
* rsync's metadata quick-check for a basis candidate (equal size is required
* separately by the caller; this adds the --size-only / mtime / --modify-window
* leg). file_basis_content_required reports whether a hit must ALSO be
* confirmed by a whole-file content digest (--verify-basis; false is the
* default rsync-parity behavior). */
bool file_basis_quick_match(const Config* config, const struct stat* st, time_t check_mtime,
long check_mtime_nsec);
bool file_basis_content_required(const Config* config);
File* receive_incremental_check(int fd, const Config* config, bool* skipped);
/* Extended variant used by the receiver. `would_transfer` (may be NULL) is set
* true only on the server-contacting --dry-run path when the file is not up to
+5
View File
@@ -57,6 +57,11 @@ typedef struct {
* equals the incoming file, and `data` is kept as the cross-filesystem
* fallback (a local copy) if the hard link cannot be created. */
char* basis_link;
/* Receiver-only, --copy-dest: when set (and basis_link is NULL), stream the
* basis file's bytes into the destination instead of `data`/`data->size`.
* This lets a basis larger than any whole-file bound materialize without
* buffering it; the source metadata on `metadata` is applied afterwards. */
char* basis_copy;
/* --hard-links (-H), sender + receiver wire state. link_group is a run-local
* id shared by every member of one source inode (0 = not part of a group).
* The FIRST member (link_first == true) carries its data on the wire and is
+1 -1
View File
@@ -53,7 +53,7 @@ typedef struct {
char* pattern; /* cleaned glob pattern (no leading '/', no trailing '/') */
} FilterRule;
typedef struct {
typedef struct FilterRuleList {
FilterRule** items; /* owned array of rule pointers */
int count;
int capacity;
+3 -1
View File
@@ -191,7 +191,9 @@ enum NET_STATUS {
* (--delete-delay). Payload: an int32 has_config flag (1 on the first plan
* of the run, 0 afterwards); when set, the three global config sections
* (protected-prefix count+paths, size-skipped count+paths, missing-args
* count+paths); then the destination-relative directory path wire string
* count+paths); then an int32 apply flag (1 for a real plan, 0 for a
* config-only carrier frame that must not walk a directory); then the
* destination-relative directory path wire string
* ("." for the receive root); then the child-directory count + names and the
* child-file count + names that must be kept. Appended after
* STATUS_DEST_INFO so no existing status is renumbered. */
+40 -14
View File
@@ -682,7 +682,8 @@ static bool is_synced_dir(const PathIndex* dirs, const char* rel) {
like any other non-directory extra (never followed). */
static bool delete_extras_fd(int dirfd, const char* rel_path, const PathIndex* keep,
const PathIndex* dirs, DeleteBudget* budget,
const DeleteSkipEntry* skips, int skip_count, bool parent_deletable,
const DeleteSkipEntry* skips, int skip_count,
const FilterRuleList* protect_rules, bool parent_deletable,
bool* all_removed, DeletePathObserver observer,
void* observer_context) {
/* openat(dirfd, ".") opens an independent file description: a dup() would
@@ -729,12 +730,23 @@ static bool delete_extras_fd(int dirfd, const char* rel_path, const PathIndex* k
free(child_rel);
continue;
}
if (S_ISDIR(st.st_mode)) {
bool is_dir = S_ISDIR(st.st_mode);
if (protect_rules && filter_rules_apply_side(protect_rules, child_rel, entry->d_name, is_dir,
FILTER_SIDE_RECEIVER) == FILTER_ACTION_PROTECT) {
/* A first-match protect rule shields the extra; for a directory the whole
subtree is shielded (rsync prunes an excluded directory), so do not
descend. */
local_survives = true;
free(child_rel);
continue;
}
if (is_dir) {
int childfd = openat(dirfd, entry->d_name, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
bool child_all_removed = false;
if (childfd >= 0) {
if (!delete_extras_fd(childfd, child_rel, keep, dirs, budget, skips, skip_count, deletable,
&child_all_removed, observer, observer_context))
if (!delete_extras_fd(childfd, child_rel, keep, dirs, budget, skips, skip_count,
protect_rules, deletable, &child_all_removed, observer,
observer_context))
operation_ok = false;
close(childfd);
} else if (errno != ENOENT) {
@@ -802,7 +814,8 @@ static bool delete_extras_fd(int dirfd, const char* rel_path, const PathIndex* k
reportable children (depth-first), matching the delete pass's ordering. */
static bool list_extras_fd(int dirfd, const char* rel_path, const PathIndex* keep,
const PathIndex* dirs, ArrayList* out, size_t* recorded,
const DeleteSkipEntry* skips, int skip_count, bool parent_deletable,
const DeleteSkipEntry* skips, int skip_count,
const FilterRuleList* protect_rules, bool parent_deletable,
bool* all_removed) {
int scanfd = openat(dirfd, ".", O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
if (scanfd < 0)
@@ -836,12 +849,21 @@ static bool list_extras_fd(int dirfd, const char* rel_path, const PathIndex* kee
free(child_rel);
continue;
}
if (S_ISDIR(st.st_mode)) {
bool is_dir = S_ISDIR(st.st_mode);
if (protect_rules && filter_rules_apply_side(protect_rules, child_rel, entry->d_name, is_dir,
FILTER_SIDE_RECEIVER) == FILTER_ACTION_PROTECT) {
/* Mirror the delete walk: a protected entry is never reported as a
would-delete and a protected directory's subtree is not enumerated. */
local_survives = true;
free(child_rel);
continue;
}
if (is_dir) {
int childfd = openat(dirfd, entry->d_name, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
bool child_all_removed = false;
if (childfd >= 0) {
if (!list_extras_fd(childfd, child_rel, keep, dirs, out, recorded, skips, skip_count,
deletable, &child_all_removed))
protect_rules, deletable, &child_all_removed))
operation_ok = false;
close(childfd);
} else if (errno != ENOENT) {
@@ -892,7 +914,7 @@ static bool list_extras_fd(int dirfd, const char* rel_path, const PathIndex* kee
bool delete_extras_list(const char* dest_root, const ArrayList* manifest,
const ArrayList* synced_dirs, const DeleteSkipEntry* skips, int skip_count,
ArrayList* out, size_t* count_out) {
const FilterRuleList* protect_rules, ArrayList* out, size_t* count_out) {
if (count_out)
*count_out = 0;
if (!manifest || !out)
@@ -928,7 +950,7 @@ bool delete_extras_list(const char* dest_root, const ArrayList* manifest,
bool all_removed = false;
size_t recorded = 0;
bool ok = list_extras_fd(rootfd, "", &keep, have_dirs ? &dirs : NULL, out, &recorded, skips,
skip_count, false, &all_removed);
skip_count, protect_rules, false, &all_removed);
if (close(rootfd) != 0)
ok = false;
path_index_free(&keep);
@@ -942,6 +964,7 @@ bool delete_extras_list(const char* dest_root, const ArrayList* manifest,
DeleteWalkResult delete_extras_limited_observed(const char* dest_root, const ArrayList* manifest,
const ArrayList* synced_dirs, size_t max_delete,
const DeleteSkipEntry* skips, int skip_count,
const FilterRuleList* protect_rules,
size_t* deleted_out, size_t* skipped_out,
DeletePathObserver observer,
void* observer_context) {
@@ -984,8 +1007,9 @@ DeleteWalkResult delete_extras_limited_observed(const char* dest_root, const Arr
}
DeleteBudget budget = {.max_delete = max_delete, .deleted = 0, .skipped = 0, .limit_hit = false};
bool all_removed = false;
bool ok = delete_extras_fd(rootfd, "", &keep, have_dirs ? &dirs : NULL, &budget, skips,
skip_count, false, &all_removed, observer, observer_context);
bool ok =
delete_extras_fd(rootfd, "", &keep, have_dirs ? &dirs : NULL, &budget, skips, skip_count,
protect_rules, false, &all_removed, observer, observer_context);
if (close(rootfd) != 0)
ok = false;
path_index_free(&keep);
@@ -1003,13 +1027,15 @@ DeleteWalkResult delete_extras_limited_observed(const char* dest_root, const Arr
DeleteWalkResult delete_extras_limited(const char* dest_root, const ArrayList* manifest,
const ArrayList* synced_dirs, size_t max_delete,
const DeleteSkipEntry* skips, int skip_count,
size_t* deleted_out, size_t* skipped_out) {
const FilterRuleList* protect_rules, size_t* deleted_out,
size_t* skipped_out) {
return delete_extras_limited_observed(dest_root, manifest, synced_dirs, max_delete, skips,
skip_count, deleted_out, skipped_out, NULL, NULL);
skip_count, protect_rules, deleted_out, skipped_out, NULL,
NULL);
}
bool delete_extras(const char* dest_root, const ArrayList* manifest) {
return delete_extras_limited(dest_root, manifest, NULL, SIZE_MAX, NULL, 0, NULL, NULL) ==
return delete_extras_limited(dest_root, manifest, NULL, SIZE_MAX, NULL, 0, NULL, NULL, NULL) ==
DELETE_WALK_OK;
}
+10 -3
View File
@@ -2,6 +2,7 @@
#define UTILS_H
#include "array_list.h"
#include "filter.h"
#include <stddef.h>
#include <stdbool.h>
#include <stdio.h>
@@ -148,7 +149,8 @@ bool path_under_skip_prefix(const char* child_rel, bool at_root, const DeleteSki
DeleteWalkResult delete_extras_limited(const char* dest_root, const ArrayList* manifest,
const ArrayList* synced_dirs, size_t max_delete,
const DeleteSkipEntry* skips, int skip_count,
size_t* deleted_out, size_t* skipped_out);
const FilterRuleList* protect_rules, size_t* deleted_out,
size_t* skipped_out);
/* Optional per-deletion observer: called for each destination-relative path
actually removed (a file, symlink, or directory), in removal order, so the
@@ -156,10 +158,15 @@ DeleteWalkResult delete_extras_limited(const char* dest_root, const ArrayList* m
typedef void (*DeletePathObserver)(void* context, const char* rel_path);
/* `delete_extras_limited_observed` is delete_extras_limited with an optional
observer; the observer is invoked only for entries truly removed. */
* observer; the observer is invoked only for entries truly removed. When
* `protect_rules` is non-NULL its receiver-side verdict is evaluated for every
* candidate extra: a first-match PROTECT leaves the entry (and, for a
* directory, its whole subtree) in place, while RISK/NONE fall through to the
* ordinary skip-prefix/keep-set logic. */
DeleteWalkResult delete_extras_limited_observed(const char* dest_root, const ArrayList* manifest,
const ArrayList* synced_dirs, size_t max_delete,
const DeleteSkipEntry* skips, int skip_count,
const FilterRuleList* protect_rules,
size_t* deleted_out, size_t* skipped_out,
DeletePathObserver observer,
void* observer_context);
@@ -170,7 +177,7 @@ DeleteWalkResult delete_extras_limited_observed(const char* dest_root, const Arr
strings appended to `out` and receives their count in *count_out. */
bool delete_extras_list(const char* dest_root, const ArrayList* manifest,
const ArrayList* synced_dirs, const DeleteSkipEntry* skips, int skip_count,
ArrayList* out, size_t* count_out);
const FilterRuleList* protect_rules, ArrayList* out, size_t* count_out);
bool delete_extras(const char* dest_root, const ArrayList* manifest);
/* Open the existing destination directory at `dest_root`, confined to the
authorized root with an O_NOFOLLOW component walk (the same confinement the
+7
View File
@@ -119,6 +119,13 @@ static void build_canonical_frame(void) {
cfg->usermap[0].to = MAP_TO;
cfg->usermap[0].to_name = NULL;
}
/* Force a non-empty receiver delete-protection block so the fuzzer mutates
* its rule count, action/sides codes and pattern strings. */
cfg->filters = array_list_create(free);
if (cfg->filters) {
array_list_add(cfg->filters, str_dup("P *.log"));
array_list_add(cfg->filters, str_dup("+r keep/**"));
}
if (!cfg->send_directory || !cfg->receive_root_directory || !cfg->usermap) {
config_delete(cfg);
return;
+13
View File
@@ -229,6 +229,18 @@ def corpus_iconv(root: str) -> None:
os.utime(full, (_SRC_MTIME, _SRC_MTIME))
# Payload for the --fuzzy basis corpus: large enough for the delta engine's
# 16 KiB minimum and with repeated content so a coinciding basis yields a
# non-zero (and identical) Matched data count in both tools.
FUZZY_PAYLOAD = (b"the quick brown fox jumps over the lazy dog\n" * 2000)[:65536]
def corpus_fuzzy(root: str) -> None:
"""A named regular file; the `fuzzy` seed adds the similar-suffix sibling."""
clean_dir(root)
_write(os.path.join(root, "report_v2.txt"), FUZZY_PAYLOAD)
CORPORA: Dict[str, Callable[[str], None]] = {
"basic": corpus_basic,
"unicode": corpus_unicode,
@@ -240,6 +252,7 @@ CORPORA: Dict[str, Callable[[str], None]] = {
"multidir": corpus_multidir,
"relative": corpus_relative,
"iconv": corpus_iconv,
"fuzzy": corpus_fuzzy,
}
+149
View File
@@ -9,6 +9,7 @@ directory when the shared test server is launched), so every scratch tree lives
under ``TEST_DATA_DIR`` rather than pytest's ``tmp_path``.
"""
import os
import random
import shutil
import subprocess
import sys
@@ -68,6 +69,84 @@ def _tree_bytes(root):
return out
_CC_DELTA_T0 = 1_600_000_000
_CC_DELTA_T1 = 1_600_000_100
_DELTA_STATS_KEYS = (
"Number of created files",
"Number of regular files transferred",
"Total transferred file size",
"Literal data",
"Matched data",
)
def _pin_tree(root, mtime):
for dirpath, dirnames, filenames in os.walk(root):
for name in dirnames + filenames:
path = os.path.join(dirpath, name)
if not os.path.islink(path):
os.utime(path, (mtime, mtime))
os.utime(root, (mtime, mtime))
def _make_delta_basis(src, size=512 * 1024):
"""Build a source and a matching pre-modification basis tree.
The source's ``big.bin`` is then modified in a few disjoint places and given
a newer mtime so both tools take the delta path. Returns the basis dir.
"""
clean_dir(src)
original = random.Random(20240101).randbytes(size)
with open(os.path.join(src, "big.bin"), "wb") as fh:
fh.write(original)
with open(os.path.join(src, "small.txt"), "wb") as fh:
fh.write(b"hello world\n")
_pin_tree(src, _CC_DELTA_T0)
basis = src.rstrip("/") + "_basis"
clean_dir(basis)
shutil.copy2(os.path.join(src, "big.bin"), os.path.join(basis, "big.bin"))
shutil.copy2(os.path.join(src, "small.txt"), os.path.join(basis, "small.txt"))
_pin_tree(basis, _CC_DELTA_T0)
modified = bytearray(original)
for off in (0, size // 3, 2 * size // 3, size - 64):
for i in range(32):
modified[off + i] ^= 0x5A
with open(os.path.join(src, "big.bin"), "wb") as fh:
fh.write(bytes(modified))
os.utime(os.path.join(src, "big.bin"), (_CC_DELTA_T1, _CC_DELTA_T1))
return basis
def _seed_from_basis(basis, target):
clean_dir(target)
for name in os.listdir(basis):
shutil.copy2(os.path.join(basis, name), os.path.join(target, name))
def _delta_stats(text):
found = {}
for line in text.splitlines():
for key in _DELTA_STATS_KEYS:
if line.startswith(key + ":"):
found[key] = line.split(":", 1)[1].strip()
return found
def _big_bin_outfmt(text):
"""The ``(c, C)`` pair from the ``big.bin`` out-format line (`%c|%C %n`)."""
for line in text.splitlines():
stripped = line.strip()
if "|" not in stripped or not stripped.endswith("big.bin"):
continue
c_field, rest = stripped.split("|", 1)
fields = rest.split()
return c_field.strip(), (fields[0] if fields else "")
return None, None
class TestCodecChoiceMatrix:
"""The CLI accept/reject set and exit codes must match rsync 3.4.1."""
@@ -191,6 +270,76 @@ class TestCodecTransferDifferential:
assert _tree_bytes(received) == _tree_bytes(rsync_dst)
class TestChecksumChoiceDeltaSurface:
"""--checksum-choice does not move the delta-transfer parity surface.
FastSync's delta BLOCK strong checksum is a fixed xxHash32, so the
negotiated algorithm only selects the whole-file comparison digest (and the
``%C`` transfer digest). A pre-seeded delta transfer must therefore land
byte-identical bytes and report the same counters for every choice, while
``%C`` -- the one token that tracks the choice -- stays byte-identical to
rsync. This pins the Track-3b reclassification in RSYNC_COMPAT.md.
"""
CHOICES = ["xxh64", "xxh128", "xxh3", "md5", "md4", "sha1",
"xxh64,sha1", "sha1,xxh64"]
@requires_rsync
@pytest.mark.ci
def test_delta_surface_invariant_to_checksum_choice(self, shared_server):
src = _scratch("ccdelta_src")
basis = _make_delta_basis(src)
source_bytes = _tree_bytes(src)
rsync_c, fastsync_c, digests = {}, {}, {}
rsync_stats, fastsync_stats = {}, {}
for choice in self.CHOICES:
tag = choice.replace(",", "_")
rsync_dst = _scratch(f"ccdelta_rs_{tag}")
fs_dst = _scratch(f"ccdelta_fs_{tag}")
fs_root = get_dest_received_dir(fs_dst, src)
_seed_from_basis(basis, rsync_dst)
_seed_from_basis(basis, fs_root)
# Pin the block size on both ends so the literal/matched split is
# comparable (rsync's adaptive default would otherwise differ from
# FastSync's 8192-byte default).
rsync_result = _rsync(["-a", "--no-whole-file", "-B8192", "--stats",
"--out-format=%c|%C %n", f"--cc={choice}",
src + "/", rsync_dst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
result, _ = run_client(
src, fs_dst,
flags=["-a", "--incremental", "--delta", "-B8192", "--stats",
"--out-format=%c|%C %n", f"--cc={choice}"],
port=shared_server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
assert _tree_bytes(rsync_dst) == source_bytes, choice
assert _tree_bytes(fs_root) == source_bytes, choice
assert _delta_stats(rsync_result.stdout) == _delta_stats(result.stdout), choice
rs_c, rs_C = _big_bin_outfmt(rsync_result.stdout)
fs_c, fs_C = _big_bin_outfmt(result.stdout)
assert rs_C == fs_C, f"{choice}: %C rsync={rs_C!r} fastsync={fs_C!r}"
rsync_c[choice] = rs_c
fastsync_c[choice] = fs_c
digests[choice] = fs_C
rsync_stats[choice] = _delta_stats(rsync_result.stdout)
fastsync_stats[choice] = _delta_stats(result.stdout)
# The choice is only observable in %C, and it is effective (the digests
# are not all the same algorithm's output).
assert len(set(digests.values())) > 1, digests
# The compared --stats counters are invariant across choices in each tool
# (and were asserted equal cross-tool inside the loop).
assert len({tuple(sorted(s.items())) for s in rsync_stats.values()}) == 1, rsync_stats
assert len({tuple(sorted(s.items())) for s in fastsync_stats.values()}) == 1, fastsync_stats
# The block-checksum token (%c) is invariant across choices in each tool.
assert len(set(rsync_c.values())) == 1, rsync_c
assert len(set(fastsync_c.values())) == 1, fastsync_c
class TestCodecNegotiationFallback:
"""FastSync's auto negotiation and deterministic fallback order."""
+45 -18
View File
@@ -217,7 +217,21 @@ class _SlicingProxy:
class TestDeleteTimingFinalStateParity:
"""On a successful transfer the per-directory timings match rsync's result."""
"""On a successful transfer the per-directory timings match rsync's result.
Plain ``--delete`` has no rsync-incompatible spelling: it defaults to
delete-during on both tools, so it is compared against rsync's own default.
``--delete-commit`` is FastSync-only and selects the late whole-tree commit,
which is rsync's ``--delete-after`` timing.
"""
# (fastsync flag, rsync flag)
PAIRS = [
("--delete", "--delete"),
("--delete-during", "--delete-during"),
("--delete-delay", "--delete-delay"),
("--delete-commit", "--delete-after"),
]
def _run_fastsync(self, tag, timing):
source, dest, received = _seed_pair(tag)
@@ -226,12 +240,12 @@ class TestDeleteTimingFinalStateParity:
result, _ = run_client(source, dest, flags=[timing], port=server.port)
return result, received
@pytest.mark.parametrize("timing", ["--delete-during", "--delete-delay"])
@pytest.mark.parametrize("fs_timing,rs_timing", PAIRS)
@requires_rsync
def test_success_final_state_matches_rsync(self, timing):
# Worker-safe names: xdist may run both parametrizations concurrently, so
# the timing is part of every fixture path.
label = timing.lstrip("-")
def test_success_final_state_matches_rsync(self, fs_timing, rs_timing):
# Worker-safe names: xdist may run the parametrizations concurrently, so
# the flags are part of every fixture path.
label = f"{fs_timing.lstrip('-')}_vs_{rs_timing.lstrip('-')}"
# Build the rsync fixture from the same seed so both sides start equal.
source, dest, received = _seed_pair(f"parity_rsync_{label}")
source2 = source
@@ -240,17 +254,18 @@ class TestDeleteTimingFinalStateParity:
# rsync mirrors src/ into dst/; seed the same extra.
_write(os.path.join(rsync_dst, "d", "old_extra"), b"stale extra\n")
rsync_result = _rsync(["-a", timing, source2 + "/", rsync_dst + "/"])
rsync_result = _rsync(["-a", rs_timing, source2 + "/", rsync_dst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
rsync_tree = _tree(rsync_dst)
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
result, _ = run_client(source, dest, flags=[timing], port=server.port)
result, _ = run_client(source, dest, flags=[fs_timing], port=server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
fastsync_tree = _tree(received)
assert fastsync_tree == rsync_tree, (
f"{timing}: fastsync tree {fastsync_tree} != rsync tree {rsync_tree}"
f"{fs_timing} vs rsync {rs_timing}: fastsync tree {fastsync_tree} != "
f"rsync tree {rsync_tree}"
)
@@ -292,7 +307,14 @@ class TestDeleteTimingTypeConflictParity:
class TestDeleteTimingFailure:
"""A mid-transfer failure distinguishes during from delay."""
"""A mid-transfer failure distinguishes the during timings from the late
commit timings.
Plain ``--delete`` must behave like ``--delete-during`` (the rsync default),
removing the extras of the directories already reached; ``--delete-commit``
must behave like ``--delete-after`` and remove nothing until the transfer
has fully succeeded.
"""
@pytest.mark.parametrize("mt", [False, True])
def test_during_removes_delay_preserves_on_failure(self, mt):
@@ -301,8 +323,12 @@ class TestDeleteTimingFailure:
assert os.path.exists(extra)
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
for timing, expect_removed in (("--delete-during", True),
("--delete-delay", False)):
for timing, expect_removed in (
("--delete-during", True),
("--delete", True),
("--delete-delay", False),
("--delete-commit", False),
("--delete-after", False)):
# Re-seed the extra before each run.
_write(extra, b"stale extra\n")
proxy = _SlicingProxy(server.port, forward_limit=MID_TRANSFER_BYTES, throttle=PROXY_THROTTLE)
@@ -449,12 +475,13 @@ class TestDeleteDelayVsAfterSnapshot:
class TestDeleteAfterThreadsKeepSet:
"""Regression: -m/--threads with the default delete-after timing (plain
--delete) must still transmit the keep-set manifest and remove destination
extras. PipelineContextSender.delete_suppressed was left uninitialized, so a
garbage true silently skipped the manifest under --threads."""
"""Regression: -j/--threads must still transmit the delete keep-set in every
timing. PipelineContextSender.delete_suppressed was left uninitialized, so a
garbage true silently skipped the late keep-set manifest under --threads.
Plain --delete now uses the per-directory plans, while --delete-commit /
--delete-after keep exercising the late whole-tree manifest."""
@pytest.mark.parametrize("delete_flag", ["--delete", "--delete-after"])
@pytest.mark.parametrize("delete_flag", ["--delete", "--delete-commit", "--delete-after"])
def test_threads_delete_after_sends_keep_set(self, delete_flag):
source, dest, received = _seed_pair("mtkeep")
extra = os.path.join(received, "d", "old_extra")
@@ -465,7 +492,7 @@ class TestDeleteAfterThreadsKeepSet:
port=server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
assert not os.path.exists(extra), (
f"{delete_flag} --threads did not remove an extra: keep-set manifest was suppressed"
f"{delete_flag} --threads did not remove an extra: delete keep-set was suppressed"
)
class TestDeleteDelayMaxDeleteRefilledDir:
+196 -4
View File
@@ -28,6 +28,7 @@ from common import ( # noqa: E402
ServerManager,
TEST_DATA_DIR,
clean_dir,
get_dest_received_dir,
)
from parity_caveats import ASPECTS, caveat_for # noqa: E402
import parity_harness as H # noqa: E402
@@ -116,12 +117,32 @@ def seed_delete_excluded(_src, rroot, froot):
_mk(os.path.join(root, "keep.txt"), b"keep\n", _OLD_MTIME)
def seed_filter_protect(_src, rroot, froot):
"""Destination-only entries, including nested ones, for the receiver-side
`protect` rule: the `.log` extras must survive --delete, the rest go."""
for root in (rroot, froot):
_mk(os.path.join(root, "extra.log"), b"dest-only log\n", _OLD_MTIME)
_mk(os.path.join(root, "other.txt"), b"dest-only other\n", _OLD_MTIME)
_mk(os.path.join(root, "sub", "extra2.log"), b"nested dest-only log\n", _OLD_MTIME)
_mk(os.path.join(root, "sub", "other2.txt"), b"nested dest-only other\n", _OLD_MTIME)
def seed_max_delete(_src, rroot, froot):
for root in (rroot, froot):
_mk(os.path.join(root, "extra1.txt"), b"e1\n", _OLD_MTIME)
_mk(os.path.join(root, "extra2.txt"), b"e2\n", _OLD_MTIME)
def fuzzy_basis_seed(_src, rroot, froot):
"""Seed a same-suffix sibling whose name is one edit from the source and
whose content matches it, with a DIFFERENT mtime so rsync's exact
size+mtime pass cannot fire: both tools must select it via the
name-distance pass. Where the two tools' basis choices coincide the
block-level results are identical when the block size is pinned."""
for root in (rroot, froot):
_mk(os.path.join(root, "report_v1.txt"), H.FUZZY_PAYLOAD, _OLD_MTIME)
def max_delete_count_check(_src, rroot, froot, _rs, _fs):
"""The exact survivor set is order-dependent; the count must still match."""
r = H.snapshot(rroot)
@@ -197,6 +218,18 @@ _CASES = [
H.Case("chmod", "basic", ["-a", "--chmod=Fu+rwx"], compare_modes=True,
ci=True, ref="--chmod"),
# --- delta / similar-file basis (--fuzzy) -----------------------------
# Basis choices coincide here (same-suffix sibling, name distance one edit,
# content identical); with the block size pinned both tools report the same
# Matched/Literal/transferred counters. The residual (FastSync's narrower
# delta size window) is covered by TestFuzzy in test_parity_quickwins.py.
H.Case("fuzzy_basis", "fuzzy",
["-a", "--no-whole-file", "--fuzzy", "--stats", "-B8192"],
fastsync_flags=["-a", "--incremental", "--delta", "--fuzzy",
"--stats", "--delta-block=8192"],
seed=fuzzy_basis_seed, stdout=H.STDOUT_STATS, ci=True,
ref="-y/--fuzzy similar-file basis"),
# --- deletion ---------------------------------------------------------
H.Case("delete", "basic", ["-a", "--delete"], seed=seed_extras,
server_args=DELETE, ci=True, ref="--delete"),
@@ -208,13 +241,36 @@ _CASES = [
server_args=DELETE, ref="--delete-delay"),
H.Case("delete_after", "basic", ["-a", "--delete-after"], seed=seed_extras,
server_args=DELETE, ref="--delete-after"),
H.Case("delete_commit", "basic", ["-a", "--delete-after"], seed=seed_extras,
fastsync_flags=["-a", "--delete-commit"], server_args=DELETE,
ref="FastSync-only --delete-commit == rsync --delete-after"),
H.Case("delete_excluded", "filters",
["-a", "--delete", "--delete-excluded", "--exclude=*.log"],
seed=seed_delete_excluded, server_args=DELETE, ref="--delete-excluded"),
H.Case("exclude_protect_dest_only", "filters",
["-a", "--delete", "--exclude=*.log"],
seed=seed_delete_excluded, server_args=DELETE, ci=True,
ref="--delete protects a destination-only excluded entry like rsync"),
H.Case("max_delete", "basic", ["-a", "--delete", "--max-delete=1"],
seed=seed_max_delete, server_args=DELETE,
extra_check=max_delete_count_check, compare_tree=False,
ref="--max-delete"),
H.Case("filter_protect", "filters",
["-a", "--delete", "--filter=P *.log"],
seed=seed_filter_protect, server_args=DELETE, ci=True,
ref="--filter P/--protect receiver-side delete protection (default during)"),
H.Case("filter_protect_during", "filters",
["-a", "--delete-during", "--filter=P *.log"],
seed=seed_filter_protect, server_args=DELETE, ci=True,
ref="--filter P/--protect under --delete-during"),
H.Case("filter_protect_delay", "filters",
["-a", "--delete-delay", "--filter=P *.log"],
seed=seed_filter_protect, server_args=DELETE, ci=True,
ref="--filter P/--protect under --delete-delay"),
H.Case("filter_protect_after", "filters",
["-a", "--delete-after", "--filter=P *.log"],
seed=seed_filter_protect, server_args=DELETE, ci=True,
ref="--filter P/--protect under the whole-tree --delete-after commit"),
# --- relative / dirs --------------------------------------------------
H.Case("relative_general", "basic", ["-a", "-R"], layout=H.MIRROR_ABS,
@@ -324,7 +380,11 @@ def _result_aspects(result):
_STANDALONE_REFS = {
"incremental_modified": "-i/--itemize-changes + incremental second run",
"compare_dest": "--compare-dest",
"copy_dest": "--copy-dest",
"link_dest": "--link-dest",
"link_dest_stats": "--link-dest + --stats",
"verify_basis": "--verify-basis (FastSync-only)",
"verify_basis_default": "--verify-basis (default quick-check vs rsync)",
"added_and_deleted": "--delete across two runs",
"added_and_deleted_seed": "--delete across two runs",
"one_file_system": "-x/--one-file-system",
@@ -385,14 +445,17 @@ def test_compare_dest_skips_basis(parity_server_factory):
fdst = os.path.join(TEST_DATA_DIR, "parity_cmpd_fdst")
clean_dir(src)
_mk(os.path.join(src, "f.txt"), b"basis-content\n")
_pin(os.path.join(src, "f.txt"), _OLD_MTIME)
server = parity_server_factory(SUPER)
rel = os.path.abspath(src).lstrip(os.sep)
# rsync resolves --compare-dest relative to the destination dir; FastSync
# resolves it under the receive root and appends the mirrored source path.
# Both rely on rsync's size+mtime quick-check, so the basis mtime is pinned
# to the source's to keep the match deterministic across a second boundary.
def seed(_src, rroot, froot):
_mk(os.path.join(rroot, "basis", "f.txt"), b"basis-content\n")
_mk(os.path.join(fdst, "basis", rel, "f.txt"), b"basis-content\n")
_mk(os.path.join(rroot, "basis", "f.txt"), b"basis-content\n", _OLD_MTIME)
_mk(os.path.join(fdst, "basis", rel, "f.txt"), b"basis-content\n", _OLD_MTIME)
def extra(_src, rroot, froot, _rs, _fs):
out = []
@@ -420,12 +483,13 @@ def test_link_dest_hardlinks_basis(parity_server_factory):
fdst = os.path.join(TEST_DATA_DIR, "parity_linkd_fdst")
clean_dir(src)
_mk(os.path.join(src, "f.txt"), b"link-basis-content\n")
_pin(os.path.join(src, "f.txt"), _OLD_MTIME)
server = parity_server_factory(SUPER)
rel = os.path.abspath(src).lstrip(os.sep)
def seed(_src, rroot, froot):
_mk(os.path.join(rroot, "basis", "f.txt"), b"link-basis-content\n")
_mk(os.path.join(fdst, "basis", rel, "f.txt"), b"link-basis-content\n")
_mk(os.path.join(rroot, "basis", "f.txt"), b"link-basis-content\n", _OLD_MTIME)
_mk(os.path.join(fdst, "basis", rel, "f.txt"), b"link-basis-content\n", _OLD_MTIME)
def extra(_src, rroot, froot, _rs, _fs):
r_basis = os.stat(os.path.join(rroot, "basis", "f.txt")).st_ino
@@ -448,6 +512,134 @@ def test_link_dest_hardlinks_basis(parity_server_factory):
_run_and_check(case_id, result)
@requires_rsync
@parity
def test_link_dest_stats_matches_rsync(parity_server_factory):
"""A basis hit must not be counted as created or literal data: rsync reports
zero for both, so FastSync's receiver tallies must too (regression for the
basis materialization over-report)."""
case_id = "link_dest_stats"
src = os.path.join(TEST_DATA_DIR, "parity_linkds_src")
rdst = os.path.join(TEST_DATA_DIR, "parity_linkds_rdst")
fdst = os.path.join(TEST_DATA_DIR, "parity_linkds_fdst")
clean_dir(src)
_mk(os.path.join(src, "f.txt"), b"link-basis-content\n")
_pin(os.path.join(src, "f.txt"), _OLD_MTIME)
server = parity_server_factory(SUPER)
rel = os.path.abspath(src).lstrip(os.sep)
def seed(_src, rroot, froot):
_mk(os.path.join(rroot, "basis", "f.txt"), b"link-basis-content\n", _OLD_MTIME)
_mk(os.path.join(fdst, "basis", rel, "f.txt"), b"link-basis-content\n", _OLD_MTIME)
result = H.run_differential(
src, rdst, fdst,
["-a", "--link-dest=basis", "--stats"],
["-a", f"--link-dest={os.path.join(fdst, 'basis')}", "--incremental", "--stats"],
server, seed=seed, ignore_paths=("basis",), stdout=H.STDOUT_STATS)
_run_and_check(case_id, result, ref="--link-dest + --stats")
@requires_rsync
@parity
def test_copy_dest_copies_basis(parity_server_factory):
"""--copy-dest: a basis match is materialized as an independent copy with the
source's attributes, matching rsync (copy then fix attributes)."""
case_id = "copy_dest"
src = os.path.join(TEST_DATA_DIR, "parity_copyd_src")
rdst = os.path.join(TEST_DATA_DIR, "parity_copyd_rdst")
fdst = os.path.join(TEST_DATA_DIR, "parity_copyd_fdst")
clean_dir(src)
_mk(os.path.join(src, "f.txt"), b"copy-basis-content\n")
_pin(os.path.join(src, "f.txt"), 1_600_000_000)
os.chmod(os.path.join(src, "f.txt"), 0o755)
server = parity_server_factory(SUPER)
rel = os.path.abspath(src).lstrip(os.sep)
def seed(_src, rroot, froot):
# Basis content matches the source; give the basis a different mode so a
# wrong "keep basis attributes" implementation is visible.
_mk(os.path.join(rroot, "basis", "f.txt"), b"copy-basis-content\n",
1_600_000_000)
os.chmod(os.path.join(rroot, "basis", "f.txt"), 0o644)
_mk(os.path.join(fdst, "basis", rel, "f.txt"), b"copy-basis-content\n",
1_600_000_000)
os.chmod(os.path.join(fdst, "basis", rel, "f.txt"), 0o644)
def extra(_src, rroot, froot, _rs, _fs):
out = []
bases = {"rsync": os.path.join(rroot, "basis", "f.txt"),
"fastsync": os.path.join(fdst, "basis", rel, "f.txt")}
for label, root in (("rsync", rroot), ("fastsync", froot)):
target = os.path.join(root, "f.txt")
if not os.path.exists(target):
out.append(f"{label}: f.txt missing")
continue
if os.stat(target).st_ino == os.stat(bases[label]).st_ino:
out.append(f"{label}: f.txt is hard-linked, not copied")
if (os.stat(target).st_mode & 0o777) != 0o755:
out.append(f"{label}: f.txt mode "
f"{oct(os.stat(target).st_mode & 0o777)} != 0o755")
return out
result = H.run_differential(
src, rdst, fdst,
["-a", "--copy-dest=basis"],
["-a", f"--copy-dest={os.path.join(fdst, 'basis')}", "--incremental"],
server, seed=seed, ignore_paths=("basis",), extra_check=extra,
compare_modes=True)
_run_and_check(case_id, result)
@requires_rsync
@parity
def test_verify_basis_restores_strict_content(parity_server_factory):
"""Default matches rsync's metadata quick-check; FastSync-only
`--verify-basis` restores strict content equality and transfers the source
when a same-size/different-content basis would otherwise be trusted."""
case_id = "verify_basis"
src = os.path.join(TEST_DATA_DIR, "parity_vbasis_src")
rdst = os.path.join(TEST_DATA_DIR, "parity_vbasis_rdst")
fdst = os.path.join(TEST_DATA_DIR, "parity_vbasis_fdst")
clean_dir(src)
_mk(os.path.join(src, "f.txt"), b"AAAA\n")
_pin(os.path.join(src, "f.txt"), _OLD_MTIME)
server = parity_server_factory(SUPER)
rel = os.path.abspath(src).lstrip(os.sep)
def seed(_src, rroot, froot):
# Same size and mtime as the source, different bytes: a metadata
# quick-check trusts it; --verify-basis must not.
for root, basis_rel in ((rroot, os.path.join("basis", "f.txt")),
(fdst, os.path.join("basis", rel, "f.txt"))):
_mk(os.path.join(root, basis_rel), b"BBBB\n", _OLD_MTIME)
# Default: both tools trust the basis (rsync's quick check), so the
# destination carries the basis bytes and the trees match.
result = H.run_differential(
src, rdst, fdst,
["-a", "--link-dest=basis"],
["-a", f"--link-dest={os.path.join(fdst, 'basis')}", "--incremental"],
server, seed=seed, ignore_paths=("basis",))
_run_and_check(case_id + "_default", result)
# --verify-basis (FastSync only): the digest mismatch rejects the basis and
# the source is transferred, so the destination is the source bytes. rsync
# has no such flag; assert the FastSync outcome directly against the source.
fdst2 = os.path.join(TEST_DATA_DIR, "parity_vbasis_fdst2")
clean_dir(fdst2)
for root, basis_rel in ((fdst2, os.path.join("basis", rel, "f.txt")),):
_mk(os.path.join(root, basis_rel), b"BBBB\n", _OLD_MTIME)
result, _ = H.run_fastsync(src, fdst2,
["-a", f"--link-dest={os.path.join(fdst2, 'basis')}",
"--incremental", "--verify-basis"], server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
target = os.path.join(get_dest_received_dir(fdst2, src), "f.txt")
with open(target, "rb") as fh:
assert fh.read() == b"AAAA\n", \
"--verify-basis must reject the same-size/different-content basis"
@requires_rsync
@parity
def test_added_and_deleted_between_runs(parity_server_factory):
+117 -53
View File
@@ -589,11 +589,10 @@ class TestRemoteDryRun:
Covers the three cases that a real run protects: the file being updated
(in the keep set), a filter-excluded source entry (protected prefix), and
a --max-size-pruned source entry (always-protected prefix). Only the
genuine destination-only extras may appear. Residual: a destination-only
entry matching an exclude pattern is still removed (FastSync derives
delete protection from the source scan, not a receiver filter engine);
that divergence is pinned by TestOptionParity.
a --max-size-pruned source entry (always-protected prefix). Track 4a
adds a fourth: a destination-only entry matching the exclude rule is
re-derived on the receiver and also protected, so only the genuine
destination-only `extra.txt` appears.
"""
source = os.path.join(TEST_DATA_DIR, "dryrep_src")
rdst = os.path.join(TEST_DATA_DIR, "dryrep_rdst")
@@ -610,7 +609,8 @@ class TestRemoteDryRun:
os.makedirs(received, exist_ok=True)
for root in (rdst, received):
for name, data in (("a.txt", b"old\n"), ("keep.log", b"log\n"),
("big.bin", b"B" * 2000), ("extra.txt", b"extra\n")):
("big.bin", b"B" * 2000), ("extra.txt", b"extra\n"),
("stray.log", b"dest only\n")):
with open(os.path.join(root, name), "wb") as fh:
fh.write(data)
os.utime(os.path.join(root, name), (1_500_000_000, 1_500_000_000))
@@ -631,6 +631,8 @@ class TestRemoteDryRun:
fs_del = sorted(l for l in (result.stdout or "").splitlines()
if l.startswith("*deleting"))
assert fs_del == rsync_del, f"rsync={rsync_del}\nfastsync={fs_del}"
assert os.path.exists(os.path.join(received, "stray.log")), \
"destination-only exclude match must be protected in the dry-run report"
@pytest.mark.ci
def test_remote_dry_run_quiet_is_silent(self, shared_server):
@@ -3980,15 +3982,16 @@ class TestDeleteTiming:
assert _read_file(os.path.join(received, "sub", "deep.txt")) == b"deeply nested file\n", \
f"{flag}: nested file was not written after the early deletion"
@pytest.mark.parametrize("flag", ["--delete", "--delete-after"])
@pytest.mark.parametrize("flag", ["--delete-commit", "--delete-after"])
@pytest.mark.parametrize("mt", [False, True])
def test_late_flags_commit_only_after_success(self, flag, mt):
"""Plain --delete/--delete-after defer deletion until the whole transfer
"""--delete-commit/--delete-after defer deletion until the whole transfer
succeeds: a mid-transfer write failure must leave every extra in place
(commit-style safety). The -m receiver must also keep the extras: the
deferred keep-set is committed by the server only after the disk-writer
thread has finished, and a failing writer means the manifest is freed,
never applied."""
(commit-style safety). Plain --delete no longer defers (it defaults to
delete-during), so only the explicitly late timings are exercised here.
The --threads receiver must also keep the extras: the deferred keep-set is
committed by the server only after the disk-writer thread has finished,
and a failing writer means the manifest is freed, never applied."""
source = self._seed("late")
dest = os.path.join(TEST_DATA_DIR, "deltiming_late_dst")
clean_dir(dest)
@@ -4693,11 +4696,11 @@ class TestDeletePolicy:
finally:
os.chmod(source, 0o755)
def test_delete_excluded_protection_is_sender_derived(self):
def test_delete_protection_reapplied_on_receiver(self):
"""Plain --delete protects destination mirrors of files the SOURCE scan
excluded, but a destination-only file that merely matches an exclude
rule is still an extra and is removed (protection never re-applies rules
to the destination)."""
excluded, and (track 4a) also protects a destination-only file matching
an exclude rule because the compiled rule set is re-applied on the
receiver, matching rsync."""
source = os.path.join(TEST_DATA_DIR, "senderderived_src")
clean_dir(source)
self._write(os.path.join(source, "keep.txt"), b"kept\n")
@@ -4717,8 +4720,8 @@ class TestDeletePolicy:
f"delete sync failed: {(result.stderr or result.stdout)[:300]}"
assert os.path.exists(os.path.join(received, "secret.log")), \
"source-excluded mirror was deleted under plain --delete"
assert not os.path.exists(os.path.join(received, "stray.log")), \
"destination-only file matching the exclude rule was left (should be deleted)"
assert os.path.exists(os.path.join(received, "stray.log")), \
"destination-only file matching the exclude rule must be protected like rsync"
def _pin_mtime(path, ts):
@@ -4738,11 +4741,12 @@ class TestBasisDestDirs:
STAGING = ".fastsync-stage"
TS = 1577836800 # 2020-01-01 00:00:00 UTC, used to pin matching mtimes
# fixture files: source and basis share the mtime pin, so a basis "match"
# is decided purely by content (xxHash). unchanged.txt is byte-identical;
# changed.txt is byte-DIFFERENT but has the SAME SIZE as the source (and
# the same pinned mtime), which is what forces the content-hash gate;
# added.txt does not exist in the basis at all.
# fixture files: source and basis share the mtime pin, so the DEFAULT
# (rsync-parity) quick-check is a size+mtime match and trusts the basis even
# when the body differs. unchanged.txt is byte-identical; changed.txt is
# byte-DIFFERENT but has the SAME SIZE as the source (and the same pinned
# mtime), which is what the FastSync-only --verify-basis content gate
# rejects; added.txt does not exist in the basis at all.
UNCHANGED = "unchanged.txt"
CHANGED = "changed.txt"
ADDED = "added.txt"
@@ -4782,18 +4786,19 @@ class TestBasisDestDirs:
}
def _basis_tree(self, prefix):
# unchanged.txt is identical to the source; changed.txt has the SAME
# byte size and pinned mtime but a different body (equal size forces
# the xxHash gate); added.txt is missing from the basis.
# unchanged.txt is identical to the source; changed.txt has a DIFFERENT
# size (and body) so the size leg of the quick-check fails and it is
# transferred normally; added.txt is missing from the basis.
return {
self.UNCHANGED: b"stable content v1\n",
self.CHANGED: b"CHANGED CONTENT NOW\n",
self.CHANGED: b"CHANGED CONTENT NOW AND LONGER\n",
}
def test_same_size_different_content_is_not_a_basis_match(self, shared_server):
# Core safety property: equal size + pinned mtime but different content
# must NEVER be hard-linked or copied from the basis -- the xxHash gate
# rejects it and the sender's data is transferred instead.
def test_same_size_different_content_default_trusts_quick_check(self, shared_server):
# Default rsync-parity behavior: equal size + pinned mtime is a basis
# match, so the basis body is materialized/linked without reading it.
# This mirrors rsync 3.4.1's quick check (differential-tested in
# test_differential_parity.py::test_verify_basis_restores_strict_content).
for flag, basis_dir in (("--link-dest", "szlb"), ("--copy-dest", "szcp"),
("--compare-dest", "szcmp")):
source = self._make_source("basis_same_size_src",
@@ -4805,7 +4810,36 @@ class TestBasisDestDirs:
result, _ = run_client(source, dest, flags=[f"{flag}={basis_dir}"],
port=shared_server.port)
assert result.returncode == 0, \
f"{flag} same-size mismatch failed: {result.stderr[:300]}"
f"{flag} same-size quick-check failed: {result.stderr[:300]}"
received = get_dest_received_dir(dest, source)
dest_file = os.path.join(received, self.UNCHANGED)
if flag == "--compare-dest":
assert not os.path.exists(dest_file), \
f"{flag}: compare-dest must leave a matching file sparse"
else:
assert _read_file(dest_file) == b"SAME LENGTH BODY!", \
f"{flag}: default quick-check did not trust the basis body"
if flag == "--link-dest":
assert os.stat(dest_file).st_ino == os.stat(basis_file).st_ino, \
f"{flag}: basis was not hard-linked"
def test_verify_basis_rejects_same_size_different_content(self, shared_server):
# FastSync-only --verify-basis: the whole-file digest gate rejects the
# same-size/different-content basis, so the source data is transferred
# instead of the wrong basis bytes.
for flag, basis_dir in (("--link-dest", "vszlb"), ("--copy-dest", "vszcp"),
("--compare-dest", "vszcmp")):
source = self._make_source("basis_verify_src",
{self.UNCHANGED: b"same length body\n"})
dest = os.path.join(TEST_DATA_DIR, f"basis_verify_dst_{basis_dir}")
clean_dir(dest)
basis_file = self._seed_basis_file(dest, source, basis_dir, self.UNCHANGED,
b"SAME LENGTH BODY!")
result, _ = run_client(source, dest,
flags=[f"{flag}={basis_dir}", "--verify-basis"],
port=shared_server.port)
assert result.returncode == 0, \
f"{flag} --verify-basis failed: {result.stderr[:300]}"
received = get_dest_received_dir(dest, source)
dest_file = os.path.join(received, self.UNCHANGED)
assert _read_file(dest_file) == b"same length body\n", \
@@ -4837,11 +4871,13 @@ class TestBasisDestDirs:
self._source_tree("c")[self.ADDED], "added file not transferred"
@pytest.mark.ci
def test_dry_run_compare_dest_does_not_read_basis(self, shared_server):
# A dry-run --compare-dest must never read/hash the basis file: doing so
# is a 1-bit content oracle against the client-supplied digest. Even a
# byte-identical basis with a matching size+mtime is therefore reported
# as would-transfer, and nothing is created.
def test_dry_run_compare_dest_quick_check_does_not_read_basis(self, shared_server):
# A dry-run --compare-dest must never read/hash the basis file. Under
# the default metadata quick-check a matching basis is reported as a
# skip (matching rsync) without reading it; nothing is created. Under
# --verify-basis, which would require hashing, the dry-run cannot
# confirm the hit (that would be a 1-bit content oracle) and reports
# would-transfer instead.
source = self._make_source("basis_dry_src", {self.UNCHANGED: b"stable content v1\n"})
dest = os.path.join(TEST_DATA_DIR, "basis_dry_dst")
clean_dir(dest)
@@ -4852,11 +4888,30 @@ class TestBasisDestDirs:
port=shared_server.port)
assert result.returncode == 0, \
f"dry-run compare-dest failed: {result.stderr[:300]}"
assert self.UNCHANGED in result.stdout, (
"dry-run compare-dest silently skipped: receiver read the basis content"
assert self.UNCHANGED not in result.stdout, (
"dry-run compare-dest did not honor the metadata quick-check "
"(reported would-transfer for a matching basis)"
)
assert _snapshot_tree(dest) == before, "dry-run compare-dest mutated the destination"
# --verify-basis: the hit needs the basis content, which a dry-run must
# not read, so the file is reported as would-transfer.
dest2 = os.path.join(TEST_DATA_DIR, "basis_dry_verify_dst")
clean_dir(dest2)
self._seed_basis(dest2, source, "drybasis", {self.UNCHANGED: b"stable content v1\n"})
before2 = _snapshot_tree(dest2)
result, _ = run_client(source, dest2,
flags=["--compare-dest=drybasis", "--dry-run",
"--verify-basis"],
port=shared_server.port)
assert result.returncode == 0, \
f"dry-run --verify-basis compare-dest failed: {result.stderr[:300]}"
assert self.UNCHANGED in result.stdout, (
"dry-run --verify-basis must not read the basis to confirm a hit"
)
assert _snapshot_tree(dest2) == before2, \
"dry-run --verify-basis compare-dest mutated the destination"
def test_compare_dest_content_mismatch_forces_transfer(self, shared_server):
# The basis holds a file with a DIFFERENT body: even though it shares
# the mtime pin, the xxHash check fails and the data must be sent.
@@ -5095,27 +5150,36 @@ class TestBasisDestDirs:
assert os.stat(dest_file).st_ino != os.stat(basis_file).st_ino, \
"--ignore-times must not hard-link to a basis file"
def test_basis_refuses_file_above_whole_file_limit(self, shared_server):
# Every whole-file payload path in FastSync (basis dirs included) is
# bounded by MAX_RECEIVE_WHOLE_FILE_SIZE. rsync supports basis dirs for
# arbitrary sizes; FastSync refuses such a run up front with a clear
# diagnostic instead of letting the receiver abort the whole transfer
# mid-stream with no client-side explanation.
def test_basis_handles_file_above_whole_file_limit(self, shared_server):
# Track 5a: a basis hit streams the copy (and the --verify-basis digest
# streams the basis), so a source larger than the whole-file payload
# bound is supported for basis dirs exactly like rsync. A basis MISS
# still falls back to the normal transfer, which keeps its own bound.
source = self._make_source("basis_oversize_src", {"small.txt": b"ok\n"})
big = os.path.join(source, "huge.bin")
with open(big, "wb") as fh:
os.ftruncate(fh.fileno(), 256 * 1024 * 1024 + 4096)
dest = os.path.join(TEST_DATA_DIR, "basis_oversize_dst")
clean_dir(dest)
result, _ = run_client(source, dest, flags=["--link-dest=nope"],
port=shared_server.port)
assert result.returncode != 0, \
"basis run with an over-limit file unexpectedly succeeded"
assert "larger than" in result.stderr, \
f"no clear over-limit diagnostic: {result.stderr[:300]}"
received = get_dest_received_dir(dest, source)
assert not os.path.exists(received), \
"over-limit basis run transferred files before failing"
rel = os.path.relpath(received, dest)
basis_big = os.path.join(dest, "ob", rel, "huge.bin")
os.makedirs(os.path.dirname(basis_big), exist_ok=True)
shutil.copyfile(big, basis_big)
os.utime(basis_big, (self.TS, self.TS))
os.utime(big, (self.TS, self.TS))
result, _ = run_client(source, dest,
flags=["--link-dest=ob", "--incremental"],
port=shared_server.port)
assert result.returncode == 0, \
f"over-limit basis run failed: {result.stderr[:300]}"
dest_big = os.path.join(received, "huge.bin")
assert os.path.exists(dest_big), "over-limit basis hit was not materialized"
assert os.path.getsize(dest_big) == 256 * 1024 * 1024 + 4096
assert os.stat(dest_big).st_ino == os.stat(basis_big).st_ino, \
"over-limit --link-dest did not hard-link to the basis"
assert _read_file(os.path.join(received, "small.txt")) == b"ok\n"
def _random_payloads(size=2 * 1024 * 1024, changed=64 * 1024, seed=1234):
+31 -10
View File
@@ -478,14 +478,14 @@ class TestRemoteOptionDaemon:
assert "remote-option" in (result.stderr + result.stdout)
class TestFilterProtectDivergence:
"""Documented residual: a protect rule that matches only a destination-only
entry is not re-derived on the receiver (FastSync derives delete protection
from the source scan), so rsync protects the extra but FastSync removes it."""
class TestFilterProtect:
"""Receiver-derived delete protection: a `protect`/`P` rule is compiled by
the sender and sent on the config frame, so the receiver shields a
destination-only entry that never appeared on the sender, matching rsync."""
@requires_rsync
@pytest.mark.ci
def test_protect_dest_only_divergence(self, shared_server):
def test_protect_dest_only_matches_rsync(self, shared_server):
source = os.path.join(TEST_DATA_DIR, "fpd_src")
dest = os.path.join(TEST_DATA_DIR, "fpd_dst")
rdst = os.path.join(TEST_DATA_DIR, "fpd_rdst")
@@ -498,6 +498,7 @@ class TestFilterProtectDivergence:
rsync_result = _rsync(["-a", "--delete", "--filter=P *.log", source + "/", rdst + "/"])
assert rsync_result.returncode == 0, rsync_result.stderr
assert os.path.exists(os.path.join(rdst, "extra.log")), "rsync did not protect extra.log"
assert not os.path.exists(os.path.join(rdst, "other.txt")), "rsync did not delete other.txt"
clean_dir(dest)
received = get_dest_received_dir(dest, source)
@@ -509,9 +510,29 @@ class TestFilterProtectDivergence:
flags=["-a", "--delete", "--filter=P *.log"],
port=server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
# Pin the known divergence: FastSync deletes the destination-only file.
assert not os.path.exists(os.path.join(received, "extra.log")), (
"FastSync now protects destination-only P matches; the --filter row may be "
"upgradable to full parity"
)
assert os.path.exists(os.path.join(received, "extra.log")), (
"FastSync must protect a destination-only P match like rsync")
assert not os.path.exists(os.path.join(received, "other.txt"))
@pytest.mark.ci
def test_protect_dest_only_dry_run_enumeration(self, shared_server):
source = os.path.join(TEST_DATA_DIR, "fpd_nd_src")
dest = os.path.join(TEST_DATA_DIR, "fpd_nd_dst")
clean_dir(source)
_write(os.path.join(source, "keep.txt"), b"keep\n")
received = get_dest_received_dir(dest, source)
clean_dir(received)
_write(os.path.join(received, "keep.txt"), b"keep\n")
_write(os.path.join(received, "extra.log"), b"extra\n")
_write(os.path.join(received, "other.txt"), b"other\n")
with ServerManager() as server:
server.start(extra_args=["--allow-delete"])
result, _ = run_client(source, dest,
flags=["-a", "-n", "--delete", "--out-format=%n",
"--filter=P *.log"],
port=server.port)
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
assert "other.txt" in result.stdout, result.stdout
assert "extra.log" not in result.stdout, result.stdout
assert os.path.exists(os.path.join(received, "extra.log"))
assert os.path.exists(os.path.join(received, "other.txt"))
+129 -7
View File
@@ -719,13 +719,18 @@ class TestVerifyAndFlip:
source = self._src("cmpd")
dest = self._dst("cmpd")
rdst = self._dst("cmpd_r")
# Pin the mtime so rsync's size+mtime quick-check (and FastSync's
# default) matches deterministically across a second boundary.
OLD = 1_500_000_000
with open(os.path.join(source, "f.txt"), "wb") as fh:
fh.write(b"basis-content\n")
os.utime(os.path.join(source, "f.txt"), (OLD, OLD))
# rsync resolves --compare-dest relative to the destination dir; its
# basis file sits at the transfer-relative path.
os.makedirs(os.path.join(rdst, "basis"), exist_ok=True)
with open(os.path.join(rdst, "basis", "f.txt"), "wb") as fh:
fh.write(b"basis-content\n")
os.utime(os.path.join(rdst, "basis", "f.txt"), (OLD, OLD))
rs = _rsync(["-a", "--compare-dest=basis", source + "/", rdst + "/"])
assert rs.returncode == 0, rs.stderr
assert not os.path.exists(os.path.join(rdst, "f.txt")), \
@@ -738,6 +743,7 @@ class TestVerifyAndFlip:
os.makedirs(basis, exist_ok=True)
with open(os.path.join(basis, "f.txt"), "wb") as fh:
fh.write(b"basis-content\n")
os.utime(os.path.join(basis, "f.txt"), (OLD, OLD))
received = get_dest_received_dir(dest, source)
result, _ = run_client(source, dest,
flags=["--compare-dest=basis", "--incremental"],
@@ -751,14 +757,17 @@ class TestVerifyAndFlip:
def test_link_dest_hardlinks_matches_rsync(self, shared_server):
source = self._src("linkd")
dest = self._dst("linkd")
OLD = 1_500_000_000
with open(os.path.join(source, "f.txt"), "wb") as fh:
fh.write(b"link-basis-content\n")
os.utime(os.path.join(source, "f.txt"), (OLD, OLD))
rel = os.path.abspath(source).lstrip(os.sep)
basis = os.path.join(dest, "basis", rel)
os.makedirs(basis, exist_ok=True)
basis_file = os.path.join(basis, "f.txt")
with open(basis_file, "wb") as fh:
fh.write(b"link-basis-content\n")
os.utime(basis_file, (OLD, OLD))
received = get_dest_received_dir(dest, source)
result, _ = run_client(source, dest,
flags=["--link-dest=basis", "--incremental"],
@@ -771,12 +780,11 @@ class TestVerifyAndFlip:
@requires_rsync
def test_basis_dir_size_only_content_residual(self, shared_server):
"""Documented residual (RSYNC_COMPAT.md basis-dir rows): FastSync
xxHash-verifies a basis hit, while rsync's `--size-only` quick check
trusts the size alone. With a same-size, different-content basis,
rsync links/copies the wrong basis content while FastSync transfers the
source. This test pins both observed behaviors (FastSync is stricter,
so the rows are reclassified Divergent)."""
"""rsync parity (default): a basis hit is decided by the metadata
quick-check alone. With `--size-only`, a same-size, different-content
basis is trusted, so rsync links the basis content and FastSync must now
do the same instead of xxHash-verifying it. `--verify-basis` restores
the stricter content equality (covered by the differential test)."""
source = self._src("basissz")
rdest = self._dst("basissz_r")
fdest = self._dst("basissz_f")
@@ -806,9 +814,123 @@ class TestVerifyAndFlip:
flags=["-a", "--size-only", "--link-dest=basis", "--incremental"],
port=shared_server.port)
assert result.returncode == 0, result.stderr[:300]
with open(os.path.join(received, "f.txt"), "rb") as fh:
assert fh.read() == b"BBBB\n", \
"FastSync must trust the metadata quick-check exactly like rsync"
@requires_rsync
def test_verify_basis_restores_content_check(self, shared_server):
"""FastSync-only `--verify-basis`: a same-size, same-mtime basis with
different content is rejected by the whole-file digest, so the source is
transferred instead of installing the wrong basis bytes. The default
(no flag) installs the basis content, matching rsync."""
source = self._src("vbasis")
fdest = self._dst("vbasis_f")
with open(os.path.join(source, "f.txt"), "wb") as fh:
fh.write(b"AAAA\n")
OLD = 1_400_000_000
os.utime(os.path.join(source, "f.txt"), (OLD, OLD))
rel = os.path.abspath(source).lstrip(os.sep)
basis = os.path.join(fdest, "basis", rel)
os.makedirs(basis, exist_ok=True)
with open(os.path.join(basis, "f.txt"), "wb") as fh:
fh.write(b"BBBB\n")
os.utime(os.path.join(basis, "f.txt"), (OLD, OLD))
received = get_dest_received_dir(fdest, source)
result, _ = run_client(source, fdest,
flags=["-a", "--link-dest=basis", "--incremental",
"--verify-basis"],
port=shared_server.port)
assert result.returncode == 0, result.stderr[:300]
with open(os.path.join(received, "f.txt"), "rb") as fh:
assert fh.read() == b"AAAA\n", \
"FastSync must verify the basis content and transfer the source"
"--verify-basis must reject the same-size/different-content basis"
def _stat_bytes(output, key):
"""Parse a --stats byte counter (e.g. ``Matched data: 65,536 bytes``)."""
for line in output.splitlines():
if line.startswith(key + ":"):
raw = line.split(":", 1)[1].strip().split()[0]
return int(raw.replace(",", ""))
return None
class TestFuzzy:
"""Track 5b: `-y`/`--fuzzy` is an internal bandwidth optimization with a
byte-exact result. FastSync ports rsync 3.4.1's weighted-Levenshtein name
heuristic, so where both delta engines admit the candidate the tools pick
the same basis (the ``fuzzy_basis`` differential asserts the tree and the
Matched/Literal counters match with the block size pinned). The residual is
candidate ELIGIBILITY: FastSync's delta size gate (both files >= 16 KiB and
a <= 10x size ratio) is narrower than rsync's, which empirically uses a
fuzzy basis well beyond 10x and below 16 KiB. These tests pin the window
boundary and prove the byte-exact fallback on both sides of it."""
_BASE = b"the quick brown fox jumps over the lazy dog\n" * 4000
def _src(self, tag):
source = os.path.join(TEST_DATA_DIR, f"fz_{tag}_src")
clean_dir(source)
return source
def _dst(self, tag):
d = os.path.join(TEST_DATA_DIR, f"fz_{tag}_dst")
clean_dir(d)
return d
def _run_both(self, shared_server, source, dest, rdst, payload, sibling,
rs_extra=(), fs_extra=()):
with open(os.path.join(source, "report_v2.txt"), "wb") as fh:
fh.write(payload)
for root in (rdst, get_dest_received_dir(dest, source)):
os.makedirs(root, exist_ok=True)
with open(os.path.join(root, "report_v1.txt"), "wb") as fh:
fh.write(sibling)
rs = _rsync(["-a", "--no-whole-file", "--fuzzy", "--stats"] +
list(rs_extra) + [source + "/", rdst + "/"])
assert rs.returncode == 0, rs.stderr[:300]
result, _ = run_client(
source, dest,
flags=["-a", "--incremental", "--delta", "--fuzzy", "--stats"] +
list(fs_extra),
port=shared_server.port)
assert result.returncode == 0, result.stderr[:300]
_assert_same_tree(rdst, get_dest_received_dir(dest, source), "(--fuzzy)")
return rs, result
@requires_rsync
def test_fuzzy_above_size_window_declines_but_tree_exact(self, shared_server):
"""A sibling >10x the source is used by rsync but declined by FastSync's
delta size-ratio gate; both destinations stay byte-identical."""
n = 65536
payload = (self._BASE * ((n // len(self._BASE)) + 1))[:n]
sibling = (self._BASE * 200)[: n * 20]
source, dest, rdst = (self._src("big"), self._dst("big"),
self._dst("big_r"))
rs, result = self._run_both(shared_server, source, dest, rdst,
payload, sibling)
assert _stat_bytes(rs.stdout, "Matched data") > 0, \
"rsync should still use a >10x fuzzy basis"
assert _stat_bytes(result.stdout, "Matched data") == 0, \
"FastSync's 10x delta size-ratio gate must decline the oversized basis"
assert _stat_bytes(result.stdout, "Literal data") == n
@requires_rsync
def test_fuzzy_below_delta_minimum_declines_but_tree_exact(self, shared_server):
"""A sibling below the 16 KiB delta minimum is used by rsync but never
enters FastSync's delta/fuzzy path; both trees stay byte-identical."""
n = 8192
payload = (self._BASE * ((n // len(self._BASE)) + 1))[:n]
source, dest, rdst = (self._src("small"), self._dst("small"),
self._dst("small_r"))
rs, result = self._run_both(shared_server, source, dest, rdst,
payload, payload)
assert _stat_bytes(rs.stdout, "Matched data") > 0, \
"rsync applies --fuzzy below 16 KiB"
assert _stat_bytes(result.stdout, "Matched data") == 0, \
"FastSync's 16 KiB delta minimum must bypass the fuzzy basis"
assert _stat_bytes(result.stdout, "Literal data") == n
class TestIgnoreExistingShortCircuit:
+2 -2
View File
@@ -118,8 +118,8 @@ class TestProtocol:
shutil.rmtree(dest, ignore_errors=True)
os.makedirs(dest)
_seed_protocol_source(source)
for bad in ("2.22.0", "2.21.0", "2.20.0", "2.19.0", "2.18.0", "2.17.0", "2.15.0", "2.16.0",
"216", "31"):
for bad in ("2.27.0", "2.26.0", "2.25.0", "2.24.0", "2.23.0", "2.22.0", "2.21.0", "2.20.0",
"2.19.0", "2.18.0", "2.17.0", "2.15.0", "2.16.0", "216", "31"):
result, _ = run_client(source, dest, flags=[f"--protocol={bad}"],
port=shared_server.port)
assert result.returncode != 0, f"--protocol={bad} should be rejected"
+77
View File
@@ -340,6 +340,7 @@ static void test_parse_args_protocol_accept_current() {
static void test_parse_args_protocol_rejects_other_versions() {
static const char* const bad_versions[] = {"2.17", "2.16", "2.15.0", "2.16.0", "2.17.0",
"2.18.0", "2.19.0", "2.20.0", "2.21.0", "2.22.0",
"2.23.0", "2.24.0", "2.25.0", "2.26.0", "2.27.0",
"216", "31", "abc", ""};
for (size_t i = 0; i < sizeof(bad_versions) / sizeof(bad_versions[0]); i++) {
Config* cfg = valid_client_config();
@@ -1045,6 +1046,23 @@ static void test_parse_args_basis_dirs() {
config_delete(cfg);
}
/* --verify-basis (FastSync-only, long-only): default off; parses on as a plain
boolean and leaves the basis implications intact. */
static void test_parse_args_verify_basis() {
Config* cfg = config_create();
EXPECT_FALSE(cfg->verify_basis);
config_delete(cfg);
cfg = config_create();
int positional_args[2];
int positional_count = 0;
char* argv[] = {"fastsync", "--link-dest=prior", "--verify-basis", "/src", "/dst"};
EXPECT_EQ_INT(parse_args(cfg, 5, argv, positional_args, &positional_count), 0);
EXPECT_TRUE(cfg->verify_basis);
EXPECT_TRUE(config_has_basis(cfg));
config_delete(cfg);
}
/* Escaping or degenerate basis-dir values must be rejected up front (they would
resolve outside the destination root on the receiver); an absolute path is
accepted (rsync parity) and canonicalized with its leading '/' preserved. */
@@ -1132,6 +1150,63 @@ static void test_parse_args_delete_timing_flags() {
config_delete(cfg);
}
/* Plain --delete with no explicit timing defaults to delete-during, matching
* rsync's --del default (progressive deletion). --delete-commit is the
* FastSync-only long spelling that selects rsync's --delete-after timing (the
* late whole-tree commit), and an explicit timing always wins over the default.
*/
static void test_parse_args_delete_default_timing_and_commit() {
Config* cfg = config_create();
char* argv[] = {"fastsync", "--delete", "/src", "/dst"};
int positional_args[2];
int positional_count = 0;
EXPECT_EQ_INT(parse_args(cfg, 4, argv, positional_args, &positional_count), 0);
EXPECT_TRUE(cfg->use_delete);
EXPECT_TRUE(cfg->delete_during);
EXPECT_FALSE(cfg->delete_before);
EXPECT_FALSE(cfg->delete_delay);
EXPECT_FALSE(cfg->delete_after);
cfg->send_directory = str_dup("/src");
cfg->receive_root_directory = str_dup("/dst");
EXPECT_TRUE(validate_config(cfg));
config_delete(cfg);
/* --delete-commit selects the late whole-tree commit (delete_after) and
implies --delete. */
cfg = config_create();
char* argv_commit[] = {"fastsync", "--delete-commit", "/src", "/dst"};
positional_count = 0;
EXPECT_EQ_INT(parse_args(cfg, 4, argv_commit, positional_args, &positional_count), 0);
EXPECT_TRUE(cfg->use_delete);
EXPECT_TRUE(cfg->delete_after);
EXPECT_FALSE(cfg->delete_before);
EXPECT_FALSE(cfg->delete_during);
EXPECT_FALSE(cfg->delete_delay);
cfg->send_directory = str_dup("/src");
cfg->receive_root_directory = str_dup("/dst");
EXPECT_TRUE(validate_config(cfg));
config_delete(cfg);
/* An explicit --delete-after alongside plain --delete keeps the late timing:
the default never overwrites an explicit timing. */
cfg = config_create();
char* argv_after[] = {"fastsync", "--delete", "--delete-after", "/src", "/dst"};
positional_count = 0;
EXPECT_EQ_INT(parse_args(cfg, 5, argv_after, positional_args, &positional_count), 0);
EXPECT_TRUE(cfg->use_delete);
EXPECT_TRUE(cfg->delete_after);
EXPECT_FALSE(cfg->delete_during);
config_delete(cfg);
/* --delete-commit conflicts with a different timing. */
cfg = config_create();
char* argv_conflict[] = {"fastsync", "--delete-commit", "--delete-during", "/src", "/dst"};
positional_count = 0;
EXPECT_EQ_INT(parse_args(cfg, 5, argv_conflict, positional_args, &positional_count), 0);
EXPECT_FALSE(validate_config(cfg));
config_delete(cfg);
}
/* Two different delete-timing flags on one command line are a conflict, not a
* silent last-one-wins choice. */
static void test_parse_args_delete_timing_conflict_rejected() {
@@ -4740,6 +4815,7 @@ void test_client_cli() {
test_parse_args_relative_no_implied_mkpath();
test_parse_args_delete_during_alias();
test_parse_args_delete_timing_flags();
test_parse_args_delete_default_timing_and_commit();
test_parse_args_delete_timing_conflict_rejected();
test_parse_args_delete_timing_without_delete_rejected();
test_parse_args_rejects_unimplemented_options();
@@ -4831,6 +4907,7 @@ void test_client_cli() {
test_parse_args_filter_rules();
test_parse_args_from0_cvs_filter_file_flags();
test_parse_args_basis_dirs();
test_parse_args_verify_basis();
test_parse_args_basis_invalid_paths();
test_validate_config_basis_rejects_chunk_serialization();
test_parse_args_delete_policy_flags();
+103 -6
View File
@@ -1,6 +1,7 @@
#include "test_config.h"
#include "config.h"
#include "delta.h"
#include "filter.h"
#include "identity.h"
#include "multiprocessing.h"
#include "protocol.h"
@@ -2119,6 +2120,48 @@ static void test_config_receive_rejects_oversized_string_budget() {
config_delete(over_bytes);
}
/* Pre-auth bounds for the receiver-side filter rule block (protocol 2.28.0).
A peer may send `protect`/`risk` rules; the receiver must reject an over-cap
count or an over-long pattern before evaluating anything, so a crafted config
cannot drive unbounded glob work or install a rule that silently never
matches. */
static void test_config_receive_rejects_bad_protect_rules() {
if (is_running_under_valgrind())
return;
/* Over-cap rule count: one more than MAX_FILTER_RULES rules. */
Config* c = config_create();
EXPECT_NOT_NULL(c);
c->send_directory = str_dup("/src");
c->receive_root_directory = str_dup("/dst");
c->filters = array_list_create(free);
EXPECT_NOT_NULL(c->filters);
for (int i = 0; i <= MAX_FILTER_RULES; i++)
EXPECT_TRUE(array_list_add(c->filters, str_dup("- *.tmp")));
EXPECT_TRUE(roundtrip_config_rejected(c));
config_delete(c);
/* A pattern longer than the receiver's evaluation bound is rejected by the
sender (mirroring the receiver's guard) instead of being sent as an inert
rule. */
c = config_create();
EXPECT_NOT_NULL(c);
c->send_directory = str_dup("/src");
c->receive_root_directory = str_dup("/dst");
c->filters = array_list_create(free);
EXPECT_NOT_NULL(c->filters);
size_t big = MAX_PROTECT_PATTERN_LEN + 1;
char* long_rule = malloc(big + 3);
EXPECT_NOT_NULL(long_rule);
long_rule[0] = '-';
long_rule[1] = ' ';
memset(long_rule + 2, 'x', big);
long_rule[big + 2] = '\0';
EXPECT_TRUE(array_list_add(c->filters, long_rule));
EXPECT_TRUE(roundtrip_config_rejected(c));
config_delete(c);
}
/* identity_copy_as_refused() is the pure, pre-snapshot refusal predicate: a
--copy-as is refused when the receiver is not root OR the effective super
mode is OFF (an operator veto), and never when --copy-as is unset. */
@@ -2590,6 +2633,46 @@ static bool basis_equal(const Config* a, const Config* b) {
return true;
}
/* The receiver reconstructs its delete-protection list from the sender's
* compiled base rules, so compare the received list against a fresh
* filter_base_build() of the sender's raw --filter texts. */
static bool filter_rules_equal(const Config* a, const FilterRuleList* got) {
int count = a->filters ? a->filters->size : 0;
const char** texts = NULL;
if (count > 0) {
texts = calloc((size_t)count, sizeof(char*));
if (!texts)
return false;
for (int i = 0; i < count; i++)
texts[i] = (const char*)a->filters->items[i];
}
char err[160];
FilterRuleList* expected =
filter_base_build(texts, count, a->cvs_exclude, a->delete_excluded, err, sizeof(err));
free(texts);
if (!expected)
return false;
bool equal = true;
int want = expected->count;
int have = got ? got->count : 0;
if (want != have) {
equal = false;
} else {
for (int i = 0; i < want; i++) {
const FilterRule* x = expected->items[i];
const FilterRule* y = got->items[i];
if (x->action != y->action || x->sides != y->sides || x->anchored != y->anchored ||
x->dir_only != y->dir_only || x->negate != y->negate ||
!str_opt_equal(x->owner, y->owner) || !str_opt_equal(x->pattern, y->pattern)) {
equal = false;
break;
}
}
}
filter_rule_list_free(expected);
return equal;
}
#define CONFIG_CMP_BOOL(a, b, name) ((a)->name == (b)->name)
#define CONFIG_CMP_INT(a, b, name) ((a)->name == (b)->name)
#define CONFIG_CMP_RAW(a, b, name) ((a)->name == (b)->name)
@@ -2615,6 +2698,7 @@ static bool basis_equal(const Config* a, const Config* b) {
#define CONFIG_CMP_COPY_AS_ID(a, b, name) (!(a)->copy_as_set || (a)->name == (b)->name)
#define CONFIG_CMP_BLOCK_SKIP_SUFFIXES(a, b, name) skip_suffixes_equal((a), (b))
#define CONFIG_CMP_BLOCK_BASIS(a, b, name) basis_equal((a), (b))
#define CONFIG_CMP_BLOCK_PROTECT_RULES(a, b, name) filter_rules_equal((a), (b)->name)
#define CONFIG_CMP_BLOCK_IDMAP(a, b, name) \
idmap_equal((a)->name, (a)->name##_count, (b)->name, (b)->name##_count)
@@ -2784,6 +2868,7 @@ static void golden_config_populate(Config* c) {
c->skip_compress_suffixes[1] = str_dup(".xz");
EXPECT_EQ_INT(config_basis_append(c, BASIS_DEST_COMPARE, "compare"), 0);
EXPECT_EQ_INT(config_basis_append(c, BASIS_DEST_LINK, "link"), 0);
c->verify_basis = true;
c->fuzzy = true;
c->checksum_algo = CHECKSUM_ALGO_MD5;
c->checksum_seed = 0x1122334455667788ULL;
@@ -2829,6 +2914,14 @@ static void golden_config_populate(Config* c) {
c->copy_as_set = true;
c->copy_as_uid = 111;
c->copy_as_gid = 222;
/* Compile-through delete-protection rules (protocol 2.28.0). The golden
* sender serializes its compiled base rules, so populate a diverse set that
* exercises both sides, negate, anchoring and dir-only. */
c->filters = array_list_create(free);
array_list_add(c->filters, str_dup("P *.log"));
array_list_add(c->filters, str_dup("+r **/*.txt"));
array_list_add(c->filters, str_dup("H,!secret"));
array_list_add(c->filters, str_dup("- /sub/dir/"));
}
/* The pinned golden frame (protocol 2.28.0). The values below are the only
@@ -2836,12 +2929,14 @@ static void golden_config_populate(Config* c) {
* them ONLY with a PROTOCOL_VERSION bump and a documented reason. The 2.24.0
* delete-plan wave changed only the version string; 2.25.0 appended the
* report_stats bool, 2.26.0 appended the compression_algo int, 2.27.0 appended
* the report_deletes bool, and 2.28.0 changed only the version string (the
* STATUS_STATS body grew, but the config frame layout is unchanged, so the
* frame length is identical). The byte-exact values are recomputed for the
* merged layout. */
#define GOLDEN_WIRE_LEN 709
#define GOLDEN_WIRE_HASH 417335736347473203ULL
* the report_deletes bool, and 2.28.0 changed only the version string and
* appended the receiver-side delete-protection rule block (the STATUS_STATS
* body also grew, but that is not part of this frame). Track 5a appends the
* FastSync-only verify_basis bool to the basis block WITHOUT a version bump
* (project decision), so the frame grew by one int to 886 bytes. The
* byte-exact values are recomputed for the merged layout. */
#define GOLDEN_WIRE_LEN 886
#define GOLDEN_WIRE_HASH 5809509022716816757ULL
static unsigned long long fnv1a_64(const unsigned char* buf, size_t len) {
unsigned long long h = 1469598103934665603ULL;
@@ -2990,6 +3085,7 @@ static void test_config_wire_golden_receive() {
recv->groupmap[0].to_name != NULL && strcmp(recv->groupmap[0].to_name, "root") == 0;
ok = ok && recv->basis_count == 2 && recv->basis_dirs[0].type == BASIS_DEST_COMPARE &&
recv->basis_dirs[1].type == BASIS_DEST_LINK;
ok = ok && recv->verify_basis;
ok = ok && recv->module != NULL && strcmp(recv->module, "goldenmod") == 0;
ok = ok && recv->copy_as_set && recv->copy_as_uid == 111 && recv->copy_as_gid == 222;
}
@@ -3250,6 +3346,7 @@ void test_config() {
test_config_wire_golden_receive();
test_config_wire_receive_bounds();
test_config_receive_rejects_overcap_counts();
test_config_receive_rejects_bad_protect_rules();
test_config_wire_roundtrip_all_fields();
test_config_preserve_attribute_wire_roundtrip();
}
+50
View File
@@ -17,6 +17,7 @@
* caller/receiver entry point) describing `dir` with no kept children. */
static void send_plan_frame(int fd, const char* dir) {
EXPECT_TRUE(send_int(fd, 0)); /* has_config */
EXPECT_TRUE(send_int(fd, 1)); /* apply: a real plan */
EXPECT_TRUE(send_wire_str(fd, dir));
EXPECT_TRUE(send_int(fd, 0)); /* kept child dirs */
EXPECT_TRUE(send_int(fd, 0)); /* kept child files */
@@ -216,9 +217,58 @@ static void test_delete_delay_actual_removal_charges_budget(void) {
config_delete(config);
}
/* Send a config-only carrier frame (apply=false): the per-run config block with
* one --delete-missing-args exact path, and no directory walk. */
static void send_config_only_frame(int fd, const char* missing_path) {
EXPECT_TRUE(send_int(fd, 1)); /* has_config */
EXPECT_TRUE(send_int(fd, 0)); /* protected prefixes */
EXPECT_TRUE(send_int(fd, 0)); /* size-skipped */
EXPECT_TRUE(send_int(fd, 1)); /* missing args */
EXPECT_TRUE(send_wire_str(fd, missing_path));
EXPECT_TRUE(send_int(fd, 0)); /* apply = false */
EXPECT_TRUE(send_wire_str(fd, "."));
EXPECT_TRUE(send_int(fd, 0));
EXPECT_TRUE(send_int(fd, 0));
}
/* The config-only carrier frame (apply=false) still applies the
* --delete-missing-args exact deletions even though it walks no directory. This
* is the fix for a --files-from list that synchronizes no directory. */
static void test_config_only_frame_applies_missing_args(void) {
char root[] = "/tmp/fastsync_dp_cfgonly_XXXXXX";
EXPECT_TRUE(mkdtemp(root) != NULL);
char gone[1024];
snprintf(gone, sizeof(gone), "%s/gone.txt", root);
int fd = open(gone, O_WRONLY | O_CREAT | O_TRUNC, 0600);
EXPECT_TRUE(fd >= 0);
close(fd);
Config* config = config_create();
EXPECT_NOT_NULL(config);
config->receive_root_directory = str_dup(root);
config->delete_missing_args = true;
int p[2];
EXPECT_EQ_INT(socketpair(AF_UNIX, SOCK_STREAM, 0, p), 0);
DeletePlanSession* session = delete_plan_session_create(config);
EXPECT_NOT_NULL(session);
send_config_only_frame(p[1], "gone.txt");
EXPECT_EQ_INT(delete_plan_session_receive(session, config, p[0]), 0);
EXPECT_EQ_INT((int)delete_plan_session_deleted(session), 1);
EXPECT_TRUE(lstat(gone, &(struct stat){0}) != 0);
delete_plan_session_destroy(session);
close(p[0]);
close(p[1]);
rmdir(root);
config_delete(config);
}
void test_delete_plan(void) {
test_delete_delay_refilled_dir_removed_recursively();
test_delete_delay_removed_file_counted();
test_delete_delay_max_delete_bounds_actual();
test_delete_delay_actual_removal_charges_budget();
test_config_only_frame_applies_missing_args();
}
+156
View File
@@ -6,6 +6,7 @@
#include "file_receive.h"
#include "data.h"
#include "config.h"
#include "charset.h"
#include "utils.h"
#include "protocol.h"
#include "test_utils.h"
@@ -1807,6 +1808,159 @@ static void test_receive_incremental_check_empty_path() {
config_delete(cfg);
}
/* Pure policy helpers behind the basis quick-check / --verify-basis decision. */
static void test_file_basis_quick_match_decision() {
Config* cfg = config_create();
EXPECT_NOT_NULL(cfg);
EXPECT_FALSE(file_basis_content_required(cfg));
cfg->verify_basis = true;
EXPECT_TRUE(file_basis_content_required(cfg));
cfg->verify_basis = false;
struct stat st;
memset(&st, 0, sizeof(st));
st.st_mtime = 1500000000;
#ifdef __linux__
st.st_mtim.tv_nsec = 500;
#endif
/* Equal size is required by the caller; this leg is the mtime / --size-only
rule. Equal mtime matches, a different mtime misses by default. */
EXPECT_TRUE(file_basis_quick_match(cfg, &st, 1500000000, 500));
EXPECT_FALSE(file_basis_quick_match(cfg, &st, 1500000001, 500));
cfg->size_only = true;
EXPECT_TRUE(file_basis_quick_match(cfg, &st, 1500000001, 500));
cfg->size_only = false;
cfg->modify_window = 2;
EXPECT_TRUE(file_basis_quick_match(cfg, &st, 1500000002, 500));
config_delete(cfg);
}
/* End-to-end handshake decision for a same-size, same-mtime, DIFFERENT-content
basis. Default (rsync parity): the metadata quick-check is trusted, the
receiver answers STATUS_OK and materializes the basis bytes. --verify-basis:
the whole-file digest is required, the basis is rejected and the receiver
asks for the source (STATUS_NEXT + full transfer). */
static void test_receive_incremental_check_basis_quick_check_and_verify() {
const char* root = "test_basis_quick_root";
const char* basis_dir = "test_basis_quick_root/basis";
const char* basis_file = "test_basis_quick_root/basis/f.txt";
unlink(basis_file);
unlink("test_basis_quick_root/f.txt");
rmdir(basis_dir);
rmdir(root);
EXPECT_EQ_INT(mkdir(root, 0755), 0);
EXPECT_EQ_INT(mkdir(basis_dir, 0755), 0);
const char* src_bytes = "AAAA";
const unsigned long long size = 4;
const time_t mtime = 1500000000;
{
FILE* fh = fopen(basis_file, "wb");
EXPECT_NOT_NULL(fh);
// cppcheck-suppress knownConditionTrueFalse
if (fh) {
EXPECT_EQ_INT((int)fwrite("BBBB", 1, (size_t)size, fh), (int)size);
fclose(fh);
}
}
struct timespec ts[2] = {{mtime, 0}, {mtime, 0}};
EXPECT_EQ_INT(utimensat(AT_FDCWD, basis_file, ts, 0), 0);
char root_abs[PATH_MAX];
EXPECT_NOT_NULL(realpath(root, root_abs));
int root_fd = open(root_abs, O_RDONLY | O_DIRECTORY | O_CLOEXEC);
EXPECT_TRUE(root_fd >= 0);
// cppcheck-suppress knownConditionTrueFalse
if (root_fd < 0) {
unlink(basis_file);
rmdir(basis_dir);
rmdir(root);
return;
}
EXPECT_TRUE(utils_set_authorized_root(root_fd, root_abs));
uint8_t digest[CHECKSUM_MAX_DIGEST_LEN];
size_t digest_len = 0;
EXPECT_TRUE(checksum_digest(CHECKSUM_ALGO_XXH64, 0, src_bytes, size, digest, sizeof(digest),
&digest_len));
/* Route protocol I/O through the explicit descriptors (a previous test group
may have left io_set_fds() bound to its own pipe). */
io_set_fds(-1, -1);
io_set_bwlimit(0);
for (int verify = 0; verify <= 1; verify++) {
Config* cfg = config_create();
EXPECT_NOT_NULL(cfg);
cfg->receive_root_directory = str_dup(root_abs);
cfg->checksum = false;
cfg->checksum_algo = CHECKSUM_ALGO_XXH64;
cfg->checksum_seed = 0;
cfg->use_incremental = true;
cfg->use_delta = false;
cfg->use_metadata = false;
cfg->verify_basis = (verify != 0);
EXPECT_EQ_INT(config_basis_append(cfg, BASIS_DEST_LINK, "basis"), 0);
int p[2];
EXPECT_EQ_INT(socketpair(AF_UNIX, SOCK_STREAM, 0, p), 0);
EXPECT_TRUE(send_wire_str(p[1], "f.txt"));
unsigned long long check_size = size;
long long check_mtime = (long long)mtime;
long long check_mtime_nsec = 0;
EXPECT_TRUE(send_n_data(p[1], &check_size, sizeof(check_size)));
EXPECT_TRUE(send_n_data(p[1], &check_mtime, sizeof(check_mtime)));
EXPECT_TRUE(send_n_data(p[1], &check_mtime_nsec, sizeof(check_mtime_nsec)));
/* Only --verify-basis needs the digest (cfg->checksum is false) and the
pre-staged fallback full transfer the receiver will request. */
if (verify) {
uint8_t wire_len = (uint8_t)digest_len;
EXPECT_TRUE(send_n_data(p[1], &wire_len, sizeof(wire_len)));
EXPECT_TRUE(send_n_data(p[1], digest, digest_len));
char* payload_bytes = str_dup(src_bytes);
EXPECT_NOT_NULL(payload_bytes);
Data* payload = data_create(payload_bytes, size);
EXPECT_NOT_NULL(payload);
// cppcheck-suppress knownConditionTrueFalse
if (payload)
EXPECT_TRUE(send_data(p[1], payload));
data_destroy(payload);
}
bool skipped = false;
File* file = receive_incremental_check(p[0], cfg, &skipped);
EXPECT_NOT_NULL(file);
// cppcheck-suppress knownConditionTrueFalse
if (file) {
EXPECT_FALSE(skipped);
Status reply = STATUS_ERROR;
EXPECT_TRUE(receive_status(p[1], &reply));
if (verify) {
EXPECT_EQ_INT((int)reply, (int)STATUS_NEXT);
EXPECT_NOT_NULL(file->data->data);
EXPECT_TRUE(file->data->data != NULL && memcmp(file->data->data, src_bytes, size) == 0);
} else {
EXPECT_EQ_INT((int)reply, (int)STATUS_OK);
EXPECT_TRUE(file->skip);
/* The default quick-check hit materializes from the basis PATH at
install time (streaming), so no content is buffered on the File. */
EXPECT_NOT_NULL(file->basis_link);
EXPECT_NULL(file->data->data);
}
file_destroy(file);
}
close(p[0]);
close(p[1]);
config_delete(cfg);
}
utils_set_authorized_root(-1, NULL);
close(root_fd);
unlink(basis_file);
rmdir(basis_dir);
rmdir(root);
}
/* -K/--keep-dirlinks secure open: with an authorized root, a destination path
* component that is a symlink to an IN-ROOT directory is used as that directory
* (its referent is opened through a relative O_NOFOLLOW walk from the root fd,
@@ -2140,6 +2294,8 @@ void test_file() {
test_dir_time_list();
test_dir_time_list_cap();
test_receive_incremental_check_empty_path();
test_file_basis_quick_match_decision();
test_receive_incremental_check_basis_quick_check_and_verify();
test_keep_dirlinks_secure_open();
test_inplace_overwrite_clears_special_mode_bits();
test_inplace_overwrite_metadata_strips_special_bits();
+5 -4
View File
@@ -18,10 +18,11 @@
/* P8 config-frame tail: super_mode (4) + copy-as presence (4) + uid (4) + gid (4). */
#define P8_TAIL_BYTES 16
/* Bytes after the P8 tail: report_dest_info (4), report_stats (4, wire-stats
* wave) and compression_algo (4, codec wave). The P8 fields sit this many
* bytes before the end of the frame. */
#define POST_P8_TAIL_BYTES 12
/* Bytes after the P8 tail: report_dest_info (4), report_stats (4),
* report_deletes (4, --info=del wave), compression_algo (4, codec wave) and the
* receiver delete-protection count (4, protocol 2.28.0). The P8 fields sit
* this many bytes before the end of the frame. */
#define POST_P8_TAIL_BYTES 20
/* Smoke test for chunk_deserialize fuzz target */
static void test_fuzz_chunk_deserialize() {
+27 -13
View File
@@ -1063,14 +1063,18 @@ static void test_incremental_check_fifo_destination_does_not_hang() {
}
/* A server-contacting --dry-run with an alternate basis dir must never read or
hash the basis file. An exact (size+mtime+content) basis match would
otherwise let a client probe the basis bytes against its own supplied digest
(a 1-bit content oracle). The dry-run decision is metadata-only, so even a
byte-identical basis is reported as would-transfer, not a compare-dest skip. */
static void test_incremental_check_dry_run_basis_does_not_read_content() {
hash the basis file. Under the default metadata quick-check a hit needs no
basis bytes, so a compare-dest match is reported as a skip (STATUS_OK) just
like a real run -- and still no content is read. Under --verify-basis a hit
would require hashing the basis against the client-supplied digest (a 1-bit
content oracle), which a dry-run must never do, so even a byte-identical
basis is reported as would-transfer. The destination is never materialized
in either arm. */
static void run_dry_run_basis_check(bool verify, Status expected) {
Config* cfg = config_create();
EXPECT_NOT_NULL(cfg);
cfg->dry_run = true;
cfg->verify_basis = verify;
char* root = make_check_root("dryb");
EXPECT_NOT_NULL(root);
cfg->receive_root_directory = str_dup(root);
@@ -1086,8 +1090,8 @@ static void test_incremental_check_dry_run_basis_does_not_read_content() {
EXPECT_EQ_INT(stat(basis_path, &bst), 0);
EXPECT_EQ_INT(config_basis_append(cfg, BASIS_DEST_COMPARE, "basis"), 0);
/* The (correct) source digest for the basis bytes: an unfixed dry-run would
read+hash the basis and treat this as an exact compare-dest hit. */
/* The (correct) source digest for the basis bytes: a buggy dry-run that read
and hashed the basis would treat this as an exact compare-dest hit. */
uint8_t digest[CHECKSUM_MAX_DIGEST_LEN];
size_t digest_len = 0;
EXPECT_TRUE(checksum_digest((ChecksumAlgo)cfg->checksum_algo, cfg->checksum_seed, content,
@@ -1106,7 +1110,8 @@ static void test_incremental_check_dry_run_basis_does_not_read_content() {
bool skipped = false;
bool would_transfer = false;
File* file = receive_incremental_check_ex(p[0], cfg, &skipped, &would_transfer);
bool ok = file == NULL && !skipped && would_transfer;
bool ok = file == NULL && skipped == (expected == STATUS_OK) &&
would_transfer == (expected != STATUS_OK);
file_destroy(file);
config_delete(cfg);
close(p[0]);
@@ -1124,13 +1129,16 @@ static void test_incremental_check_dry_run_basis_does_not_read_content() {
EXPECT_TRUE(send_n_data(p[1], &size, sizeof(size)));
EXPECT_TRUE(send_n_data(p[1], &mtime, sizeof(mtime)));
EXPECT_TRUE(send_n_data(p[1], &mtime_nsec, sizeof(mtime_nsec)));
uint8_t wire_len = (uint8_t)digest_len;
EXPECT_TRUE(send_n_data(p[1], &wire_len, sizeof(wire_len)));
EXPECT_TRUE(send_n_data(p[1], digest, digest_len));
/* The digest is only on the wire when --checksum or --verify-basis needs it
(cfg->checksum is false here); the default quick-check arm sends none. */
if (verify) {
uint8_t wire_len = (uint8_t)digest_len;
EXPECT_TRUE(send_n_data(p[1], &wire_len, sizeof(wire_len)));
EXPECT_TRUE(send_n_data(p[1], digest, digest_len));
}
Status s;
EXPECT_TRUE(receive_status(p[1], &s));
/* A skip here would mean the receiver read+hashed the basis file. */
EXPECT_EQ_INT(s, STATUS_DRY_RUN_TRANSFER);
EXPECT_EQ_INT(s, expected);
int status;
waitpid(pid, &status, 0);
@@ -1148,6 +1156,11 @@ static void test_incremental_check_dry_run_basis_does_not_read_content() {
}
}
static void test_incremental_check_dry_run_basis_does_not_read_content() {
run_dry_run_basis_check(false, STATUS_OK);
run_dry_run_basis_check(true, STATUS_DRY_RUN_TRANSFER);
}
/* B1: a FIFO planted in a --link-dest basis directory must not block
* basis_open_regular() either; the basis match is simply declined. */
static void test_incremental_check_basis_fifo_does_not_hang() {
@@ -1278,6 +1291,7 @@ static void test_dry_run_delete_plan_commit_does_not_delete() {
EXPECT_TRUE(send_int(p[1], 0)); /* size-skipped prefixes */
EXPECT_TRUE(send_int(p[1], 1)); /* missing-args exact deletions */
EXPECT_TRUE(send_str(p[1], "victim.txt"));
EXPECT_TRUE(send_int(p[1], 1)); /* apply: a real plan */
EXPECT_TRUE(send_str(p[1], ".")); /* receive root plan */
EXPECT_TRUE(send_int(p[1], 0)); /* kept child directories */
EXPECT_TRUE(send_int(p[1], 0)); /* kept child files */
+47 -6
View File
@@ -138,7 +138,7 @@ static void test_walker_removes_extras_keeps_manifest_and_protected() {
DeleteSkipEntry skip = {"prot", false};
size_t deleted = 0;
DeleteWalkResult result =
delete_extras_limited(root, manifest, NULL, 100000, &skip, 1, &deleted, NULL);
delete_extras_limited(root, manifest, NULL, 100000, &skip, 1, NULL, &deleted, NULL);
EXPECT_EQ_INT((int)result, (int)DELETE_WALK_OK);
EXPECT_FALSE(file_exists(root, "a.txt"));
EXPECT_TRUE(file_exists(root, "keep.txt"));
@@ -173,7 +173,7 @@ static void test_walker_keeps_nested_manifest_dirs() {
EXPECT_NOT_NULL(manifest);
size_t deleted = 0;
DeleteWalkResult result =
delete_extras_limited(root, manifest, NULL, 100000, NULL, 0, &deleted, NULL);
delete_extras_limited(root, manifest, NULL, 100000, NULL, 0, NULL, &deleted, NULL);
EXPECT_EQ_INT((int)result, (int)DELETE_WALK_OK);
EXPECT_FALSE(file_exists(root, "extra.txt"));
EXPECT_TRUE(file_exists(root, "keepdir/deep/keep.txt"));
@@ -204,7 +204,7 @@ static void test_walker_max_delete_partial_deletes_up_to_cap() {
size_t deleted = 999;
size_t skipped = 0;
DeleteWalkResult result =
delete_extras_limited(root, manifest, NULL, 2, NULL, 0, &deleted, &skipped);
delete_extras_limited(root, manifest, NULL, 2, NULL, 0, NULL, &deleted, &skipped);
EXPECT_EQ_INT((int)result, (int)DELETE_WALK_LIMIT_REACHED);
EXPECT_EQ_INT((int)deleted, 2);
EXPECT_EQ_INT((int)skipped, 1);
@@ -225,7 +225,8 @@ static void test_walker_max_delete_exact_bound_deletes() {
ArrayList* manifest = make_manifest_strings(keeps, 0);
EXPECT_NOT_NULL(manifest);
size_t deleted = 0;
DeleteWalkResult result = delete_extras_limited(root, manifest, NULL, 2, NULL, 0, &deleted, NULL);
DeleteWalkResult result =
delete_extras_limited(root, manifest, NULL, 2, NULL, 0, NULL, &deleted, NULL);
EXPECT_EQ_INT((int)result, (int)DELETE_WALK_OK);
EXPECT_EQ_INT((int)deleted, 2);
EXPECT_FALSE(file_exists(root, "a.txt"));
@@ -258,7 +259,7 @@ static void test_walker_removes_extraneous_symlinks() {
EXPECT_NOT_NULL(manifest);
size_t deleted = 0;
DeleteWalkResult result =
delete_extras_limited(root, manifest, NULL, 100000, NULL, 0, &deleted, NULL);
delete_extras_limited(root, manifest, NULL, 100000, NULL, 0, NULL, &deleted, NULL);
EXPECT_EQ_INT((int)result, (int)DELETE_WALK_OK);
EXPECT_FALSE(file_exists(root, "link_file"));
EXPECT_FALSE(file_exists(root, "link_dir"));
@@ -294,7 +295,7 @@ static void test_walker_confines_deletion_to_synced_dirs() {
EXPECT_TRUE(array_list_add(dirs, str_dup("inscope")));
size_t deleted = 0;
DeleteWalkResult result =
delete_extras_limited(root, manifest, dirs, 100000, NULL, 0, &deleted, NULL);
delete_extras_limited(root, manifest, dirs, 100000, NULL, 0, NULL, &deleted, NULL);
EXPECT_EQ_INT((int)result, (int)DELETE_WALK_OK);
EXPECT_TRUE(file_exists(root, "rootextra.txt"));
EXPECT_FALSE(file_exists(root, "inscope/extra.txt"));
@@ -324,6 +325,45 @@ static void test_walker_unlimited_deletes_all() {
free(root);
}
/* Receiver-side filter protection (protocol 2.28.0): a compiled protect rule
shields a DESTINATION-ONLY extra that never appeared on the sender, a risk
rule cancels an earlier/later protect (first match wins), and a dir-only
protect rule shields the whole subtree. */
static void test_walker_protect_rules_shield_dest_only() {
char* root = make_walk_root("protectrules");
EXPECT_NOT_NULL(root);
EXPECT_TRUE(write_file_at(root, "keep.txt", "kept"));
EXPECT_TRUE(write_file_at(root, "extra.log", "risk cancels protect"));
EXPECT_TRUE(write_file_at(root, "safe.log", "protected"));
EXPECT_TRUE(write_file_at(root, "other.txt", "deleted"));
EXPECT_EQ_INT(make_subdir(root, "prot"), 0);
EXPECT_TRUE(write_file_at(root, "prot/inside.txt", "shielded subtree"));
EXPECT_TRUE(write_file_at(root, "prot/deep.log", "shielded subtree"));
const char* keeps[] = {"keep.txt"};
ArrayList* manifest = make_manifest_strings(keeps, 1);
EXPECT_NOT_NULL(manifest);
const char* rule_text[] = {"R extra.log", "P *.log", "P prot/"};
char err[160];
FilterRuleList* rules = filter_base_build(rule_text, 3, false, false, err, sizeof(err));
EXPECT_NOT_NULL(rules);
size_t deleted = 0;
DeleteWalkResult result =
delete_extras_limited(root, manifest, NULL, 100000, NULL, 0, rules, &deleted, NULL);
EXPECT_EQ_INT((int)result, (int)DELETE_WALK_OK);
EXPECT_TRUE(file_exists(root, "keep.txt"));
EXPECT_FALSE(file_exists(root, "extra.log")); /* risk wins the first match */
EXPECT_TRUE(file_exists(root, "safe.log")); /* protect shields the extra */
EXPECT_FALSE(file_exists(root, "other.txt"));
EXPECT_TRUE(dir_exists(root, "prot"));
EXPECT_TRUE(file_exists(root, "prot/inside.txt"));
EXPECT_TRUE(file_exists(root, "prot/deep.log"));
filter_rule_list_free(rules);
array_list_delete(manifest);
remove_walk_tree(root);
free(root);
}
typedef struct {
bool eight_bit_output;
const char* expected;
@@ -626,6 +666,7 @@ void test_shared_utils() {
test_walker_removes_extraneous_symlinks();
test_walker_confines_deletion_to_synced_dirs();
test_walker_unlimited_deletes_all();
test_walker_protect_rules_shield_dest_only();
test_loopback_helpers();
test_fd_peer_ip();
+2 -1
View File
@@ -210,7 +210,8 @@ static void test_xattr_receive_drops_acl_without_preserve_acls() {
/* MINOR-2: a --link-dest / -H copy fallback (linkat refused) must still apply
* the per-file xattrs and --fake-super stat. A DIRECTORY basis forces linkat
* to fail with EPERM, exercising the byte-copy fallback deterministically.
* to fail with EPERM, exercising the byte-copy fallback deterministically (the
* basis is not a regular file, so the fallback uses the caller's bytes).
* Guarded on filesystem xattr support. */
static void test_link_copy_fallback_preserves_xattrs() {
const char* dest = "test_link_xattr_dest.txt";