Compare commits
120
Commits
v2.26.0
...
b6b5eee20e
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
b6b5eee20e | ||
|
|
8dcd87609d | ||
|
|
19fe63bd59 | ||
|
|
1f5f8dc5a7 | ||
|
|
3787695ba6 | ||
|
|
b7cb213c8c | ||
|
|
8cc3dd993b | ||
|
|
1167e7970b | ||
|
|
e98729f00e | ||
|
|
c0020364b2 | ||
|
|
d7ac940a6b | ||
|
|
be20e836de | ||
|
|
b549887138 | ||
|
|
1ffd4744c6 | ||
|
|
494cef2a0b | ||
|
|
ed7527cc2c | ||
|
|
934defa965 | ||
|
|
ade9be8600 | ||
|
|
707bb659e8 | ||
|
|
33980bc4c8 | ||
|
|
6a9372f4f5 | ||
|
|
5e1d6b6e10 | ||
|
|
d2d1b63f44 | ||
|
|
a6659472fe | ||
|
|
4638030288 | ||
|
|
07dfec629f | ||
|
|
c062a0762e | ||
|
|
283f9f0823 | ||
|
|
5754b9a952 | ||
|
|
b8a0efef7b | ||
|
|
2e77c09447 | ||
|
|
06c4026b74 | ||
|
|
9f47b13712 | ||
|
|
b1eddf0133 | ||
|
|
338c27db73 | ||
|
|
8d46a26c04 | ||
|
|
707ba272df | ||
|
|
bc71a3c0a5 | ||
|
|
cee9b7647c | ||
|
|
3adb6dddb5 | ||
|
|
d77849774e | ||
|
|
b5c6f8b60f | ||
|
|
134dcd74cc | ||
|
|
42f2f845ac | ||
|
|
0669335ca5 | ||
|
|
391f76cd56 | ||
|
|
2021afe613 | ||
|
|
ba914e8ab3 | ||
|
|
df8fe1ae4e | ||
|
|
2c490d58b7 | ||
|
|
d55dabff2e | ||
|
|
b478a59a81 | ||
|
|
865941f850 | ||
|
|
221cefa7cc | ||
|
|
bb9b59024a | ||
|
|
5baf120243 | ||
|
|
84b7e1fd2f | ||
|
|
800a978e15 | ||
|
|
0f40e747f0 | ||
|
|
b07306d5bc | ||
|
|
f2c89b6e7c | ||
|
|
379f127859 | ||
|
|
3799200f71 | ||
|
|
91c4a967e6 | ||
|
|
88c8968b4c | ||
|
|
082b31886a | ||
|
|
423a62e691 | ||
|
|
bc18ae205b | ||
|
|
10a61c6101 | ||
|
|
bc1e1191af | ||
|
|
0fbb9de915 | ||
|
|
9b05972375 | ||
|
|
d119f35066 | ||
|
|
00829fd265 | ||
|
|
ff261bc38a | ||
|
|
b235721f8b | ||
|
|
79a28cdb96 | ||
|
|
402cae80ad | ||
|
|
9691dba6f0 | ||
|
|
558782d339 | ||
|
|
5597e74f6a | ||
|
|
ee6523afac | ||
|
|
4163caa1d3 | ||
|
|
10159dc120 | ||
|
|
cd7b96d0bb | ||
|
|
f6f49d536e | ||
|
|
38d304103c | ||
|
|
4b09213b88 | ||
|
|
36fd0774e8 | ||
|
|
711b7e50b3 | ||
|
|
67076bf218 | ||
|
|
8ec8cb7203 | ||
|
|
82959395fb | ||
|
|
c9f94ea46e | ||
|
|
eff9852038 | ||
|
|
c80098623f | ||
|
|
6119e1e75c | ||
|
|
25062352f6 | ||
|
|
00d628d4ba | ||
|
|
3545d88905 | ||
|
|
596a039454 | ||
|
|
ae037cc27b | ||
|
|
3d0672a721 | ||
|
|
b82aab72c5 | ||
|
|
8e7764007d | ||
|
|
1042d15db7 | ||
|
|
cbe37a77dd | ||
|
|
a5d45ef266 | ||
|
|
a960391b34 | ||
|
|
134f8b027a | ||
|
|
eb7e3fd2e0 | ||
|
|
6a129b54d4 | ||
|
|
c7b2c7eb2b | ||
|
|
e77dfbec70 | ||
|
|
f3ac4df4d0 | ||
|
|
91197fd7cf | ||
|
|
13708352ec | ||
|
|
d162d93570 | ||
|
|
9fa1696eff | ||
|
|
b705fb807f |
@@ -49,6 +49,47 @@ jobs:
|
||||
if: github.event_name == 'push'
|
||||
run: python3 -m pytest tests/integration/ -n 4 --dist=load -m "not setpriv" --durations=25 --tb=short -q
|
||||
|
||||
# Differential rsync-parity gate: runs real rsync 3.4.1 and FastSync over the
|
||||
# same corpora and compares destinations + normalized output. The fast subset
|
||||
# guards the ✅ surface on every PR; the full set (with FASTSYNC_PARITY_STRICT
|
||||
# so a fixed caveat must be removed from the allowlist) burns the documented
|
||||
# ⚠️/❌ residuals down on push. See tests/integration/README.md.
|
||||
parity-fast:
|
||||
runs-on: ubuntu-latest
|
||||
container: gitea.tap-tap.win/taptap/fastsync-ci:v11
|
||||
needs: lint
|
||||
if: github.event_name == 'pull_request'
|
||||
steps:
|
||||
- name: Checkout
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
|
||||
- name: Configure
|
||||
run: cmake -B build -S . -DSTRICT_WARNINGS=ON
|
||||
|
||||
- name: Build
|
||||
run: cmake --build build -j$(nproc)
|
||||
|
||||
- name: Differential parity (fast subset)
|
||||
run: python3 -m pytest tests/integration/test_differential_parity.py -n 4 --dist=load -m parity_ci -q
|
||||
|
||||
parity-full:
|
||||
runs-on: ubuntu-latest
|
||||
container: gitea.tap-tap.win/taptap/fastsync-ci:v11
|
||||
needs: lint
|
||||
if: github.event_name == 'push'
|
||||
steps:
|
||||
- name: Checkout
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
|
||||
- name: Configure
|
||||
run: cmake -B build -S . -DSTRICT_WARNINGS=ON
|
||||
|
||||
- name: Build
|
||||
run: cmake --build build -j$(nproc)
|
||||
|
||||
- name: Differential parity (full set)
|
||||
run: FASTSYNC_PARITY_STRICT=1 python3 -m pytest tests/integration/test_differential_parity.py -n 4 --dist=load -m parity -q
|
||||
|
||||
sanitizers:
|
||||
runs-on: ubuntu-latest
|
||||
container: gitea.tap-tap.win/taptap/fastsync-ci:v11
|
||||
|
||||
@@ -12,3 +12,12 @@ build_docker2/
|
||||
# Test/run artifacts
|
||||
root/
|
||||
test_partial_install_tmp/
|
||||
|
||||
# Editor/tooling + test caches/artifacts
|
||||
.pytest_cache/
|
||||
*.gcda
|
||||
*.gcno
|
||||
*.gcov
|
||||
di/
|
||||
test_data-manual/
|
||||
*.log
|
||||
|
||||
@@ -4,9 +4,9 @@ FastSync is a high-performance file synchronization system written in C11. It su
|
||||
|
||||
## Dependency installation
|
||||
|
||||
**CI rule:** never add `apt-get install` / `pip install` steps to CI workflows — use the custom Docker image instead. The image is built from the repo-root `Dockerfile` and is the same image CI uses: `gitea.tap-tap.win/taptap/fastsync-ci:v11`. It contains the full toolchain: gcc/g++, CMake, libzstd-dev, libssl-dev, make, git, cppcheck, clang-format, python3 + pytest + pytest-xdist, openssh-client, Node.js, plus `rsync` 3.4.1 (with zstd/xxhash/lz4), `acl` and `attr` (setfacl/getfacl, setfattr/getfattr) for drop-in parity tests.
|
||||
**CI rule:** never add `apt-get install` / `pip install` steps to CI workflows — use the custom Docker image instead. The image is built from the repo-root `Dockerfile` and is the same image CI uses: `gitea.tap-tap.win/taptap/fastsync-ci:v11`. It contains the full toolchain: gcc/g++, CMake, libzstd-dev, zlib1g-dev, liblz4-dev, libxxhash-dev, libssl-dev, make, git, cppcheck, clang-format, python3 + pytest + pytest-xdist, openssh-client, Node.js, plus `rsync` 3.4.1 (with zstd/xxhash/lz4), `acl` and `attr` (setfacl/getfacl, setfattr/getfattr) for drop-in parity tests. (CMake hard-requires zstd, zlib, and lz4; xxHash is fetched via `FetchContent`.)
|
||||
|
||||
**Host rule:** for local development, use `nix-shell` (see `README.md`) which provides zstd, OpenSSL, CMake, and gcc. The Docker image can also be used locally for CI parity.
|
||||
**Host rule:** for local development, use `nix-shell` (see `README.md`) which provides zstd, zlib, lz4, OpenSSL, CMake, and gcc. The Docker image can also be used locally for CI parity.
|
||||
|
||||
```bash
|
||||
# Use the prebuilt CI image directly (faster, guaranteed CI parity)
|
||||
@@ -38,11 +38,12 @@ If a dependency is missing from the CI image, add it to the `Dockerfile` (and re
|
||||
When configuring for CI parity, use:
|
||||
```bash
|
||||
cmake -B build -S . -DSTRICT_WARNINGS=ON # -Wextra -Wpedantic -Werror
|
||||
cmake -B build -S . -DSANITIZER=address # AddressSanitizer (ASan)
|
||||
cmake -B build -S . -DSANITIZER=thread # ThreadSanitizer (TSan)
|
||||
cmake -B build -S . -DSANITIZER=address # AddressSanitizer (ASan); in the CI matrix
|
||||
cmake -B build -S . -DSANITIZER=undefined # UndefinedBehaviorSanitizer (UBSan); in the CI matrix
|
||||
cmake -B build -S . -DSANITIZER=thread # ThreadSanitizer (TSan); local-only, NOT in CI
|
||||
```
|
||||
|
||||
The CI workflow (`.gitea/workflows/ci.yaml`) runs lint (clang-format, cppcheck), then a **fast PR gate** — build + unit + a representative subset of integration tests marked `@pytest.mark.ci`, parallelized with pytest-xdist (`-n 4 --dist=load`). The full coverage jobs (full integration suite as `-m "not setpriv"`, sanitizer, fuzz, coverage, valgrind) run **only on push to `dev`/`main`**; pull requests skip them to keep PR CI under ~3 minutes. The two `setpriv` privilege tests are excluded from CI via a marker because their result depends on the runner/container uid and host mount permissions.
|
||||
The CI workflow (`.gitea/workflows/ci.yaml`) runs lint (clang-format, cppcheck), then a **fast PR gate** — build + unit + a representative subset of integration tests marked `@pytest.mark.ci`, parallelized with pytest-xdist (`-n 4 --dist=load`). The full coverage jobs (full integration suite as `-m "not setpriv"`, the `address`+`undefined` sanitizer matrix, fuzz, coverage, valgrind) run **only on push to `dev`/`main`**; pull requests skip them to keep PR CI under ~3 minutes. TSan is not part of the CI matrix and is a local-only configuration. The `setpriv`-marked privilege tests (four decorated functions, collecting to eight instances because two are parametrized) are excluded from CI via a marker because their result depends on the runner/container uid and host mount permissions.
|
||||
|
||||
## Build
|
||||
|
||||
@@ -56,6 +57,21 @@ cmake -B build -S . && cmake --build build -j$(nproc)
|
||||
./build/tests # unit tests
|
||||
python3 -m pytest tests/integration/ -n 4 --dist=load -m "not setpriv" # full integration suite (CI excludes env-dependent privilege tests)
|
||||
python3 -m pytest tests/integration/ -n 4 --dist=load -m ci # PR-gate subset only
|
||||
|
||||
# Differential rsync-parity gate (real rsync 3.4.1 vs FastSync)
|
||||
python3 -m pytest tests/integration/test_differential_parity.py -n 4 --dist=load -m parity_ci # fast PR subset
|
||||
python3 -m pytest tests/integration/test_differential_parity.py -n 4 --dist=load -m parity # full set
|
||||
```
|
||||
|
||||
See `tests/integration/README.md` for the differential parity gate and its
|
||||
`parity_caveats.py` allowlist (the residual burn-down mechanism).
|
||||
|
||||
Unit tests under valgrind must set `FASTSYNC_UNDER_VALGRIND=1` (CI does): the
|
||||
tests use it to skip fork-based tests, because valgrind 3.22 does not expose
|
||||
`vgpreload` in the guest's `/proc/self/maps`.
|
||||
|
||||
```bash
|
||||
FASTSYNC_UNDER_VALGRIND=1 valgrind --leak-check=full --show-leak-kinds=definite --error-exitcode=1 ./build/tests
|
||||
```
|
||||
|
||||
## CI Workflow — Waiting for Results
|
||||
@@ -90,7 +106,7 @@ Two main branches: `dev` (integration) and `main` (stable releases).
|
||||
|
||||
### Rules
|
||||
- **All PRs target `dev`** — never target `main` directly
|
||||
- **`dev` is the default branch** in Gitea repo settings
|
||||
- **`dev` is intended to be the default branch** in Gitea repo settings — verify in the repo settings, since this clone's `origin/HEAD` still points at `main`
|
||||
- **`main` is protected** — only merged from `dev` via PR with 2 approvals + full CI pass
|
||||
- **Feature/bug branches** branch from `dev`, PR back to `dev`
|
||||
- **`dev` → `main` merges** happen on-demand or weekly, requiring full CI + review
|
||||
|
||||
+214
@@ -4,8 +4,222 @@ All notable changes to FastSync are documented here. Versions match
|
||||
`PROTOCOL_VERSION` (printed by `fastsync --version`); the client and server must
|
||||
run the same version because the handshake is strict.
|
||||
|
||||
## [Unreleased]
|
||||
|
||||
The rsync-parity cycle 2.29 (no wire change; `PROTOCOL_VERSION` stays 2.28.0).
|
||||
`RSYNC_COMPAT.md` moves from **116 ✅ / 14 ⚠️ / 27 ❌** to
|
||||
**120 ✅ / 10 ⚠️ / 27 ❌** of 157 rows.
|
||||
|
||||
An audit cycle follows on the same wire version (`PROTOCOL_VERSION` stays
|
||||
2.28.0): a security-and-correctness pass over the parity-2.29 baseline, plus a
|
||||
set of audit follow-ups (filter merge modifiers, the `--inplace`/`--partial-dir`
|
||||
conflict, credential-file hardening, and small leak/log/test fixes). It fixes
|
||||
a `--temp-dir` symlink escape, gates client-controlled special permission bits,
|
||||
corrects `--partial-dir`/`--bwlimit`/`-z` behavior, handles unsupported filter
|
||||
modifiers, and tightens client and wire validation. The only parity
|
||||
reclassification is `--filter=RULE` moving ✅ → ⚠️, because its merge-only
|
||||
`e`/`n`/`w`/`-` modifiers are now accepted and consumed but their semantics
|
||||
remain unimplemented (accepted-but-ignored); the matrix is therefore **119 ✅ /
|
||||
11 ⚠️ / 27 ❌** of 157 rows. The affected rows' notes and the summary tally in
|
||||
`RSYNC_COMPAT.md` were updated.
|
||||
|
||||
### Changed
|
||||
|
||||
- **rsync-exact traversal order.** The sequential scanner now walks each
|
||||
directory's entries in rsync 3.4.1's flist order (non-directories ascending,
|
||||
then directories ascending, depth-first), so `--info=name`, the
|
||||
`--delete-during`/`--delete-delay`/`-n` would-delete order and the partial
|
||||
`--max-delete` survivor set match rsync byte-for-byte. `--threads` has no
|
||||
rsync analogue and stays unordered.
|
||||
- **Delete timing.** The complete `--delete-during`/`--delete-delay`
|
||||
per-directory plan set is transmitted before the first data frame, so a
|
||||
mid-transfer abort has already removed every planned extra like rsync's
|
||||
generator; `-d/--dirs` uses per-directory plans (shielded untraversed
|
||||
subdirectories) instead of the end-of-transfer commit. `-n`, `--delete`,
|
||||
`--del`/`--delete-during` and `--delete-delay` are now ✅ Parity.
|
||||
- **Basis directories.** A relative `--compare-dest`/`--copy-dest`/`--link-dest`
|
||||
DIR resolves against the destination directory with the transfer-relative
|
||||
name appended, exactly like rsync 3.4.1.
|
||||
- **`-y`/`--fuzzy`.** The candidate search no longer inherits the ordinary delta
|
||||
engine's 16 KiB minimum or 10× size-ratio bound, so an oversized or
|
||||
sub-16-KiB sibling is reused exactly as rsync reuses it.
|
||||
- `--info=mount` prints rsync's mount-point skip line (repeated `-xx` drops the
|
||||
mount-point directory); `--info=stats` enables the `--stats` block; `-x` is
|
||||
repeatable. `--stats` counts traversed directories for the `Number of files`
|
||||
breakdown under a plain `-r` scan. `--debug` emits real output for
|
||||
`flist`/`del`/`hash`/`deltasum`/`recv`/`filter`/`send`.
|
||||
|
||||
### Known residuals
|
||||
|
||||
- `--progress` and `--info` still need a receiver→sender event channel for the
|
||||
root `./` line, ancestor-directory suppression, receiver-side `skip`/`backup`
|
||||
wording, and symlink/empty-directory quick-checks.
|
||||
- `--delete-before`'s phase-0 late-file divergence remains (rsync's pre-scan
|
||||
fixes the file list before the data pass).
|
||||
- A single file larger than 256 MiB cannot be streamed in the default path
|
||||
(a general whole-file limit, not basis-specific).
|
||||
- `--stats` byte totals and `--msgs2stderr` stay documented divergences.
|
||||
|
||||
### Security
|
||||
|
||||
- **`--temp-dir` symlink escape fixed.** The receiver's scratch directory was
|
||||
opened with a bare `open()`, so a symlink planted under the receive root could
|
||||
redirect receiver scratch files outside the authorized root. The opened
|
||||
directory is now judged by the real path of its fd (`/proc/self/fd` via
|
||||
`realpath`) and an escaping target is refused (`EACCES`, logged); an in-root
|
||||
link to another filesystem (the `EXDEV` fallback case) still works.
|
||||
- **Client-controlled special bits masked when super-user activities are not
|
||||
permitted.** Setuid/setgid/sticky bits (`--perms`, `--chmod`, the symlink and
|
||||
special-node paths, and deferred directory modes) are now stripped when the
|
||||
connection forbids super activities (`--no-super`, a non-opted daemon module,
|
||||
a privileged listener without `--allow-super`); exact rsync semantics are
|
||||
preserved wherever super activities are permitted.
|
||||
- **Daemon umask no longer forced to `0`.** `daemonize()` now sets the
|
||||
conventional `022`, so implied parent directories created without `-p` are no
|
||||
longer world-writable `0777`.
|
||||
- **Credentials and signal handling hardened.** Secret files are opened with
|
||||
`O_NOFOLLOW|O_NONBLOCK` (while allowing fd-backed store paths and bound-waiting
|
||||
a FIFO read for ~3 s so a slow process substitution works but a connected-but-
|
||||
silent FIFO cannot hang), and signal handlers use `sigaction` with
|
||||
async-signal-safe bodies.
|
||||
|
||||
### Fixed
|
||||
|
||||
- **`-z` on 100–256 MiB files.** The decompressor's internal ceiling was 100 MiB
|
||||
while the receiver advertises and the sender compresses whole files up to
|
||||
`MAX_RECEIVE_WHOLE_FILE_SIZE` (256 MiB), so `-z` on a 100–256 MiB regular file
|
||||
failed with `Declared decompressed size exceeds 104857600 bytes`. The ceiling
|
||||
is now defined in terms of the protocol whole-file bound (still an
|
||||
allocation-clamped bomb guard).
|
||||
- **`--bwlimit` now paces `--sendfile`.** The plaintext-TCP `--sendfile` fast
|
||||
path bypassed the protocol's token bucket, so the limit was ignored there. It
|
||||
now throttles through the same per-session leaky bucket as the TLS path.
|
||||
- **`--partial-dir` implies `--partial`.** Matching rsync 3.4.1 (which sets
|
||||
`keep_partial` after option parsing), `--partial-dir=DIR` alone retains an
|
||||
interrupted transfer's partial and wins over an explicit `--no-partial`;
|
||||
`--inplace` still bypasses the partial machinery, and combining `--inplace`
|
||||
with `--partial-dir` is now rejected up front with rsync's message
|
||||
(`--inplace cannot be used with --partial-dir`).
|
||||
- **Filter modifiers handled.** The `x` xattr-name modifier is rejected with a
|
||||
clear error everywhere. The merge-only `e`/`n`/`w` and `-` modifiers are now
|
||||
accepted and consumed on `merge`/`dir-merge` rules (so they no longer leak
|
||||
into the merge filename) while still being rejected on non-merge rules,
|
||||
matching rsync; their semantics remain unimplemented (accepted-but-ignored).
|
||||
Glued patterns (`-newfile`, `-e2e`) and mixed tokens (`H,!secret`) keep their
|
||||
historical parsing.
|
||||
- **Credential-file reads hardened.** Secret files (`--password-file`/
|
||||
`--early-input`/`--hash-credentials` input) are opened with `O_NOFOLLOW`, so a
|
||||
symlinked credential path now fails closed (`ELOOP`) instead of being followed
|
||||
before the owner/mode gate; literal fd-backed paths (`/dev/fd/<digits>`,
|
||||
`/proc/self/fd/<digits>`) are exempt so process substitution still works. A
|
||||
FIFO/process-substitution read now waits under a bounded ~3 s deadline for its
|
||||
writer, so a slow producer works while a connected-but-silent FIFO fails
|
||||
instead of hanging.
|
||||
- **Miscellaneous correctness fixes:** `--filter` rule count is checked
|
||||
client-side against `MAX_FILTER_RULES` before any network I/O (the receiver
|
||||
still re-checks the expanded count); unknown wire `Status` values are rejected
|
||||
as protocol errors; a mutex leak on an init-failure path, an `errno` read
|
||||
after `free()` in deferred delete application, `log_perror` misuse for
|
||||
non-`errno` conditions, and a `NULL` `server_host`/`ssh_destination`
|
||||
allocation path were fixed (the `config_create` failure now releases through
|
||||
`config_delete`); the decompression-limit log now prints the effective bound
|
||||
rather than the compile-time ceiling; the daemon umask and root test fixtures
|
||||
were hardened; `SSL_read` length is clamped and `sendfile` `poll()` retries on
|
||||
`EINTR`.
|
||||
|
||||
### Refactored / Docs
|
||||
|
||||
- Dropped dead `filter_rules_apply` and dead `--old-args` plumbing, unified
|
||||
`set_error`, deduplicated `path_is_within` and shared constants, and added
|
||||
printf format attributes (fixing format mismatches). `RSYNC_COMPAT.md`,
|
||||
`CHANGELOG.md` and `HANDOFF.md` were updated for the audit cycle; the
|
||||
`RSYNC_COMPAT.md` summary tally was corrected to match the rows.
|
||||
|
||||
## [2.28.0] - 2026-09-20
|
||||
|
||||
The rsync-parity cycle. `PROTOCOL_VERSION` moves `2.26.0 → 2.27.0 → 2.28.0`;
|
||||
client and server must run the same version (the handshake is strict). See
|
||||
`RSYNC_COMPAT.md` for the per-option matrix, now **116 ✅ / 14 ⚠️ / 27 ❌** of
|
||||
157 rows.
|
||||
|
||||
### Added
|
||||
|
||||
- **Differential rsync 3.4.1 parity gate** (`tests/integration/
|
||||
test_differential_parity.py`, `parity_harness.py`, `parity_caveats.py`): runs
|
||||
real `rsync` and FastSync over generated corpora and diffs the destination
|
||||
tree, normalized stdout and exit code. A fast subset runs on pull requests and
|
||||
the full strict set on push; the residual allowlist is empty.
|
||||
- FastSync-only long option **`--verify-basis`**: require a
|
||||
`--compare-dest`/`--copy-dest`/`--link-dest` hit to match the source by
|
||||
whole-file digest instead of trusting the size+mtime quick-check.
|
||||
- FastSync-only long option **`--delete-commit`** (implies `--delete`): the old
|
||||
atomic late whole-tree commit.
|
||||
- `--bwlimit` now parses rsync's units exactly and paces like rsync's leaky
|
||||
bucket; `--ignore-errors` reproduces rsync's skip-unreadable-subdir and
|
||||
IO-error-suppressed deletion (exit 23).
|
||||
- `--info=name/flist/del/remove/nonreg/progress` emit rsync's line format,
|
||||
including real-run `deleting`/`*deleting` lines carried by a new
|
||||
`report_deletes` wire bool.
|
||||
- Receiver-observed `--stats` counters: `Number of created files` now carries
|
||||
rsync's `(reg/dir/link/special)` breakdown and `Literal data` is exact for a
|
||||
delta transfer (extended `STATUS_STATS`).
|
||||
- `--progress` uses an opt-in paths-only pre-count so the `to-chk` denominator
|
||||
counts every entry like rsync, and emits per-directory/symlink/special names.
|
||||
- Receiver-side `protect`/`risk` filter engine (new bounded filter-rule wire
|
||||
block): `--filter='P ...'` now shields a destination-only entry like rsync.
|
||||
- `auto` for `--compress-choice`/`--checksum-choice` honors
|
||||
`RSYNC_COMPRESS_LIST`/`RSYNC_CHECKSUM_LIST`, and per-codec compression-level
|
||||
defaults match rsync.
|
||||
- Empty source directories are recreated recursively; `-R --no-implied-dirs
|
||||
--files-from` places listed files under missing implied parents; `--iconv`
|
||||
matches rsync's push direction; `--delete-delay` reports actual removals and
|
||||
recursively removes a refilled deferred directory.
|
||||
|
||||
### Changed
|
||||
|
||||
- **`--delete` now defaults to delete-during (rsync `--del`) timing.** With no
|
||||
explicit timing flag, a plain `--delete` removes each directory's extras as
|
||||
that directory is processed instead of committing one whole-tree deletion only
|
||||
after the entire transfer succeeds. This matches rsync, frees destination
|
||||
space progressively, and avoids the whole-old+new-tree peak that could
|
||||
`ENOSPC` a tight destination. The client maps the default onto the existing
|
||||
`delete_during` wire boolean, so `PROTOCOL_VERSION` stays `2.28.0`.
|
||||
- Basis directories (`--compare-dest`/`--copy-dest`/`--link-dest`) now default
|
||||
to rsync's metadata quick-check (equal size and mtime; `--size-only` drops the
|
||||
mtime leg) instead of FastSync's historical always-verify content hash.
|
||||
`--copy-dest` re-applies the source attributes, and basis materialization is
|
||||
streamed so the 256 MiB whole-file cap no longer applies to a basis hit.
|
||||
- The per-directory `STATUS_DELETE_PLAN` frame gained a one-int `apply` flag:
|
||||
the one-shot per-run config block (protected prefixes, size-pruned mirrors,
|
||||
`--delete-missing-args` exact paths) is now always transmitted first on a
|
||||
config-only carrier (`apply=false`), fixing a latent bug where a
|
||||
`--delete-missing-args` run whose `--files-from` list synchronized no directory
|
||||
never sent its exact deletions.
|
||||
|
||||
### Notes
|
||||
|
||||
- `--delete`/`--delete-during` remain caveats for the mid-transfer abort
|
||||
boundary (rsync's generator removes all planned extras ahead of its throttled
|
||||
sender; FastSync removes only reached directories — final trees agree).
|
||||
`--delete-before`, `--progress`, `--stats`, `--fuzzy` and the basis rows keep
|
||||
their documented residuals in `RSYNC_COMPAT.md`; `--filter` and
|
||||
`--delete-excluded` are now parity, including protection of a destination-only
|
||||
excluded entry under default `--delete`.
|
||||
|
||||
### Migration
|
||||
|
||||
- Scripts that relied on plain `--delete` deleting nothing until the transfer
|
||||
fully succeeded must pass **`--delete-commit`** (or `--delete-after`) to keep
|
||||
that behavior. Plain `--delete` now removes reached directories' extras during
|
||||
the transfer, exactly like rsync's default; on a completed run the final tree
|
||||
is unchanged.
|
||||
- Deployments that relied on FastSync's stricter basis verification should pass
|
||||
**`--verify-basis`**; the default now trusts the size+mtime quick-check like
|
||||
rsync.
|
||||
|
||||
## [2.26.0] - 2026-09-17
|
||||
|
||||
|
||||
### Added
|
||||
|
||||
- **Parity-completion wave.** Closed the remaining rsync-parity gaps against
|
||||
|
||||
+14
-3
@@ -1,6 +1,6 @@
|
||||
cmake_minimum_required(VERSION 3.22)
|
||||
|
||||
project(FastFileTransfer VERSION 2.26.0)
|
||||
project(FastFileTransfer VERSION 2.28.0)
|
||||
|
||||
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
|
||||
set(CMAKE_C_STANDARD 11)
|
||||
@@ -26,9 +26,9 @@ elseif(NOT SANITIZER STREQUAL "none")
|
||||
endif()
|
||||
|
||||
# --- Strict warnings option ---
|
||||
option(STRICT_WARNINGS "Enable strict warnings (Wextra, Wpedantic, Werror)" OFF)
|
||||
option(STRICT_WARNINGS "Enable strict warnings (Wextra, Wpedantic, Wformat-signedness, Werror)" OFF)
|
||||
if(STRICT_WARNINGS)
|
||||
add_compile_options(-Wextra -Wpedantic -Werror)
|
||||
add_compile_options(-Wextra -Wpedantic -Wformat-signedness -Werror)
|
||||
endif()
|
||||
|
||||
# --- Coverage option ---
|
||||
@@ -99,17 +99,21 @@ set(SHARED_SRCS
|
||||
src/shared/daemon_limits.c
|
||||
src/shared/data.c
|
||||
src/shared/delay_updates.c
|
||||
src/shared/delete.c
|
||||
src/shared/delete_commit.c
|
||||
src/shared/delete_plan.c
|
||||
src/shared/delta.c
|
||||
src/shared/file.c
|
||||
src/shared/file_list.c
|
||||
src/shared/file_receive.c
|
||||
src/shared/file_save.c
|
||||
src/shared/file_send.c
|
||||
src/shared/file_store.c
|
||||
src/shared/filter.c
|
||||
src/shared/format.c
|
||||
src/shared/hardlink.c
|
||||
src/shared/identity.c
|
||||
src/shared/incremental_check.c
|
||||
src/shared/log.c
|
||||
src/shared/metadata.c
|
||||
src/shared/motd.c
|
||||
@@ -136,9 +140,14 @@ set(SERVER_MAIN_SRCS src/server/server.c)
|
||||
# Client implementation (no main): everything except the CLI entry point.
|
||||
set(CLIENT_CORE_SRCS
|
||||
src/client/change_list.c
|
||||
src/client/client_manifest.c
|
||||
src/client/client_report.c
|
||||
src/client/client_scan.c
|
||||
src/client/client_send.c
|
||||
src/client/client_validation.c
|
||||
src/client/scanner.c
|
||||
src/client/scanner_filter.c
|
||||
src/client/scanner_parallel.c
|
||||
src/client/usage.c
|
||||
)
|
||||
set(CLIENT_MAIN_SRCS src/client/client_cli.c)
|
||||
@@ -221,10 +230,12 @@ set(TEST_SRCS
|
||||
tests/test_daemon_limits.c
|
||||
tests/test_data.c
|
||||
tests/test_delay_updates.c
|
||||
tests/test_delete_plan.c
|
||||
tests/test_delta.c
|
||||
tests/test_file.c
|
||||
tests/test_file_list.c
|
||||
tests/test_file_sendfile.c
|
||||
tests/test_filter.c
|
||||
tests/test_format.c
|
||||
tests/test_fuzz_smoke.c
|
||||
tests/test_glob.c
|
||||
|
||||
+215
-22
@@ -1,16 +1,31 @@
|
||||
# FastSync — Session Handoff (2026-09-17)
|
||||
# FastSync — Session Handoff (2026-09-21)
|
||||
|
||||
## Current status
|
||||
- **Release `v2.21.0`** tagged (`919a729`, "Release v2.21.0"); full CI green
|
||||
(run 552: lint, build-and-test, ASan, UBSan, fuzz-build, coverage, valgrind).
|
||||
`dev` has the release commit plus later doc-only merges (a README refresh and
|
||||
this handoff).
|
||||
- **Release PR #284 (`dev` -> `main`)** open, CI green (run 553).
|
||||
`main` is protected: it needs review/approval to merge.
|
||||
https://gitea.tap-tap.win/TapTap/FastSync/pulls/284
|
||||
- **`PROTOCOL_VERSION` = `"2.26.0"`** (`src/shared/config.h`); CMake
|
||||
`project(FastFileTransfer VERSION 2.26.0)`.
|
||||
- Working tree clean; no wave worktrees remain.
|
||||
- **Release `v2.28.0`** is tagged and merged to `main`: tag `v2.28.0` points at
|
||||
`ee6523a`, and the PR #304 merge commit `b4d54504` is on `main`.
|
||||
- **`dev` is at `0fbb9de`** — the merge of parity cycle 2.29 (PR #305). The old
|
||||
`558782d` (incremental-check flake fix) is an ancestor.
|
||||
- **`PROTOCOL_VERSION` = `"2.28.0"`** (`src/shared/config.h`); CMake
|
||||
`project(FastFileTransfer VERSION 2.28.0)`.
|
||||
- **Parity cycle 2.29 is merged to `dev`** (PR #305), no wire change. It closed
|
||||
the scanner-order, delete-timing, relative-basis and fuzzy-eligibility
|
||||
residuals and improved the `--info`/`--stats`/`--debug` partials. Parity
|
||||
matrix: **120 ✅ / 10 ⚠️ / 27 ❌ = 157**. Remaining ⚠️ rows: `--info`,
|
||||
`--debug`, `--msgs2stderr`, `--stats`, `--progress`, `--delete-before`, the
|
||||
three basis-dir options, and `-y`/`--fuzzy`.
|
||||
- **Audit cycle complete on branch `fix/audit-cycle`** (branched from `dev` @
|
||||
`0fbb9de`), integration PR to `dev` pending. No wire change
|
||||
(`PROTOCOL_VERSION` stays 2.28.0). It lands the receiver/client security and
|
||||
correctness fixes — `--temp-dir` symlink-escape confinement, special-bit
|
||||
masking under a super-off policy, daemon `umask(022)`, the `-z` decompression
|
||||
ceiling raised to the 256 MiB whole-file bound, `--bwlimit` pacing the
|
||||
plaintext `--sendfile` path, `--partial-dir` implying `--partial`, rejection
|
||||
of unsupported filter modifiers (`x`/`e`/`n`/`w`), client-side
|
||||
`MAX_FILTER_RULES` enforcement, unknown wire `Status` rejection, and the
|
||||
accompanying refactors/docs. The parity matrix is unchanged at
|
||||
**120 ✅ / 10 ⚠️ / 27 ❌ = 157**; this docs pass (worktree `fix/audit-docs2`)
|
||||
corrects the `RSYNC_COMPAT.md` summary tally to match the rows.
|
||||
|
||||
|
||||
## What landed this session
|
||||
1. **Wave 8 (refactors):** Config X-macro wire table; single-owner `authorized_root`;
|
||||
@@ -39,19 +54,197 @@
|
||||
docs state push-only / remote-source unsupported.
|
||||
5. **Preserve-attribute split (protocol 2.22.0)** landed on `feat/preserve-attr-split`: per-attribute `-p/-t/-o/-g` + `--no-*` negations, `-a` = `-rlptgoD`, and the 2.21.0 → 2.22.0 wire bump.
|
||||
6. **Rsync-parity wave (protocol 2.23.0)** on `feat/rsync-parity`: rsync short options/clustering/attached values (`-r`/`-b`/`-L`/`-B`, `-av`, `-aAX`, `-B1000`, `-essh`, `-MOPT`), `-c` checksum quick-check, `--checksum-choice`/`--compress-choice` validation and seed randomization, rsync timeout/max-alloc defaults, temp-dir confinement + `EXDEV` fallback, ownership/mapping parity (numeric-ids modifier, map ranges/`*`/empty-FROM, `--chown`+map conflicts, fake-super resolved-owner record), verbatim symlink storage with rsync `--safe-links`/`--munge-links`, socket recreation under `--specials`, `--chmod` 3.4.1 semantics, and delete scoping + `--max-delete` partial/exit-25. Wire: appended delete-manifest synchronized-directory section and `STATUS_DELETE_LIMIT`.
|
||||
7. **Parity-completion wave (protocol 2.24.0 → 2.26.0)** on `feat/parity-completion`: per-directory delete plans (`STATUS_DELETE_PLAN`) for `--delete-during`/`--delete-delay`; receiver `STATUS_STATS` counters feeding `--stats`/`--progress` and `--out-format %b/%c/%C`, plus `-n --delete` lines; `lz4`/`zlib`/`zlibx` compression and `md4`/`sha1`/`none` checksums with `auto` negotiation (default `xxh128`/`zstd`); general `-R`/`--no-implied-dirs`/`-d`; the full filter grammar (`merge`/`dir-merge`/`hide`/`show`/`protect`/`risk`/`clear` + modifiers) and corrected `-F`/`-FF`; receiver-side `--chown`/map TO-name resolution; absolute basis dirs + `--link-dest` relink; receiver-side `--ignore-existing` short-circuit; `--preallocate` over `--sparse` via `fallocate(2)`; `--iconv=.`/`-`/`--no-iconv`; lone `-h` help; aliases `--ignore-non-existing`/`--protect-args`/`--msgs2stderr`; and the full `--info`/`--debug` vocabulary. `RSYNC_COMPAT.md` reclassifies the matrix to 106 ✅ / 27 ⚠️ / 23 ❌.
|
||||
7. **Parity-completion wave (protocol 2.24.0 → 2.26.0)** on `feat/parity-completion`: per-directory delete plans (`STATUS_DELETE_PLAN`) for `--delete-during`/`--delete-delay`; receiver `STATUS_STATS` counters feeding `--stats`/`--progress` and `--out-format %b/%c/%C`, plus `-n --delete` lines; `lz4`/`zlib`/`zlibx` compression and `md4`/`sha1`/`none` checksums with `auto` negotiation (default `xxh128`/`zstd`); general `-R`/`--no-implied-dirs`/`-d`; the full filter grammar (`merge`/`dir-merge`/`hide`/`show`/`protect`/`risk`/`clear` + modifiers) and corrected `-F`/`-FF`; receiver-side `--chown`/map TO-name resolution; absolute basis dirs + `--link-dest` relink; receiver-side `--ignore-existing` short-circuit; `--preallocate` over `--sparse` via `fallocate(2)`; `--iconv=.`/`-`/`--no-iconv`; lone `-h` help; aliases `--ignore-non-existing`/`--protect-args`/`--msgs2stderr`; and the full `--info`/`--debug` vocabulary. `RSYNC_COMPAT.md` reclassifies the matrix to 106 ✅ / 27 ⚠️ / 23 ❌; the later rsync-parity-stats pass (`fix/parity-stats`) moves it to 107 ✅ / 25 ⚠️ / 24 ❌ (see item 8).
|
||||
8. **rsync-parity-stats pass** on `fix/parity-stats` (no wire change, `PROTOCOL_VERSION` stays `2.26.0`): `--delete-delay` now reports only entries it actually removes, while the `--max-delete` budget is charged at plan/snapshot time (`planned`, via `defer_add`) to bound the deferred list (a refilled deferred directory that survives `ENOTEMPTY` is not reported but still consumes budget); `--stats` gained the `(reg/dir/link/special)` `Number of files` breakdown and now counts only regular files actually stored for `Number of regular files transferred`/transferred size/literal data (up-to-date re-runs report 0); `Total file size` includes symlink target lengths; `--progress` prints the leading `./` root line and counts it in `to-chk` so a single-file transfer matches rsync; and `%C` uses the selected transfer checksum with `checksum_digest_file` supporting md4/sha1/none, byte-identical to rsync for every algorithm. `--out-format` reclassified ❌ (`%b`/delta-`%c` are protocol-specific). Differential + regression tests added; full suite + ASan + clang-format + cppcheck clean.
|
||||
9. **Option-parity wave (protocol 2.26.0 → 2.27.0, on `fix/parity-options`):**
|
||||
`--bwlimit` now ports rsync 3.4.1's units/quantization and paces like its
|
||||
leaky bucket; `--ignore-errors` reproduces rsync's default (an I/O error
|
||||
skips deletion unless the flag is set; the readable tree still transfers and
|
||||
the run exits 23) across every delete timing; the `--info` categories with a
|
||||
FastSync event (`name`/`flist`/`del`/`remove`/`nonreg`/`progress`) emit
|
||||
rsync's line format, with real-run `deleting`/`*deleting` lines carried over
|
||||
the new trailing config bool `report_deletes` (golden wire updated by
|
||||
`tests/test_config.c`). Two residuals were reclassified **divergent**: `-M`
|
||||
over daemon/TCP (no argv channel in FastSync's binary config handshake;
|
||||
rsync-daemon differential pins the rsync behavior) and receiver-side
|
||||
`protect`/`risk` re-derivation for destination-only entries (would need a
|
||||
receiver filter engine; differential pins the divergence — **reversed by
|
||||
track 4a below**, which adds that engine). The options pass
|
||||
stands at **110 ✅ / 21 ⚠️ / 26 ❌**. New `tests/integration/test_option_parity.py`
|
||||
holds the rsync differentials (bwlimit parse+rate, info lines, real-setpriv
|
||||
`--ignore-errors`, rsync-daemon `-M`, filter-protect pin).
|
||||
|
||||
10. **rsync-parity-fs pass** on `fix/parity-fs` (no wire change of its own; integrated
|
||||
on top of the 2.27.0 options wave): recursive transfers now recreate empty source directories (and
|
||||
`-m/--prune-empty-dirs` still suppresses them), a directory entry replaces a
|
||||
blocking destination regular file, and `-R --no-implied-dirs --files-from`
|
||||
places a listed file under a missing implied parent with default attributes
|
||||
instead of refusing (real rsync 3.4.1 parity, differential-tested). `--iconv`
|
||||
now reproduces rsync's push direction (destination charset = the spec's REMOTE
|
||||
half; a server `--iconv` overrides), and `-T/--temp-dir` relative semantics are
|
||||
confirmed identical while the absolute-path confinement is a deliberate
|
||||
divergence. The basis-dir options, `--delay-updates` and `--dry-run` were
|
||||
reclassified to ❌ after a differential test reproduced each exact residual
|
||||
(basis content verification, fixed staging-name collision, and dry-run
|
||||
would-delete over-report). `--fuzzy` was also reclassified to ❌ (deterministic
|
||||
heuristic with a 10× size window, not rsync's matcher), but its residual is the
|
||||
candidate-selection heuristic itself: the final tree is byte-exact by design, so
|
||||
it is pinned by the `TestFuzzy` threshold suite rather than a byte-level rsync
|
||||
differential. (Track 5b later found the name heuristic is rsync's own and moved
|
||||
the row ❌ → ⚠️, leaving only the narrower delta size window; see entry 15.) The parity-review pass then moved `--delete-delay` to ⚠️ (the
|
||||
plan-time `--max-delete` charge and non-recursive deferred removal differ from
|
||||
rsync when a snapshotted entry fails removal). Differential-gate allowlist
|
||||
entries `min_size`/`empty_dirs_recursive`/`dirs_plain` were removed. The
|
||||
integrated stats+options+fs branch stands at **111 ✅ / 13 ⚠️ / 33 ❌ = 157**;
|
||||
full suite + ASan + clang-format + cppcheck clean.
|
||||
|
||||
11. **No-wire parity track 1** on `feat/parity-2.28` (no protocol change):
|
||||
`-n --delete` now sends the same filter-excluded + size-pruned protected
|
||||
prefixes and synchronized-directory scope as a real run (dry-run would-delete
|
||||
matches rsync for source-derived protections; the destination-only exclude
|
||||
residual was later closed by track 4a, readdir ordering remains);
|
||||
`--delete-delay` now charges
|
||||
`--max-delete` on actual removals and re-scans a queued directory at commit
|
||||
to remove content created after the plan, with an independent deferred-list
|
||||
cap (only partial-delete ordering remains); and `--info=name2` emits `NAME is
|
||||
uptodate` plus the leading `./` root name line for `--info=name` (only the
|
||||
root-line trigger condition and receiver-side `skip` wording remain). Matrix
|
||||
now **111 ✅ / 14 ⚠️ / 32 ❌ = 157**; differential + unit tests added in
|
||||
`test_features.py`, `test_option_parity.py`, the unit test
|
||||
`tests/test_delete_plan.c`,
|
||||
`test_delete_delay_budget_parity.py`, `test_delete_timing_parity.py`.
|
||||
12. **No-wire parity track 2b** on `feat/parity-2.28` (no protocol change):
|
||||
`--progress`/`-P`/`--info=progress` (when not `--quiet`) now run an opt-in
|
||||
paths-only metadata pre-count (no file reads/hashing) that supplies rsync's
|
||||
full file-list total for the `to-chk` denominator and the directory names,
|
||||
and emits per-directory/symlink/special name lines, in both the sequential
|
||||
and `--threads` paths. `--delete-during`/`--delete-delay` reuse their
|
||||
keep-set pre-scan instead of a second walk; non-progress runs are
|
||||
unaffected. Differential tests (`progress`/`progress_threads` over a new
|
||||
`multidir` corpus) match rsync's name set and `to-chk` denominator on a
|
||||
fresh transfer, and the single-file byte-identical test still passes;
|
||||
emission order (rsync's sorted depth-first vs FastSync's readdir/BFS stream)
|
||||
plus re-run over-naming (unconditional `./`, ancestor dirs named with a
|
||||
transferred child, and no quick-check for symlinks/empty dirs) remain the
|
||||
caveats, so the row stays ⚠️ and the matrix is unchanged at
|
||||
**111 ✅ / 14 ⚠️ / 32 ❌ = 157**.
|
||||
|
||||
13. **Wire parity track 4a** on `feat/parity-2.28` (`PROTOCOL_VERSION` stays
|
||||
`2.28.0`): the receiver now has a delete-time filter engine. The sender
|
||||
compiles its root-level selection rules exactly as the scanner does
|
||||
(`filter_base_build`) and streams them as one bounded, self-describing
|
||||
config-frame block (action, sides, anchored, dir-only, negate, owner,
|
||||
pattern; bounded rule count and pattern bytes, unknown action/sides is a
|
||||
protocol error). The receiver reconstructs `protect_rules` and applies them
|
||||
first-match-wins to each extraneous destination path in every delete timing
|
||||
(the whole-tree commit walker, the `--delete-during`/`--delete-delay`
|
||||
per-directory plans, and the `-n` would-delete enumeration), so a
|
||||
`P *.log` rule protects a destination-only `extra.log` like rsync (with
|
||||
`risk` cancelling); the sender-derived protected-prefix behavior is
|
||||
preserved when no rules are sent and `--delete-excluded` semantics are
|
||||
unchanged. Per-directory merge (`:`/`.`) receiver re-derivation remains the
|
||||
residual. `TestFilterProtect` (real + dry-run) plus differential cases
|
||||
`filter_protect`, `filter_protect_during`, `filter_protect_delay` added and
|
||||
the `--filter=RULE` row moves ❌ → ✅: matrix now
|
||||
**115 ✅ / 11 ⚠️ / 31 ❌ = 157**; unit tests, the three named integration
|
||||
files, clang-format and cppcheck clean.
|
||||
|
||||
14. **Wire parity track 5a** on `feat/parity-2.28` (`PROTOCOL_VERSION` stays
|
||||
`2.28.0` by project decision): the three basis-dir options now default to
|
||||
rsync's metadata quick-check (equal size + equal mtime, or size alone under
|
||||
`--size-only`; `-I` disables matching) instead of FastSync's historical
|
||||
xxHash64 content equality, so a same-size/different-content basis is trusted
|
||||
exactly as rsync trusts it. A new FastSync-only, long-only `--verify-basis`
|
||||
flag restores the strict whole-file content equality; its bool is appended to
|
||||
the basis block of the config frame (golden wire frame 882 → 886 bytes).
|
||||
`--verify-basis` streams the confined basis descriptor to hash it, and a
|
||||
basis hit is no longer capped at the 256 MiB whole-file payload bound:
|
||||
`--copy-dest` streams the basis through a bounded buffer and `--link-dest`'s
|
||||
copy fallback streams from the basis, so an over-limit hit materializes (a
|
||||
basis MISS still falls back to the normal transfer and keeps its own bound).
|
||||
A `--copy-dest` hit re-applies the SOURCE attributes (the sender transmits
|
||||
the source metadata with the basis check frame), matching rsync's
|
||||
"copy then fix attributes"; a `--link-dest` success keeps the shared inode's
|
||||
attributes (writing through it would mutate the basis). Differential cases
|
||||
`copy_dest` and `verify_basis` added; `test_basis_dir_size_only_content_residual`
|
||||
converted to a passing parity assertion; `TestBasisDestDirs` updated for the
|
||||
new default + `--verify-basis`; unit tests cover the quick-check/verify
|
||||
decision and the same-size/different-content handshake. The
|
||||
`--compare-dest`/`--copy-dest`/`--link-dest` rows move ❌ → ⚠️ (relative-DIR
|
||||
resolution base and over-limit MISS refusal): matrix now
|
||||
**116 ✅ / 13 ⚠️ / 28 ❌ = 157**.
|
||||
|
||||
15. **No-wire parity track 5b** on `feat/parity-2.28` (`PROTOCOL_VERSION` stays
|
||||
`2.28.0` by project decision): `-y`/`--fuzzy` reclassified ❌ → ⚠️. A probe
|
||||
against real rsync 3.4.1 (pinned `-B8192`, repeated-content 64 KiB corpus)
|
||||
showed the name heuristic is already rsync's (`util1.c fuzzy_distance` /
|
||||
`find_filename_suffix` + the exact size+mtime pass) and the output is always
|
||||
byte-exact; the only residual is candidate ELIGIBILITY, because FastSync's
|
||||
`delta_should_attempt` gate caps the size ratio at 10× and requires both
|
||||
files ≥ 16 KiB while rsync will reuse a basis from 0.25× to 10000× and below
|
||||
16 KiB. The choice is observable only as `--stats` bandwidth counters. Added
|
||||
differential case `fuzzy_basis` (same-suffix sibling, one name edit,
|
||||
identical content, block size pinned) asserting tree **and** normalized
|
||||
`--stats` parity where the choices coincide, plus `TestFuzzy` pinning the
|
||||
window boundary on both sides (>10× and <16 KiB siblings declined by
|
||||
FastSync while rsync uses them, both trees byte-identical). Matrix now
|
||||
**116 ✅ / 14 ⚠️ / 27 ❌ = 157**.
|
||||
|
||||
16. **Lockstep delete-default track 6** on `feat/parity-2.28` (`PROTOCOL_VERSION`
|
||||
stays `2.28.0`): plain `--delete` now defaults to rsync's delete-during
|
||||
(`--del`) timing, normalized on the client onto the existing `delete_during`
|
||||
wire bool. The old late whole-tree commit is opt-in via `--delete-after` or
|
||||
the FastSync-only long `--delete-commit` (identical `delete_after` timing).
|
||||
`-d/--dirs` still falls back to the end commit, `--delay-updates` still
|
||||
deletes before publication, and `--files-from`/`-R` scope is unchanged. The
|
||||
`STATUS_DELETE_PLAN` frame gained a one-int `apply` flag so the per-run
|
||||
config block (including `--delete-missing-args` exact paths) is always
|
||||
transmitted, on a config-only carrier when the scope allows no directory
|
||||
plan — fixing a latent bug with a file-only `--files-from` list. Differential
|
||||
cases `delete`/`delete_commit`/`filter_protect_after` plus the extended
|
||||
`test_delete_timing_parity.py` (plain `--delete` mid-abort removes reached
|
||||
extras, `--delete-commit` defers) pass; full `-m "not setpriv"` suite,
|
||||
clang-format and cppcheck clean. Matrix unchanged at
|
||||
**116 ✅ / 14 ⚠️ / 27 ❌ = 157** (the `--delete`/`--delete-during` rows stay
|
||||
⚠️ for the abort boundary; `--delete-after` stays ✅).
|
||||
|
||||
17. **Audit cycle** on `fix/audit-cycle` (from `dev` @ `0fbb9de`;
|
||||
`PROTOCOL_VERSION` stays `2.28.0`): a security/correctness pass over the
|
||||
parity-2.29 baseline. It raises the decompression ceiling to the 256 MiB
|
||||
protocol whole-file bound (`-z` on 100–256 MiB files now works), paces the
|
||||
plaintext-TCP `--sendfile` path with `--bwlimit`, confines the `--temp-dir`
|
||||
scratch dir by the fd's real path (symlink escape refused), masks
|
||||
client-controlled setuid/setgid/sticky bits when super activities are not
|
||||
permitted, sets the daemon umask to `022`, makes `--partial-dir` imply
|
||||
`--partial`, rejects the unsupported filter modifiers (`x`/`e`/`n`/`w`),
|
||||
enforces `MAX_FILTER_RULES` client-side, rejects unknown wire `Status`
|
||||
values, and hardens credentials/signal handling (with the accompanying
|
||||
refactors and docs). No row changes classification, so the matrix stays
|
||||
**120 ✅ / 10 ⚠️ / 27 ❌ = 157**. This docs pass is on `fix/audit-docs2`.
|
||||
|
||||
## Next steps
|
||||
1. **Merge PR #284** (`dev` -> `main`) once reviewed (protected branch).
|
||||
2. **Deferred security items** (documented, not implemented):
|
||||
- Pre-auth config/daemon-auth handshake has no aggregate wall-clock deadline
|
||||
(per-message timeout only) — slowloris holds connection slots.
|
||||
- Per-source registry fails open when the shared table is full (per-module/global
|
||||
caps and host ACLs still apply); consider fail-closed or larger/evicting table.
|
||||
- SCRAM-like daemon auth has no TLS channel binding (and is not RFC 5802).
|
||||
- `cleanup()` signal handler calls non-async-signal-safe teardown; daemon `umask(0)`.
|
||||
- Wire protocol assumes homogeneous word size/endianness (lengths are native
|
||||
`size_t`) — document or move to fixed-width framing.
|
||||
1. **Open and merge the audit-cycle PR** (`fix/audit-cycle`, including this
|
||||
`fix/audit-docs2` docs pass) into `dev` once reviewed. `dev` is the default
|
||||
branch; all PRs target `dev`, never `main` directly.
|
||||
2. **Remaining deferred items:**
|
||||
- **Large structural refactors:** delete-engine consolidation
|
||||
(`delete_extras_fd`/`manifest_delete_extras`/the delete-plan path),
|
||||
god-function splits, and translation-unit splits.
|
||||
- **`--progress`/`--info` receiver→sender event channel:** the root `./`
|
||||
line, ancestor-directory suppression, receiver-side `skip`/`backup` echo,
|
||||
and symlink/empty-dir quick-check feedback.
|
||||
- **`--delete-before` phase-0 keep-set** (rsync fixes the file list before
|
||||
the data pass; FastSync keeps its pre-scan snapshot race).
|
||||
- **>256 MiB single-file streaming** (B4, the general whole-file limit).
|
||||
- **Wire native-size framing:** lengths are native `size_t` and the protocol
|
||||
assumes homogeneous word size/endianness — document or move to fixed-width
|
||||
framing.
|
||||
- **SCRAM-like daemon auth channel binding:** no TLS channel binding today
|
||||
(and it is not RFC 5802).
|
||||
- Still-open security nits: the pre-auth config/daemon-auth handshake has no
|
||||
aggregate wall-clock deadline (per-message timeout only — slowloris holds
|
||||
connection slots); the per-source registry fails open when the shared table
|
||||
is full (per-module/global caps and host ACLs still apply).
|
||||
3. **Out of scope / intentional:** pull (remote source) mode is **not** planned —
|
||||
FastSync is push-only; see `RSYNC_COMPAT.md#direction`.
|
||||
|
||||
|
||||
@@ -75,8 +75,9 @@ matrix is classified as parity, caveat, or divergent in
|
||||
owner, group, devices, and special files — and does not imply compression or
|
||||
multithreading (see [Client](#client)). Ownership application is still
|
||||
privilege-gated: a receiver that cannot `chown` logs a warning and skips it.
|
||||
Under `-p` the source mode is copied exactly, including setuid/setgid/sticky
|
||||
and group/other-write bits (strict rsync parity; see
|
||||
Under `-p` the source mode is copied exactly, including group/other-write
|
||||
bits; setuid/setgid/sticky bits are copied only when super-user activities are
|
||||
permitted, and are masked under `SUPER_MODE_OFF`/`--no-super` (see
|
||||
[`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)).
|
||||
- Symlink transfer stores targets **verbatim** (`-l`/`--links`), including
|
||||
absolute and `..`-bearing targets, matching rsync. The receiver does not
|
||||
@@ -113,13 +114,18 @@ matrix is classified as parity, caveat, or divergent in
|
||||
(`-B1000`, `-essh`, `-MOPT`, `--opt=value`) are accepted, matching rsync.
|
||||
- `-r`, `-b`, `-L`, and `-B` are parsed with the rsync short names.
|
||||
- `--stats` prints the counters FastSync can observe plus the receiver-only
|
||||
counters (`Matched data`, deleted files) reported over the wire; rsync's
|
||||
per-type `Number of files` breakdown is not reproduced. `--progress` prints
|
||||
rsync-style per-file blocks (without rsync's leading `./` line).
|
||||
counters reported over the wire (`Matched data`, deleted files, and the
|
||||
created/literal counters); `Number of files` and `Number of created files`
|
||||
carry rsync's per-type breakdown. `--progress` prints rsync-style per-file
|
||||
blocks including the leading `./` line, and (when progress is requested) a
|
||||
paths-only pre-count supplies rsync's `to-chk` denominator.
|
||||
- Codecs match rsync 3.4.1: `zstd`/`lz4`/`zlib`/`zlibx` compression and
|
||||
`xxh128`/`xxh3`/`xxh64`/`md5`/`md4`/`sha1`/`none` checksums, negotiated with
|
||||
`auto`; `zlibx` behaves as `zlib`, and the transfer checksum is not separately
|
||||
selectable.
|
||||
`xxh128`/`xxh3`/`xxh64`/`md5`/`md4`/`sha1`/`none` checksums. `auto` honors
|
||||
`RSYNC_COMPRESS_LIST`/`RSYNC_CHECKSUM_LIST` and otherwise follows rsync's
|
||||
compiled-in order. An omitted `--compress-level` uses the codec's rsync
|
||||
default (zstd 3, zlib/zlibx 6, lz4 ignored); `zlib`/`zlibx` share the
|
||||
literal-only zlib path (rsync's zlibx semantics), and the transfer checksum is
|
||||
not separately selectable.
|
||||
|
||||
The detailed flag matrix is maintained in
|
||||
[`RSYNC_COMPAT.md`](RSYNC_COMPAT.md). It reports each row as **parity**,
|
||||
@@ -167,7 +173,7 @@ This produces `./build/client` and `./build/server`. `compile_commands.json` is
|
||||
| `--preserve` | Preserve mode and mtime (`-p` + `-t`; add `-o`/`-g` for owner/group or `-U`/`--atimes` for atime; `-N`/`--crtimes` captures birth time but cannot apply it) |
|
||||
| `-U, --atimes` | Preserve access times. Captured with the metadata payload; does not enable ownership. |
|
||||
| `-N, --crtimes` | Capture birth time; cannot be applied (documented divergence) |
|
||||
| `-p, --perms` | Preserve permission bits. Strict rsync parity: the source mode is copied exactly, including setuid/setgid/sticky and group/other-write bits |
|
||||
| `-p, --perms` | Preserve permission bits. The source mode is copied exactly, including group/other-write bits; setuid/setgid/sticky are copied only when super-user activities are permitted (`SUPER_MODE_OFF`/`--no-super` masks them) |
|
||||
| `-t, --times` | Preserve modification times |
|
||||
| `-o, --owner` | Preserve the source owner (privilege-gated; mapped by name on the receiver with a numeric fallback) |
|
||||
| `-g, --group` | Preserve the source group (privilege-gated; mapped by name on the receiver with a numeric fallback) |
|
||||
@@ -199,26 +205,29 @@ This produces `./build/client` and `./build/server`. `compile_commands.json` is
|
||||
| `-u, --update` | Skip files newer than the source on the receiver |
|
||||
| `--incremental` | Skip files unchanged since last transfer (size + mtime). Auto-enables `--preserve`. Incompatible with `--chunk-serialization`. |
|
||||
| `--existing` | Skip files not already present at the destination; update existing files normally. |
|
||||
| `--compare-dest <dir>` | Extra comparison basis: unchanged files are not transferred (requires/implies `--incremental`) |
|
||||
| `--copy-dest <dir>` | Like `--compare-dest`, but copies the unchanged file from DIR into the destination |
|
||||
| `--link-dest <dir>` | Like `--copy-dest`, but hard-links the unchanged file from DIR (repeatable; earlier DIRs win) |
|
||||
| `--delete` | Delete files on receiver not present in source (default timing: delete-after, i.e. only after the whole transfer succeeded). Scoped to the synchronized directories, so `--files-from` subsets are safe |
|
||||
| `--ignore-existing` | Skip files that already exist on the receiver; like rsync it does not apply to directories or symlinks. |
|
||||
| `--compare-dest <dir>` | Extra comparison basis: unchanged files are not transferred (requires/implies `--incremental`; a basis MISS above the 256 MiB whole-file payload bound is refused — see [`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)) |
|
||||
| `--copy-dest <dir>` | Like `--compare-dest`, but copies the unchanged file from DIR into the destination (same basis-size caveat; see [`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)) |
|
||||
| `--link-dest <dir>` | Like `--copy-dest`, but hard-links the unchanged file from DIR (repeatable; earlier DIRs win; same basis-size caveat; see [`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)) |
|
||||
| `--verify-basis` | FastSync-only: require a basis hit (`--compare-dest`/`--copy-dest`/`--link-dest`) to match the source by whole-file digest instead of trusting the size+mtime quick-check (default matches rsync) |
|
||||
| `--delete` | Delete files on receiver not present in source (default timing: delete-during, matching rsync, so destination space is freed progressively). Scoped to the synchronized directories, so `--files-from` subsets are safe |
|
||||
| `--delete-before` | Delete extras before the transfer starts (implies `--delete`) |
|
||||
| `--delete-during`, `--del` | Delete extras once the keep-set is known, before data is applied (implies `--delete`) |
|
||||
| `--delete-delay` | Delete extras only after a successful transfer (implies `--delete`) |
|
||||
| `--delete-after` | Explicit delete-after timing (implies `--delete`) |
|
||||
| `--delete-commit` | FastSync-only: keep the pre-2.28 atomic timing — delete only after the whole transfer succeeded (identical timing to `--delete-after`) |
|
||||
| `--delete-excluded` | Also delete filter-excluded destination mirrors (size-pruned mirrors stay protected) |
|
||||
| `--max-delete <n>` | Delete at most n destination entries; the rest are skipped and the run exits 25 (partial), matching rsync |
|
||||
| `--delay-updates` | Put updated files into place only at the end of the transfer (`--force` is honored at publication) |
|
||||
| `--delay-updates` | Put updated files into place only at the end of the transfer (`--force` is honored at publication; the fixed `.fastsync-stage` staging name diverges from rsync — see [`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)) |
|
||||
| `-T, --temp-dir <dir>` | Scratch directory for temp files before the atomic install; confined to the receive root (relative only), with an `EXDEV` non-atomic copy fallback |
|
||||
| `-n, --dry-run` | Report what would be transferred without mutating the destination. Since protocol 2.21.0 a server-routed target contacts the receiver and reports would-transfer based on receiver state; a plain local destination keeps the client-side scan. Never mutates or deletes. |
|
||||
| `-v, --verbose` | Enable debug logging |
|
||||
| `-q, --quiet` | Suppress non-error output |
|
||||
| `--progress` | Show rsync-style per-file progress blocks from the receiver's wire counters (FastSync does not print rsync's leading `./` line) |
|
||||
| `--progress` | Show rsync-style per-file progress blocks from the receiver's wire counters; the root `./` line is printed whenever progress is active (rsync prints it only when the transfer root is created) |
|
||||
| `-P` | Enables partial-transfer mode + progress output; interrupted writes retain the already-written temp for resumption |
|
||||
| `--stats` | Print transfer statistics at end (bytes, files, timing), including the receiver-only counters reported over the wire; rsync's per-type `Number of files` breakdown is not reproduced |
|
||||
| `--stats` | Print transfer statistics at end (bytes, files, timing), including the receiver-only counters reported over the wire; `Number of files` and `Number of created files` carry rsync's per-type breakdown (deleted files are reported as a single total) |
|
||||
| `-i, --itemize-changes` | Print an rsync-style per-file change line |
|
||||
| `--out-format=FORMAT` | Output format for changed files (`%f %n %l %b %M %%`) |
|
||||
| `--out-format=FORMAT` | Output format for changed files (`%f %n %l %b %c %C %i %M %%`) |
|
||||
| `--list-only` | List source files instead of transferring |
|
||||
| `--fsync` | Fsync every written file before publication |
|
||||
| `-h, --human-readable` | Format transfer byte/rate counts with rsync's decimal (base-1000) units |
|
||||
@@ -239,7 +248,7 @@ This produces `./build/client` and `./build/server`. `compile_commands.json` is
|
||||
| `-4, --ipv4` | Force IPv4 for destination resolution |
|
||||
| `-6, --ipv6` | Force IPv6 for destination resolution |
|
||||
| `--sockopts=OPTS` | Comma-separated OPT=VAL socket options applied before connect (`TCP_NODELAY`, `SO_KEEPALIVE`, `SO_RCVBUF`, `SO_SNDBUF`, `SO_REUSEADDR`) |
|
||||
| `--bwlimit <KB/s>` | Bandwidth limit in kilobytes per second |
|
||||
| `--bwlimit <KB/s>` | Bandwidth limit in kilobytes per second; also paces `--sendfile` transfers |
|
||||
| `--chunk-size <n>` | Chunk size in bytes (default: 10485760) |
|
||||
| `--timeout <sec>` | I/O timeout in seconds, applied to both the socket (`SO_RCVTIMEO`/`SO_SNDTIMEO`) and the per-message protocol poll deadline. Default `0` = disabled (matching rsync); `0` disables it. `--no-timeout` is the negation. The value is not sent on the wire; the server side keeps its own safe floor. |
|
||||
| `--contimeout <sec>` | Connection timeout in seconds (default: 60, matching rsync); `0` disables it (`--no-contimeout` is the negation) |
|
||||
@@ -307,17 +316,22 @@ transfer is never aborted.
|
||||
`timeout`, `contimeout`, `quiet`, `stats`, `max_depth`, and `log_file` are
|
||||
client-only.
|
||||
5. **Queue** — thread-safe bounded queue with condition variables.
|
||||
6. **DirectoryScanner** — recursive BFS traversal with exclude and include
|
||||
pattern support, max-depth enforcement.
|
||||
6. **DirectoryScanner** — recursive traversal that buffers and sorts each
|
||||
directory (non-directories ascending, then directories ascending) and walks
|
||||
depth-first in rsync flist order, with exclude and include pattern support and
|
||||
max-depth enforcement.
|
||||
|
||||
### Key Algorithms
|
||||
|
||||
1. **File scanning** — BFS directory traversal; entries matched against exclude
|
||||
and include patterns, with max-depth enforced.
|
||||
1. **File scanning** — sorted depth-first traversal in rsync flist order (each
|
||||
directory's non-directories ascending, then its directories ascending);
|
||||
entries matched against exclude and include patterns, with max-depth
|
||||
enforced. The `--threads` parallel scanner remains unordered.
|
||||
2. **Chunking** — files accumulated until the `chunk_size` threshold (default
|
||||
10 MiB) is reached, then flushed.
|
||||
3. **Compression** — streaming zstd via `ZSTD_compressStream2()` /
|
||||
`ZSTD_decompressStream()`.
|
||||
`ZSTD_decompressStream()`, with lz4 and zlib/zlibx codecs also supported
|
||||
(selectable with `--compress-choice`).
|
||||
4. **Network protocol** — status-code-driven exchange with metadata packing,
|
||||
keep-alive, and abort support.
|
||||
5. **Incremental check** — the client sends `STATUS_CHECK` + path + size +
|
||||
@@ -372,6 +386,8 @@ Received files are written to a temporary path (suffixed with `.tmp`) and then a
|
||||
- C11 compiler
|
||||
- CMake >= 3.22
|
||||
- zstd library
|
||||
- zlib library
|
||||
- lz4 library
|
||||
- OpenSSL (development headers and libraries)
|
||||
- pthreads
|
||||
- SSH client (for SSH transport mode only)
|
||||
@@ -380,12 +396,12 @@ Received files are written to a temporary path (suffixed with `.tmp`) and then a
|
||||
|
||||
**Ubuntu/Debian:**
|
||||
```bash
|
||||
sudo apt install cmake build-essential libzstd-dev libssl-dev openssh-client
|
||||
sudo apt install cmake build-essential libzstd-dev zlib1g-dev liblz4-dev libssl-dev openssh-client
|
||||
```
|
||||
|
||||
**Nix:**
|
||||
```bash
|
||||
nix-shell # provides zstd, openssl, cmake, gcc
|
||||
nix-shell # provides zstd, zlib, lz4, openssl, cmake, gcc
|
||||
```
|
||||
|
||||
## Building
|
||||
@@ -505,8 +521,8 @@ features without changing the meaning of ordinary compatibility options.
|
||||
|---|---|
|
||||
| `-j`, `--threads[=N]` | Enable the multithreaded scanner/loader/sender pipeline. `N` (1–256) sets the parallel scanner worker count; bare `-j`/`--threads` uses the default. |
|
||||
| `-z [level]`, `--compress [level]` | Enable streaming compression (default `zstd`), levels 1-22. |
|
||||
| `--compress-level <n>` | Set the compression level. |
|
||||
| `--zc <alg>` | Alias for `--compress-choice`. FastSync supports `zstd` (default), `lz4`, `zlib`, `zlibx`, `none`, and `auto`; `zlibx` behaves as `zlib`. |
|
||||
| `--compress-level <n>` | Set the compression level (1-22). Omitted, each codec uses its rsync default: zstd 3, zlib/zlibx 6, lz4 ignored. |
|
||||
| `--zc <alg>` | Alias for `--compress-choice`. FastSync supports `zstd` (default), `lz4`, `zlib`, `zlibx`, `none`, and `auto`; `zlib`/`zlibx` share the same literal-only zlib path. |
|
||||
| `--zl <n>` | Alias for `--compress-level`. |
|
||||
| `--skip-compress <list>` | Skip compression for `/`- or `,`-separated suffixes; defaults to rsync 3.4.1's built-in list. Incompatible with `--chunk-serialization`. |
|
||||
| `--compress-threads <n>` | Use `n` zstd compression workers. Requires compression and a zstd build with threaded support; the setting affects sender CPU work only. |
|
||||
@@ -519,9 +535,9 @@ features without changing the meaning of ordinary compatibility options.
|
||||
| `--server-host <host>` | Select the TCP server host. |
|
||||
| `--server-port <port>` | Select the TCP server port (`--port <port>` and `--port=<port>` are rsync-friendly aliases). |
|
||||
| `--tls` | Enable TLS for TCP transport. |
|
||||
| `--bwlimit <KB/s>` | Apply token-bucket bandwidth limiting. |
|
||||
| `--progress` | Show rsync-style per-file progress blocks from the receiver's wire counters (FastSync omits rsync's leading `./` line). |
|
||||
| `--stats` | Print transfer statistics, including the receiver-only counters reported over the wire; rsync's per-type `Number of files` breakdown is not reproduced. |
|
||||
| `--bwlimit <KB/s>` | Apply token-bucket bandwidth limiting (also paces `--sendfile` transfers). |
|
||||
| `--progress` | Show rsync-style per-file progress blocks from the receiver's wire counters; the root `./` line is printed whenever progress is active (rsync prints it only when the transfer root is created). |
|
||||
| `--stats` | Print transfer statistics, including the receiver-only counters reported over the wire; `Number of files`/`Number of created files` carry rsync's per-type breakdown (deleted files are a single total). |
|
||||
| `--timeout <seconds>` | Set the socket **and** per-message protocol I/O timeout. Default `0` = disabled (matching rsync); `0` disables it. |
|
||||
| `--contimeout <seconds>` | Connection timeout (default 60, matching rsync); `0` disables it. |
|
||||
|
||||
@@ -555,22 +571,27 @@ remote SSH argv is already built injection-safe.
|
||||
| `--size-only` | Skip incremental files matching in size, ignoring mtime. |
|
||||
| `-I, --ignore-times` | Transfer files even when size and mtime match. |
|
||||
| `-u, --update` | Skip files newer than the source on the receiver. |
|
||||
| `--ignore-existing` | Skip files that already exist on the receiver; like rsync it does not apply to directories or symlinks. |
|
||||
| `-@, --modify-window <sec>` | Modification-time tolerance (seconds) for the incremental/basis quick-check; `0` requires an exact mtime match. |
|
||||
| `-W, --whole-file` | Transfer changed files without delta processing (`--no-whole-file` clears it). |
|
||||
| `-B <n>, --block-size <n>` | Delta block size in bytes (alias `--delta-block`). |
|
||||
| `-d, --dirs` | Transfer the named directory entries without recursing into their contents (aliases `--old-dirs`/`--old-d`). |
|
||||
| `-R, --relative` | Use rsync's relative path semantics (including the `/./` cut); with `--files-from`, preserve each listed entry's relative path below the destination root. |
|
||||
| `--files-from <file>` | Read the source file list from FILE (paths relative to the source root). |
|
||||
| `--delay-updates` | Put updated files into place only at the end of the transfer. |
|
||||
| `--compare-dest <dir>` | Extra comparison basis: unchanged files are not transferred (requires/implies `--incremental`). |
|
||||
| `--copy-dest <dir>` | Like `--compare-dest`, but copies the unchanged file from DIR into the destination. |
|
||||
| `--link-dest <dir>` | Like `--copy-dest`, but hard-links the unchanged file from DIR (repeatable; earlier DIRs win). |
|
||||
| `-0, --from0` | Treat entries in `--files-from` files as NUL-delimited instead of newline-delimited. |
|
||||
| `--delay-updates` | Put updated files into place only at the end of the transfer (the fixed `.fastsync-stage` staging name diverges from rsync; see [`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)). |
|
||||
| `--compare-dest <dir>` | Extra comparison basis: unchanged files are not transferred (requires/implies `--incremental`; a basis MISS above the 256 MiB whole-file payload bound is refused — see [`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)). |
|
||||
| `--copy-dest <dir>` | Like `--compare-dest`, but copies the unchanged file from DIR into the destination (same basis-size caveat; see [`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)). |
|
||||
| `--link-dest <dir>` | Like `--copy-dest`, but hard-links the unchanged file from DIR (repeatable; earlier DIRs win; same basis-size caveat; see [`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)). |
|
||||
| `--verify-basis` | FastSync-only: require a basis hit to match the source by whole-file digest instead of trusting the size+mtime quick-check (default matches rsync). |
|
||||
| `--preallocate` | Allocate destination file space up front (fail-fast on a full disk). |
|
||||
| `--append` | Resume a shorter destination by appending only its tail (prefix not verified; requires `--incremental`). |
|
||||
| `--append-verify` | Like `--append`, but verifies the retained prefix checksum first (falls back to a full transfer on mismatch). |
|
||||
| `--delete` | Request removal of destination entries absent from the source. The server must allow deletion. Default timing is delete-after: extras are removed only after the whole transfer succeeded. Scoped to the synchronized directories, so `--files-from` subsets are safe. |
|
||||
| `--delete` | Request removal of destination entries absent from the source. The server must allow deletion. Default timing is delete-during (matching rsync's `--del`): extras are removed per directory as the transfer proceeds, so destination space is freed progressively. Scoped to the synchronized directories, so `--files-from` subsets are safe. |
|
||||
| `--delete-before` | Delete extras before the transfer starts (implies `--delete`). |
|
||||
| `--delete-during`, `--del` | Delete extras once the keep-set manifest is known, before data is applied (implies `--delete`; early mode, same engine behaviour as `--delete-before`). |
|
||||
| `--delete-delay` | Delete extras only after a successful transfer (implies `--delete`; commit mode, same behaviour as `--delete-after`). |
|
||||
| `--delete-during`, `--del` | Delete each directory's extras as that directory is processed (implies `--delete`). Since protocol 2.24.0 the sender streams a per-directory `STATUS_DELETE_PLAN` frame as it reaches each source directory; this is also the default timing of a plain `--delete`. |
|
||||
| `--delete-delay` | Record extras per directory during the scan but remove them only after a successful transfer (implies `--delete`). Uses the same per-directory `STATUS_DELETE_PLAN` frames as `--delete-during`, applied late. |
|
||||
| `--delete-commit` | FastSync-only: atomic delete-after timing (only after the whole transfer succeeded). |
|
||||
| `--delete-after` | Explicit delete-after timing: delete only after the transfer succeeded (implies `--delete`). |
|
||||
| `--delete-excluded` | Also delete filter-excluded destination mirrors (size-pruned mirrors stay protected). |
|
||||
| `--max-delete <n>` | Delete at most n destination entries; the rest are skipped and the run exits 25 (partial), matching rsync. |
|
||||
@@ -589,8 +610,8 @@ remote SSH argv is already built injection-safe.
|
||||
| `--backup-dir <dir>` | Store backups under a separate directory (requires `--backup`). |
|
||||
| `--suffix <suffix>` | Set the backup filename suffix (default: `~`). |
|
||||
| `--partial` | Select partial-transfer handling. On failed/interrupted writes the already-written temp file is retained (best-effort) for resumption. With `--partial --partial-dir <dir>`, completed files are written under the partial directory and installed atomically. |
|
||||
| `--partial-dir <dir>` | Set a relative partial-transfer directory below the server destination root. Use with `--partial`. |
|
||||
| `--inplace` | Write directly to the destination instead of using a temporary file. |
|
||||
| `--partial-dir <dir>` | Set a relative partial-transfer directory below the server destination root. Implies `--partial`. Rejected together with `--inplace` (`--inplace cannot be used with --partial-dir`, matching rsync), because the inplace path bypasses partial/temp staging. |
|
||||
| `--inplace` | Write directly to the destination instead of using a temporary file. Cannot be combined with `--partial-dir`. |
|
||||
| `--fsync` | Fsync every written file before publication. |
|
||||
| `--write-batch=FILE` | Run the normal live transfer and also emit a self-contained batch file of the source tree. |
|
||||
| `--only-write-batch=FILE` | Emit the batch file only (no destination, no server). |
|
||||
@@ -605,8 +626,11 @@ remote SSH argv is already built injection-safe.
|
||||
| `--preserve` | Preserve mode and mtime (long form only; equivalent to `-p` + `-t`). Add `-o`/`-g` for owner/group, `-U`/`--atimes` for atime, or an identity flag (`--chown`/`--usermap`/`--groupmap`/`--numeric-ids`/`--copy-as`) for mapped ownership. |
|
||||
| `-U`, `--atimes` | Preserve access times. Captured with the metadata payload; does not enable ownership. |
|
||||
| `-N`, `--crtimes` | Capture birth time and transmit it; it cannot be applied because no portable filesystem call can set a birth time (documented divergence). |
|
||||
| `-p`, `--perms` | Preserve permission bits. One of the four per-attribute preserve flags (with `-t`/`-o`/`-g`); under `-p` the source mode is copied exactly (setuid/setgid/sticky and group/other-write included), matching rsync. |
|
||||
| `-p`, `--perms` | Preserve permission bits. One of the four per-attribute preserve flags (with `-t`/`-o`/`-g`); under `-p` the source mode is copied exactly (group/other-write included; setuid/setgid/sticky included only when super-user activities are permitted, masked under `SUPER_MODE_OFF`/`--no-super`), matching rsync otherwise. |
|
||||
| `-t`, `--times` | Preserve modification times. Independent of the other attributes; `-O`/`--omit-dir-times` suppresses directories only. |
|
||||
| `-O`, `--omit-dir-times` | Do not apply modification times to directories. |
|
||||
| `-J`, `--omit-link-times` | Do not apply times to symlinks. |
|
||||
| `--open-noatime` | Open source files with `O_NOATIME` so reading for a transfer does not update their access time (client-only). |
|
||||
| `-o`, `--owner` | Preserve the source owner (uid). Mapped by name on the receiver with a raw-numeric fallback (only numeric ids cross the wire); application is privilege-gated. |
|
||||
| `-g`, `--group` | Preserve the source group (gid). Same name-mapping/numeric-fallback and privilege gating as `-o`. |
|
||||
| `--no-perms`, `--no-times`, `--no-owner`, `--no-group` | Negate each per-attribute flag (also `--no-p`/`--no-t`/`--no-o`/`--no-g`); `--no-preserve` clears all four. |
|
||||
@@ -633,6 +657,7 @@ remote SSH argv is already built injection-safe.
|
||||
| `-D` | Preserve device and special files (implies `--devices --specials`). |
|
||||
| `--devices` | Recreate device nodes on the destination (privileged; skipped without `CAP_MKNOD`). |
|
||||
| `--specials` | Recreate special files: FIFOs and unix sockets. |
|
||||
| `--copy-devices` | Copy a source device's content as an ordinary regular file on the destination (rsync's non-privileged safe mode) instead of recreating the device node. |
|
||||
| `-S`, `--sparse` | Sparse-file handling: receiver preserves holes (zero runs are written as holes; no wire change). |
|
||||
|
||||
### Output and logging
|
||||
@@ -641,12 +666,17 @@ remote SSH argv is already built injection-safe.
|
||||
|---|---|
|
||||
| `-v`, `--verbose` | Enable debug logging. |
|
||||
| `-q`, `--quiet` | Suppress non-error output. |
|
||||
| `--progress` | Show rsync-style per-file progress blocks (not rsync's leading `./` line). |
|
||||
| `--stats` | Print transfer statistics, including the receiver-only counters reported over the wire. |
|
||||
| `-i`, `--itemize-changes` | Print an rsync-style per-file change line. |
|
||||
| `--out-format=FORMAT` | Output format for changed files (`%f %n %l %b %M %%`). |
|
||||
| `--progress` | Show rsync-style per-file progress blocks; the root `./` line is printed whenever progress is active (rsync prints it only when the transfer root is created). |
|
||||
| `--stats` | Print transfer statistics, including the receiver-only counters reported over the wire; `Number of files`/`Number of created files` carry rsync's per-type breakdown (deleted files are a single total). |
|
||||
| `-i, --itemize-changes` | Print an rsync-style per-file change line. |
|
||||
| `--out-format=FORMAT` | Output format for changed files (`%f %n %l %b %c %C %i %M %%`). |
|
||||
| `--list-only` | List source files instead of transferring. |
|
||||
| `--outbuf=MODE` | stdout/stderr buffering: `N` (none/unbuffered), `L` (line-buffered), or `B` (block-buffered, default). |
|
||||
| `--log-file <path>` | Write log output to a file. |
|
||||
| `--log-file-format=FORMAT` | Per-file log-line format (requires `--log-file`). |
|
||||
| `--stderr=MODE` | Route logging to stderr: `errors` or `all`. |
|
||||
| `--msgs2stderr` | Route all messages to stderr (deprecated spelling of `--stderr=all`). |
|
||||
| `--no-msgs2stderr` | Select errors-only stderr (deprecated spelling; the default). |
|
||||
| `-V`, `--version` | Print the FastSync protocol version. |
|
||||
| `--help` | Print command usage. |
|
||||
|
||||
@@ -671,6 +701,11 @@ remote SSH argv is already built injection-safe.
|
||||
| `-4`, `--ipv4` | Force IPv4 for destination resolution. |
|
||||
| `-6`, `--ipv6` | Force IPv6 for destination resolution. |
|
||||
| `--sockopts=OPTS` | Comma-separated OPT=VAL socket options applied before connect. |
|
||||
| `--blocking-io` | SSH transport only: leave the socket without read/write timeouts so it blocks naturally (no effect on TCP). |
|
||||
| `--protocol=NUM` | Force the wire protocol version; must equal the current `PROTOCOL_VERSION` (FastSync cannot speak older/virtual wire formats). |
|
||||
| `--old-args` | Accepted for rsync CLI compatibility; no effect (the remote server path is always safely quoted). |
|
||||
| `--iconv=LOCAL[,REMOTE]` | Convert file-name charsets at the wire boundary (`LOCAL` is our names' charset, `REMOTE` the peer's, defaulting to `LOCAL`). |
|
||||
| `--no-iconv` | Disable `--iconv` charset conversion (same as `--iconv=-`). |
|
||||
| `--tls` | Enable TLS. Requires `--cert`, `--key`, and `--ca`. |
|
||||
| `--cert <path>` | TLS certificate file. |
|
||||
| `--key <path>` | TLS private key file. |
|
||||
@@ -789,7 +824,7 @@ before the module list, before authentication, and the connecting peer address
|
||||
|
||||
## Protocol and Security
|
||||
|
||||
FastSync protocol version `2.26.0` is shared by the client and server. The
|
||||
FastSync protocol version `2.28.0` is shared by the client and server. The
|
||||
current protocol is sender-driven and includes configuration negotiation,
|
||||
including the maximum allocation limit, incremental checks, checksums,
|
||||
manifests, keep-alives, abort handling, per-file remove-source results, and
|
||||
|
||||
+168
-77
File diff suppressed because one or more lines are too long
@@ -8,3 +8,6 @@ markers =
|
||||
daemon_detach: real double-fork backgrounding path (--daemon without
|
||||
--no-detach); slower/fragile, so it runs in the full suite but not the
|
||||
fast PR gate
|
||||
parity: differential rsync-parity case (full set; runs on push to
|
||||
dev/main)
|
||||
parity_ci: fast differential rsync-parity subset (runs on the PR gate)
|
||||
|
||||
+133
-27
@@ -1,5 +1,6 @@
|
||||
#include "change_list.h"
|
||||
#include "checksum.h"
|
||||
#include "log.h"
|
||||
#include "utils.h"
|
||||
#include <fcntl.h>
|
||||
#include <limits.h>
|
||||
@@ -70,7 +71,16 @@ static bool strbuf_append(StrBuf* buf, const char* text) {
|
||||
|
||||
bool change_list_enabled(const Config* config) {
|
||||
return config != NULL && (config->itemize_changes || config->out_format != NULL ||
|
||||
(config->log_file != NULL && config->log_file_format != NULL));
|
||||
(config->log_file != NULL && config->log_file_format != NULL) ||
|
||||
(config->info_level & LOG_INFO_NAME) != 0);
|
||||
}
|
||||
|
||||
/* Emitted once, lazily, ahead of the first --info=name entry: rsync prints the
|
||||
* transfer-root `./` name line when the root directory is (re)created. */
|
||||
static bool name_root_printed = false;
|
||||
|
||||
void change_reset_name_root(void) {
|
||||
name_root_printed = false;
|
||||
}
|
||||
|
||||
/* ---- Itemize code ---- */
|
||||
@@ -131,6 +141,10 @@ static void itemize_code(const Config* config, const ChangeEvent* event, char co
|
||||
update = 'h';
|
||||
else if (created)
|
||||
update = (event->is_directory || event->is_symlink || event->is_special) ? 'c' : '>';
|
||||
else if (event->is_directory)
|
||||
/* rsync: an existing directory that only has attribute changes carries no
|
||||
transfer, so the update column is `.` rather than `>`. */
|
||||
update = '.';
|
||||
else
|
||||
update = '>';
|
||||
code[0] = update;
|
||||
@@ -158,12 +172,15 @@ static void itemize_code(const Config* config, const ChangeEvent* event, char co
|
||||
code[11] = '\0';
|
||||
}
|
||||
|
||||
/* rsync %n: the transfer-relative name, with a trailing slash for directories. */
|
||||
/* rsync %n: the transfer-relative name, with a trailing slash for directories.
|
||||
* The transfer root is `.` (so `%n` renders `./`), matching rsync's root entry. */
|
||||
static bool append_name(StrBuf* buf, const ChangeEvent* event) {
|
||||
if (!strbuf_append(buf, event->name != NULL ? event->name : ""))
|
||||
const char* name = event->name != NULL ? event->name : "";
|
||||
if (event->is_directory && name[0] == '\0')
|
||||
return strbuf_append(buf, "./");
|
||||
if (!strbuf_append(buf, name))
|
||||
return false;
|
||||
if (event->is_directory && (event->name == NULL || event->name[0] == '\0' ||
|
||||
event->name[strlen(event->name) - 1] != '/'))
|
||||
if (event->is_directory && name[strlen(name) - 1] != '/')
|
||||
return strbuf_append_char(buf, '/');
|
||||
return true;
|
||||
}
|
||||
@@ -192,27 +209,55 @@ char* change_render_itemize(const Config* config, const ChangeEvent* event) {
|
||||
return line.data;
|
||||
}
|
||||
|
||||
/* ---- --out-format / --log-file-format ---- */
|
||||
|
||||
/* rsync 3.4.1's `%C` uses the negotiated transfer checksum; with the default
|
||||
* "auto" choice on both ends that is xxh128. FastSync's internal XXH64 default
|
||||
* is not an rsync algorithm, so map it to xxh128 for parity. */
|
||||
static ChecksumAlgo out_format_checksum_algo(const Config* config) {
|
||||
switch ((ChecksumAlgo)config->checksum_algo) {
|
||||
case CHECKSUM_ALGO_MD5:
|
||||
return CHECKSUM_ALGO_MD5;
|
||||
case CHECKSUM_ALGO_XXH3:
|
||||
return CHECKSUM_ALGO_XXH3;
|
||||
case CHECKSUM_ALGO_XXH128:
|
||||
return CHECKSUM_ALGO_XXH128;
|
||||
case CHECKSUM_ALGO_XXH64:
|
||||
default:
|
||||
return CHECKSUM_ALGO_XXH128;
|
||||
/* rsync's `--info=name` line for an updated entry: the transfer-relative name
|
||||
* (trailing slash for directories) plus the ` -> target` / ` => target` link
|
||||
* suffix. `--info=name` does not alter an itemize/out-format run. */
|
||||
static char* change_render_name(const ChangeEvent* event) {
|
||||
StrBuf line = {0};
|
||||
bool ok = append_name(&line, event) && append_link_suffix(&line, event);
|
||||
if (!ok) {
|
||||
strbuf_free(&line);
|
||||
return NULL;
|
||||
}
|
||||
if (line.data == NULL) {
|
||||
line.data = str_dup("");
|
||||
if (!line.data)
|
||||
return NULL;
|
||||
}
|
||||
return line.data;
|
||||
}
|
||||
|
||||
/* Render a digest as rsync's sum_as_hex: for xxh128 the HIGH 64-bit half is
|
||||
* printed before the low half; every other algorithm prints its bytes in order. */
|
||||
/* rsync's `--info=name2` line for an unchanged entry: `NAME is uptodate`. */
|
||||
static char* change_render_name_uptodate(const ChangeEvent* event) {
|
||||
char* name = change_render_name(event);
|
||||
if (name == NULL)
|
||||
return NULL;
|
||||
size_t length = strlen(name);
|
||||
char* line = malloc(length + sizeof(" is uptodate"));
|
||||
if (line == NULL) {
|
||||
free(name);
|
||||
return NULL;
|
||||
}
|
||||
memcpy(line, name, length);
|
||||
memcpy(line + length, " is uptodate", sizeof(" is uptodate"));
|
||||
free(name);
|
||||
return line;
|
||||
}
|
||||
|
||||
/* ---- --out-format / --log-file-format ---- */
|
||||
|
||||
/* rsync 3.4.1's `%C` uses the negotiated TRANSFER checksum (the first name of a
|
||||
* two-name "transfer,pre-transfer" --checksum-choice), not the pre-transfer
|
||||
* whole-file digest FastSync compares against on the wire. The default "auto"
|
||||
* resolves to xxh128, so an explicit selection and the default both render the
|
||||
* selected algorithm's digest. */
|
||||
static ChecksumAlgo out_format_checksum_algo(const Config* config) {
|
||||
return (ChecksumAlgo)config->cli.checksum_transfer_algo;
|
||||
}
|
||||
|
||||
/* Render a digest as rsync's sum_as_hex: xxh128 prints the HIGH 64-bit half
|
||||
* before the low half, and xxh64/xxh3 print their 64-bit value big-endian; every
|
||||
* other algorithm prints its bytes in order. */
|
||||
static void digest_to_hex(ChecksumAlgo algo, const uint8_t* digest, size_t len, char* out) {
|
||||
if (algo == CHECKSUM_ALGO_XXH128 && len == 16) {
|
||||
uint64_t low = 0;
|
||||
@@ -222,6 +267,12 @@ static void digest_to_hex(ChecksumAlgo algo, const uint8_t* digest, size_t len,
|
||||
snprintf(out, len * 2 + 1, "%016llx%016llx", (unsigned long long)high, (unsigned long long)low);
|
||||
return;
|
||||
}
|
||||
if ((algo == CHECKSUM_ALGO_XXH64 || algo == CHECKSUM_ALGO_XXH3) && len == 8) {
|
||||
uint64_t value = 0;
|
||||
memcpy(&value, digest, sizeof(value));
|
||||
snprintf(out, len * 2 + 1, "%016llx", (unsigned long long)value);
|
||||
return;
|
||||
}
|
||||
static const char hex[] = "0123456789abcdef";
|
||||
for (size_t i = 0; i < len; i++) {
|
||||
out[i * 2] = hex[(digest[i] >> 4) & 0xf];
|
||||
@@ -260,6 +311,9 @@ static void fill_event_checksum(const Config* config, const File* file, ChangeEv
|
||||
if (file->path == NULL)
|
||||
return;
|
||||
ChecksumAlgo algo = out_format_checksum_algo(config);
|
||||
/* rsync renders `--checksum-choice=none` as a blank 2-character column. */
|
||||
if (algo == CHECKSUM_ALGO_NONE)
|
||||
return;
|
||||
uint8_t digest[CHECKSUM_MAX_DIGEST_LEN];
|
||||
size_t len = 0;
|
||||
/* rsync's %C is the transfer checksum, which is always seeded with 0 (it is
|
||||
@@ -329,9 +383,10 @@ char* change_render_format(const char* format, const Config* config, const Chang
|
||||
if (event->checksum_known) {
|
||||
ok = strbuf_append(&line, event->checksum);
|
||||
} else {
|
||||
/* rsync pads a non-regular / untransferred entry with spaces. */
|
||||
/* rsync pads a non-regular / untransferred / `none` entry with spaces;
|
||||
`none` renders as a blank 2-character column. */
|
||||
ChecksumAlgo algo = out_format_checksum_algo(config);
|
||||
int width = checksum_digest_len(algo) * 2;
|
||||
int width = algo == CHECKSUM_ALGO_NONE ? 2 : checksum_digest_len(algo) * 2;
|
||||
for (int i = 0; i < width && ok; i++)
|
||||
ok = strbuf_append_char(&line, ' ');
|
||||
}
|
||||
@@ -438,10 +493,24 @@ static void print_escaped_line(FILE* stream, const char* line, bool eight_bit_ou
|
||||
void change_emit(const Config* config, const ChangeEvent* event) {
|
||||
if (event == NULL || !change_list_enabled(config))
|
||||
return;
|
||||
if (event->decision == CHANGE_UP_TO_DATE)
|
||||
return;
|
||||
bool to_stdout = config->itemize_changes || config->out_format != NULL;
|
||||
bool to_log = config->log_file != NULL && config->log_file_format != NULL;
|
||||
bool progress_active = config->show_progress || (config->info_level & LOG_INFO_PROGRESS);
|
||||
if (event->decision == CHANGE_UP_TO_DATE) {
|
||||
/* --info=name2 prints `NAME is uptodate` for entries the receiver already
|
||||
had. An itemize/out-format run reports them through its own format (or
|
||||
not at all), the progress stream has no frame for them, and neither the
|
||||
itemize nor the log-file stream previously reported an up-to-date entry,
|
||||
so nothing else here changes. */
|
||||
if (!to_stdout && (config->info_level & LOG_INFO_NAME_UPTODATE) != 0 && !progress_active) {
|
||||
char* line = change_render_name_uptodate(event);
|
||||
if (line != NULL) {
|
||||
print_escaped_line(stdout, line, config->eight_bit_output);
|
||||
free(line);
|
||||
}
|
||||
}
|
||||
return;
|
||||
}
|
||||
if (to_stdout) {
|
||||
char* line = config->out_format != NULL
|
||||
? change_render_format(config->out_format, config, event)
|
||||
@@ -450,6 +519,20 @@ void change_emit(const Config* config, const ChangeEvent* event) {
|
||||
print_escaped_line(stdout, line, config->eight_bit_output);
|
||||
free(line);
|
||||
}
|
||||
} else if ((config->info_level & LOG_INFO_NAME) != 0 && !progress_active) {
|
||||
/* --info=name without -i/--out-format: print the updated entry's name. The
|
||||
--progress path owns the name line when progress output is active (it
|
||||
emits the same names before the progress frames), so do not duplicate.
|
||||
The transfer-root `./` line precedes the first such name. */
|
||||
if (!name_root_printed) {
|
||||
name_root_printed = true;
|
||||
fputs("./\n", stdout);
|
||||
}
|
||||
char* line = change_render_name(event);
|
||||
if (line != NULL) {
|
||||
print_escaped_line(stdout, line, config->eight_bit_output);
|
||||
free(line);
|
||||
}
|
||||
}
|
||||
if (to_log) {
|
||||
char* line = change_render_format(config->log_file_format, config, event);
|
||||
@@ -610,6 +693,29 @@ void change_emit_file_sent(const Config* config, const File* file) {
|
||||
change_emit_file_sent_bytes(config, file, payload, 0);
|
||||
}
|
||||
|
||||
void change_emit_file_uptodate(const Config* config, const File* file) {
|
||||
if (file == NULL || !change_list_enabled(config))
|
||||
return;
|
||||
ChangeEvent event;
|
||||
memset(&event, 0, sizeof(event));
|
||||
event.decision = CHANGE_UP_TO_DATE;
|
||||
event.is_directory = false;
|
||||
event.is_symlink = file->is_symlink;
|
||||
event.is_special = file->is_special;
|
||||
event.is_hardlink = file->link_group != 0 && !file->link_first;
|
||||
event.symlink_target = file->symlink_target;
|
||||
event.hardlink_target = file->hardlink_target;
|
||||
event.size = file->data != NULL ? file->data->size : 0;
|
||||
event.dest = file->dest_state;
|
||||
char* name = NULL;
|
||||
char* path = NULL;
|
||||
fill_event_from_file(config, file, &event, &name, &path);
|
||||
if (name != NULL && path != NULL)
|
||||
change_emit(config, &event);
|
||||
free(name);
|
||||
free(path);
|
||||
}
|
||||
|
||||
void change_emit_dir_sent(const Config* config, const File* file) {
|
||||
if (file == NULL || !change_list_enabled(config))
|
||||
return;
|
||||
|
||||
@@ -102,4 +102,13 @@ void change_emit_file_sent(const Config* config, const File* file);
|
||||
/* Build and emit a CHANGE_SENT event for an explicit directory entry (-d). */
|
||||
void change_emit_dir_sent(const Config* config, const File* file);
|
||||
|
||||
/* Build and emit a CHANGE_UP_TO_DATE event for a file the receiver already had.
|
||||
* With --info=name2 it renders rsync's "NAME is uptodate" line (no output
|
||||
* otherwise). */
|
||||
void change_emit_file_uptodate(const Config* config, const File* file);
|
||||
|
||||
/* Reset the lazy transfer-root `./` line emitted ahead of the first
|
||||
* --info=name entry. Call once at the start of a transfer. */
|
||||
void change_reset_name_root(void);
|
||||
|
||||
#endif
|
||||
|
||||
+331
-59
@@ -23,6 +23,7 @@
|
||||
#include <langinfo.h>
|
||||
#include <limits.h>
|
||||
#include <locale.h>
|
||||
#include <math.h>
|
||||
#include <time.h>
|
||||
#include <signal.h>
|
||||
#include <stdbool.h>
|
||||
@@ -53,13 +54,26 @@ bool client_abort_pending(void) {
|
||||
}
|
||||
|
||||
#ifndef FASTSYNC_TEST_BUILD
|
||||
/* SIG_DFL disposition used by the handler's "not armed" fallback. It is built
|
||||
* once at load time so the handler can restore the default action with
|
||||
* sigaction(2) -- which is async-signal-safe -- instead of signal(3), which is
|
||||
* not. The zero-initialized sa_mask is the empty set. */
|
||||
static const struct sigaction client_default_action = {
|
||||
.sa_handler = SIG_DFL,
|
||||
.sa_flags = 0,
|
||||
};
|
||||
|
||||
/* Signal handler: perform NO work beyond storing the flag. Logging, protocol
|
||||
* I/O and the STATUS_ABORT frame are all done later on the normal send path,
|
||||
* which is not async-signal-safe. When no transfer is armed, fall back to the
|
||||
* default action so local-only modes remain interruptible. */
|
||||
* which is not async-signal-safe. When no transfer is armed, restore the
|
||||
* default disposition (async-signal-safe sigaction) and re-raise so local-only
|
||||
* modes remain interruptible. The handler deliberately stays installed while a
|
||||
* transfer is armed -- rather than using SA_RESETHAND -- so a second Ctrl-C
|
||||
* during the graceful abort keeps setting the flag instead of hard-killing the
|
||||
* process mid-cleanup. */
|
||||
static void client_signal_handler(int signo) {
|
||||
if (!client_abort_armed) {
|
||||
signal(signo, SIG_DFL);
|
||||
sigaction(signo, &client_default_action, NULL);
|
||||
raise(signo);
|
||||
return;
|
||||
}
|
||||
@@ -156,20 +170,26 @@ static int set_positive_int_option(int* dest, const char* value, const char* opt
|
||||
* name is a hard error with rsync's exit code 4, never a silent no-op. */
|
||||
static int set_compression_choice(Config* config, const char* value) {
|
||||
if (!value) {
|
||||
config->cli_exit_code = 4;
|
||||
config->cli.cli_exit_code = 4;
|
||||
return -1;
|
||||
}
|
||||
int algo;
|
||||
if (strcasecmp(value, "auto") == 0)
|
||||
algo = (int)compression_negotiate_default();
|
||||
else
|
||||
if (strcasecmp(value, "auto") == 0) {
|
||||
algo = compression_choice_resolve();
|
||||
if (algo < 0) {
|
||||
log_message(LOG_LEVEL_ERROR, "RSYNC_COMPRESS_LIST names no supported compression algorithm");
|
||||
config->cli.cli_exit_code = 4;
|
||||
return -1;
|
||||
}
|
||||
} else {
|
||||
algo = compression_algo_from_name(value);
|
||||
}
|
||||
if (algo < 0) {
|
||||
log_message(LOG_LEVEL_ERROR,
|
||||
"--compress-choice '%s' is not a supported algorithm; FastSync supports zstd, "
|
||||
"lz4, zlib, zlibx, none or auto",
|
||||
value);
|
||||
config->cli_exit_code = 4;
|
||||
config->cli.cli_exit_code = 4;
|
||||
return -1;
|
||||
}
|
||||
const char* canonical = compression_algo_name((CompressionAlgo)algo);
|
||||
@@ -206,7 +226,7 @@ static int resolve_checksum_name(const char* name, size_t len, int* out) {
|
||||
* resolves to FastSync's negotiated default (xxh128). */
|
||||
static int set_checksum_choice(Config* config, const char* value) {
|
||||
if (!value) {
|
||||
config->cli_exit_code = 4;
|
||||
config->cli.cli_exit_code = 4;
|
||||
return -1;
|
||||
}
|
||||
const char* comma = strchr(value, ',');
|
||||
@@ -224,19 +244,28 @@ static int set_checksum_choice(Config* config, const char* value) {
|
||||
"--checksum-choice '%s' is invalid; FastSync supports xxh64 (or xxhash), xxh128, "
|
||||
"xxh3, md5, md4, sha1, none or auto, optionally as 'transfer,pre-transfer'",
|
||||
value);
|
||||
config->cli_exit_code = 4;
|
||||
config->cli.cli_exit_code = 4;
|
||||
return -1;
|
||||
}
|
||||
ChecksumAlgo negotiated = checksum_negotiate_default();
|
||||
int negotiated = -1;
|
||||
if (rc1 == 1 || rc2 == 1) {
|
||||
negotiated = checksum_choice_resolve();
|
||||
if (negotiated < 0) {
|
||||
log_message(LOG_LEVEL_ERROR, "RSYNC_CHECKSUM_LIST names no supported checksum algorithm");
|
||||
config->cli.cli_exit_code = 4;
|
||||
return -1;
|
||||
}
|
||||
}
|
||||
if (rc1 == 1)
|
||||
transfer = (int)negotiated;
|
||||
transfer = negotiated;
|
||||
if (!name2)
|
||||
pre = transfer;
|
||||
else if (rc2 == 1)
|
||||
pre = (int)negotiated;
|
||||
pre = negotiated;
|
||||
|
||||
config->checksum_algo = pre;
|
||||
config->checksum_transfer_algo = transfer;
|
||||
config->cli.checksum_transfer_algo = transfer;
|
||||
config->cli.checksum_choice_set = true;
|
||||
/* rsync: "none" for the transfer checksum forces --whole-file. */
|
||||
if (transfer == (int)CHECKSUM_ALGO_NONE)
|
||||
config->whole_file = true;
|
||||
@@ -488,9 +517,8 @@ static bool split_flag_level(const char* token, char* name, size_t name_size, in
|
||||
* of rsync's `symsafe`, `hlink`, and `own`. */
|
||||
static bool is_accepted_debug_category(const char* name) {
|
||||
static const char* const categories[] = {
|
||||
"acl", "backup", "bind", "chdir", "cmd", "connect", "del", "deltasum",
|
||||
"dup", "exit", "filter", "flist", "fuzzy", "genr", "hash", "hl",
|
||||
"hlink", "iconv", "nstr", "own", "owner", "recv", "send", "time",
|
||||
"acl", "backup", "bind", "chdir", "cmd", "connect", "dup", "exit", "fuzzy",
|
||||
"genr", "hl", "hlink", "iconv", "nstr", "own", "owner", "time",
|
||||
};
|
||||
for (size_t i = 0; i < sizeof(categories) / sizeof(categories[0]); i++) {
|
||||
if (strcmp(name, categories[i]) == 0)
|
||||
@@ -501,7 +529,9 @@ static bool is_accepted_debug_category(const char* name) {
|
||||
|
||||
static bool is_accepted_info_category(const char* name) {
|
||||
static const char* const categories[] = {
|
||||
"backup", "del", "flist", "mount", "nonreg", "progress", "remove", "syms", "symsafe",
|
||||
"backup",
|
||||
"syms",
|
||||
"symsafe",
|
||||
};
|
||||
for (size_t i = 0; i < sizeof(categories) / sizeof(categories[0]); i++) {
|
||||
if (strcmp(name, categories[i]) == 0)
|
||||
@@ -552,6 +582,18 @@ static int parse_debug_flags(const char* value, Config* config) {
|
||||
flag = LOG_DEBUG_PACK;
|
||||
} else if (strcmp(name, "util") == 0) {
|
||||
flag = LOG_DEBUG_UTIL;
|
||||
} else if (strcmp(name, "flist") == 0) {
|
||||
flag = LOG_DEBUG_FLIST;
|
||||
} else if (strcmp(name, "del") == 0) {
|
||||
flag = LOG_DEBUG_DEL;
|
||||
} else if (strcmp(name, "hash") == 0 || strcmp(name, "deltasum") == 0) {
|
||||
flag = LOG_DEBUG_HASH;
|
||||
} else if (strcmp(name, "recv") == 0) {
|
||||
flag = LOG_DEBUG_RECV;
|
||||
} else if (strcmp(name, "filter") == 0) {
|
||||
flag = LOG_DEBUG_FILTER;
|
||||
} else if (strcmp(name, "send") == 0) {
|
||||
flag = LOG_DEBUG_SEND;
|
||||
} else if (is_accepted_debug_category(name)) {
|
||||
continue;
|
||||
} else {
|
||||
@@ -608,14 +650,41 @@ static int parse_info_flags(const char* value, Config* config) {
|
||||
free(flags);
|
||||
return 1;
|
||||
}
|
||||
if (strcmp(name, "copy") == 0 || strcmp(name, "name") == 0)
|
||||
if (strcmp(name, "copy") == 0)
|
||||
flag = LOG_INFO_COPY;
|
||||
else if (strcmp(name, "misc") == 0)
|
||||
else if (strcmp(name, "name") == 0) {
|
||||
/* name level 2 adds rsync's "is uptodate" lines. */
|
||||
if (level == 0)
|
||||
parsed &= ~(uint32_t)(LOG_INFO_NAME | LOG_INFO_NAME_UPTODATE);
|
||||
else {
|
||||
parsed |= LOG_INFO_NAME;
|
||||
if (level >= 2)
|
||||
parsed |= LOG_INFO_NAME_UPTODATE;
|
||||
else
|
||||
parsed &= ~(uint32_t)LOG_INFO_NAME_UPTODATE;
|
||||
}
|
||||
continue;
|
||||
} else if (strcmp(name, "misc") == 0)
|
||||
flag = LOG_INFO_MISC;
|
||||
else if (strcmp(name, "skip") == 0)
|
||||
flag = LOG_INFO_SKIP;
|
||||
else if (strcmp(name, "stats") == 0)
|
||||
else if (strcmp(name, "stats") == 0) {
|
||||
flag = LOG_INFO_STATS;
|
||||
/* `--info=stats` requests the same transfer-statistics block as
|
||||
`--stats`; `--info=stats0` turns it back off. */
|
||||
config->stats = level > 0;
|
||||
} else if (strcmp(name, "del") == 0)
|
||||
flag = LOG_INFO_DEL;
|
||||
else if (strcmp(name, "remove") == 0)
|
||||
flag = LOG_INFO_REMOVE;
|
||||
else if (strcmp(name, "flist") == 0)
|
||||
flag = LOG_INFO_FLIST;
|
||||
else if (strcmp(name, "nonreg") == 0)
|
||||
flag = LOG_INFO_NONREG;
|
||||
else if (strcmp(name, "mount") == 0)
|
||||
flag = LOG_INFO_MOUNT;
|
||||
else if (strcmp(name, "progress") == 0)
|
||||
flag = LOG_INFO_PROGRESS;
|
||||
else if (is_accepted_info_category(name))
|
||||
continue;
|
||||
else {
|
||||
@@ -894,7 +963,10 @@ static const OptionEntry OPTION_TABLE[] = {
|
||||
* faithful no-op (accepted silently, never consumes an argument). */
|
||||
{"--recursive", "-r", OPT_NOOP, 0},
|
||||
{"--update", "-u", OPT_FLAG, offsetof(Config, update)},
|
||||
{"--old-args", NULL, OPT_FLAG, offsetof(Config, old_args)},
|
||||
/* rsync's --old-args: accepted for CLI compatibility as a documented no-op
|
||||
* (the remote server path is always safely quoted; see usage.c). It is
|
||||
* recognized but stores no Config field. */
|
||||
{"--old-args", NULL, OPT_NOOP, 0},
|
||||
{"--rsh", "-e", OPT_STRING, offsetof(Config, rsh_command)},
|
||||
{"--blocking-io", NULL, OPT_FLAG, offsetof(Config, blocking_io)},
|
||||
{"--links", "-l", OPT_FLAG, offsetof(Config, follow_symlinks)},
|
||||
@@ -956,6 +1028,12 @@ static const OptionEntry OPTION_TABLE[] = {
|
||||
{"--delete-during", "--del", OPT_FLAG, offsetof(Config, delete_during)},
|
||||
{"--delete-delay", NULL, OPT_FLAG, offsetof(Config, delete_delay)},
|
||||
{"--delete-after", NULL, OPT_FLAG, offsetof(Config, delete_after)},
|
||||
/* FastSync-only long spelling of the late whole-tree commit, which selects
|
||||
the same timing as rsync's --delete-after in FastSync (the whole-tree
|
||||
keep-set manifest is committed only after the entire transfer succeeded).
|
||||
Plain --delete now defaults to delete-during, so this restores the old
|
||||
FastSync behavior; it maps onto the same delete_after wire field. */
|
||||
{"--delete-commit", NULL, OPT_FLAG, offsetof(Config, delete_after)},
|
||||
{"--delete-excluded", NULL, OPT_FLAG, offsetof(Config, delete_excluded)},
|
||||
{"--max-delete", NULL, OPT_SIGNED_INT, offsetof(Config, max_delete)},
|
||||
{"--ignore-errors", NULL, OPT_FLAG, offsetof(Config, ignore_errors)},
|
||||
@@ -1013,6 +1091,11 @@ static const OptionEntry OPTION_TABLE[] = {
|
||||
* --remote-option is parsed. --trust-sender is a local receiver policy and
|
||||
* never travels to the remote peer. */
|
||||
{"--trust-sender", NULL, OPT_FLAG, offsetof(Config, trust_sender)},
|
||||
/* FastSync-only (not an rsync option): require a basis-hit's content to
|
||||
* match the source by whole-file digest instead of trusting rsync's
|
||||
* size+mtime quick-check. Long-only; crosses the wire so the receiver
|
||||
* performs the extra read/hash. */
|
||||
{"--verify-basis", NULL, OPT_FLAG, offsetof(Config, verify_basis)},
|
||||
};
|
||||
|
||||
/* Only boolean options with no required argument are safe to negate. */
|
||||
@@ -1055,6 +1138,7 @@ static const NegatableOption NEGATABLE_OPTIONS[] = {
|
||||
{"xattrs", "X", offsetof(Config, preserve_xattrs)},
|
||||
{"acls", "A", offsetof(Config, preserve_acls)},
|
||||
{"fake-super", NULL, offsetof(Config, fake_super)},
|
||||
{"verify-basis", NULL, offsetof(Config, verify_basis)},
|
||||
};
|
||||
|
||||
static bool opt_is(const char* arg, const char* name, const char* alias) {
|
||||
@@ -1115,12 +1199,12 @@ static int apply_negation(Config* config, const char* arg) {
|
||||
config->preserve_times = false;
|
||||
config->preserve_owner = false;
|
||||
config->preserve_group = false;
|
||||
config->metadata_explicitly_disabled = true;
|
||||
config->cli.metadata_explicitly_disabled = true;
|
||||
/* --no-preserve is an explicit opt-out of the whole bundle: record it so
|
||||
* the --incremental/--delta auto-preserve in cli_finalize_config does not
|
||||
* silently re-enable perms/times. */
|
||||
config->preserve_perms_explicit_off = true;
|
||||
config->preserve_times_explicit_off = true;
|
||||
config->cli.preserve_perms_explicit_off = true;
|
||||
config->cli.preserve_times_explicit_off = true;
|
||||
return 0;
|
||||
}
|
||||
*(bool*)((char*)config + entry->offset) = false;
|
||||
@@ -1128,9 +1212,9 @@ static int apply_negation(Config* config, const char* arg) {
|
||||
* auto-preserve the OTHER attribute without undoing this one. A later
|
||||
* -p/-t sets the attribute directly; this flag only gates the implication. */
|
||||
if (entry->offset == offsetof(Config, preserve_perms))
|
||||
config->preserve_perms_explicit_off = true;
|
||||
config->cli.preserve_perms_explicit_off = true;
|
||||
else if (entry->offset == offsetof(Config, preserve_times))
|
||||
config->preserve_times_explicit_off = true;
|
||||
config->cli.preserve_times_explicit_off = true;
|
||||
return 0;
|
||||
}
|
||||
|
||||
@@ -1160,7 +1244,13 @@ static int apply_table_option(Config* config, const OptionEntry* entry, const ch
|
||||
void* field = (char*)config + entry->offset;
|
||||
switch (entry->kind) {
|
||||
case OPT_FLAG:
|
||||
*(bool*)field = true;
|
||||
/* -x/--one-file-system is repeatable in rsync: `-xx` increments the level so
|
||||
the scanner drops mount-point directories instead of recreating them
|
||||
empty. Everything else is a plain boolean. */
|
||||
if (entry->offset == offsetof(Config, one_file_system))
|
||||
(*(int*)field)++;
|
||||
else
|
||||
*(bool*)field = true;
|
||||
return 0;
|
||||
case OPT_NOOP:
|
||||
return 0;
|
||||
@@ -1353,7 +1443,7 @@ static bool cli_handle_range_time_options(CliParseCtx* ctx) {
|
||||
ctx->exit_code = -1;
|
||||
return true;
|
||||
}
|
||||
config->stop_at_set = true;
|
||||
config->cli.stop_at_set = true;
|
||||
return true;
|
||||
}
|
||||
if (strcmp(arg, "--stop-at") == 0) {
|
||||
@@ -1368,7 +1458,7 @@ static bool cli_handle_range_time_options(CliParseCtx* ctx) {
|
||||
ctx->exit_code = -1;
|
||||
return true;
|
||||
}
|
||||
config->stop_at_set = true;
|
||||
config->cli.stop_at_set = true;
|
||||
return true;
|
||||
}
|
||||
const char* threads_prefix = "--compress-threads=";
|
||||
@@ -1438,6 +1528,8 @@ static bool cli_handle_table_option(CliParseCtx* ctx) {
|
||||
ctx->exit_code = -1;
|
||||
return true;
|
||||
}
|
||||
if (entry->offset == offsetof(Config, compression_level))
|
||||
config->cli.compression_level_set = true;
|
||||
if (entry->offset == offsetof(Config, chmod_spec)) {
|
||||
mode_t ignored;
|
||||
if (!chmod_apply(0, config->chmod_spec, &ignored)) {
|
||||
@@ -1450,7 +1542,7 @@ static bool cli_handle_table_option(CliParseCtx* ctx) {
|
||||
defaults to 127.0.0.1, so a value check cannot distinguish it). Used
|
||||
by --dry-run to route an explicit remote target to the server. */
|
||||
if (entry->offset == offsetof(Config, server_host))
|
||||
config->server_host_set = true;
|
||||
config->cli.server_host_set = true;
|
||||
}
|
||||
} else if (apply_table_option(config, entry, NULL) != 0) {
|
||||
ctx->exit_code = -1;
|
||||
@@ -1464,7 +1556,9 @@ static bool cli_handle_table_option(CliParseCtx* ctx) {
|
||||
if (entry->offset == offsetof(Config, per_dir_filter) && config->per_dir_filter_count < INT_MAX)
|
||||
config->per_dir_filter_count++;
|
||||
/* A delete-timing flag selects when --delete removes extras, so it
|
||||
implies --delete exactly like the rsync options do. */
|
||||
implies --delete exactly like the rsync options do. --delete-commit (the
|
||||
FastSync-only late-commit spelling) is mapped onto delete_after and so is
|
||||
covered here too. */
|
||||
if (entry->offset == offsetof(Config, delete_before) ||
|
||||
entry->offset == offsetof(Config, delete_during) ||
|
||||
entry->offset == offsetof(Config, delete_delay) ||
|
||||
@@ -1722,6 +1816,7 @@ static bool cli_handle_transfer_flags(CliParseCtx* ctx) {
|
||||
return true;
|
||||
}
|
||||
config->compression_level = (int)level;
|
||||
config->cli.compression_level_set = true;
|
||||
log_info_message(LOG_INFO_MISC, "Set Compression level to %ld", level);
|
||||
ctx->i++;
|
||||
}
|
||||
@@ -1785,7 +1880,7 @@ static int set_server_port_option(Config* config, const char* value, const char*
|
||||
return -1;
|
||||
}
|
||||
config->server_port = port;
|
||||
config->server_port_set = true;
|
||||
config->cli.server_port_set = true;
|
||||
return 0;
|
||||
}
|
||||
|
||||
@@ -1819,22 +1914,130 @@ static int set_log_file_option(Config* config, const char* log_path) {
|
||||
return 0;
|
||||
}
|
||||
|
||||
/* Apply a --bwlimit value (kilobytes per second). Returns 0 on success, -1 on
|
||||
* error. */
|
||||
/* Faithful port of rsync 3.4.1's `parse_size_arg(bwlimit_arg, 'K', "bwlimit",
|
||||
* 512, -1, True)`: a default KiB suffix, binary (1024) multipliers unless a
|
||||
* `b`/`B` decimal suffix or explicit `iB` is given, an optional decimal
|
||||
* fraction, the P/T/G/M/K suffixes, and the special rules that a value of 0
|
||||
* means "no limit" while any other value below 512 bytes is rejected. The
|
||||
* parsed byte count is then quantized to whole KiB exactly like rsync's
|
||||
* `bwlimit = (size + 512) / 1024`. Returns 0 on success, -1 on a parse error. */
|
||||
static int parse_bwlimit_value(const char* value, unsigned long long* bytes_per_sec_out) {
|
||||
if (!value || !bytes_per_sec_out)
|
||||
return -1;
|
||||
const char* arg = value;
|
||||
int reps;
|
||||
long long mult;
|
||||
while (*arg >= '0' && *arg <= '9')
|
||||
arg++;
|
||||
if (*arg != '\0' && (*arg == '.' || *arg == localeconv()->decimal_point[0]))
|
||||
for (arg++; *arg >= '0' && *arg <= '9'; arg++) {
|
||||
}
|
||||
|
||||
char suffix = *arg && *arg != '+' && *arg != '-' ? *arg++ : 'K';
|
||||
switch (suffix) {
|
||||
case 'b':
|
||||
case 'B':
|
||||
reps = 0;
|
||||
break;
|
||||
case 'k':
|
||||
case 'K':
|
||||
reps = 1;
|
||||
break;
|
||||
case 'm':
|
||||
case 'M':
|
||||
reps = 2;
|
||||
break;
|
||||
case 'g':
|
||||
case 'G':
|
||||
reps = 3;
|
||||
break;
|
||||
case 't':
|
||||
case 'T':
|
||||
reps = 4;
|
||||
break;
|
||||
case 'p':
|
||||
case 'P':
|
||||
reps = 5;
|
||||
break;
|
||||
default:
|
||||
log_message(LOG_LEVEL_ERROR, "--bwlimit=%s is invalid", value);
|
||||
return -1;
|
||||
}
|
||||
if (*arg == 'b' || *arg == 'B') {
|
||||
mult = 1000;
|
||||
arg++;
|
||||
} else if (*arg == '\0' || *arg == '+' || *arg == '-') {
|
||||
mult = 1024;
|
||||
} else if ((arg[0] == 'i' || arg[0] == 'I') && (arg[1] == 'b' || arg[1] == 'B')) {
|
||||
mult = 1024;
|
||||
arg += 2;
|
||||
} else {
|
||||
log_message(LOG_LEVEL_ERROR, "--bwlimit=%s is invalid", value);
|
||||
return -1;
|
||||
}
|
||||
|
||||
long long base = 1;
|
||||
for (int i = 0; i < reps; i++) {
|
||||
if (base > LLONG_MAX / mult) {
|
||||
log_message(LOG_LEVEL_ERROR, "--bwlimit=%s is too large", value);
|
||||
return -1;
|
||||
}
|
||||
base *= mult;
|
||||
}
|
||||
/* rsync multiplies the numeric prefix (atof) by mult^reps in a signed
|
||||
* ssize_t, which is undefined on overflow. Scale in double and range-check
|
||||
* before converting, so a huge value is rejected as "too large" (where
|
||||
* rsync's overflow happens to land on a negative result) without invoking
|
||||
* signed-overflow UB. */
|
||||
double scaled = (double)base * strtod(value, NULL);
|
||||
/* (double)LLONG_MAX rounds up to 2^63, which is itself out of range for the
|
||||
* cast, so reject at >= that bound; LLONG_MIN == -2^63 is exactly
|
||||
* representable and thus castable, so the lower bound stays strict. */
|
||||
if (!isfinite(scaled) || scaled >= (double)LLONG_MAX || scaled < (double)LLONG_MIN) {
|
||||
log_message(LOG_LEVEL_ERROR, "--bwlimit=%s is too large", value);
|
||||
return -1;
|
||||
}
|
||||
long long size = (long long)scaled;
|
||||
if ((*arg == '+' || *arg == '-') && arg[1] == '1' && arg != value) {
|
||||
/* The only form accepted here is "+1"/"-1" (a longer number leaves a
|
||||
trailing byte and is rejected below), so apply the delta directly and
|
||||
guard the one overflow direction. */
|
||||
if (*arg == '+') {
|
||||
if (size == LLONG_MAX) {
|
||||
log_message(LOG_LEVEL_ERROR, "--bwlimit=%s is too large", value);
|
||||
return -1;
|
||||
}
|
||||
size += 1;
|
||||
} else {
|
||||
size -= 1;
|
||||
}
|
||||
arg += 2;
|
||||
}
|
||||
if (*arg != '\0' || size < 0) {
|
||||
log_message(LOG_LEVEL_ERROR, "--bwlimit=%s is %s", value, size < 0 ? "too large" : "invalid");
|
||||
return -1;
|
||||
}
|
||||
if (size != 0 && size < 512) {
|
||||
log_message(LOG_LEVEL_ERROR, "--bwlimit=%s is too small (min: 512 or 0 for unlimited)", value);
|
||||
return -1;
|
||||
}
|
||||
long long kib = size == 0 ? 0 : (size + 512) / 1024;
|
||||
if (kib > (long long)(ULLONG_MAX / 1024)) {
|
||||
log_message(LOG_LEVEL_ERROR, "--bwlimit=%s is too large", value);
|
||||
return -1;
|
||||
}
|
||||
*bytes_per_sec_out = (unsigned long long)kib * 1024;
|
||||
return 0;
|
||||
}
|
||||
|
||||
/* Apply a --bwlimit value using rsync 3.4.1's units/semantics. Returns 0 on
|
||||
* success, -1 on error. */
|
||||
static int set_bwlimit_option(const char* value) {
|
||||
unsigned long long kbps;
|
||||
if (parse_ull_arg(value, &kbps, "--bwlimit") != 0)
|
||||
unsigned long long bytes_per_sec;
|
||||
if (parse_bwlimit_value(value, &bytes_per_sec) != 0)
|
||||
return -1;
|
||||
if (kbps == 0) {
|
||||
log_message(LOG_LEVEL_ERROR, "--bwlimit must be a positive integer");
|
||||
return -1;
|
||||
}
|
||||
if (kbps > ULLONG_MAX / 1024) {
|
||||
log_message(LOG_LEVEL_ERROR, "--bwlimit value too large");
|
||||
return -1;
|
||||
}
|
||||
io_set_bwlimit(kbps * 1024);
|
||||
log_info_message(LOG_INFO_MISC, "Set bandwidth limit to %llu KB/s", kbps);
|
||||
io_set_bwlimit(bytes_per_sec);
|
||||
log_info_message(LOG_INFO_MISC, "Set bandwidth limit to %llu KB/s", bytes_per_sec / 1024);
|
||||
return 0;
|
||||
}
|
||||
|
||||
@@ -2408,7 +2611,32 @@ static bool cli_handle_outbuf_option(CliParseCtx* ctx) {
|
||||
* load --files-from once every argument has been seen. Returns 0 on success,
|
||||
* -1 on error. */
|
||||
static int cli_finalize_config(Config* config, bool verbose, bool no_delta, bool no_incremental) {
|
||||
set_log_level(config->quiet ? LOG_LEVEL_ERROR : (verbose ? LOG_LEVEL_DEBUG : LOG_LEVEL_WARNING));
|
||||
/* An explicit --debug=FLAGS enables the debug log level by itself (rsync
|
||||
behaviour); -v enables every other INFO-level message. */
|
||||
bool debug_enabled = verbose || config->debug_level != 0;
|
||||
set_log_level(config->quiet ? LOG_LEVEL_ERROR
|
||||
: (debug_enabled ? LOG_LEVEL_DEBUG : LOG_LEVEL_WARNING));
|
||||
/* rsync's plain --delete defaults to delete-during (--del): each directory's
|
||||
extras are removed as that directory is processed, so space is freed
|
||||
progressively and a tight destination never has to hold the whole old+new
|
||||
tree at once. The late whole-tree commit FastSync historically used is
|
||||
still selected explicitly by --delete-after or by the FastSync-only long
|
||||
spelling --delete-commit (an exact alias for --delete-after). Resolve the
|
||||
default on the client, before validation and before the config crosses the
|
||||
wire, so exactly one timing flag is ever set; an explicit timing (including
|
||||
--delete-commit) always wins. */
|
||||
if (config->use_delete && !config->delete_before && !config->delete_during &&
|
||||
!config->delete_delay && !config->delete_after)
|
||||
config->delete_during = true;
|
||||
/* rsync parity: --partial-dir=DIR chooses where an interrupted transfer's
|
||||
partial file is kept, so it implies --partial. rsync applies the
|
||||
implication after option parsing, so it wins over an explicit --no-partial
|
||||
regardless of the order the two options appear in (verified on rsync
|
||||
3.4.1). --inplace is the exception: the destination file is written in
|
||||
place with no partial/temp staging, so the partial machinery is bypassed
|
||||
and the implication is skipped to leave --inplace behavior untouched. */
|
||||
if (config->partial_dir && !config->inplace)
|
||||
config->partial = true;
|
||||
if (config->compress_choice) {
|
||||
int algo = compression_algo_from_name(config->compress_choice);
|
||||
if (algo >= 0) {
|
||||
@@ -2416,14 +2644,49 @@ static int cli_finalize_config(Config* config, bool verbose, bool no_delta, bool
|
||||
config->use_compression = (algo != (int)COMPRESSION_ALGO_NONE);
|
||||
}
|
||||
}
|
||||
if (config->use_compression && config->compression_algo == (int)COMPRESSION_ALGO_NONE)
|
||||
config->compression_algo = (int)compression_negotiate_default();
|
||||
/* A bare -z (no --compress-choice) resolves like rsync's "auto": the
|
||||
* RSYNC_COMPRESS_LIST preference list first, then the compiled-in order. A
|
||||
* list that names no supported codec is rsync's failed negotiation (exit 4). */
|
||||
if (config->use_compression && !config->compress_choice) {
|
||||
int resolved = compression_choice_resolve();
|
||||
if (resolved < 0) {
|
||||
log_message(LOG_LEVEL_ERROR, "RSYNC_COMPRESS_LIST names no supported compression algorithm");
|
||||
config->cli.cli_exit_code = 4;
|
||||
return -1;
|
||||
}
|
||||
config->compression_algo = resolved;
|
||||
if (resolved == (int)COMPRESSION_ALGO_NONE)
|
||||
config->use_compression = false;
|
||||
}
|
||||
/* Apply rsync's per-codec compression level: an explicit --compress-level is
|
||||
* clamped to the codec's range, otherwise the codec's own default is used. */
|
||||
if (config->use_compression) {
|
||||
CompressionAlgo algo = (CompressionAlgo)config->compression_algo;
|
||||
config->compression_level = config->cli.compression_level_set
|
||||
? compression_clamp_level(algo, config->compression_level)
|
||||
: compression_default_level(algo);
|
||||
log_debug_message(LOG_DEBUG_UTIL, "Client compression: %s (level %d)",
|
||||
compression_algo_name(algo), config->compression_level);
|
||||
}
|
||||
/* The negotiated checksum is always resolved (rsync negotiates one for the
|
||||
* delta strong sum even without --checksum): RSYNC_CHECKSUM_LIST first, then
|
||||
* the compiled-in order. An explicit --checksum-choice already set it. */
|
||||
if (!config->cli.checksum_choice_set) {
|
||||
int resolved = checksum_choice_resolve();
|
||||
if (resolved < 0) {
|
||||
log_message(LOG_LEVEL_ERROR, "RSYNC_CHECKSUM_LIST names no supported checksum algorithm");
|
||||
config->cli.cli_exit_code = 4;
|
||||
return -1;
|
||||
}
|
||||
config->checksum_algo = resolved;
|
||||
config->cli.checksum_transfer_algo = resolved;
|
||||
}
|
||||
/* rsync parity: "none" as the pre-transfer checksum cannot be combined with
|
||||
* --checksum (exit 4). The check runs here because --checksum may appear on
|
||||
* either side of --checksum-choice. */
|
||||
if (config->checksum && config->checksum_algo == (int)CHECKSUM_ALGO_NONE) {
|
||||
log_message(LOG_LEVEL_ERROR, "Invalid checksum-choice for --checksum: none");
|
||||
config->cli_exit_code = 4;
|
||||
config->cli.cli_exit_code = 4;
|
||||
return -1;
|
||||
}
|
||||
|
||||
@@ -2502,11 +2765,11 @@ static int cli_finalize_config(Config* config, bool verbose, bool no_delta, bool
|
||||
* explicitly negated them (--no-perms/--no-times/--no-preserve). This runs
|
||||
* BEFORE the derived use_metadata bit so the transport frame is still sent
|
||||
* for the incremental/delta handshake even when both attributes were negated
|
||||
* via --no-preserve (metadata_explicitly_disabled handles that opt-out). */
|
||||
if (preserve_implied && !config->metadata_explicitly_disabled) {
|
||||
if (!config->preserve_perms_explicit_off)
|
||||
* via --no-preserve (cli.metadata_explicitly_disabled handles that opt-out). */
|
||||
if (preserve_implied && !config->cli.metadata_explicitly_disabled) {
|
||||
if (!config->cli.preserve_perms_explicit_off)
|
||||
config->preserve_perms = true;
|
||||
if (!config->preserve_times_explicit_off)
|
||||
if (!config->cli.preserve_times_explicit_off)
|
||||
config->preserve_times = true;
|
||||
}
|
||||
|
||||
@@ -2549,8 +2812,17 @@ static int cli_finalize_config(Config* config, bool verbose, bool no_delta, bool
|
||||
}
|
||||
}
|
||||
}
|
||||
config->report_stats = config->stats || config->show_progress || format_needs_wire ||
|
||||
(config->dry_run && config->use_delete);
|
||||
/* --info=del on a real --delete run asks the receiver to report the paths it
|
||||
actually removed; the report rides the STATUS_STATS path list, so the wire
|
||||
stats frame must be negotiated too. --debug=del needs the same paths, so
|
||||
it opts into the existing report (no new wire field). */
|
||||
config->report_deletes =
|
||||
config->use_delete && !config->dry_run &&
|
||||
((config->info_level & LOG_INFO_DEL) != 0 || config->itemize_changes ||
|
||||
config->out_format != NULL || (config->debug_level & LOG_DEBUG_DEL) != 0);
|
||||
config->report_stats = config->stats || config->show_progress ||
|
||||
(config->info_level & LOG_INFO_PROGRESS) || format_needs_wire ||
|
||||
config->report_deletes || (config->dry_run && config->use_delete);
|
||||
return 0;
|
||||
}
|
||||
|
||||
@@ -2884,7 +3156,7 @@ int main(int argc, char* argv[]) {
|
||||
int parse_ret = parse_args(config, argc, argv, positional_args, &positional_count);
|
||||
if (parse_ret != 0) {
|
||||
if (parse_ret < 0)
|
||||
exit_code = config->cli_exit_code ? config->cli_exit_code : 1;
|
||||
exit_code = config->cli.cli_exit_code ? config->cli.cli_exit_code : 1;
|
||||
goto cleanup;
|
||||
}
|
||||
|
||||
@@ -3027,7 +3299,7 @@ int main(int argc, char* argv[]) {
|
||||
exit_code = 1;
|
||||
}
|
||||
} else if (config->use_multithreading) {
|
||||
exit_code = send_files_multithreaded(&config);
|
||||
exit_code = send_files_multithreaded(config);
|
||||
} else {
|
||||
exit_code = send_files(config);
|
||||
}
|
||||
|
||||
@@ -0,0 +1,682 @@
|
||||
#include "client_send_internal.h"
|
||||
#include "array_list.h"
|
||||
#include "change_list.h"
|
||||
#include "charset.h"
|
||||
#include "config.h"
|
||||
#include "data.h"
|
||||
#include "delta.h"
|
||||
#include "file.h"
|
||||
#include "format.h"
|
||||
#include "log.h"
|
||||
#include "protocol.h"
|
||||
#include "scanner.h"
|
||||
#include "transport_tls.h"
|
||||
#include "utils.h"
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <sys/stat.h>
|
||||
#include <time.h>
|
||||
|
||||
/* True when --dry-run should contact a receiver rather than running the
|
||||
* client-side local manifest. Any target a real run would reach over the wire
|
||||
* selects the server-contacting path: a remote (SSH host:path), a daemon
|
||||
* (host::module/path), an explicit --server-host, --server-port/--port, TLS, or
|
||||
* a source-bind --address. A plain local destination (none of these) keeps the
|
||||
* original client-side behavior, which never dials the default 127.0.0.1:8080. */
|
||||
bool dry_run_targets_server(const Config* config) {
|
||||
if (!config)
|
||||
return false;
|
||||
if (config->transport == TRANSPORT_SSH)
|
||||
return true;
|
||||
if (config->module && config->module[0] != '\0')
|
||||
return true;
|
||||
if (config->cli.server_host_set || config->cli.server_port_set)
|
||||
return true;
|
||||
if (config->use_tls)
|
||||
return true;
|
||||
if (config->address != NULL)
|
||||
return true;
|
||||
return false;
|
||||
}
|
||||
|
||||
bool add_chunk_to_manifest(ArrayList* manifest, const Chunk* chunk) {
|
||||
if (!manifest)
|
||||
return true;
|
||||
for (int i = 0; i < chunk->element_count; i++) {
|
||||
const char* path = file_wire_path(chunk->items[i]);
|
||||
if (*path == '/')
|
||||
path++;
|
||||
char* entry = str_dup(path);
|
||||
if (!entry) {
|
||||
log_message(LOG_LEVEL_ERROR, "Failed to allocate manifest entry");
|
||||
return false;
|
||||
}
|
||||
if (!array_list_add(manifest, entry)) {
|
||||
free(entry);
|
||||
return false;
|
||||
}
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
/* Print dry-run manifest showing files that would be transferred. Returns 0 on success. */
|
||||
int send_dry_run_manifest(const Config* config) {
|
||||
int skipped = 0;
|
||||
ArrayList* missing_dest = NULL;
|
||||
if (config->delete_missing_args) {
|
||||
missing_dest = array_list_create(free);
|
||||
if (!missing_dest)
|
||||
return -1;
|
||||
}
|
||||
if (!files_from_list_check(config, missing_dest, &skipped)) {
|
||||
if (missing_dest)
|
||||
array_list_delete(missing_dest);
|
||||
return -1;
|
||||
}
|
||||
PreparedScanner prepared;
|
||||
if (!prepare_scanner(config, 0, &prepared)) {
|
||||
if (missing_dest)
|
||||
array_list_delete(missing_dest);
|
||||
return -1;
|
||||
}
|
||||
DirectoryScanner* scanner =
|
||||
directory_scanner_create_with_options(config->send_directory, &prepared.options);
|
||||
if (!scanner) {
|
||||
prepared_scanner_destroy(&prepared);
|
||||
if (missing_dest)
|
||||
array_list_delete(missing_dest);
|
||||
return -1;
|
||||
}
|
||||
Chunk* chunk;
|
||||
int file_count = 0;
|
||||
unsigned long long total_bytes = 0;
|
||||
char size_buffer[32];
|
||||
if (!config->quiet)
|
||||
printf("Dry run: files to be transferred\n");
|
||||
while ((chunk = directory_scanner_next(scanner)) != NULL) {
|
||||
for (int i = 0; i < chunk->element_count; i++) {
|
||||
if (!config->quiet) {
|
||||
char* escaped_path =
|
||||
output_escape(file_wire_path(chunk->items[i]), config->eight_bit_output);
|
||||
if (!escaped_path) {
|
||||
chunk_destroy(chunk);
|
||||
directory_scanner_destroy(scanner);
|
||||
prepared_scanner_destroy(&prepared);
|
||||
if (missing_dest)
|
||||
array_list_delete(missing_dest);
|
||||
return -1;
|
||||
}
|
||||
if (config->human_readable)
|
||||
printf(
|
||||
" %s (%s)\n", escaped_path,
|
||||
display_bytes(chunk->items[i]->data->size, true, size_buffer, sizeof(size_buffer)));
|
||||
else
|
||||
printf(" %s (%zu bytes)\n", escaped_path, chunk->items[i]->data->size);
|
||||
free(escaped_path);
|
||||
}
|
||||
total_bytes += chunk->items[i]->data->size;
|
||||
file_count++;
|
||||
}
|
||||
chunk_destroy(chunk);
|
||||
}
|
||||
directory_scanner_destroy(scanner);
|
||||
prepared_scanner_destroy(&prepared);
|
||||
/* --delete-missing-args: the missing entries' destination mirrors render as
|
||||
would-be deletions (rsync's dry-run also lists its *deleting lines). */
|
||||
if (missing_dest && !config->quiet) {
|
||||
for (int i = 0; i < missing_dest->size; i++) {
|
||||
char* escaped = output_escape((char*)missing_dest->items[i], config->eight_bit_output);
|
||||
printf(" %s (missing; would be deleted)\n", escaped ? escaped : "<allocation failed>");
|
||||
free(escaped);
|
||||
}
|
||||
}
|
||||
if (missing_dest)
|
||||
array_list_delete(missing_dest);
|
||||
if (!config->quiet) {
|
||||
if (config->human_readable)
|
||||
printf("Total: %d files, %s\n", file_count,
|
||||
display_bytes(total_bytes, true, size_buffer, sizeof(size_buffer)));
|
||||
else
|
||||
printf("Total: %d files, %.1f MB\n", file_count, (double)total_bytes / (double)BYTES_PER_MIB);
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
typedef struct {
|
||||
char* name; /* transfer-relative name ("" == the source root) */
|
||||
mode_t mode;
|
||||
unsigned long long size;
|
||||
time_t mtime;
|
||||
long mtime_nsec;
|
||||
bool is_dir;
|
||||
bool is_symlink;
|
||||
char* link_target;
|
||||
} ListEntry;
|
||||
|
||||
static void list_entries_destroy(ListEntry* entries, size_t count) {
|
||||
if (entries == NULL)
|
||||
return;
|
||||
for (size_t i = 0; i < count; i++) {
|
||||
free(entries[i].name);
|
||||
free(entries[i].link_target);
|
||||
}
|
||||
free(entries);
|
||||
}
|
||||
|
||||
static int compare_list_entries(const void* left, const void* right) {
|
||||
const ListEntry* a = (const ListEntry*)left;
|
||||
const ListEntry* b = (const ListEntry*)right;
|
||||
return strcmp(a->name, b->name);
|
||||
}
|
||||
|
||||
/* Relative path of an entry below `root` ("" for the root itself). Mirrors
|
||||
* change_list's relative_name for list-only rendering. */
|
||||
static char* list_relative_name(const char* root, const char* full) {
|
||||
if (root == NULL || full == NULL)
|
||||
return str_dup(full != NULL ? full : "");
|
||||
size_t root_len = strlen(root);
|
||||
while (root_len > 1 && root[root_len - 1] == '/')
|
||||
root_len--;
|
||||
if (strncmp(root, full, root_len) == 0) {
|
||||
if (full[root_len] == '\0')
|
||||
return str_dup("");
|
||||
if (full[root_len] == '/')
|
||||
return str_dup(full + root_len + 1);
|
||||
}
|
||||
return str_dup(full);
|
||||
}
|
||||
|
||||
/* --list-only: print an ls-style listing of the entries that WOULD be
|
||||
* transferred and exit without contacting the server or writing anything.
|
||||
* Names are transfer-relative (rsync prints `a.txt`, `sub/b.txt`, `.`) and
|
||||
* directory entries are included. Returns 0 on success, 1 on error. */
|
||||
int send_list_only(const Config* config) {
|
||||
int skipped = 0;
|
||||
if (!files_from_list_check(config, NULL, &skipped))
|
||||
return 1;
|
||||
PreparedScanner prepared;
|
||||
if (!prepare_scanner(config, 0, &prepared))
|
||||
return 1;
|
||||
prepared.options.use_metadata = true; /* capture mode + mtime for the listing */
|
||||
prepared.options.list_dirs = true;
|
||||
DirectoryScanner* scanner =
|
||||
directory_scanner_create_with_options(config->send_directory, &prepared.options);
|
||||
if (!scanner) {
|
||||
prepared_scanner_destroy(&prepared);
|
||||
return 1;
|
||||
}
|
||||
ListEntry* entries = NULL;
|
||||
size_t count = 0;
|
||||
size_t capacity = 0;
|
||||
bool oom = false;
|
||||
|
||||
/* rsync lists the source root itself (as "."). Only when the source is a
|
||||
* directory and no --files-from subset is in effect. */
|
||||
if (config->files_from_set == NULL && config->send_directory != NULL) {
|
||||
struct stat st;
|
||||
if (stat(config->send_directory, &st) == 0 && S_ISDIR(st.st_mode)) {
|
||||
capacity = 64;
|
||||
entries = calloc(capacity, sizeof(ListEntry));
|
||||
if (entries == NULL) {
|
||||
oom = true;
|
||||
} else if ((entries[0].name = str_dup("")) == NULL) {
|
||||
/* A NULL name would be dereferenced by qsort/render: fail the listing. */
|
||||
oom = true;
|
||||
} else {
|
||||
entries[0].mode = st.st_mode;
|
||||
entries[0].mtime = st.st_mtime;
|
||||
entries[0].mtime_nsec = st.st_mtim.tv_nsec;
|
||||
entries[0].size = (unsigned long long)st.st_size;
|
||||
entries[0].is_dir = true;
|
||||
count = 1;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
Chunk* chunk;
|
||||
while (!oom && (chunk = directory_scanner_next(scanner)) != NULL) {
|
||||
for (int i = 0; i < chunk->element_count; i++) {
|
||||
File* f = chunk->items[i];
|
||||
if (f == NULL)
|
||||
continue;
|
||||
if (count == capacity) {
|
||||
size_t new_capacity = capacity > 0 ? capacity * 2 : 64;
|
||||
if (new_capacity <= capacity) {
|
||||
oom = true;
|
||||
break;
|
||||
}
|
||||
ListEntry* grown = realloc(entries, new_capacity * sizeof(ListEntry));
|
||||
if (!grown) {
|
||||
oom = true;
|
||||
break;
|
||||
}
|
||||
entries = grown;
|
||||
memset(entries + capacity, 0, (new_capacity - capacity) * sizeof(ListEntry));
|
||||
capacity = new_capacity;
|
||||
}
|
||||
char* name = list_relative_name(config->send_directory, file_wire_path(f));
|
||||
if (!name) {
|
||||
oom = true;
|
||||
break;
|
||||
}
|
||||
mode_t mode = 0;
|
||||
time_t mtime = 0;
|
||||
long mtime_nsec = 0;
|
||||
if (f->metadata != NULL) {
|
||||
mode = f->metadata->mode;
|
||||
mtime = f->metadata->mtime_sec;
|
||||
mtime_nsec = f->metadata->mtime_nsec;
|
||||
} else {
|
||||
struct stat st;
|
||||
if (lstat(f->path, &st) == 0) {
|
||||
mode = st.st_mode;
|
||||
mtime = st.st_mtime;
|
||||
mtime_nsec = st.st_mtim.tv_nsec;
|
||||
}
|
||||
}
|
||||
entries[count].name = name;
|
||||
entries[count].mode = mode;
|
||||
entries[count].mtime = mtime;
|
||||
entries[count].mtime_nsec = mtime_nsec;
|
||||
if (f->is_symlink)
|
||||
entries[count].size = f->symlink_target != NULL ? strlen(f->symlink_target) : 0;
|
||||
else if (f->is_dir) {
|
||||
struct stat dir_st;
|
||||
entries[count].size = stat(f->path, &dir_st) == 0 ? (unsigned long long)dir_st.st_size : 0;
|
||||
} else
|
||||
entries[count].size = f->data != NULL ? f->data->size : 0;
|
||||
entries[count].is_dir = f->is_dir;
|
||||
entries[count].is_symlink = f->is_symlink;
|
||||
entries[count].link_target =
|
||||
f->is_symlink && f->symlink_target ? str_dup(f->symlink_target) : NULL;
|
||||
count++;
|
||||
}
|
||||
chunk_destroy(chunk);
|
||||
}
|
||||
bool failed = oom || directory_scanner_failed(scanner) || directory_scanner_had_io_error(scanner);
|
||||
directory_scanner_destroy(scanner);
|
||||
prepared_scanner_destroy(&prepared);
|
||||
if (failed) {
|
||||
list_entries_destroy(entries, count);
|
||||
if (oom)
|
||||
log_message(LOG_LEVEL_ERROR, "memory allocation failed while listing");
|
||||
return 1;
|
||||
}
|
||||
if (count > 1)
|
||||
qsort(entries, count, sizeof(ListEntry), compare_list_entries);
|
||||
for (size_t i = 0; i < count; i++) {
|
||||
ChangeEvent event;
|
||||
memset(&event, 0, sizeof(event));
|
||||
event.name = entries[i].name;
|
||||
event.path = entries[i].name;
|
||||
event.mode = entries[i].mode;
|
||||
event.size = entries[i].size;
|
||||
event.mtime_sec = entries[i].mtime;
|
||||
event.mtime_nsec = entries[i].mtime_nsec;
|
||||
event.is_directory = entries[i].is_dir;
|
||||
event.is_symlink = entries[i].is_symlink;
|
||||
event.symlink_target = entries[i].link_target;
|
||||
char* line = change_render_list_line(config, &event);
|
||||
if (line != NULL) {
|
||||
char* escaped = output_escape(line, config->eight_bit_output);
|
||||
printf("%s\n", escaped != NULL ? escaped : line);
|
||||
free(escaped);
|
||||
free(line);
|
||||
}
|
||||
}
|
||||
list_entries_destroy(entries, count);
|
||||
return 0;
|
||||
}
|
||||
|
||||
/* Send the delete manifest to the server. Returns 0 on success, -1 on
|
||||
failure. It carries FOUR sections: the keep-set paths, the protected
|
||||
excluded prefixes, the --delete-missing-args exact-delete paths, and the
|
||||
destination-relative directories the sender synchronized this run.
|
||||
When --delete-excluded is given `protected` is empty: excluded destination
|
||||
mirrors are then ordinary extras and are removed. When
|
||||
--delete-missing-args is active `missing_args` holds the destination mirrors
|
||||
of missing --files-from entries: each is an explicit receiver-side deletion
|
||||
request, independent of the extras walk. `synced_dirs` confines the extras
|
||||
walk to entries directly inside a synchronized directory. A NULL
|
||||
keep-set / protected / missing / dirs list transmits an empty section. All
|
||||
four sections are unbounded on the sender; the receiver enforces
|
||||
MAX_MANIFEST_ENTRIES per section and a single MAX_MANIFEST_BYTES budget
|
||||
shared across the sections, rejecting (with STATUS_ERROR) an over-budget
|
||||
frame. A heavily filtered source whose exclusion list is large therefore
|
||||
fails the run cleanly on the receiver rather than being truncated. */
|
||||
int send_delete_manifest(int fd, ArrayList* manifest, ArrayList* protected_prefixes,
|
||||
ArrayList* size_skipped, ArrayList* missing_args, ArrayList* synced_dirs) {
|
||||
if (!send_status(fd, STATUS_MANIFEST))
|
||||
return -1;
|
||||
int keep_count = manifest ? manifest->size : 0;
|
||||
if (!send_int(fd, keep_count))
|
||||
return -1;
|
||||
for (int i = 0; i < keep_count; i++) {
|
||||
if (!send_wire_str(fd, (char*)manifest->items[i]))
|
||||
return -1;
|
||||
}
|
||||
/* The receiver has ONE protected-prefix section; filter-excluded prefixes
|
||||
(dropped under --delete-excluded) and size-pruned prefixes (always
|
||||
protected) are concatenated into it. */
|
||||
int protected_count =
|
||||
(protected_prefixes ? protected_prefixes->size : 0) + (size_skipped ? size_skipped->size : 0);
|
||||
if (!send_int(fd, protected_count))
|
||||
return -1;
|
||||
if (protected_prefixes) {
|
||||
for (int i = 0; i < protected_prefixes->size; i++) {
|
||||
if (!send_wire_str(fd, (char*)protected_prefixes->items[i]))
|
||||
return -1;
|
||||
}
|
||||
}
|
||||
if (size_skipped) {
|
||||
for (int i = 0; i < size_skipped->size; i++) {
|
||||
if (!send_wire_str(fd, (char*)size_skipped->items[i]))
|
||||
return -1;
|
||||
}
|
||||
}
|
||||
int missing_count = missing_args ? missing_args->size : 0;
|
||||
if (!send_int(fd, missing_count))
|
||||
return -1;
|
||||
for (int i = 0; i < missing_count; i++) {
|
||||
if (!send_wire_str(fd, (char*)missing_args->items[i]))
|
||||
return -1;
|
||||
}
|
||||
int dirs_count = synced_dirs ? synced_dirs->size : 0;
|
||||
if (!send_int(fd, dirs_count))
|
||||
return -1;
|
||||
for (int i = 0; i < dirs_count; i++) {
|
||||
if (!send_wire_str(fd, (char*)synced_dirs->items[i]))
|
||||
return -1;
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
/* Transmit the keep-set manifest and wait for the receiver's verdict. Used by
|
||||
--delete-before/--delete-during, where the extras are removed on the receiver
|
||||
BEFORE the first byte of file data is sent: the receiver acknowledges with
|
||||
STATUS_OK once the bounded delete committed, or STATUS_ERROR if it could not
|
||||
(in which case the sender aborts without streaming any data). The ACK may
|
||||
take much longer than an ordinary per-message round trip because the receiver
|
||||
performs the whole bounded deletion walk (up to MAX_SERVER_DELETE_COUNT
|
||||
unlinks) before replying, so the wait uses a generous explicit deadline
|
||||
instead of the default 60 s receive window. */
|
||||
#define DELETE_ACK_TIMEOUT_SEC 3600
|
||||
/* While waiting for the (potentially slow) receiver-side deletion, send a
|
||||
* STATUS_KEEPALIVE at most this often so the connection is demonstrably alive
|
||||
* and neither side's per-message timeout trips. */
|
||||
#define DELETE_ACK_KEEPALIVE_SEC 10
|
||||
|
||||
bool send_delete_manifest_early(Client* client, ArrayList* manifest, ArrayList* protected_prefixes,
|
||||
ArrayList* size_skipped, ArrayList* missing_args,
|
||||
ArrayList* synced_dirs) {
|
||||
if (!client || !manifest)
|
||||
return false;
|
||||
if (send_delete_manifest(client->file_descriptor, manifest, protected_prefixes, size_skipped,
|
||||
missing_args, synced_dirs) != 0)
|
||||
return false;
|
||||
Status ack;
|
||||
/* The wait is long (up to an hour) and runs inline on this thread: a helper
|
||||
* thread would race the non-thread-safe protocol send path, so keepalives are
|
||||
* emitted from this wait loop itself. A Ctrl-C/SIGTERM abort flag also ends
|
||||
* the wait; the caller then best-effort sends STATUS_ABORT. */
|
||||
if (!receive_status_keepalive(client->file_descriptor, &ack, DELETE_ACK_TIMEOUT_SEC,
|
||||
DELETE_ACK_KEEPALIVE_SEC, client_abort_pending)) {
|
||||
/* A Ctrl-C/SIGTERM abort ends the wait above; tell the receiver before the
|
||||
caller tears the connection down (best-effort). */
|
||||
if (client_abort_pending()) {
|
||||
log_info_message(LOG_INFO_MISC,
|
||||
"Abort requested while awaiting delete ack; sending STATUS_ABORT");
|
||||
send_status(client->file_descriptor, STATUS_ABORT);
|
||||
}
|
||||
return false;
|
||||
}
|
||||
if (ack != STATUS_OK) {
|
||||
log_server_rejection("Server failed to delete files before the transfer");
|
||||
return false;
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
/* Server-contacting --dry-run. Connects to the configured remote/daemon and
|
||||
* runs the normal per-file incremental decision WITHOUT transmitting any file
|
||||
* data: the receiver (which also sees dry_run=true on the wire) answers
|
||||
* STATUS_OK for an up-to-date file and STATUS_DRY_RUN_TRANSFER for a file it
|
||||
* would otherwise write, mutating nothing on either side. The would-transfer
|
||||
* set and the same trailer as the local dry-run are printed. A
|
||||
* --compare-dest exact basis hit with no destination copy is reported as a
|
||||
* skip by the receiver.
|
||||
*
|
||||
* Only regular files take the receiver-consulted check; directory / symlink /
|
||||
* special / hard-link-sibling entries have no per-file content check, so they
|
||||
* are reported conservatively as would-transfer and their frames are never
|
||||
* sent (which is what keeps the receiver mutation-free). --delete* is
|
||||
* deliberately NOT transmitted in dry-run, so no deletion can occur; the
|
||||
* would-delete manifest report is a documented follow-up.
|
||||
*
|
||||
* Returns 0 on success, 1 on error. */
|
||||
int send_dry_run_remote(Config* config) {
|
||||
int from_skipped = 0;
|
||||
ArrayList* missing_args = NULL;
|
||||
if (config->delete_missing_args) {
|
||||
missing_args = array_list_create(free);
|
||||
if (!missing_args)
|
||||
return 1;
|
||||
}
|
||||
if (!files_from_list_check(config, missing_args, &from_skipped)) {
|
||||
if (missing_args)
|
||||
array_list_delete(missing_args);
|
||||
return 1;
|
||||
}
|
||||
if (missing_args)
|
||||
array_list_delete(missing_args);
|
||||
/* A live session may follow, so arm graceful abort handling. */
|
||||
client_set_abort_armed(true);
|
||||
Client* client = connect_transfer_client(config);
|
||||
if (!client) {
|
||||
if (config->transport == TRANSPORT_TCP)
|
||||
log_message(LOG_LEVEL_ERROR, "could not connect to server%s",
|
||||
config->use_tls ? " via TLS" : "");
|
||||
client_set_abort_armed(false);
|
||||
return 1;
|
||||
}
|
||||
ProtocolSession session;
|
||||
protocol_session_init(&session, client->file_descriptor, client->file_descriptor);
|
||||
protocol_session_set_io_timeout(&session, config->timeout);
|
||||
protocol_session_set_ssl(&session, (SSL*)client->ssl);
|
||||
protocol_session_bind(&session);
|
||||
|
||||
int ret = 1;
|
||||
time_t dry_start = time(NULL);
|
||||
ReceiverStats dry_stats;
|
||||
memset(&dry_stats, 0, sizeof(dry_stats));
|
||||
PreparedScanner prepared;
|
||||
memset(&prepared, 0, sizeof(prepared));
|
||||
DirectoryScanner* scanner = NULL;
|
||||
ArrayList* dry_manifest = NULL;
|
||||
ArrayList* dry_dirs = NULL;
|
||||
ArrayList* dry_excluded = NULL;
|
||||
ArrayList* dry_size_skipped = NULL;
|
||||
if (!config_send(client->file_descriptor, config))
|
||||
goto dry_fail;
|
||||
receive_daemon_motd(client, config);
|
||||
if (!prepare_scanner(config, 0, &prepared))
|
||||
goto dry_fail;
|
||||
/* -n --delete: build the same keep-set manifest, protected prefixes, and
|
||||
synchronized-directory scope a real run would send, so the receiver's
|
||||
read-only extras walk enumerates exactly the deletions a real run makes. */
|
||||
if (config->use_delete) {
|
||||
dry_manifest = array_list_create(free);
|
||||
dry_dirs = array_list_create(free);
|
||||
dry_size_skipped = array_list_create(free);
|
||||
if (!dry_manifest || !dry_dirs || !dry_size_skipped)
|
||||
goto dry_fail;
|
||||
if (!config->delete_excluded) {
|
||||
dry_excluded = array_list_create(free);
|
||||
if (!dry_excluded)
|
||||
goto dry_fail;
|
||||
prepared.options.excluded_paths = dry_excluded;
|
||||
}
|
||||
prepared.options.size_skipped_paths = dry_size_skipped;
|
||||
/* A --files-from subset confines the extras walk to the directories the
|
||||
scan synchronized; a full recursive transfer marks the root itself. */
|
||||
if (config->files_from_set == NULL) {
|
||||
char* root_marker = delete_scope_root_marker(config);
|
||||
if (!root_marker || !array_list_add(dry_dirs, root_marker)) {
|
||||
free(root_marker);
|
||||
goto dry_fail;
|
||||
}
|
||||
} else {
|
||||
prepared.options.synced_dirs = dry_dirs;
|
||||
}
|
||||
}
|
||||
scanner = directory_scanner_create_with_options(config->send_directory, &prepared.options);
|
||||
if (!scanner)
|
||||
goto dry_fail;
|
||||
|
||||
int file_count = 0;
|
||||
unsigned long long total_bytes = 0;
|
||||
char size_buffer[32];
|
||||
if (!config->quiet)
|
||||
printf("Dry run: files to be transferred\n");
|
||||
Chunk* chunk;
|
||||
while ((chunk = directory_scanner_next(scanner)) != NULL) {
|
||||
if (dry_manifest && !add_chunk_to_manifest(dry_manifest, chunk)) {
|
||||
chunk_destroy(chunk);
|
||||
goto dry_fail;
|
||||
}
|
||||
for (int i = 0; i < chunk->element_count; i++) {
|
||||
File* f = chunk->items[i];
|
||||
if (!f)
|
||||
continue;
|
||||
unsigned long long fsize = f->data ? f->data->size : 0;
|
||||
bool would;
|
||||
if (f->is_dir || f->is_symlink || f->is_special ||
|
||||
(f->link_group != 0 && !f->link_first && f->hardlink_target != NULL)) {
|
||||
/* No receiver-side content check exists for these frame types; a real
|
||||
run would (re)create them, so report would-transfer and send no
|
||||
frame (the receiver must stay mutation-free). */
|
||||
would = true;
|
||||
} else if (fsize > MAX_RECEIVE_WHOLE_FILE_SIZE && !config->use_incremental &&
|
||||
!config_has_basis(config)) {
|
||||
/* A non-incremental run streams a >whole-file-limit source without the
|
||||
STATUS_CHECK handshake, so no read-only receiver decision is possible
|
||||
(and none is needed: a real run would transfer it). */
|
||||
would = true;
|
||||
} else {
|
||||
DeltaSignature* sig = NULL;
|
||||
unsigned long long resume_offset = 0;
|
||||
int rc = incremental_check(client, f, config, &sig, &resume_offset);
|
||||
delta_signature_destroy(sig);
|
||||
if (rc < 0) {
|
||||
chunk_destroy(chunk);
|
||||
goto dry_fail;
|
||||
}
|
||||
if (rc == 1)
|
||||
continue; /* up to date; nothing to report */
|
||||
if (rc != 4) {
|
||||
log_message(LOG_LEVEL_ERROR, "Unexpected receiver reply during dry-run");
|
||||
chunk_destroy(chunk);
|
||||
goto dry_fail;
|
||||
}
|
||||
would = true;
|
||||
}
|
||||
if (would) {
|
||||
if (!config->quiet) {
|
||||
char* escaped_path = output_escape(file_wire_path(f), config->eight_bit_output);
|
||||
if (!escaped_path) {
|
||||
chunk_destroy(chunk);
|
||||
goto dry_fail;
|
||||
}
|
||||
if (config->human_readable)
|
||||
printf(" %s (%s)\n", escaped_path,
|
||||
display_bytes(fsize, true, size_buffer, sizeof(size_buffer)));
|
||||
else
|
||||
printf(" %s (%llu bytes)\n", escaped_path, fsize);
|
||||
free(escaped_path);
|
||||
}
|
||||
total_bytes += fsize;
|
||||
file_count++;
|
||||
}
|
||||
}
|
||||
chunk_destroy(chunk);
|
||||
}
|
||||
bool io_error = directory_scanner_had_io_error(scanner);
|
||||
if (directory_scanner_failed(scanner))
|
||||
goto dry_fail;
|
||||
if (io_error)
|
||||
log_message(LOG_LEVEL_WARNING, "source scan hit an unreadable directory");
|
||||
/* Send the keep-set manifest (no data frames) so the receiver can enumerate
|
||||
the destination extras; an early-timing delete ACKs before it will accept
|
||||
the terminal FINISHED. */
|
||||
bool early_delete = config->use_delete && config_delete_timing_early(config);
|
||||
if (dry_manifest) {
|
||||
if (send_delete_manifest(client->file_descriptor, dry_manifest, dry_excluded, dry_size_skipped,
|
||||
NULL, dry_dirs) != 0)
|
||||
goto dry_fail;
|
||||
if (early_delete) {
|
||||
Status ack;
|
||||
if (!receive_status_keepalive(client->file_descriptor, &ack, DELETE_ACK_TIMEOUT_SEC,
|
||||
DELETE_ACK_KEEPALIVE_SEC, client_abort_pending) ||
|
||||
ack != STATUS_OK)
|
||||
goto dry_fail;
|
||||
}
|
||||
}
|
||||
/* Terminate the stream so the receiver emits its success frame; no data frame
|
||||
is ever sent in dry-run. */
|
||||
if (!send_status(client->file_descriptor, STATUS_FINISHED))
|
||||
goto dry_fail;
|
||||
Status status;
|
||||
if (!receive_status(client->file_descriptor, &status))
|
||||
goto dry_fail;
|
||||
if (status == STATUS_STATS) {
|
||||
ArrayList* would_delete = array_list_create(free);
|
||||
if (!would_delete)
|
||||
goto dry_fail;
|
||||
if (!receive_stats_record(client->file_descriptor, &dry_stats, would_delete)) {
|
||||
array_list_delete(would_delete);
|
||||
goto dry_fail;
|
||||
}
|
||||
print_delete_reports(config, would_delete);
|
||||
array_list_delete(would_delete);
|
||||
if (!receive_status(client->file_descriptor, &status))
|
||||
goto dry_fail;
|
||||
}
|
||||
if (status != STATUS_OK)
|
||||
goto dry_fail;
|
||||
if (!config->quiet) {
|
||||
if (config->human_readable)
|
||||
printf("Total: %d files, %s\n", file_count,
|
||||
display_bytes(total_bytes, true, size_buffer, sizeof(size_buffer)));
|
||||
else
|
||||
printf("Total: %d files, %.1f MB\n", file_count, (double)total_bytes / (double)BYTES_PER_MIB);
|
||||
}
|
||||
{
|
||||
TransferStats dry_transfer;
|
||||
memset(&dry_transfer, 0, sizeof(dry_transfer));
|
||||
dry_transfer.flist_reg = (unsigned long long)file_count;
|
||||
dry_transfer.total_file_size = total_bytes;
|
||||
dry_transfer.transferred_regular = (unsigned long long)file_count;
|
||||
dry_transfer.transferred_file_size = total_bytes;
|
||||
dry_transfer.literal_data = total_bytes;
|
||||
report_transfer_stats(config, &dry_transfer, dry_start, &dry_stats);
|
||||
}
|
||||
ret = io_error ? 1 : 0;
|
||||
|
||||
dry_fail:
|
||||
if (dry_manifest)
|
||||
array_list_delete(dry_manifest);
|
||||
if (dry_dirs)
|
||||
array_list_delete(dry_dirs);
|
||||
if (dry_excluded)
|
||||
array_list_delete(dry_excluded);
|
||||
if (dry_size_skipped)
|
||||
array_list_delete(dry_size_skipped);
|
||||
if (scanner)
|
||||
directory_scanner_destroy(scanner);
|
||||
prepared_scanner_destroy(&prepared);
|
||||
disconnect_transfer_client(client);
|
||||
protocol_session_unbind();
|
||||
client_set_abort_armed(false);
|
||||
return ret;
|
||||
}
|
||||
@@ -0,0 +1,890 @@
|
||||
#include "client_send_internal.h"
|
||||
#include "array_list.h"
|
||||
#include "change_list.h"
|
||||
#include "charset.h"
|
||||
#include "config.h"
|
||||
#include "file.h"
|
||||
#include "format.h"
|
||||
#include "log.h"
|
||||
#include "protocol.h"
|
||||
#include "utils.h"
|
||||
#include <stdatomic.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <time.h>
|
||||
|
||||
/* Surface a server rejection to the user. When the last status exchange
|
||||
carried a STATUS_ERROR_DETAIL reason (protocol 2.21.0) it is appended to the
|
||||
client-side context; a bare STATUS_ERROR still logs the context alone. */
|
||||
void log_server_rejection(const char* context) {
|
||||
const char* detail = protocol_last_error();
|
||||
if (detail && detail[0] != '\0') {
|
||||
/* The detail is peer-controlled: escape it so terminal/log-format
|
||||
* metacharacters cannot be injected into the client's output. */
|
||||
char* escaped = output_escape(detail, log_get_8_bit_output());
|
||||
log_message(LOG_LEVEL_ERROR, "%s: %s", context, escaped ? escaped : "<allocation failed>");
|
||||
free(escaped);
|
||||
} else {
|
||||
log_message(LOG_LEVEL_ERROR, "%s", context);
|
||||
}
|
||||
}
|
||||
|
||||
const char* display_bytes(unsigned long long bytes, bool human_readable, char* buffer,
|
||||
size_t buffer_size) {
|
||||
if (human_readable && format_human_size_decimal(bytes, buffer, buffer_size))
|
||||
return buffer;
|
||||
snprintf(buffer, buffer_size, "%.1f MB", (double)bytes / (double)BYTES_PER_MIB);
|
||||
return buffer;
|
||||
}
|
||||
|
||||
/* rsync byte count: human-readable decimal when -h was given, otherwise a
|
||||
* comma-grouped integer (rsync's big_num in the C locale). */
|
||||
static const char* stats_bytes(const Config* config, unsigned long long bytes, char* buffer,
|
||||
size_t buffer_size) {
|
||||
if (!format_big_num(bytes, config->human_readable, buffer, buffer_size))
|
||||
snprintf(buffer, buffer_size, "%llu", bytes);
|
||||
return buffer;
|
||||
}
|
||||
|
||||
/* Build rsync's per-type parenthetical: each non-zero category, in
|
||||
reg/dir/link/special order. Empty when every count is zero. */
|
||||
static void type_breakdown(unsigned long long reg, unsigned long long dir, unsigned long long link,
|
||||
unsigned long long special, char* out, size_t out_size) {
|
||||
if (reg + dir + link + special == 0) {
|
||||
out[0] = '\0';
|
||||
return;
|
||||
}
|
||||
out[0] = '\0';
|
||||
size_t used = 0;
|
||||
const struct {
|
||||
const char* name;
|
||||
unsigned long long count;
|
||||
} parts[4] = {{"reg", reg}, {"dir", dir}, {"link", link}, {"special", special}};
|
||||
bool first = true;
|
||||
for (size_t i = 0; i < 4; i++) {
|
||||
if (parts[i].count == 0)
|
||||
continue;
|
||||
int written = snprintf(out + used, out_size - used, "%s%s: %llu", first ? "(" : ", ",
|
||||
parts[i].name, parts[i].count);
|
||||
if (written < 0 || (size_t)written >= out_size - used)
|
||||
break;
|
||||
used += (size_t)written;
|
||||
first = false;
|
||||
}
|
||||
if (!first && used + 1 < out_size)
|
||||
out[used++] = ')';
|
||||
out[used] = '\0';
|
||||
}
|
||||
|
||||
/* Build rsync's `Number of files` parenthetical from the scan's flist counts. */
|
||||
static void stats_type_breakdown(const TransferStats* stats, char* out, size_t out_size) {
|
||||
type_breakdown(stats->flist_reg, stats->flist_dir, stats->flist_link, stats->flist_special, out,
|
||||
out_size);
|
||||
}
|
||||
|
||||
/* rsync's `Number of files` counts every directory. A recursive scan that
|
||||
preserves a directory attribute captures them in `dir_entries`; a `-r` scan
|
||||
(no -t/-p) captures nothing, so fall back to the scanner's shared counter of
|
||||
traversed directories that are not already represented by an inline
|
||||
directory entry. The -d generator counts its explicit directory entries
|
||||
inline and does not traverse, so it is excluded here. */
|
||||
unsigned long long dir_count_for_stats(const Config* config, const ArrayList* dir_entries,
|
||||
atomic_ullong* counter) {
|
||||
if (config == NULL || config->dirs || config->list_only)
|
||||
return 0;
|
||||
if (dir_metadata_should_capture(config))
|
||||
return dir_entries != NULL ? (unsigned long long)dir_entries->size : 0;
|
||||
return counter != NULL ? (unsigned long long)atomic_load(counter) : 0;
|
||||
}
|
||||
|
||||
/* Print the rsync `--stats` block on stdout. The source-side flist and
|
||||
transferred counters come from `stats` (filled while scanning/sending), the
|
||||
receiver-only counters from the STATUS_STATS frame, and the wire byte totals
|
||||
from the process-wide protocol counters. The labels, layout and
|
||||
rate/speedup formulas match rsync 3.4.1. Shared by the single-threaded and
|
||||
multithreaded send paths. */
|
||||
void report_transfer_stats(const Config* config, const TransferStats* stats, time_t start,
|
||||
const ReceiverStats* recv) {
|
||||
if (!config->stats || config->quiet)
|
||||
return;
|
||||
TransferStats empty = {0};
|
||||
if (stats == NULL)
|
||||
stats = ∅
|
||||
ReceiverStats none = {0};
|
||||
if (recv == NULL)
|
||||
recv = &none;
|
||||
unsigned long long sent = protocol_bytes_written();
|
||||
unsigned long long received = protocol_bytes_read();
|
||||
/* rsync: bytes_per_sec = (written + read) / (0.5 + (end - start)). */
|
||||
double elapsed = difftime(time(NULL), start);
|
||||
double rate = (double)(sent + received) / (0.5 + elapsed);
|
||||
char total_buffer[32];
|
||||
char transferred_buffer[32];
|
||||
char literal_buffer[32];
|
||||
char matched_buffer[32];
|
||||
char sent_buffer[32];
|
||||
char recv_buffer[32];
|
||||
char rate_buffer[32] = {0};
|
||||
char human_rate[32] = {0};
|
||||
const char* total =
|
||||
stats_bytes(config, stats->total_file_size, total_buffer, sizeof(total_buffer));
|
||||
const char* transferred = stats_bytes(config, stats->transferred_file_size, transferred_buffer,
|
||||
sizeof(transferred_buffer));
|
||||
/* Protocol 2.28.0: the receiver reports the bytes it literally stored, which
|
||||
is exact for a delta transfer (the sender's own literal_data counts each
|
||||
stored file's whole source size and is only an upper bound). Fall back to
|
||||
the sender total when the receiver reported no delta/literal accounting
|
||||
(e.g. a local no-server path). */
|
||||
unsigned long long literal_bytes = (recv->literal_bytes != 0 || recv->matched_data != 0)
|
||||
? recv->literal_bytes
|
||||
: stats->literal_data;
|
||||
const char* literal = stats_bytes(config, literal_bytes, literal_buffer, sizeof(literal_buffer));
|
||||
const char* sent_s = stats_bytes(config, sent, sent_buffer, sizeof(sent_buffer));
|
||||
const char* recv_s = stats_bytes(config, received, recv_buffer, sizeof(recv_buffer));
|
||||
const char* rate_str = rate_buffer;
|
||||
if (config->human_readable) {
|
||||
if (!format_human_size_decimal((unsigned long long)rate, human_rate, sizeof(human_rate)))
|
||||
snprintf(human_rate, sizeof(human_rate), "0");
|
||||
rate_str = human_rate;
|
||||
} else {
|
||||
snprintf(rate_buffer, sizeof(rate_buffer), "%.2f", rate);
|
||||
}
|
||||
double speedup =
|
||||
(sent + received) > 0 ? (double)stats->total_file_size / (double)(sent + received) : 0.0;
|
||||
char breakdown[128];
|
||||
stats_type_breakdown(stats, breakdown, sizeof(breakdown));
|
||||
unsigned long long flist_total =
|
||||
stats->flist_reg + stats->flist_dir + stats->flist_link + stats->flist_special;
|
||||
char created_breakdown[128];
|
||||
type_breakdown(recv->created_reg, recv->created_dir, recv->created_link, recv->created_special,
|
||||
created_breakdown, sizeof(created_breakdown));
|
||||
unsigned long long created_total =
|
||||
recv->created_reg + recv->created_dir + recv->created_link + recv->created_special;
|
||||
printf("\n");
|
||||
if (breakdown[0] != '\0')
|
||||
printf("Number of files: %llu %s\n", flist_total, breakdown);
|
||||
else
|
||||
printf("Number of files: %llu\n", flist_total);
|
||||
/* Protocol 2.28.0: the receiver reports which destination entries it newly
|
||||
created, split by type, so this line matches rsync exactly. */
|
||||
if (created_breakdown[0] != '\0')
|
||||
printf("Number of created files: %llu %s\n", created_total, created_breakdown);
|
||||
else
|
||||
printf("Number of created files: %llu\n", created_total);
|
||||
printf("Number of deleted files: %llu\n", recv->deleted_files);
|
||||
printf("Number of regular files transferred: %llu\n", stats->transferred_regular);
|
||||
printf("Total file size: %s bytes\n", total);
|
||||
printf("Total transferred file size: %s bytes\n", transferred);
|
||||
printf("Literal data: %s bytes\n", literal);
|
||||
const char* matched =
|
||||
stats_bytes(config, recv->matched_data, matched_buffer, sizeof(matched_buffer));
|
||||
printf("Matched data: %s bytes\n", matched);
|
||||
printf("File list size: 0\n");
|
||||
printf("File list generation time: 0.000 seconds\n");
|
||||
printf("File list transfer time: 0.000 seconds\n");
|
||||
printf("Total bytes sent: %s\n", sent_s);
|
||||
printf("Total bytes received: %s\n", recv_s);
|
||||
printf("\n");
|
||||
printf("sent %s bytes received %s bytes %s bytes/sec\n", sent_s, recv_s, rate_str);
|
||||
printf("total size is %s speedup is %.2f%s\n", total, speedup,
|
||||
config->dry_run ? " (DRY RUN)" : "");
|
||||
fflush(stdout);
|
||||
}
|
||||
|
||||
/* Classify one scanned source entry into the rsync flist counters. Called for
|
||||
every entry the sender walks, transferred or skipped. Directory entries are
|
||||
counted here only for the explicit -d/--dirs generator; a recursive scan's
|
||||
directories are accounted from the scanner's dir_entries list at report time. */
|
||||
void transfer_stats_note_entry(TransferStats* stats, const File* file) {
|
||||
if (stats == NULL || file == NULL)
|
||||
return;
|
||||
if (file->is_dir) {
|
||||
stats->flist_dir++;
|
||||
return;
|
||||
}
|
||||
if (file->is_symlink) {
|
||||
stats->flist_link++;
|
||||
stats->total_file_size += file->symlink_target ? strlen(file->symlink_target) : 0;
|
||||
return;
|
||||
}
|
||||
if (file->is_special) {
|
||||
stats->flist_special++;
|
||||
return;
|
||||
}
|
||||
stats->flist_reg++;
|
||||
stats->total_file_size += file->data ? file->data->size : 0;
|
||||
}
|
||||
|
||||
/* Account for a regular file (or a whole-file append) the receiver actually
|
||||
stored: rsync's transferred-file count and transferred/literal byte totals.
|
||||
`literal_data` counts the whole source size, which is exact for a whole-file
|
||||
send but an upper bound for a delta send (the receiver reuses basis blocks
|
||||
the sender never ships); see TransferStats.literal_data in format.h. */
|
||||
void transfer_stats_note_transferred(TransferStats* stats, const File* file) {
|
||||
if (stats == NULL || file == NULL)
|
||||
return;
|
||||
if (file->is_dir || file->is_symlink || file->is_special)
|
||||
return;
|
||||
if (file->link_group != 0 && !file->link_first)
|
||||
return;
|
||||
unsigned long long size = file->data ? file->data->size : 0;
|
||||
stats->transferred_regular++;
|
||||
stats->transferred_file_size += size;
|
||||
stats->literal_data += size;
|
||||
}
|
||||
|
||||
/* ---- rsync-style per-file --progress ------------------------------------
|
||||
* rsync prints, for each transferred regular file, the file name followed by a
|
||||
* two-frame progress line: the first at the initial 32 KiB read window (always
|
||||
* 0.00 kB/s / 0:00:00 on a sub-second transfer) and a final 100% frame carrying
|
||||
* `(xfr#N, to-chk=X/Y)`. Rates are wall-clock dependent, so only the final
|
||||
* rate is measured here; the layout matches rsync 3.4.1's progress.c. */
|
||||
#define RSYNC_PROGRESS_IO_WINDOW (32ULL * 1024ULL)
|
||||
|
||||
/* Paths-only pre-count of the source file list, built once at transfer start
|
||||
* when progress output or -i/--out-format needs it. rsync's `to-chk`
|
||||
* denominator is the whole file list -- every regular file, directory, symlink
|
||||
* and special plus the transfer root -- while the streaming scan only emits
|
||||
* empty directories. A metadata-only walk (no file reads, no hashing) supplies
|
||||
* that total and a metadata-bearing File for every directory, so --progress can
|
||||
* name them and -i/--out-format can itemize them without a second full scan. */
|
||||
/* One directory in the pre-count, keyed by its transfer-relative display name
|
||||
* ("" is the transfer root). `file` is owned by ProgressPrecount.dir_files and
|
||||
* carries the source metadata needed by -i/--out-format (%M/%B/%U/%G). */
|
||||
typedef struct {
|
||||
char* name; /* owned */
|
||||
File* file;
|
||||
} DirRef;
|
||||
|
||||
typedef struct {
|
||||
unsigned long long total;
|
||||
ArrayList* dir_paths; /* owned char* in transfer-relative display form */
|
||||
ArrayList* dir_files; /* owned File* captured during the metadata walk */
|
||||
ArrayList* dir_refs; /* owned DirRef*, sorted by name for prefix lookup */
|
||||
} ProgressPrecount;
|
||||
|
||||
static bool g_progress_active;
|
||||
/* True when -i/--out-format need the pre-counted directory entries fed into the
|
||||
* change-event stream (independent of --progress). */
|
||||
static bool g_change_dirs_active;
|
||||
static unsigned long long g_progress_xferred;
|
||||
static unsigned long long g_progress_index;
|
||||
static unsigned long long g_progress_total;
|
||||
static struct timespec g_progress_file_start;
|
||||
static ProgressPrecount g_progress_precount;
|
||||
static PathIndex g_progress_dir_index;
|
||||
static bool g_progress_dir_index_valid;
|
||||
static StrHashSet g_progress_emitted;
|
||||
static bool g_progress_emitted_valid;
|
||||
static ArrayList* g_progress_emitted_keys;
|
||||
|
||||
bool progress_requested(const Config* config) {
|
||||
return config != NULL && !config->quiet &&
|
||||
(config->show_progress || (config->info_level & LOG_INFO_PROGRESS) != 0);
|
||||
}
|
||||
|
||||
static void dir_ref_destroy(void* item) {
|
||||
DirRef* ref = (DirRef*)item;
|
||||
if (ref == NULL)
|
||||
return;
|
||||
free(ref->name);
|
||||
free(ref);
|
||||
}
|
||||
|
||||
/* Sort DirRef pointers by their transfer-relative name for binary search. */
|
||||
static int dir_ref_compare(const void* left, const void* right) {
|
||||
const DirRef* a = *(const DirRef* const*)left;
|
||||
const DirRef* b = *(const DirRef* const*)right;
|
||||
return strcmp(a->name, b->name);
|
||||
}
|
||||
|
||||
/* Look up the pre-counted directory File for a transfer-relative name ("" is
|
||||
* the transfer root). Returns NULL when no pre-count was built or the name is
|
||||
* not a known directory. */
|
||||
static File* progress_dir_lookup(const char* name) {
|
||||
if (name == NULL || g_progress_precount.dir_refs == NULL)
|
||||
return NULL;
|
||||
ArrayList* refs = g_progress_precount.dir_refs;
|
||||
size_t lo = 0;
|
||||
size_t hi = (size_t)refs->size;
|
||||
while (lo < hi) {
|
||||
size_t mid = lo + (hi - lo) / 2;
|
||||
DirRef* ref = (DirRef*)refs->items[mid];
|
||||
int cmp = strcmp(ref->name, name);
|
||||
if (cmp < 0)
|
||||
lo = mid + 1;
|
||||
else if (cmp > 0)
|
||||
hi = mid;
|
||||
else
|
||||
return ref->file;
|
||||
}
|
||||
return NULL;
|
||||
}
|
||||
|
||||
static void progress_precount_dispose(ProgressPrecount* p) {
|
||||
if (p->dir_paths != NULL) {
|
||||
array_list_delete(p->dir_paths);
|
||||
p->dir_paths = NULL;
|
||||
}
|
||||
if (p->dir_files != NULL) {
|
||||
array_list_delete(p->dir_files);
|
||||
p->dir_files = NULL;
|
||||
}
|
||||
if (p->dir_refs != NULL) {
|
||||
array_list_delete(p->dir_refs);
|
||||
p->dir_refs = NULL;
|
||||
}
|
||||
p->total = 0;
|
||||
}
|
||||
|
||||
/* Record the transfer root's pre-transfer state for -i/--out-format. The
|
||||
* receive root always exists, so rsync never marks it `cd`; its only observable
|
||||
* change is its timestamp, which FastSync cannot observe remotely. Force a time
|
||||
* mismatch so the root renders rsync's `.d..t...... ./` rather than the `cd`
|
||||
* a zeroed destination state would produce. */
|
||||
static void progress_precount_mark_root(File* root) {
|
||||
if (root == NULL)
|
||||
return;
|
||||
root->dest_state.known = true;
|
||||
root->dest_state.existed = true;
|
||||
root->dest_state.mode = root->metadata != NULL ? root->metadata->mode : 0;
|
||||
root->dest_state.uid = root->metadata != NULL ? root->metadata->uid : 0;
|
||||
root->dest_state.gid = root->metadata != NULL ? root->metadata->gid : 0;
|
||||
root->dest_state.size = 0;
|
||||
root->dest_state.mtime_sec = (root->metadata != NULL ? root->metadata->mtime_sec : 0) - 3600;
|
||||
root->dest_state.mtime_nsec = root->metadata != NULL ? root->metadata->mtime_nsec : 0;
|
||||
}
|
||||
|
||||
/* Append one DirRef (name -> file) to the pre-count, marking the transfer
|
||||
* root's destination state. Returns false on allocation failure. */
|
||||
static bool progress_precount_add_ref(ProgressPrecount* p, const Config* config, File* file) {
|
||||
const char* rel = delete_display_path(config, file_wire_path(file));
|
||||
char* name = rel != NULL ? str_dup(rel) : NULL;
|
||||
if (name == NULL)
|
||||
return false;
|
||||
DirRef* ref = malloc(sizeof(*ref));
|
||||
if (ref == NULL) {
|
||||
free(name);
|
||||
return false;
|
||||
}
|
||||
ref->name = name;
|
||||
ref->file = file;
|
||||
if (name[0] == '\0')
|
||||
progress_precount_mark_root(file);
|
||||
if (!array_list_add(p->dir_refs, ref)) {
|
||||
dir_ref_destroy(ref);
|
||||
return false;
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
void client_progress_cleanup(void) {
|
||||
if (g_progress_dir_index_valid) {
|
||||
path_index_free(&g_progress_dir_index);
|
||||
g_progress_dir_index_valid = false;
|
||||
}
|
||||
if (g_progress_emitted_valid) {
|
||||
str_hash_set_free(&g_progress_emitted);
|
||||
g_progress_emitted_valid = false;
|
||||
}
|
||||
if (g_progress_emitted_keys != NULL) {
|
||||
array_list_delete(g_progress_emitted_keys);
|
||||
g_progress_emitted_keys = NULL;
|
||||
}
|
||||
progress_precount_dispose(&g_progress_precount);
|
||||
g_progress_active = false;
|
||||
g_change_dirs_active = false;
|
||||
g_progress_total = 0;
|
||||
g_progress_index = 0;
|
||||
g_progress_xferred = 0;
|
||||
}
|
||||
|
||||
static void progress_first_frame(unsigned long long size, char* out, size_t out_size) {
|
||||
char ofs_buf[32];
|
||||
unsigned long long ofs = size < RSYNC_PROGRESS_IO_WINDOW ? size : RSYNC_PROGRESS_IO_WINDOW;
|
||||
if (!format_big_num(ofs, false, ofs_buf, sizeof(ofs_buf)))
|
||||
snprintf(ofs_buf, sizeof(ofs_buf), "%llu", ofs);
|
||||
int pct = size == 0 ? 100 : (ofs == size ? 100 : (int)(100.0 * (double)ofs / (double)size));
|
||||
snprintf(out, out_size, "\r%15s %3d%% %7.2f%s %s%s", ofs_buf, pct, 0.0, "kB/s", " 0:00:00",
|
||||
" ");
|
||||
}
|
||||
|
||||
static void progress_final_frame(unsigned long long size, char* out, size_t out_size) {
|
||||
char ofs_buf[32];
|
||||
char rembuf[32];
|
||||
unsigned long long last_ofs = size < RSYNC_PROGRESS_IO_WINDOW ? size : RSYNC_PROGRESS_IO_WINDOW;
|
||||
if (!format_big_num(size, false, ofs_buf, sizeof(ofs_buf)))
|
||||
snprintf(ofs_buf, sizeof(ofs_buf), "%llu", size);
|
||||
struct timespec now;
|
||||
clock_gettime(CLOCK_MONOTONIC, &now);
|
||||
long long diff_ms = (long long)(now.tv_sec - g_progress_file_start.tv_sec) * 1000 +
|
||||
(now.tv_nsec - g_progress_file_start.tv_nsec) / 1000000;
|
||||
if (diff_ms <= 0)
|
||||
diff_ms = 1;
|
||||
double rate =
|
||||
size > last_ofs ? (double)(size - last_ofs) * 1000.0 / (double)diff_ms / 1024.0 : 0.0;
|
||||
const char* units = "kB/s";
|
||||
if (rate > 1024.0 * 1024.0) {
|
||||
rate /= 1024.0 * 1024.0;
|
||||
units = "GB/s";
|
||||
} else if (rate > 1024.0) {
|
||||
rate /= 1024.0;
|
||||
units = "MB/s";
|
||||
}
|
||||
unsigned long long remain = (unsigned long long)(diff_ms / 1000);
|
||||
snprintf(rembuf, sizeof(rembuf), "%4u:%02u:%02u", (unsigned)(remain / 3600),
|
||||
(unsigned)((remain / 60) % 60), (unsigned)(remain % 60));
|
||||
/* rsync's `to-chk` denominator is the whole file list (the pre-count); the
|
||||
numerator falls as each entry is processed, root first. Without a
|
||||
pre-count (the paths-only walk failed) fall back to the transferred-file
|
||||
count so the single-file layout stays intact. */
|
||||
unsigned long long total = g_progress_total > 0 ? g_progress_total : g_progress_xferred + 1;
|
||||
unsigned long long to_chk = total > g_progress_index ? total - g_progress_index - 1 : 0;
|
||||
snprintf(out, out_size, "\r%15s %3d%% %7.2f%s %s (xfr#%llu, to-chk=%llu/%llu)\n", ofs_buf, 100,
|
||||
rate, units, rembuf, g_progress_xferred, to_chk, total);
|
||||
}
|
||||
|
||||
bool info_flag_enabled(const Config* config, LogInfoFlag flag) {
|
||||
return config != NULL && (config->info_level & flag) != 0;
|
||||
}
|
||||
|
||||
/* Print rsync's deletion lines for a received list of destination-relative
|
||||
* paths: `*deleting PATH` when itemizing, the --out-format expansion when a
|
||||
* format is set, else `deleting PATH` for --info=del. Used by both the dry-run
|
||||
* would-delete report and the real --info=del report. */
|
||||
void print_delete_reports(const Config* config, const ArrayList* paths) {
|
||||
if (!config || !paths || config->quiet)
|
||||
return;
|
||||
/* --debug=del is independent of the --info=del/itemize/out-format display:
|
||||
emit the debug trace even when no deletion line would be printed. */
|
||||
if (log_debug_enabled(LOG_DEBUG_DEL)) {
|
||||
for (int i = 0; i < paths->size; i++) {
|
||||
const char* raw = (const char*)paths->items[i];
|
||||
const char* path = delete_display_path(config, raw);
|
||||
log_debug_message(LOG_DEBUG_DEL, "del: %s", path ? path : raw);
|
||||
}
|
||||
}
|
||||
if (!(config->itemize_changes || config->out_format != NULL ||
|
||||
info_flag_enabled(config, LOG_INFO_DEL)))
|
||||
return;
|
||||
for (int i = 0; i < paths->size; i++) {
|
||||
const char* raw = (const char*)paths->items[i];
|
||||
const char* path = delete_display_path(config, raw);
|
||||
if (config->out_format != NULL) {
|
||||
ChangeEvent event;
|
||||
memset(&event, 0, sizeof(event));
|
||||
event.decision = CHANGE_SENT;
|
||||
event.deleted = true;
|
||||
event.name = path;
|
||||
event.path = path;
|
||||
char* line = change_render_format(config->out_format, config, &event);
|
||||
if (line) {
|
||||
char* escaped = output_escape(line, config->eight_bit_output);
|
||||
printf("%s\n", escaped ? escaped : line);
|
||||
free(escaped);
|
||||
free(line);
|
||||
}
|
||||
} else {
|
||||
char* escaped = output_escape(path, config->eight_bit_output);
|
||||
if (config->itemize_changes)
|
||||
printf("*deleting %s\n", escaped ? escaped : path);
|
||||
else
|
||||
printf("deleting %s\n", escaped ? escaped : path);
|
||||
free(escaped);
|
||||
}
|
||||
}
|
||||
fflush(stdout);
|
||||
}
|
||||
|
||||
/* Emit every not-yet-seen ancestor directory of `rel`, outermost first, in the
|
||||
* order rsync's depth-first flist walk visits them. With -i/--out-format each
|
||||
* ancestor becomes a real change line (`cd+++++++++ sub/`, `.d..t...... ./`)
|
||||
* rendered by the shared itemize code; otherwise it is the `--info=name` /
|
||||
* --progress directory name line. */
|
||||
static void client_progress_emit_ancestors(const Config* config, const char* rel) {
|
||||
if (!g_progress_dir_index_valid || !g_progress_emitted_valid || g_progress_emitted_keys == NULL ||
|
||||
rel == NULL)
|
||||
return;
|
||||
size_t rel_len = strlen(rel);
|
||||
for (size_t i = 0; i < rel_len; i++) {
|
||||
if (rel[i] != '/')
|
||||
continue;
|
||||
char* prefix = malloc(i + 1);
|
||||
if (prefix == NULL)
|
||||
return;
|
||||
memcpy(prefix, rel, i);
|
||||
prefix[i] = '\0';
|
||||
if (path_index_contains(&g_progress_dir_index, prefix) &&
|
||||
!str_hash_set_lookup(&g_progress_emitted, prefix)) {
|
||||
char* key = str_dup(prefix);
|
||||
if (key != NULL && array_list_add(g_progress_emitted_keys, key)) {
|
||||
str_hash_set_insert_ref(&g_progress_emitted, key);
|
||||
if (g_change_dirs_active) {
|
||||
File* dir = progress_dir_lookup(prefix);
|
||||
if (dir != NULL)
|
||||
change_emit_dir_sent(config, dir);
|
||||
} else {
|
||||
char* escaped = output_escape(prefix, config->eight_bit_output);
|
||||
printf("%s/\n", escaped ? escaped : prefix);
|
||||
free(escaped);
|
||||
}
|
||||
g_progress_index++;
|
||||
} else {
|
||||
free(key);
|
||||
}
|
||||
}
|
||||
free(prefix);
|
||||
}
|
||||
}
|
||||
|
||||
/* Feed a transferred entry's ancestor directories into the change-event stream
|
||||
* before the entry's own line, so -i/--out-format and --progress report
|
||||
* directories in rsync's depth-first order. Every directory is an ancestor of
|
||||
* some emitted entry (a file, symlink, special, hard link or the empty-directory
|
||||
* entry the scanner emits for a leaf), so this covers the whole tree. */
|
||||
void client_change_emit_ancestors(const Config* config, const File* file) {
|
||||
if (config == NULL || file == NULL)
|
||||
return;
|
||||
if (!g_progress_active && !g_change_dirs_active)
|
||||
return;
|
||||
const char* rel = delete_display_path(config, file_wire_path(file));
|
||||
client_progress_emit_ancestors(config, rel);
|
||||
}
|
||||
|
||||
/* rsync's --info=name/progress line for one entry: transfer-relative name (a
|
||||
* trailing slash for directories) plus the ` -> target` symlink suffix. */
|
||||
static char* progress_entry_line(const File* file, const char* rel) {
|
||||
const char* arrow = NULL;
|
||||
const char* target = NULL;
|
||||
if (file->is_symlink && file->symlink_target != NULL) {
|
||||
arrow = " -> ";
|
||||
target = file->symlink_target;
|
||||
} else if (file->link_group != 0 && !file->link_first && file->hardlink_target != NULL) {
|
||||
arrow = " => ";
|
||||
target = file->hardlink_target;
|
||||
}
|
||||
size_t rel_len = strlen(rel);
|
||||
bool dir_slash = file->is_dir && (rel_len == 0 || rel[rel_len - 1] != '/');
|
||||
size_t extra = (dir_slash ? 1u : 0u) + (target != NULL ? 4u + strlen(target) : 0u);
|
||||
char* line = malloc(rel_len + extra + 1);
|
||||
if (line == NULL)
|
||||
return NULL;
|
||||
memcpy(line, rel, rel_len);
|
||||
size_t off = rel_len;
|
||||
if (dir_slash)
|
||||
line[off++] = '/';
|
||||
if (target != NULL) {
|
||||
memcpy(line + off, arrow, 4);
|
||||
off += 4;
|
||||
memcpy(line + off, target, strlen(target));
|
||||
off += strlen(target);
|
||||
}
|
||||
line[off] = '\0';
|
||||
return line;
|
||||
}
|
||||
|
||||
void client_progress_begin(const Config* config) {
|
||||
change_reset_name_root();
|
||||
g_progress_active = progress_requested(config);
|
||||
g_progress_xferred = 0;
|
||||
g_progress_index = 1; /* the transfer root is file-list entry #0 */
|
||||
if (!g_progress_active && !g_change_dirs_active) {
|
||||
/* `--info=flist` prints rsync's file-list header even without progress. */
|
||||
if (!config->quiet && info_flag_enabled(config, LOG_INFO_FLIST)) {
|
||||
printf("sending incremental file list\n");
|
||||
fflush(stdout);
|
||||
}
|
||||
return;
|
||||
}
|
||||
if (g_progress_active)
|
||||
printf("sending incremental file list\n");
|
||||
/* rsync prints the transfer-root directory before the first entry. Under
|
||||
-i/--out-format it is the root change line (`.d..t...... ./`); otherwise it
|
||||
is the plain --info=name / --progress name line. */
|
||||
if (g_change_dirs_active) {
|
||||
File* root = progress_dir_lookup("");
|
||||
if (root != NULL)
|
||||
change_emit_dir_sent(config, root);
|
||||
} else {
|
||||
printf("./\n");
|
||||
}
|
||||
fflush(stdout);
|
||||
}
|
||||
|
||||
/* Emit the name (unless itemize/out-format already did) and the two progress
|
||||
* frames for one transferred regular file. */
|
||||
void client_progress_file(const Config* config, const File* file) {
|
||||
if (!g_progress_active || file == NULL || !file->data)
|
||||
return;
|
||||
g_progress_xferred++;
|
||||
unsigned long long size = file->data->size;
|
||||
if (!config->itemize_changes && config->out_format == NULL) {
|
||||
const char* rel = delete_display_path(config, file_wire_path(file));
|
||||
char* escaped = output_escape(rel, config->eight_bit_output);
|
||||
printf("%s\n", escaped ? escaped : (rel ? rel : ""));
|
||||
free(escaped);
|
||||
}
|
||||
clock_gettime(CLOCK_MONOTONIC, &g_progress_file_start);
|
||||
char frame[160];
|
||||
progress_first_frame(size, frame, sizeof(frame));
|
||||
fputs(frame, stdout);
|
||||
progress_final_frame(size, frame, sizeof(frame));
|
||||
fputs(frame, stdout);
|
||||
g_progress_index++;
|
||||
fflush(stdout);
|
||||
}
|
||||
|
||||
/* Emit the name line for a transferred non-regular entry (directory, symlink,
|
||||
* special or hard-link sibling): rsync prints these in the file list but has no
|
||||
* progress frame for them. */
|
||||
void client_progress_name(const Config* config, const File* file) {
|
||||
if (!g_progress_active || file == NULL)
|
||||
return;
|
||||
const char* rel = delete_display_path(config, file_wire_path(file));
|
||||
if (!config->itemize_changes && config->out_format == NULL) {
|
||||
char* line = progress_entry_line(file, rel ? rel : "");
|
||||
if (line != NULL) {
|
||||
char* escaped = output_escape(line, config->eight_bit_output);
|
||||
printf("%s\n", escaped ? escaped : line);
|
||||
free(escaped);
|
||||
free(line);
|
||||
fflush(stdout);
|
||||
}
|
||||
}
|
||||
g_progress_index++;
|
||||
}
|
||||
|
||||
/* An entry the receiver already had prints no name under --progress but still
|
||||
* occupies a file-list slot in the `to-chk` numerator. */
|
||||
void client_progress_uptodate(const Config* config, const File* file) {
|
||||
(void)config;
|
||||
(void)file;
|
||||
if (!g_progress_active)
|
||||
return;
|
||||
g_progress_index++;
|
||||
}
|
||||
|
||||
static bool progress_precount_add_dir(ProgressPrecount* p, const char* path) {
|
||||
if (path == NULL || path[0] == '\0')
|
||||
return true;
|
||||
char* dup = str_dup(path);
|
||||
if (dup == NULL)
|
||||
return false;
|
||||
if (array_list_add(p->dir_paths, dup))
|
||||
return true;
|
||||
free(dup);
|
||||
return false;
|
||||
}
|
||||
|
||||
/* Metadata-only walk collecting the full file-list total and every directory
|
||||
* name. It uses its own scanner (fresh filter compilation and hard-link table)
|
||||
* so the data pass's link-group state is never perturbed. */
|
||||
static bool progress_precount_scan(const Config* config, ProgressPrecount* out) {
|
||||
out->dir_paths = array_list_create(free);
|
||||
out->dir_files = array_list_create(file_destroy);
|
||||
out->dir_refs = array_list_create(dir_ref_destroy);
|
||||
if (out->dir_paths == NULL || out->dir_files == NULL || out->dir_refs == NULL) {
|
||||
progress_precount_dispose(out);
|
||||
return false;
|
||||
}
|
||||
out->total = 0;
|
||||
PreparedScanner prepared;
|
||||
memset(&prepared, 0, sizeof(prepared));
|
||||
if (!prepare_scanner(config, 0, &prepared)) {
|
||||
progress_precount_dispose(out);
|
||||
return false;
|
||||
}
|
||||
ScannerOptions local = prepared.options;
|
||||
local.list_dirs = true;
|
||||
local.note_nonreg = false;
|
||||
local.note_mount = false;
|
||||
local.dir_count = NULL;
|
||||
local.use_metadata = false;
|
||||
local.preserve_xattrs = false;
|
||||
local.preserve_acls = false;
|
||||
local.checksum = false;
|
||||
/* Capture one metadata-bearing File per traversed directory (including the
|
||||
transfer root) so -i/--out-format can render %M/%B/%U/%G for directories. */
|
||||
local.capture_dir_times = true;
|
||||
local.excluded_paths = NULL;
|
||||
local.size_skipped_paths = NULL;
|
||||
local.synced_dirs = NULL;
|
||||
local.plan_dirs = NULL;
|
||||
local.dir_entries = out->dir_files;
|
||||
local.dir_entries_mutex = NULL;
|
||||
local.hardlinks = NULL;
|
||||
DirectoryScanner* scanner = directory_scanner_create_with_options(config->send_directory, &local);
|
||||
bool ok = scanner != NULL;
|
||||
if (scanner != NULL) {
|
||||
Chunk* chunk;
|
||||
while (ok && (chunk = directory_scanner_next(scanner)) != NULL) {
|
||||
out->total += (unsigned long long)chunk->element_count;
|
||||
for (int i = 0; i < chunk->element_count && ok; i++) {
|
||||
const File* f = chunk->items[i];
|
||||
if (f != NULL && f->is_dir)
|
||||
ok = progress_precount_add_dir(out, delete_display_path(config, file_wire_path(f)));
|
||||
}
|
||||
chunk_destroy(chunk);
|
||||
}
|
||||
if (ok && directory_scanner_failed(scanner))
|
||||
ok = false;
|
||||
directory_scanner_destroy(scanner);
|
||||
}
|
||||
prepared_scanner_destroy(&prepared);
|
||||
if (!ok) {
|
||||
progress_precount_dispose(out);
|
||||
return false;
|
||||
}
|
||||
/* Build the name -> File lookup from the captured directory Files. */
|
||||
for (int i = 0; i < out->dir_files->size; i++) {
|
||||
File* f = (File*)out->dir_files->items[i];
|
||||
if (f == NULL)
|
||||
continue;
|
||||
if (!progress_precount_add_ref(out, config, f)) {
|
||||
progress_precount_dispose(out);
|
||||
return false;
|
||||
}
|
||||
}
|
||||
if (out->dir_refs->size > 1)
|
||||
qsort(out->dir_refs->items, (size_t)out->dir_refs->size, sizeof(DirRef*), dir_ref_compare);
|
||||
out->total += 1; /* the transfer root "." */
|
||||
return true;
|
||||
}
|
||||
|
||||
/* Reuse the --delete-during/--delete-delay keep-set pre-scan: its traversed
|
||||
* directory list already holds every directory and `non_dir_count` the entries
|
||||
* counted during that same pass, so progress costs no second walk. */
|
||||
static bool progress_precount_from_plan_dirs(const Config* config, const ArrayList* plan_dirs,
|
||||
unsigned long long non_dir_count,
|
||||
ProgressPrecount* out) {
|
||||
out->dir_paths = array_list_create(free);
|
||||
out->dir_files = array_list_create(file_destroy);
|
||||
out->dir_refs = array_list_create(dir_ref_destroy);
|
||||
if (out->dir_paths == NULL || out->dir_files == NULL || out->dir_refs == NULL) {
|
||||
progress_precount_dispose(out);
|
||||
return false;
|
||||
}
|
||||
out->total = non_dir_count + 1;
|
||||
/* The delete pre-scan's plan list omits the transfer root, so synthesize its
|
||||
entry here; it is only used for the root change line. */
|
||||
File* root = file_create("");
|
||||
if (root == NULL || !array_list_add(out->dir_files, root)) {
|
||||
file_destroy(root);
|
||||
progress_precount_dispose(out);
|
||||
return false;
|
||||
}
|
||||
root->is_dir = true;
|
||||
if (!progress_precount_add_ref(out, config, root)) {
|
||||
progress_precount_dispose(out);
|
||||
return false;
|
||||
}
|
||||
for (int i = 0; i < plan_dirs->size; i++) {
|
||||
const char* path = (const char*)plan_dirs->items[i];
|
||||
const char* rel = config->send_directory != NULL
|
||||
? utils_strip_transfer_root(path, config->send_directory)
|
||||
: path;
|
||||
if (!progress_precount_add_dir(out, rel)) {
|
||||
progress_precount_dispose(out);
|
||||
return false;
|
||||
}
|
||||
File* dir = file_create("");
|
||||
if (dir == NULL) {
|
||||
progress_precount_dispose(out);
|
||||
return false;
|
||||
}
|
||||
dir->is_dir = true;
|
||||
dir->send_path = str_dup(rel != NULL ? rel : "");
|
||||
if (dir->send_path == NULL || !array_list_add(out->dir_files, dir)) {
|
||||
file_destroy(dir);
|
||||
progress_precount_dispose(out);
|
||||
return false;
|
||||
}
|
||||
if (!progress_precount_add_ref(out, config, dir)) {
|
||||
progress_precount_dispose(out);
|
||||
return false;
|
||||
}
|
||||
}
|
||||
if (out->dir_refs->size > 1)
|
||||
qsort(out->dir_refs->items, (size_t)out->dir_refs->size, sizeof(DirRef*), dir_ref_compare);
|
||||
out->total += (unsigned long long)out->dir_paths->size;
|
||||
return true;
|
||||
}
|
||||
|
||||
/* Build the optional progress pre-count. A failed pre-count is non-fatal: the
|
||||
* transfer proceeds and the progress denominator falls back to the transferred
|
||||
* file count. */
|
||||
void client_progress_prepare(const Config* config, const ArrayList* plan_dirs,
|
||||
unsigned long long plan_non_dir_count) {
|
||||
client_progress_cleanup();
|
||||
g_progress_active = progress_requested(config);
|
||||
g_change_dirs_active = config->itemize_changes || config->out_format != NULL;
|
||||
if (!g_progress_active && !g_change_dirs_active)
|
||||
return;
|
||||
bool ok = plan_dirs != NULL ? progress_precount_from_plan_dirs(
|
||||
config, plan_dirs, plan_non_dir_count, &g_progress_precount)
|
||||
: progress_precount_scan(config, &g_progress_precount);
|
||||
if (!ok) {
|
||||
g_progress_total = 0;
|
||||
return;
|
||||
}
|
||||
g_progress_total = g_progress_precount.total;
|
||||
if (g_progress_precount.dir_paths != NULL && g_progress_precount.dir_paths->size > 0 &&
|
||||
path_index_build(&g_progress_dir_index,
|
||||
(const char* const*)g_progress_precount.dir_paths->items,
|
||||
(size_t)g_progress_precount.dir_paths->size))
|
||||
g_progress_dir_index_valid = true;
|
||||
if (str_hash_set_init(&g_progress_emitted, (size_t)(g_progress_precount.dir_paths != NULL
|
||||
? g_progress_precount.dir_paths->size + 1
|
||||
: 1)))
|
||||
g_progress_emitted_valid = true;
|
||||
g_progress_emitted_keys = array_list_create(free);
|
||||
}
|
||||
|
||||
/* Read the optional STATUS_STATS record (protocol 2.25.0) that the receiver
|
||||
* sends just before its terminal status when report_stats was negotiated.
|
||||
* Consumes the would-delete path list into `would_delete` (optional). */
|
||||
bool receive_stats_record(int fd, ReceiverStats* stats, ArrayList* would_delete) {
|
||||
if (!format_stats_receive(fd, stats))
|
||||
return false;
|
||||
int count = 0;
|
||||
if (!receive_int(fd, &count) || count < 0 || count > MAX_MANIFEST_ENTRIES)
|
||||
return false;
|
||||
/* Mirror the delete-plan parser: every retained path must be a valid
|
||||
destination-relative path, and the whole list shares one MAX_MANIFEST_BYTES
|
||||
budget so a hostile peer cannot make the client retain unbounded memory. */
|
||||
size_t bytes = 0;
|
||||
for (int i = 0; i < count; i++) {
|
||||
char* path = receive_wire_str(fd);
|
||||
if (!path)
|
||||
return false;
|
||||
if (path[0] == '\0' || path[0] == '/' || has_path_traversal(path)) {
|
||||
free(path);
|
||||
return false;
|
||||
}
|
||||
if (would_delete) {
|
||||
size_t entry_size = strlen(path) + sizeof(char*) + 16;
|
||||
if (entry_size > MAX_MANIFEST_BYTES - bytes) {
|
||||
free(path);
|
||||
return false;
|
||||
}
|
||||
bytes += entry_size;
|
||||
if (!array_list_add(would_delete, path)) {
|
||||
free(path);
|
||||
return false;
|
||||
}
|
||||
} else {
|
||||
free(path);
|
||||
}
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
/* Strip the transfer-root prefix from a receiver-reported destination-relative
|
||||
* delete path so a `*deleting` line matches rsync's transfer-relative name
|
||||
* (FastSync's destination mirror includes the source's absolute path). */
|
||||
const char* delete_display_path(const Config* config, const char* path) {
|
||||
if (!config || !path || !config->send_directory)
|
||||
return path;
|
||||
return utils_strip_transfer_root(path, config->send_directory);
|
||||
}
|
||||
@@ -0,0 +1,454 @@
|
||||
#include "client_send_internal.h"
|
||||
#include "array_list.h"
|
||||
#include "charset.h"
|
||||
#include "config.h"
|
||||
#include "delete_plan.h"
|
||||
#include "file.h"
|
||||
#include "file_list.h"
|
||||
#include "filter.h"
|
||||
#include "hardlink.h"
|
||||
#include "log.h"
|
||||
#include "scanner.h"
|
||||
#include "utils.h"
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <sys/stat.h>
|
||||
|
||||
/* Build the scanner options for one scan. Returns false and logs on failure. */
|
||||
bool prepare_scanner(const Config* config, int num_threads, PreparedScanner* out) {
|
||||
if (!out)
|
||||
return false;
|
||||
out->base_filters = NULL;
|
||||
out->hardlinks = NULL;
|
||||
out->relative_prefix = NULL;
|
||||
memset(&out->options, 0, sizeof(out->options));
|
||||
|
||||
int rule_count = config->filters ? config->filters->size : 0;
|
||||
const char** texts = NULL;
|
||||
if (rule_count > 0) {
|
||||
texts = malloc((size_t)rule_count * sizeof(char*));
|
||||
if (!texts) {
|
||||
log_message(LOG_LEVEL_ERROR, "memory allocation failed for filter rules");
|
||||
return false;
|
||||
}
|
||||
for (int i = 0; i < rule_count; i++)
|
||||
texts[i] = (const char*)config->filters->items[i];
|
||||
}
|
||||
if (rule_count > 0 || config->cvs_exclude) {
|
||||
char err[160];
|
||||
out->base_filters = filter_base_build(texts, rule_count, config->cvs_exclude,
|
||||
config->delete_excluded, err, sizeof(err));
|
||||
free(texts);
|
||||
if (!out->base_filters) {
|
||||
log_message(LOG_LEVEL_ERROR, "invalid filter rule: %s", err);
|
||||
return false;
|
||||
}
|
||||
} else {
|
||||
free(texts);
|
||||
}
|
||||
|
||||
ScannerOptions* options = &out->options;
|
||||
options->use_metadata = config->use_metadata;
|
||||
options->preserve_atimes = config->preserve_atimes;
|
||||
options->preserve_crtimes = config->preserve_crtimes;
|
||||
options->preserve_xattrs = config->preserve_xattrs;
|
||||
options->preserve_acls = config->preserve_acls;
|
||||
options->chunk_size = config->chunk_size;
|
||||
/* --exclude/--include are compiled, in command-line order, into the SAME
|
||||
* ordered filter rule list as --filter/-f (see config_add_selection_rule), so
|
||||
* the legacy per-kind arrays are deliberately NOT passed to the scanner:
|
||||
* doing so would re-apply them with the old "excludes first, then includes as
|
||||
* a mandatory whitelist" precedence and defeat rsync's first-match-wins
|
||||
* ordering. The arrays remain populated purely for the Config API surface. */
|
||||
options->exclude_patterns = NULL;
|
||||
options->exclude_count = 0;
|
||||
options->include_patterns = NULL;
|
||||
options->include_count = 0;
|
||||
options->max_size = config->max_size;
|
||||
options->min_size = config->min_size;
|
||||
options->max_depth = config->max_depth;
|
||||
options->num_threads = num_threads;
|
||||
options->follow_symlinks = config->follow_symlinks;
|
||||
options->copy_links = config->copy_links;
|
||||
options->safe_links = config->safe_links;
|
||||
options->copy_unsafe_links = config->copy_unsafe_links;
|
||||
options->copy_dirlinks = config->copy_dirlinks;
|
||||
options->munge_links = config->munge_links;
|
||||
options->checksum = config->checksum;
|
||||
options->one_file_system = config->one_file_system;
|
||||
options->preserve_devices = config->preserve_devices;
|
||||
options->preserve_specials = config->preserve_specials;
|
||||
options->copy_devices = config->copy_devices;
|
||||
options->file_list = (const FileListSet*)config->files_from_set;
|
||||
options->base_filters = out->base_filters;
|
||||
options->per_dir_filters = config->per_dir_filter;
|
||||
options->delete_excluded = config->delete_excluded;
|
||||
options->exclude_per_dir_filter_files = config->per_dir_filter_count >= 2;
|
||||
options->dirs = config->dirs;
|
||||
options->relative = config->relative;
|
||||
/* A real recursive transfer recreates empty source directories (rsync
|
||||
parity); low-level scanner users leave this off. */
|
||||
options->emit_empty_dirs = true;
|
||||
/* --no-implied-dirs only has meaning with -R (rsync): without it the option
|
||||
is a documented no-op, so the scanner must not suppress directory
|
||||
metadata. */
|
||||
options->no_implied_dirs = config->no_implied_dirs && config->relative;
|
||||
/* -R/--relative outside --files-from reconstructs every destination path from
|
||||
* the source spec (rsync's '/./' cut point). With --files-from the listed
|
||||
* entry already supplies the bare relative path, so no prefix is built. */
|
||||
if (config->relative && config->files_from_set == NULL && config->send_directory) {
|
||||
out->relative_prefix = scanner_relative_prefix(config->send_directory);
|
||||
if (!out->relative_prefix) {
|
||||
log_message(LOG_LEVEL_ERROR, "memory allocation failed building --relative path prefix");
|
||||
filter_rule_list_free(out->base_filters);
|
||||
out->base_filters = NULL;
|
||||
return false;
|
||||
}
|
||||
options->relative_prefix = out->relative_prefix;
|
||||
}
|
||||
options->prune_empty_dirs = config->prune_empty_dirs;
|
||||
options->ignore_io_errors = config->ignore_errors;
|
||||
options->ignore_missing_args = config->ignore_missing_args || config->delete_missing_args;
|
||||
options->note_nonreg = (config->info_level & LOG_INFO_NONREG) != 0 && !config->quiet;
|
||||
options->note_mount = (config->info_level & LOG_INFO_MOUNT) != 0 && !config->quiet;
|
||||
options->send_directory = config->send_directory;
|
||||
options->eight_bit_output = config->eight_bit_output;
|
||||
options->excluded_paths = NULL;
|
||||
options->excluded_mutex = NULL;
|
||||
options->size_skipped_paths = NULL;
|
||||
options->synced_dirs = NULL;
|
||||
options->hardlinks = NULL;
|
||||
/* Set by the real send paths; NULL for the metadata-only scans (progress
|
||||
pre-count, batch) that must not perturb the sender's --stats counter. */
|
||||
options->dir_count = NULL;
|
||||
/* P7 Wave D: capture source directory metadata when a directory attribute is
|
||||
requested (-p for modes, -t for times unless -O omits them). Whether they
|
||||
are APPLIED is decided receiver-side. */
|
||||
options->capture_dir_times = dir_metadata_should_capture(config);
|
||||
options->dir_entries = NULL;
|
||||
options->dir_entries_mutex = NULL;
|
||||
if (config->preserve_hard_links) {
|
||||
out->hardlinks = hardlink_table_create();
|
||||
if (!out->hardlinks) {
|
||||
filter_rule_list_free(out->base_filters);
|
||||
out->base_filters = NULL;
|
||||
return false;
|
||||
}
|
||||
options->hardlinks = out->hardlinks;
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
void prepared_scanner_destroy(PreparedScanner* prepared) {
|
||||
if (!prepared)
|
||||
return;
|
||||
filter_rule_list_free(prepared->base_filters);
|
||||
prepared->base_filters = NULL;
|
||||
hardlink_table_destroy(prepared->hardlinks);
|
||||
prepared->hardlinks = NULL;
|
||||
free(prepared->relative_prefix);
|
||||
prepared->relative_prefix = NULL;
|
||||
}
|
||||
|
||||
/* -R/--relative implied directories: rsync transmits the metadata of the
|
||||
* parent directories implied by the source path (every prefix component above
|
||||
* the source root) so the receiver applies their attributes to the created
|
||||
* parents. FastSync's scan only covers the source root and below, so append
|
||||
* one metadata-only directory entry per implied ancestor. --no-implied-dirs
|
||||
* suppresses this exactly like rsync. A missing ancestor is never fatal. */
|
||||
bool append_implied_dir_times(const Config* config, ArrayList* dir_entries) {
|
||||
if (!dir_entries || !config->relative || config->files_from_set != NULL ||
|
||||
config->no_implied_dirs || !config->send_directory)
|
||||
return true;
|
||||
char* prefix = scanner_relative_prefix(config->send_directory);
|
||||
if (!prefix)
|
||||
return true;
|
||||
int ncomp = 0;
|
||||
for (const char* s = prefix; *s;) {
|
||||
while (*s == '/')
|
||||
s++;
|
||||
if (!*s)
|
||||
break;
|
||||
while (*s && *s != '/')
|
||||
s++;
|
||||
ncomp++;
|
||||
}
|
||||
if (ncomp <= 1) {
|
||||
free(prefix);
|
||||
return true;
|
||||
}
|
||||
char* fs = str_dup(config->send_directory);
|
||||
if (!fs) {
|
||||
free(prefix);
|
||||
return true;
|
||||
}
|
||||
size_t flen = strlen(fs);
|
||||
while (flen > 1 && fs[flen - 1] == '/')
|
||||
fs[--flen] = '\0';
|
||||
bool ok = true;
|
||||
/* Walk the source path upwards one component at a time (fs is truncated in
|
||||
place, so each step targets the next implied ancestor). */
|
||||
for (int depth = ncomp - 2; depth >= 0 && ok; depth--) {
|
||||
char* slash = strrchr(fs, '/');
|
||||
if (!slash || slash == fs)
|
||||
break;
|
||||
*slash = '\0';
|
||||
char* p = prefix;
|
||||
int c = 0;
|
||||
while (c <= depth) {
|
||||
while (*p == '/')
|
||||
p++;
|
||||
while (*p && *p != '/')
|
||||
p++;
|
||||
c++;
|
||||
}
|
||||
char saved = *p;
|
||||
*p = '\0';
|
||||
struct stat st;
|
||||
if (stat(fs, &st) == 0 && S_ISDIR(st.st_mode)) {
|
||||
File* file = file_create(fs);
|
||||
if (!file) {
|
||||
ok = false;
|
||||
} else {
|
||||
file->is_dir = true;
|
||||
file->metadata =
|
||||
file_metadata_create(fs, &st, config->preserve_atimes, config->preserve_crtimes);
|
||||
file->send_path = str_dup(prefix);
|
||||
if (!file->metadata || !file->send_path || !array_list_add(dir_entries, file)) {
|
||||
file_destroy(file);
|
||||
ok = false;
|
||||
}
|
||||
}
|
||||
}
|
||||
*p = saved;
|
||||
}
|
||||
free(fs);
|
||||
free(prefix);
|
||||
return ok;
|
||||
}
|
||||
|
||||
/* The delete-walk root scope for a full (non---files-from) transfer: rsync
|
||||
* confines --delete to the directories it actually transferred. A plain
|
||||
* recursive run mirrors the source under the receive root, so "." (the whole
|
||||
* tree) is correct; an -R run transfers only the reconstructed prefix subtree,
|
||||
* so the walk is scoped to that prefix instead. Returns a malloc'd wire path
|
||||
* (or "."), or NULL on allocation failure. */
|
||||
char* delete_scope_root_marker(const Config* config) {
|
||||
if (config->relative && config->files_from_set == NULL && config->send_directory) {
|
||||
char* prefix = scanner_relative_prefix(config->send_directory);
|
||||
if (!prefix)
|
||||
return NULL;
|
||||
if (prefix[0] != '\0')
|
||||
return prefix;
|
||||
free(prefix);
|
||||
}
|
||||
return str_dup(".");
|
||||
}
|
||||
|
||||
/* The -R destination prefix that confines a per-directory delete walk, or NULL
|
||||
* when the whole receive root is in scope. The marker was installed into
|
||||
* `synced_dirs` by delete_scope_root_marker(); for a plain recursive transfer
|
||||
* it is "." (whole root) and for --files-from the list is not a single prefix. */
|
||||
const char* delete_plan_walk_root(const Config* config, const ArrayList* synced_dirs) {
|
||||
if (!config || config->files_from_set != NULL || !config->relative || !config->send_directory)
|
||||
return NULL;
|
||||
if (!synced_dirs || synced_dirs->size != 1)
|
||||
return NULL;
|
||||
const char* marker = (const char*)synced_dirs->items[0];
|
||||
if (marker[0] == '\0' || strcmp(marker, ".") == 0)
|
||||
return NULL;
|
||||
return marker;
|
||||
}
|
||||
|
||||
/* The destination-relative mirror path for a missing --files-from entry: where
|
||||
a PRESENT entry with the same name would have been written. With -R that is
|
||||
the entry's bare relative path (the bare wire path the receiver uses);
|
||||
otherwise it is the full source mirror below the destination root
|
||||
(`send_directory` joined to the entry, leading '/' stripped), exactly the
|
||||
path the manifest records for a present sibling. Returns an owned string, or
|
||||
NULL on allocation failure. */
|
||||
static char* files_from_missing_dest_path(const Config* config, const char* entry) {
|
||||
if (config->relative)
|
||||
return str_dup(entry);
|
||||
char* joined = path_cat(config->send_directory, entry);
|
||||
if (!joined)
|
||||
return NULL;
|
||||
const char* rel = *joined == '/' ? joined + 1 : joined;
|
||||
char* dup = str_dup(rel);
|
||||
free(joined);
|
||||
return dup;
|
||||
}
|
||||
|
||||
/* --files-from semantics: every listed entry must resolve under the source
|
||||
* root, otherwise rsync reports a hard error instead of silently transferring
|
||||
* nothing. An entry of "." (the whole tree) and listed-but-empty directories
|
||||
* are valid. An empty list is valid too: rsync transfers nothing and exits 0.
|
||||
* With --ignore-missing-args
|
||||
* (implied by --delete-missing-args) a listed-but-missing entry is instead
|
||||
* skipped: nothing is transferred for it, it never enters the keep-set and the
|
||||
* run succeeds for the rest (an all-missing non-empty list succeeds
|
||||
* transferring nothing, matching rsync). With --delete-missing-args
|
||||
* `missing_dest` (when non-NULL) collects the entry's destination-relative
|
||||
* mirror for the receiver's exact-deletion request. Runs before any
|
||||
* transfer so the failure/skip is surfaced uniformly in the single-threaded,
|
||||
* -m, dry-run and --list-only paths. */
|
||||
bool files_from_list_check(const Config* config, ArrayList* missing_dest, int* skipped_out) {
|
||||
*skipped_out = 0;
|
||||
const FileListSet* set = (const FileListSet*)config->files_from_set;
|
||||
if (!set)
|
||||
return true;
|
||||
if (!config->send_directory) {
|
||||
log_message(LOG_LEVEL_ERROR, "--files-from requires a source directory");
|
||||
return false;
|
||||
}
|
||||
if (set->count == 0) {
|
||||
/* rsync treats an empty --files-from list as "nothing to transfer" and
|
||||
exits 0 (the source directory is still a valid source arg), so this is
|
||||
not an error. Nothing passes the (empty) allow-set, so no file is sent
|
||||
and no keep-set entry is produced. */
|
||||
return true;
|
||||
}
|
||||
bool ignore = config->ignore_missing_args || config->delete_missing_args;
|
||||
for (int i = 0; i < set->count; i++) {
|
||||
const char* entry = set->entries[i];
|
||||
if (entry[0] == '\0')
|
||||
continue; /* "." == list the whole tree */
|
||||
char* full = path_cat(config->send_directory, entry);
|
||||
if (!full) {
|
||||
log_message(LOG_LEVEL_ERROR, "memory allocation failed while validating --files-from");
|
||||
return false;
|
||||
}
|
||||
struct stat st;
|
||||
if (lstat(full, &st) != 0) {
|
||||
free(full);
|
||||
if (ignore) {
|
||||
(*skipped_out)++;
|
||||
char* escaped_entry = output_escape(entry, log_get_8_bit_output());
|
||||
log_info_message(LOG_INFO_MISC, "skipping missing --files-from entry '%s'",
|
||||
escaped_entry ? escaped_entry : "<allocation failed>");
|
||||
free(escaped_entry);
|
||||
if (config->delete_missing_args && missing_dest) {
|
||||
char* mirror = files_from_missing_dest_path(config, entry);
|
||||
if (!mirror || !array_list_add(missing_dest, mirror)) {
|
||||
free(mirror);
|
||||
log_message(LOG_LEVEL_ERROR, "memory allocation failed while validating --files-from");
|
||||
return false;
|
||||
}
|
||||
}
|
||||
continue;
|
||||
}
|
||||
char* escaped_entry = output_escape(entry, log_get_8_bit_output());
|
||||
char* escaped_src = output_escape(config->send_directory, log_get_8_bit_output());
|
||||
log_message(LOG_LEVEL_ERROR, "--files-from entry '%s' not found in source '%s'",
|
||||
escaped_entry ? escaped_entry : "<allocation failed>",
|
||||
escaped_src ? escaped_src : "<allocation failed>");
|
||||
free(escaped_entry);
|
||||
free(escaped_src);
|
||||
return false;
|
||||
}
|
||||
free(full);
|
||||
}
|
||||
if (*skipped_out > 0) {
|
||||
if (config->delete_missing_args) {
|
||||
/* --list-only never deletes and a --dry-run only shows intent, so the
|
||||
summary must not claim a real deletion happened in those modes. */
|
||||
if (config->list_only)
|
||||
log_message(LOG_LEVEL_WARNING,
|
||||
"--delete-missing-args: %d missing --files-from entr%s skipped (--list-only "
|
||||
"never deletes)",
|
||||
*skipped_out, *skipped_out == 1 ? "y" : "ies");
|
||||
else if (config->dry_run)
|
||||
log_message(LOG_LEVEL_WARNING,
|
||||
"--delete-missing-args: %d missing --files-from entr%s would be deleted from "
|
||||
"the destination (dry run)",
|
||||
*skipped_out, *skipped_out == 1 ? "y" : "ies");
|
||||
else
|
||||
log_message(
|
||||
LOG_LEVEL_WARNING,
|
||||
"--delete-missing-args: %d missing --files-from entr%s will be deleted from the "
|
||||
"destination",
|
||||
*skipped_out, *skipped_out == 1 ? "y" : "ies");
|
||||
} else if (config->ignore_missing_args)
|
||||
log_message(LOG_LEVEL_WARNING,
|
||||
"--ignore-missing-args: ignored %d missing --files-from entr%s", *skipped_out,
|
||||
*skipped_out == 1 ? "y" : "ies");
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
/* Walk the whole source tree once collecting only destination-relative wire
|
||||
paths, loading and sending nothing. --delete-before/--delete-during need the
|
||||
complete keep-set manifest before the first data byte, so it is built by a
|
||||
dedicated pre-scan pass and transmitted early; the data pass then re-scans
|
||||
with a fresh scanner. A source I/O error is fatal unless the options carry
|
||||
--ignore-errors, in which case the scan continues past the unreadable
|
||||
directory and *io_error_out reports it (the caller still performs the
|
||||
deletion but reports the run as errored). */
|
||||
bool scan_paths_only(const Config* config, const ScannerOptions* options, ArrayList* manifest,
|
||||
DeletePlanSender* plans, bool* io_error_out,
|
||||
unsigned long long* non_dir_count_out) {
|
||||
if (io_error_out)
|
||||
*io_error_out = false;
|
||||
if (non_dir_count_out)
|
||||
*non_dir_count_out = 0;
|
||||
ScannerOptions local = *options;
|
||||
/* The pre-scan is a paths-only pass with no client output; it must not emit
|
||||
--info=nonreg lines (the data pass does that once). */
|
||||
local.note_nonreg = false;
|
||||
DirectoryScanner* scanner = directory_scanner_create_with_options(config->send_directory, &local);
|
||||
if (!scanner)
|
||||
return false;
|
||||
bool ok = true;
|
||||
Chunk* chunk;
|
||||
while ((chunk = directory_scanner_next(scanner)) != NULL) {
|
||||
if (non_dir_count_out) {
|
||||
for (int i = 0; i < chunk->element_count; i++) {
|
||||
const File* f = chunk->items[i];
|
||||
if (f && !f->is_dir)
|
||||
(*non_dir_count_out)++;
|
||||
}
|
||||
}
|
||||
if (manifest && !add_chunk_to_manifest(manifest, chunk)) {
|
||||
ok = false;
|
||||
chunk_destroy(chunk);
|
||||
break;
|
||||
}
|
||||
if (plans) {
|
||||
for (int i = 0; i < chunk->element_count; i++) {
|
||||
File* f = chunk->items[i];
|
||||
if (!f)
|
||||
continue;
|
||||
const char* path = file_wire_path(f);
|
||||
if (!delete_plan_sender_add(plans, path, f->is_dir)) {
|
||||
ok = false;
|
||||
break;
|
||||
}
|
||||
}
|
||||
if (!ok) {
|
||||
chunk_destroy(chunk);
|
||||
break;
|
||||
}
|
||||
}
|
||||
chunk_destroy(chunk);
|
||||
}
|
||||
if (ok) {
|
||||
/* Keep every traversed source directory, including empty ones, so a plan
|
||||
no longer removes the destination directory itself. Their own plans are
|
||||
emitted after the data stream (no file frame triggers them). */
|
||||
if (plans && options->plan_dirs) {
|
||||
for (int i = 0; i < options->plan_dirs->size; i++) {
|
||||
if (!delete_plan_sender_add(plans, (const char*)options->plan_dirs->items[i], true)) {
|
||||
ok = false;
|
||||
break;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
if (ok && directory_scanner_failed(scanner))
|
||||
ok = false;
|
||||
if (io_error_out)
|
||||
*io_error_out = directory_scanner_had_io_error(scanner);
|
||||
directory_scanner_destroy(scanner);
|
||||
return ok;
|
||||
}
|
||||
+453
-1764
File diff suppressed because it is too large
Load Diff
@@ -22,7 +22,13 @@ void client_set_abort_armed(bool armed);
|
||||
* never free it, and the caller retains ownership (freeing it with
|
||||
* config_delete() once the call returns). */
|
||||
int send_files(Config* config);
|
||||
int send_files_multithreaded(Config** config);
|
||||
int send_files_multithreaded(Config* config);
|
||||
/* rsync's --ignore-errors deletion gate: with no I/O error during the scan the
|
||||
* deletion phase always proceeds; with one it is suppressed unless
|
||||
* `--ignore-errors` was given. Exposed so the decision can be unit-tested
|
||||
* without a privileged (mode-000) source directory. See client_send.c. */
|
||||
bool ignore_errors_allows_delete(const Config* config, bool had_io_error);
|
||||
|
||||
/* Phase 6 residual-batch (client-only). See client_send.c. */
|
||||
int write_batch_from_source(const Config* config, const char* batch_path);
|
||||
int apply_batch_to_dest(const Config* config, const char* batch_path, const char* dest_root);
|
||||
|
||||
@@ -0,0 +1,94 @@
|
||||
#ifndef CLIENT_SEND_INTERNAL_H
|
||||
#define CLIENT_SEND_INTERNAL_H
|
||||
|
||||
/* Declarations shared between the client_send.c transfer orchestration and the
|
||||
* reporting (client_report.c), scanner-preparation (client_scan.c) and
|
||||
* manifest/list/dry-run (client_manifest.c) translation units that were split
|
||||
* out of it. Nothing here is part of the public client_send.h facade. */
|
||||
|
||||
#include "array_list.h"
|
||||
#include "client_send.h"
|
||||
#include "config.h"
|
||||
#include "delete_plan.h"
|
||||
#include "delta.h"
|
||||
#include "format.h"
|
||||
#include "log.h"
|
||||
#include "scanner.h"
|
||||
#include <stdatomic.h>
|
||||
#include <stdbool.h>
|
||||
#include <stddef.h>
|
||||
#include <time.h>
|
||||
|
||||
/* One mebibyte in bytes; the unit used by the --stats/--progress lines.
|
||||
Always cast to double when dividing so the output stays fractional. */
|
||||
#define BYTES_PER_MIB (1024ULL * 1024ULL)
|
||||
|
||||
/* Compiled scanner inputs that are shared read-only across scanner instances
|
||||
* and, in -m mode, across worker threads. `base_filters` owns the compiled
|
||||
* command-line + -C rules; the FileListSet allow-set lives in the Config.
|
||||
* `hardlinks` owns the --hard-links/-H link-group detection table (NULL when
|
||||
* off) and is shared (mutex-guarded) across every scanner/worker of one scan. */
|
||||
typedef struct {
|
||||
ScannerOptions options;
|
||||
FilterRuleList* base_filters; /* owned; may be NULL */
|
||||
HardLinkTable* hardlinks; /* owned; may be NULL */
|
||||
char* relative_prefix; /* owned -R prefix; may be NULL */
|
||||
} PreparedScanner;
|
||||
|
||||
/* client_scan.c */
|
||||
bool prepare_scanner(const Config* config, int num_threads, PreparedScanner* out);
|
||||
void prepared_scanner_destroy(PreparedScanner* prepared);
|
||||
bool append_implied_dir_times(const Config* config, ArrayList* dir_entries);
|
||||
char* delete_scope_root_marker(const Config* config);
|
||||
const char* delete_plan_walk_root(const Config* config, const ArrayList* synced_dirs);
|
||||
bool files_from_list_check(const Config* config, ArrayList* missing_dest, int* skipped_out);
|
||||
bool scan_paths_only(const Config* config, const ScannerOptions* options, ArrayList* manifest,
|
||||
DeletePlanSender* plans, bool* io_error_out,
|
||||
unsigned long long* non_dir_count_out);
|
||||
|
||||
/* client_report.c */
|
||||
void log_server_rejection(const char* context);
|
||||
const char* display_bytes(unsigned long long bytes, bool human_readable, char* buffer,
|
||||
size_t buffer_size);
|
||||
unsigned long long dir_count_for_stats(const Config* config, const ArrayList* dir_entries,
|
||||
atomic_ullong* counter);
|
||||
void report_transfer_stats(const Config* config, const TransferStats* stats, time_t start,
|
||||
const ReceiverStats* recv);
|
||||
void transfer_stats_note_entry(TransferStats* stats, const File* file);
|
||||
void transfer_stats_note_transferred(TransferStats* stats, const File* file);
|
||||
bool info_flag_enabled(const Config* config, LogInfoFlag flag);
|
||||
void print_delete_reports(const Config* config, const ArrayList* paths);
|
||||
const char* delete_display_path(const Config* config, const char* path);
|
||||
bool progress_requested(const Config* config);
|
||||
void client_progress_cleanup(void);
|
||||
void client_progress_begin(const Config* config);
|
||||
void client_progress_file(const Config* config, const File* file);
|
||||
void client_progress_name(const Config* config, const File* file);
|
||||
/* Emit a transferred entry's ancestor directories (as -i/--out-format change
|
||||
* lines or --progress name lines) before the entry's own line. */
|
||||
void client_change_emit_ancestors(const Config* config, const File* file);
|
||||
void client_progress_uptodate(const Config* config, const File* file);
|
||||
void client_progress_prepare(const Config* config, const ArrayList* plan_dirs,
|
||||
unsigned long long plan_non_dir_count);
|
||||
bool receive_stats_record(int fd, ReceiverStats* stats, ArrayList* would_delete);
|
||||
|
||||
/* client_send.c */
|
||||
void receive_daemon_motd(Client* client, const Config* config);
|
||||
Client* connect_transfer_client(const Config* config);
|
||||
void disconnect_transfer_client(Client* client);
|
||||
int incremental_check(Client* client, File* file, const Config* config, DeltaSignature** out_sig,
|
||||
unsigned long long* resume_offset);
|
||||
|
||||
/* client_manifest.c */
|
||||
bool dry_run_targets_server(const Config* config);
|
||||
bool add_chunk_to_manifest(ArrayList* manifest, const Chunk* chunk);
|
||||
int send_dry_run_manifest(const Config* config);
|
||||
int send_list_only(const Config* config);
|
||||
int send_dry_run_remote(Config* config);
|
||||
int send_delete_manifest(int fd, ArrayList* manifest, ArrayList* protected_prefixes,
|
||||
ArrayList* size_skipped, ArrayList* missing_args, ArrayList* synced_dirs);
|
||||
bool send_delete_manifest_early(Client* client, ArrayList* manifest, ArrayList* protected_prefixes,
|
||||
ArrayList* size_skipped, ArrayList* missing_args,
|
||||
ArrayList* synced_dirs);
|
||||
|
||||
#endif
|
||||
@@ -53,7 +53,7 @@ bool validate_config(const Config* config) {
|
||||
return false;
|
||||
}
|
||||
if (config->compression_threads > 0 && !config->use_compression) {
|
||||
log_message(LOG_LEVEL_ERROR, "--compress-threads requires compression (-c or -z)");
|
||||
log_message(LOG_LEVEL_ERROR, "--compress-threads requires compression (-z/--compress)");
|
||||
return false;
|
||||
}
|
||||
if (config->transport == TRANSPORT_SSH && config->use_sendfile) {
|
||||
@@ -75,6 +75,13 @@ bool validate_config(const Config* config) {
|
||||
log_message(LOG_LEVEL_ERROR, "-4/--ipv4 and -6/--ipv6 are mutually exclusive");
|
||||
return false;
|
||||
}
|
||||
/* rsync 3.4.1 rejects --inplace together with --partial-dir (exit 1): the
|
||||
inplace write path bypasses partial staging, so a partial-dir name would be
|
||||
silently ignored. Match rsync's message and refuse before any I/O. */
|
||||
if (config->inplace && config->partial_dir) {
|
||||
log_message(LOG_LEVEL_ERROR, "--inplace cannot be used with --partial-dir");
|
||||
return false;
|
||||
}
|
||||
if (config->log_file_format && !config->log_file) {
|
||||
log_message(LOG_LEVEL_ERROR, "--log-file-format requires --log-file");
|
||||
return false;
|
||||
@@ -102,6 +109,16 @@ bool validate_config(const Config* config) {
|
||||
log_message(LOG_LEVEL_ERROR, "%s", invariants_error);
|
||||
return false;
|
||||
}
|
||||
/* The receiver rejects a protect-rule block with more than MAX_FILTER_RULES
|
||||
entries as an opaque protocol error; reject an over-limit --filter set here,
|
||||
before any network I/O, with an actionable message. send_protect_entries()
|
||||
re-checks the final built count because cvs-exclude / merge rules can
|
||||
expand it beyond config->filters->size. */
|
||||
if (config->filters && config->filters->size > MAX_FILTER_RULES) {
|
||||
log_message(LOG_LEVEL_ERROR, "too many filter rules: %d (maximum %d)", config->filters->size,
|
||||
MAX_FILTER_RULES);
|
||||
return false;
|
||||
}
|
||||
/* --protocol: FastSync has exactly one wire format, so the forced version
|
||||
must equal the current PROTOCOL_VERSION exactly. Rejected here, before any
|
||||
network I/O, rather than letting the server hit its own mismatch check. */
|
||||
|
||||
+516
-1549
File diff suppressed because it is too large
Load Diff
+58
-5
@@ -10,6 +10,7 @@
|
||||
#include "stop_condition.h"
|
||||
#include <dirent.h>
|
||||
#include <stdbool.h>
|
||||
#include <stddef.h>
|
||||
#include <stdatomic.h>
|
||||
#include <sys/types.h>
|
||||
#include <threads.h>
|
||||
@@ -50,7 +51,7 @@ typedef struct {
|
||||
bool copy_dirlinks;
|
||||
bool munge_links;
|
||||
bool checksum;
|
||||
bool one_file_system;
|
||||
int one_file_system;
|
||||
/* Phase 4 special/devices: whether device nodes (--devices) and special files
|
||||
* (--specials) are preserved via recreation, and whether --copy-devices
|
||||
* copies a device's content as an ordinary regular file. */
|
||||
@@ -121,9 +122,29 @@ typedef struct {
|
||||
* it) and to emit its plan after the data stream, when no file frame would
|
||||
* otherwise trigger it. Guarded by `excluded_mutex`. */
|
||||
ArrayList* plan_dirs;
|
||||
/* --ignore-errors: an unreadable directory during the scan is recorded as an
|
||||
* I/O error and skipped instead of aborting the scan. Client-only. */
|
||||
/* --ignore-errors: an unreadable subdirectory no longer aborts the scan (it
|
||||
* is always skipped so the rest of the tree transfers); this flag is kept so
|
||||
* the client can distinguish the option state when deciding deletion policy.
|
||||
* Client-only. */
|
||||
bool ignore_io_errors;
|
||||
/* --info=nonreg: print rsync's `skipping non-regular file "NAME"` line for a
|
||||
* non-regular entry that is not being preserved. Client-only. */
|
||||
bool note_nonreg;
|
||||
/* --info=mount: print rsync's `[sender] skipping mount-point dir NAME` when
|
||||
* -xx drops a mount-point directory. Client-only. */
|
||||
bool note_mount;
|
||||
/* --stats directory accounting for a `-r` run (no -t/-p): a shared counter of
|
||||
* traversed directories that are NOT otherwise represented by an inline
|
||||
* directory entry (rsync still counts every directory in `Number of files`).
|
||||
* Incremented when a directory is opened and decremented when an empty
|
||||
* directory is emitted inline (so it is counted exactly once). Atomic
|
||||
* because the parallel scanner's workers share it; NULL disables the
|
||||
* accounting. Client-only. */
|
||||
atomic_ullong* dir_count;
|
||||
/* Source root and 8-bit-output policy used to render a `--info=nonreg` name
|
||||
* relative to the transfer root. Borrowed read-only. */
|
||||
const char* send_directory;
|
||||
bool eight_bit_output;
|
||||
/* --ignore-missing-args (implied by --delete-missing-args): an explicitly
|
||||
* --files-from-listed entry that does not exist under the source is skipped
|
||||
* instead of failing (the --dirs generator is the only scanner path that
|
||||
@@ -151,6 +172,17 @@ typedef struct {
|
||||
bool capture_dir_times;
|
||||
ArrayList* dir_entries;
|
||||
mtx_t* dir_entries_mutex;
|
||||
/* Recreate empty source directories on a recursive transfer: emit a
|
||||
* payload-less directory entry for every traversed directory that produced
|
||||
* no transferred/descended child. Off by default so low-level scanner users
|
||||
* (unit helpers, --list-only) see only the historical file list; the real
|
||||
* sender sets it in prepare_scanner. */
|
||||
bool emit_empty_dirs;
|
||||
/* --no-implied-dirs with -R + --files-from: a directory that is only an
|
||||
* implied parent of a listed entry (not itself listed, nor below a listed
|
||||
* directory) must not carry source metadata; it is created with default
|
||||
* attributes at the destination, matching rsync. */
|
||||
bool no_implied_dirs;
|
||||
} ScannerOptions;
|
||||
|
||||
/* Internal per-scanner filter state. FilterNode chains represent the ordered
|
||||
@@ -168,6 +200,22 @@ typedef struct {
|
||||
int current_depth;
|
||||
dev_t root_dev;
|
||||
bool failed;
|
||||
/* rsync-order traversal: each opened directory's entries are inspected once
|
||||
and buffered (an internal SortedEntry[] owned here) sorted as rsync's flist
|
||||
orders them -- non-directories ascending, then directories ascending. The
|
||||
entries are walked in order and child directories are collected in
|
||||
`pending_dirs` (an ArrayList of DirEntry*, owned here) and pushed onto the
|
||||
LIFO `directories` stack in reverse at directory exhaustion, so the emitted
|
||||
stream is depth-first like rsync. `sorted_*` are reset per directory. */
|
||||
void* sorted_entries;
|
||||
size_t sorted_count;
|
||||
size_t sorted_index;
|
||||
void* pending_dirs;
|
||||
/* Recursive scan: whether the open directory yielded any transferred or
|
||||
descended entry. When it did not, closing it emits a directory entry so
|
||||
the empty source directory is recreated at the destination (rsync
|
||||
parity). */
|
||||
bool current_dir_produced;
|
||||
/* Phase 2 (files-from / filter layer). */
|
||||
char* root_path; /* transfer root (fs path) for rel computation */
|
||||
char* current_rel; /* rel path of the open directory ("" == root) */
|
||||
@@ -186,6 +234,10 @@ typedef struct {
|
||||
--ignore-errors the scan continues past it and the caller decides what to
|
||||
do; `failed` is reserved for fatal errors that always abort the scan. */
|
||||
bool io_error;
|
||||
/* The transfer ROOT could not be opened. It is always fatal, even under
|
||||
--ignore-errors, but the client still maps it to rsync's partial-transfer
|
||||
exit (23) rather than a generic failure. */
|
||||
bool root_io_error;
|
||||
} DirectoryScanner;
|
||||
|
||||
typedef struct {
|
||||
@@ -205,7 +257,8 @@ typedef struct {
|
||||
int completed;
|
||||
Chunk* initial_chunk;
|
||||
ProtocolSession* allocation_session;
|
||||
FilterNode* root_filter_node; /* root .rsync-filter context (owned by ps) */
|
||||
FilterNode* root_filter_node; /* root .rsync-filter context (owned by ps) */
|
||||
const ScannerOptions* options; /* borrowed scan options (--info=nonreg output) */
|
||||
} ParallelScanner;
|
||||
|
||||
DirectoryScanner* directory_scanner_create(const char* root_directory, bool use_metadata,
|
||||
@@ -224,7 +277,7 @@ void directory_scanner_destroy(DirectoryScanner* scanner);
|
||||
/* --one-file-system (-x) decision: a directory entry may be descended into
|
||||
* only when the option is disabled or the entry lives on the same device as
|
||||
* the transfer root. Exposed so tests can exercise the rule directly. */
|
||||
bool scanner_same_filesystem(bool one_file_system, dev_t root_device, dev_t entry_device);
|
||||
bool scanner_same_filesystem(int one_file_system, dev_t root_device, dev_t entry_device);
|
||||
|
||||
/* Relative path of an on-disk path below `root` ("" == the root itself, NULL
|
||||
* when `fs_path` is not under `root`). Handles trailing slashes and a root of
|
||||
|
||||
@@ -0,0 +1,671 @@
|
||||
#include "log.h"
|
||||
#include "scanner.h"
|
||||
#include "scanner_internal.h"
|
||||
#include "array_list.h"
|
||||
#include "chunk.h"
|
||||
#include "file.h"
|
||||
#include "queue.h"
|
||||
#include "utils.h"
|
||||
#include <dirent.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <sys/stat.h>
|
||||
#include <sys/sysmacros.h>
|
||||
#include <threads.h>
|
||||
#include <unistd.h>
|
||||
#include <limits.h>
|
||||
|
||||
#include "xattr.h"
|
||||
|
||||
/* A chain node: `own` holds the .rsync-filter rules of one directory, `parent`
|
||||
* the context that directory inherited (nearest ancestor with a filter file).
|
||||
* The chain for a directory's contents runs from that directory's own node up
|
||||
* to the root; the command-line base rules are evaluated after the whole
|
||||
* chain. */
|
||||
struct FilterNode {
|
||||
FilterNode* parent;
|
||||
FilterRuleList* own;
|
||||
};
|
||||
|
||||
void filter_node_destroy(void* item) {
|
||||
if (item) {
|
||||
FilterNode* node = (FilterNode*)item;
|
||||
if (node->own)
|
||||
filter_rule_list_free(node->own);
|
||||
free(node);
|
||||
}
|
||||
}
|
||||
|
||||
FilterNode* filter_node_alloc(FilterNode* parent, FilterRuleList* own) {
|
||||
FilterNode* node = malloc(sizeof(FilterNode));
|
||||
if (!node)
|
||||
return NULL;
|
||||
node->parent = parent;
|
||||
node->own = own;
|
||||
return node;
|
||||
}
|
||||
|
||||
/* Evaluate a rule chain for one entry. rsync precedence, highest first: the
|
||||
* innermost (current) directory's .rsync-filter rules, then each ancestor's,
|
||||
* then the root's, and finally the command-line base rules (--filter/-C). The
|
||||
* sender-side verdict decides whether the entry is hidden from the transfer;
|
||||
* the receiver-side verdict decides whether its destination mirror is protected
|
||||
* from --delete. Each side takes the FIRST matching rule independently. */
|
||||
typedef struct {
|
||||
bool hide; /* sender-side exclude matched */
|
||||
bool protect; /* receiver-side exclude matched */
|
||||
} FilterOutcome;
|
||||
|
||||
static void chain_rules_outcome(const FilterRuleList* base, const FilterNode* node, const char* rel,
|
||||
const char* leaf, bool is_dir, FilterOutcome* out) {
|
||||
memset(out, 0, sizeof(*out));
|
||||
bool sender_decided = false;
|
||||
bool receiver_decided = false;
|
||||
const FilterNode* n = node;
|
||||
while (!sender_decided || !receiver_decided) {
|
||||
const FilterRuleList* list = n ? n->own : base;
|
||||
if (list) {
|
||||
if (!sender_decided) {
|
||||
FilterAction action = filter_rules_apply_side(list, rel, leaf, is_dir, FILTER_SIDE_SENDER);
|
||||
if (action != FILTER_ACTION_NONE) {
|
||||
out->hide = action == FILTER_ACTION_EXCLUDE;
|
||||
sender_decided = true;
|
||||
}
|
||||
}
|
||||
if (!receiver_decided) {
|
||||
FilterAction action =
|
||||
filter_rules_apply_side(list, rel, leaf, is_dir, FILTER_SIDE_RECEIVER);
|
||||
if (action != FILTER_ACTION_NONE) {
|
||||
out->protect = action == FILTER_ACTION_PROTECT;
|
||||
receiver_decided = true;
|
||||
}
|
||||
}
|
||||
}
|
||||
if (!n)
|
||||
break;
|
||||
n = n->parent;
|
||||
}
|
||||
}
|
||||
|
||||
static bool entry_allowed(const FilterRuleList* base, const FilterNode* node, const char* rel,
|
||||
const char* leaf, bool is_dir, bool exclude_filter_files,
|
||||
bool* protect_out) {
|
||||
/* -FF: per-directory .rsync-filter files are never transferred (single -F
|
||||
transfers them, matching rsync). */
|
||||
if (exclude_filter_files && !is_dir && strcmp(leaf, ".rsync-filter") == 0) {
|
||||
if (protect_out)
|
||||
*protect_out = false;
|
||||
return false;
|
||||
}
|
||||
FilterOutcome outcome;
|
||||
chain_rules_outcome(base, node, rel, leaf, is_dir, &outcome);
|
||||
if (protect_out)
|
||||
*protect_out = outcome.protect;
|
||||
return !outcome.hide;
|
||||
}
|
||||
|
||||
void dir_entry_destroy(void* item) {
|
||||
if (item) {
|
||||
DirEntry* de = (DirEntry*)item;
|
||||
free(de->path);
|
||||
free(de);
|
||||
}
|
||||
}
|
||||
|
||||
DirEntry* dir_entry_create(const char* path, int depth, FilterNode* context) {
|
||||
DirEntry* de = malloc(sizeof(DirEntry));
|
||||
if (!de)
|
||||
return NULL;
|
||||
de->path = str_dup(path);
|
||||
if (!de->path) {
|
||||
free(de);
|
||||
return NULL;
|
||||
}
|
||||
de->depth = depth;
|
||||
de->context = context;
|
||||
return de;
|
||||
}
|
||||
|
||||
/* Apply rsync's symlink-resolution precedence to one S_ISLNK entry:
|
||||
* --copy-links dereferences every symlink;
|
||||
* --copy-unsafe-links dereferences only targets unsafe_symlink() flags;
|
||||
* -k/--copy-dirlinks dereferences only a symlink whose referent is a dir;
|
||||
* --safe-links (receiver-side in rsync; modelled here) ignores an unsafe
|
||||
* target that would otherwise be carried; with --munge-links
|
||||
* every stored target becomes absolute, so --safe-links then
|
||||
* ignores every symlink, exactly as rsync documents;
|
||||
* -l/--links carries the link.
|
||||
* `link_rel` is the symlink's transfer-relative path (incl. name) and is used
|
||||
* only for the lexical unsafe test. `target` receives the raw link value. */
|
||||
LinkAction scanner_link_action(const ScannerOptions* options, const char* path,
|
||||
const char* link_rel, char* target, size_t target_size) {
|
||||
if (!options->follow_symlinks && !options->copy_links && !options->safe_links &&
|
||||
!options->copy_unsafe_links && !options->copy_dirlinks)
|
||||
return LINK_ACTION_SKIP;
|
||||
ssize_t length = readlink(path, target, target_size - 1);
|
||||
if (length < 0)
|
||||
return LINK_ACTION_SKIP;
|
||||
target[length] = '\0';
|
||||
|
||||
bool unsafe = file_symlink_unsafe(target, link_rel);
|
||||
if (options->copy_links || (options->copy_unsafe_links && unsafe))
|
||||
return LINK_ACTION_DEREF;
|
||||
if (options->copy_dirlinks) {
|
||||
struct stat ref;
|
||||
if (stat(path, &ref) == 0 && S_ISDIR(ref.st_mode))
|
||||
return LINK_ACTION_DEREF;
|
||||
}
|
||||
if (options->safe_links && (unsafe || options->munge_links))
|
||||
return LINK_ACTION_SKIP_PROTECTED;
|
||||
if (!options->follow_symlinks || target[0] == '\0')
|
||||
return LINK_ACTION_SKIP;
|
||||
return LINK_ACTION_CARRY;
|
||||
}
|
||||
|
||||
/* --one-file-system (-x) decision. Only directories can carry a different
|
||||
* device than their parent (mount points), so this is checked when a child
|
||||
* directory is about to be descended into. */
|
||||
bool scanner_same_filesystem(int one_file_system, dev_t root_device, dev_t entry_device) {
|
||||
return one_file_system <= 0 || entry_device == root_device;
|
||||
}
|
||||
|
||||
/* Build a payload-less directory File carrying the captured metadata (when
|
||||
* requested). Used by -x mount-point emission and --list-only directory
|
||||
* entries. Returns NULL on allocation failure. */
|
||||
File* scanner_build_dir_file(const char* path, const struct stat* stats,
|
||||
const ScannerOptions* options) {
|
||||
File* dir = file_create(path);
|
||||
if (dir == NULL)
|
||||
return NULL;
|
||||
dir->is_dir = true;
|
||||
if (options->use_metadata) {
|
||||
dir->metadata =
|
||||
file_metadata_create(dir->path, stats, options->preserve_atimes, options->preserve_crtimes);
|
||||
if (!dir->metadata) {
|
||||
file_destroy(dir);
|
||||
return NULL;
|
||||
}
|
||||
}
|
||||
return dir;
|
||||
}
|
||||
|
||||
/* Relative path of an on-disk path below `root`. The transfer root may be
|
||||
* given with a trailing slash; the returned rel path never has one and is ""
|
||||
* for the root itself. A root of "/" is handled (its children start at "/").
|
||||
* Exposed so tests can exercise the mapping directly. */
|
||||
char* scanner_path_relative(const char* root, const char* fs_path) {
|
||||
size_t root_len = strlen(root);
|
||||
while (root_len > 1 && root[root_len - 1] == '/')
|
||||
root_len--;
|
||||
if (strncmp(root, fs_path, root_len) != 0)
|
||||
return NULL;
|
||||
if (root_len == 1 && root[0] == '/') {
|
||||
if (fs_path[1] == '\0')
|
||||
return str_dup("");
|
||||
return str_dup(fs_path + 1);
|
||||
}
|
||||
if (fs_path[root_len] == '\0')
|
||||
return str_dup("");
|
||||
if (fs_path[root_len] != '/')
|
||||
return NULL;
|
||||
return str_dup(fs_path + root_len + 1);
|
||||
}
|
||||
|
||||
/* -R/--relative destination-relative prefix reconstructed from a source spec:
|
||||
* everything after the first '.' path component (rsync's '/./' cut point),
|
||||
* with leading/trailing slashes removed; or the whole spec (normalized) when
|
||||
* there is no cut. Returns "" for the receive root. Exposed for tests. */
|
||||
char* scanner_relative_prefix(const char* spec) {
|
||||
if (!spec || spec[0] == '\0')
|
||||
return NULL;
|
||||
const char* after = spec;
|
||||
if (spec[0] == '.' && spec[1] == '/') {
|
||||
after = spec + 2;
|
||||
} else {
|
||||
const char* cut = strstr(spec, "/./");
|
||||
if (cut)
|
||||
after = cut + 3;
|
||||
}
|
||||
size_t cap = strlen(spec) + 1;
|
||||
char* out = malloc(cap);
|
||||
if (!out)
|
||||
return NULL;
|
||||
size_t len = 0;
|
||||
for (const char* s = after; *s;) {
|
||||
while (*s == '/')
|
||||
s++;
|
||||
const char* comp = s;
|
||||
while (*s && *s != '/')
|
||||
s++;
|
||||
size_t clen = (size_t)(s - comp);
|
||||
if (clen == 0 || (clen == 1 && comp[0] == '.'))
|
||||
continue;
|
||||
if (len)
|
||||
out[len++] = '/';
|
||||
memcpy(out + len, comp, clen);
|
||||
len += clen;
|
||||
}
|
||||
out[len] = '\0';
|
||||
return out;
|
||||
}
|
||||
|
||||
/* Relative path of a child entry below the current directory. */
|
||||
char* child_rel_path(const char* parent_rel, const char* name) {
|
||||
if (!parent_rel || parent_rel[0] == '\0')
|
||||
return str_dup(name);
|
||||
return path_cat(parent_rel, name);
|
||||
}
|
||||
|
||||
/* Destination-relative wire path for an entry under an -R prefix. */
|
||||
char* scanner_prefix_send_path(const char* prefix, const char* rel) {
|
||||
if (prefix[0] == '\0')
|
||||
return str_dup(rel);
|
||||
if (rel[0] == '\0')
|
||||
return str_dup(prefix);
|
||||
return path_cat(prefix, rel);
|
||||
}
|
||||
|
||||
/* Apply the --files-from allow-set and the filter layer to one entry. On
|
||||
* return `*protect_out` is true when a receiver-side rule protects the entry's
|
||||
* destination mirror from deletion. */
|
||||
bool entry_passes_selection(const FileListSet* file_list, const FilterRuleList* base,
|
||||
const FilterNode* node, const char* rel, const char* leaf, bool is_dir,
|
||||
bool per_dir_filters, bool exclude_filter_files, bool* protect_out) {
|
||||
if (protect_out)
|
||||
*protect_out = false;
|
||||
if (file_list && !file_list_affects(file_list, rel))
|
||||
return false;
|
||||
if (base || per_dir_filters)
|
||||
return entry_allowed(base, node, rel, leaf, is_dir, exclude_filter_files, protect_out);
|
||||
return true;
|
||||
}
|
||||
|
||||
/* Best-effort capture of the file's whitelisted xattrs (-X/-A). A failure to
|
||||
* read xattrs is non-fatal: the file is transferred without them. */
|
||||
void scanner_capture_xattrs(const DirectoryScanner* scanner, File* file) {
|
||||
if (!scanner || !file || !(scanner->options.preserve_xattrs || scanner->options.preserve_acls))
|
||||
return;
|
||||
file->xattrs = xattr_capture_path(file->path, scanner->options.preserve_acls);
|
||||
}
|
||||
|
||||
/* Apply --hard-links (-H) detection to one regular File. On a sibling (a
|
||||
* later member of an already-seen source inode) the File keeps the group id
|
||||
* and the first member's wire path but carries NO data payload (size 0); the
|
||||
* first member is left untouched (data present, link_first). Allocation
|
||||
* failure is fatal: the scanner is marked failed. */
|
||||
void scanner_assign_hardlink(DirectoryScanner* scanner, HardLinkTable* table, File* file,
|
||||
const struct stat* stats) {
|
||||
if (!table || !file || !stats)
|
||||
return;
|
||||
int gid;
|
||||
bool is_first;
|
||||
char* first_path = NULL;
|
||||
if (!hardlink_table_assign(table, file_wire_path(file), stats->st_dev, stats->st_ino, &gid,
|
||||
&is_first, &first_path)) {
|
||||
if (scanner)
|
||||
scanner->failed = true;
|
||||
return;
|
||||
}
|
||||
file->link_group = gid;
|
||||
file->link_first = is_first;
|
||||
if (!is_first) {
|
||||
file->hardlink_target = first_path;
|
||||
file->data->size = 0;
|
||||
} else {
|
||||
free(first_path);
|
||||
}
|
||||
}
|
||||
|
||||
/* Phase 4 special/devices decision for one non-regular entry, matching rsync:
|
||||
- a char/block device is RECREATED as a node under -D/--devices, unless
|
||||
--copy-devices asks for its content to be copied into a regular file;
|
||||
- a FIFO/socket is RECREATED under --specials;
|
||||
- when the matching flag is absent the entry is SKIPPED ("skipping
|
||||
non-regular file"), exactly like rsync's default, instead of being
|
||||
silently copied as a zero-length regular file;
|
||||
- anything else (regular/directory) is left to the normal data path. */
|
||||
ScannerSpecial scanner_prepare_special(bool preserve_devices, bool preserve_specials,
|
||||
bool copy_devices, File* file, const struct stat* stats) {
|
||||
if (!file || !stats)
|
||||
return SCANNER_SPECIAL_REGULAR;
|
||||
bool is_device = S_ISCHR(stats->st_mode) || S_ISBLK(stats->st_mode);
|
||||
bool is_fifo = S_ISFIFO(stats->st_mode);
|
||||
bool is_socket = S_ISSOCK(stats->st_mode);
|
||||
if (!is_device && !is_fifo && !is_socket)
|
||||
return SCANNER_SPECIAL_REGULAR;
|
||||
if (is_device && copy_devices)
|
||||
return SCANNER_SPECIAL_REGULAR; /* copy device content as a regular file */
|
||||
bool preserve = is_device ? preserve_devices : preserve_specials;
|
||||
if (!preserve)
|
||||
return SCANNER_SPECIAL_SKIP;
|
||||
file->is_special = true;
|
||||
file->data->size = 0;
|
||||
file->data->data = NULL;
|
||||
if (is_device) {
|
||||
file->rdev_major = (int32_t)major(stats->st_rdev);
|
||||
file->rdev_minor = (int32_t)minor(stats->st_rdev);
|
||||
}
|
||||
return SCANNER_SPECIAL_RECREATE;
|
||||
}
|
||||
|
||||
/* Append `rel` to the caller's exclusion sink, taking `mtx` when shared across
|
||||
parallel worker threads. Returns false on allocation failure (list left
|
||||
unchanged). */
|
||||
bool excluded_sink_append(ArrayList* list, mtx_t* mtx, const char* rel) {
|
||||
if (!list)
|
||||
return true;
|
||||
char* dup = str_dup(rel);
|
||||
if (!dup)
|
||||
return false;
|
||||
if (mtx)
|
||||
mtx_lock(mtx);
|
||||
bool ok = array_list_add(list, dup);
|
||||
if (mtx)
|
||||
mtx_unlock(mtx);
|
||||
if (!ok)
|
||||
free(dup);
|
||||
return ok;
|
||||
}
|
||||
|
||||
/* Record one pruned filesystem path in a delete-protection sink. The stored
|
||||
form is the entry's wire/destination-relative path (a single leading '/'
|
||||
removed, exactly how manifest keep entries are stored), so the receiver's
|
||||
walker prefixes match the destination layout. An allocation failure is a
|
||||
fatal scan error. */
|
||||
static void scanner_record_protected(DirectoryScanner* scanner, const char* fs_path,
|
||||
ArrayList* sink) {
|
||||
if (!sink || !fs_path)
|
||||
return;
|
||||
const char* rel = *fs_path == '/' ? fs_path + 1 : fs_path;
|
||||
if (!excluded_sink_append(sink, scanner->options.excluded_mutex, rel))
|
||||
scanner->failed = true;
|
||||
}
|
||||
|
||||
/* rsync's `--info=nonreg` line for a non-regular entry that is not being
|
||||
* preserved: `skipping non-regular file "NAME"`. The name is the path relative
|
||||
* to the transfer root, so it matches rsync's displayed name. */
|
||||
void scanner_note_nonreg(const ScannerOptions* options, const char* fs_path) {
|
||||
if (!options || !options->note_nonreg || !fs_path)
|
||||
return;
|
||||
const char* rel = utils_strip_transfer_root(fs_path, options->send_directory);
|
||||
char* escaped = output_escape(rel, options->eight_bit_output);
|
||||
printf("skipping non-regular file \"%s\"\n", escaped ? escaped : rel);
|
||||
free(escaped);
|
||||
fflush(stdout);
|
||||
}
|
||||
|
||||
/* rsync 3.4.1's `--info=mount` line, emitted when `-xx` drops a mount-point
|
||||
* directory: `[sender] skipping mount-point dir NAME` (the client is the
|
||||
* sender). Plain `-x` keeps the empty directory and prints nothing, matching
|
||||
* rsync. */
|
||||
void scanner_note_mount(const ScannerOptions* options, const char* fs_path) {
|
||||
if (!options || !options->note_mount || !fs_path)
|
||||
return;
|
||||
const char* rel = utils_strip_transfer_root(fs_path, options->send_directory);
|
||||
char* escaped = output_escape(rel, options->eight_bit_output);
|
||||
printf("[sender] skipping mount-point dir %s\n", escaped ? escaped : rel);
|
||||
free(escaped);
|
||||
fflush(stdout);
|
||||
}
|
||||
|
||||
/* --debug=filter: a selection/filter decision dropped an entry. */
|
||||
void scanner_note_filter(const ScannerOptions* options, const char* name) {
|
||||
if (!options || !log_debug_enabled(LOG_DEBUG_FILTER) || !name)
|
||||
return;
|
||||
log_debug_message(LOG_DEBUG_FILTER, "filter: excluded %s", name);
|
||||
}
|
||||
|
||||
/* Account for a directory that will not be represented by an inline directory
|
||||
* entry. Paired with scanner_dir_count_uncount for empty directories that are
|
||||
* emitted inline, so every traversed directory is counted exactly once. */
|
||||
void scanner_dir_count_count(const ScannerOptions* options) {
|
||||
if (options && options->dir_count)
|
||||
atomic_fetch_add(options->dir_count, 1);
|
||||
}
|
||||
|
||||
void scanner_dir_count_uncount(const ScannerOptions* options) {
|
||||
if (options && options->dir_count)
|
||||
atomic_fetch_sub(options->dir_count, 1);
|
||||
}
|
||||
|
||||
/* A user-selection exclusion (--filter/-C/per-dir or --exclude/--include). */
|
||||
void scanner_record_excluded(DirectoryScanner* scanner, const char* fs_path) {
|
||||
scanner_record_protected(scanner, fs_path, scanner->options.excluded_paths);
|
||||
}
|
||||
|
||||
/* A --max-size/--min-size prune (always protected, even under --delete-excluded). */
|
||||
void scanner_record_size_skipped(DirectoryScanner* scanner, const char* fs_path) {
|
||||
scanner_record_protected(scanner, fs_path, scanner->options.size_skipped_paths);
|
||||
}
|
||||
|
||||
/* Record a directory the scan synchronized. `fs_path` is its absolute path and
|
||||
`rel` its path relative to the transfer root ("" for the root); the stored
|
||||
form matches the wire layout (the bare relative path in -R+--files-from, else
|
||||
the source path with a leading '/' removed, with "." for the receive root).
|
||||
Returns false on allocation failure. */
|
||||
bool scanner_record_synced_dir(const ScannerOptions* options, const char* fs_path, const char* rel,
|
||||
bool relative_mode) {
|
||||
if (!options->synced_dirs && !options->plan_dirs)
|
||||
return true;
|
||||
if (!file_list_dir_in_scope(options->file_list, rel))
|
||||
return true;
|
||||
char* prefixed = NULL;
|
||||
const char* dest;
|
||||
if (relative_mode) {
|
||||
dest = rel;
|
||||
} else if (options->relative_prefix) {
|
||||
prefixed = scanner_prefix_send_path(options->relative_prefix, rel);
|
||||
if (!prefixed)
|
||||
return false;
|
||||
dest = prefixed;
|
||||
} else {
|
||||
dest = fs_path;
|
||||
}
|
||||
if (dest[0] == '/')
|
||||
dest++;
|
||||
if (dest[0] == '\0')
|
||||
dest = ".";
|
||||
bool ok = true;
|
||||
if (options->synced_dirs)
|
||||
ok = excluded_sink_append(options->synced_dirs, options->excluded_mutex, dest);
|
||||
/* The delete-plan keep set needs an entry for every traversed source
|
||||
directory, including empty ones, so its destination mirror is kept rather
|
||||
than deleted as an extra; the receive root (".") is implicit. */
|
||||
if (ok && options->plan_dirs && strcmp(dest, ".") != 0)
|
||||
ok = excluded_sink_append(options->plan_dirs, options->excluded_mutex, dest);
|
||||
free(prefixed);
|
||||
return ok;
|
||||
}
|
||||
|
||||
/* Read every per-directory filter file that applies to `dir_path` (its
|
||||
* .rsync-filter when -F is active, plus each registered "dir-merge NAME") into a
|
||||
* fresh list. Returns NULL on allocation/parse failure (message in `err`);
|
||||
* returns an empty list (and *any_exists=false) when no file exists. */
|
||||
FilterRuleList* read_dir_filters(const ScannerOptions* options, const char* dir_path,
|
||||
const char* rel, bool* any_exists, char* err, size_t err_size) {
|
||||
if (err && err_size > 0)
|
||||
err[0] = '\0';
|
||||
const FilterRuleList* base = options->base_filters;
|
||||
bool have_names = options->per_dir_filters || (base && base->dir_merge_count > 0);
|
||||
if (any_exists)
|
||||
*any_exists = false;
|
||||
if (!have_names)
|
||||
return NULL;
|
||||
FilterRuleList* own = filter_rule_list_create();
|
||||
if (!own) {
|
||||
snprintf(err, err_size, "memory allocation failed");
|
||||
return NULL;
|
||||
}
|
||||
FilterParseOptions opts = {.delete_excluded = options->delete_excluded, .cvs_exclude = false};
|
||||
bool exists = false;
|
||||
if (options->per_dir_filters) {
|
||||
if (!filter_file_append(own, dir_path, ".rsync-filter", rel, &opts, &exists, err, err_size))
|
||||
goto fail;
|
||||
if (exists && any_exists)
|
||||
*any_exists = true;
|
||||
}
|
||||
if (base) {
|
||||
for (int i = 0; i < base->dir_merge_count; i++) {
|
||||
if (!filter_file_append(own, dir_path, base->dir_merge_names[i], rel, &opts, &exists, err,
|
||||
err_size))
|
||||
goto fail;
|
||||
if (exists && any_exists)
|
||||
*any_exists = true;
|
||||
}
|
||||
}
|
||||
return own;
|
||||
fail:
|
||||
filter_rule_list_free(own);
|
||||
return NULL;
|
||||
}
|
||||
|
||||
/* Merge the open directory's own per-directory filter files (the default
|
||||
* .rsync-filter when -F is active, plus every "dir-merge NAME" registered on the
|
||||
* base rule list) into the inherited context, returning the context used for
|
||||
* this directory's entries. On a parse error the scanner is marked failed.
|
||||
* Returns 0 on success, -1 on failure. */
|
||||
int open_directory_filter_context(DirectoryScanner* scanner, const FilterNode* inherited) {
|
||||
char err[256];
|
||||
bool any_exists = false;
|
||||
FilterRuleList* own = read_dir_filters(&scanner->options, scanner->current_path,
|
||||
scanner->current_rel ? scanner->current_rel : "",
|
||||
&any_exists, err, sizeof(err));
|
||||
if (!own) {
|
||||
/* read_dir_filters() leaves `err` set on a parse/allocation failure even
|
||||
when an earlier merge file in the same directory existed (any_exists true);
|
||||
key off the error text rather than any_exists so an invalid per-directory
|
||||
filter file can never be silently ignored. */
|
||||
if (err[0] == '\0') {
|
||||
scanner->current_node = (FilterNode*)inherited;
|
||||
return 0;
|
||||
}
|
||||
char* escaped_path = output_escape(scanner->current_path, log_get_8_bit_output());
|
||||
log_message(LOG_LEVEL_ERROR, "invalid per-directory filter in %s: %s",
|
||||
escaped_path ? escaped_path : "<allocation failed>", err);
|
||||
free(escaped_path);
|
||||
scanner->failed = true;
|
||||
return -1;
|
||||
}
|
||||
if (any_exists && (own->count > 0 || own->dir_merge_count > 0)) {
|
||||
FilterNode* node = filter_node_alloc((FilterNode*)inherited, own);
|
||||
if (!node || !array_list_add(scanner->filter_nodes, node)) {
|
||||
filter_node_destroy(node);
|
||||
scanner->failed = true;
|
||||
return -1;
|
||||
}
|
||||
scanner->current_node = node;
|
||||
} else {
|
||||
filter_rule_list_free(own);
|
||||
scanner->current_node = (FilterNode*)inherited;
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
/* Inspect symlinks, resolve the entry type, and apply file filters once for both scanners.
|
||||
* `link_rel` is the entry's path relative to the transfer root (including its
|
||||
* name), used for the lexical rsync unsafe-symlink test. */
|
||||
int scanner_inspect_entry(const ScannerOptions* options, const char* containing_dir,
|
||||
const char* link_rel, const char* name, ScannerEntry* entry) {
|
||||
entry->excluded = false;
|
||||
entry->size_excluded = false;
|
||||
entry->referent_error = false;
|
||||
entry->is_symlink = false;
|
||||
entry->link_target = NULL;
|
||||
entry->path = path_cat(containing_dir, name);
|
||||
if (!entry->path)
|
||||
return -1;
|
||||
|
||||
struct stat link_stats;
|
||||
if (lstat(entry->path, &link_stats) != 0) {
|
||||
free(entry->path);
|
||||
return 0;
|
||||
}
|
||||
if (!S_ISLNK(link_stats.st_mode))
|
||||
goto regular;
|
||||
|
||||
char link_target[4096];
|
||||
switch (scanner_link_action(options, entry->path, link_rel, link_target, sizeof(link_target))) {
|
||||
case LINK_ACTION_SKIP:
|
||||
goto skip;
|
||||
case LINK_ACTION_SKIP_PROTECTED:
|
||||
/* --safe-links ignored the link, but rsync still counts it as present in
|
||||
the transfer, so its destination mirror survives --delete. Record it as
|
||||
an excluded path (the same delete-protection channel as a filter prune). */
|
||||
entry->excluded = true;
|
||||
goto skip;
|
||||
case LINK_ACTION_DEREF:
|
||||
if (stat(entry->path, &entry->stats) != 0) {
|
||||
/* rsync reports "symlink has no referent" and continues with a partial
|
||||
transfer (exit 23); record the error so the run exits 23 too. */
|
||||
char* escaped = output_escape(entry->path, log_get_8_bit_output());
|
||||
log_message(LOG_LEVEL_WARNING, "symlink has no referent: %s",
|
||||
escaped ? escaped : "<allocation failed>");
|
||||
free(escaped);
|
||||
entry->referent_error = true;
|
||||
goto skip;
|
||||
}
|
||||
entry->is_directory = S_ISDIR(entry->stats.st_mode);
|
||||
if (entry->is_directory)
|
||||
return 1;
|
||||
goto apply_filters;
|
||||
case LINK_ACTION_CARRY:
|
||||
break;
|
||||
}
|
||||
|
||||
/* Carry the link as a symlink. --munge-links is applied by the RECEIVER (it
|
||||
prefixes every stored target with /rsyncd-munged/); when the SOURCE already
|
||||
holds a munged value the sender strips it so the receiver re-munges a clean
|
||||
target, round-tripping a munged tree exactly like rsync. */
|
||||
entry->is_symlink = true;
|
||||
entry->stats = link_stats;
|
||||
entry->is_directory = false;
|
||||
entry->link_target = str_dup(link_target);
|
||||
if (!entry->link_target)
|
||||
goto skip;
|
||||
if (options->munge_links)
|
||||
file_symlink_unmunge(entry->link_target);
|
||||
goto apply_filters;
|
||||
|
||||
regular:
|
||||
/* Not a symlink: the lstat() above already described this entry, and lstat
|
||||
and stat are identical for every non-symlink, so reuse that result instead
|
||||
of issuing a redundant stat() on the scanner hot path. stat() is still
|
||||
used on the dereference paths above/below for actual symlinks (copy-links,
|
||||
safe/copy-unsafe links, and -k symlinks-to-directories). */
|
||||
entry->stats = link_stats;
|
||||
entry->is_directory = S_ISDIR(link_stats.st_mode);
|
||||
if (entry->is_directory)
|
||||
return 1;
|
||||
|
||||
apply_filters:
|
||||
for (int i = 0; i < options->exclude_count; i++)
|
||||
if (glob_match(options->exclude_patterns[i], name)) {
|
||||
entry->excluded = true;
|
||||
goto skip;
|
||||
}
|
||||
if (options->include_count > 0) {
|
||||
bool included = false;
|
||||
for (int i = 0; i < options->include_count; i++)
|
||||
if (glob_match(options->include_patterns[i], name))
|
||||
included = true;
|
||||
if (!included) {
|
||||
entry->excluded = true;
|
||||
goto skip;
|
||||
}
|
||||
}
|
||||
if ((options->max_size > 0 && (unsigned long long)entry->stats.st_size > options->max_size) ||
|
||||
(options->min_size > 0 && (unsigned long long)entry->stats.st_size < options->min_size)) {
|
||||
entry->excluded = true;
|
||||
entry->size_excluded = true;
|
||||
goto skip;
|
||||
}
|
||||
return 1;
|
||||
|
||||
skip:
|
||||
free(entry->path);
|
||||
entry->path = NULL;
|
||||
free(entry->link_target);
|
||||
entry->link_target = NULL;
|
||||
return 0;
|
||||
}
|
||||
@@ -0,0 +1,107 @@
|
||||
#ifndef SCANNER_INTERNAL_H
|
||||
#define SCANNER_INTERNAL_H
|
||||
|
||||
/* Internal declarations shared between the scanner translation units
|
||||
* (scanner_filter.c, scanner.c, scanner_parallel.c). Nothing here is part of
|
||||
* the public scanner façade (scanner.h); every symbol stays internal to the
|
||||
* client module. */
|
||||
|
||||
#include "array_list.h"
|
||||
#include "file.h"
|
||||
#include "scanner.h"
|
||||
#include <stdbool.h>
|
||||
#include <stddef.h>
|
||||
#include <sys/stat.h>
|
||||
|
||||
typedef struct {
|
||||
char* path;
|
||||
int depth;
|
||||
FilterNode* context; /* inherited per-directory filter context */
|
||||
} DirEntry;
|
||||
|
||||
/* How rsync's readlink_stat()/generator resolves one source symlink. */
|
||||
typedef enum {
|
||||
LINK_ACTION_SKIP, /* not transferred (no link option) */
|
||||
LINK_ACTION_SKIP_PROTECTED, /* ignored as unsafe by --safe-links; rsync keeps
|
||||
it in the transfer, so its destination mirror
|
||||
must be protected from --delete */
|
||||
LINK_ACTION_DEREF, /* follow the referent (--copy-links, an unsafe
|
||||
target under --copy-unsafe-links, or -k dir) */
|
||||
LINK_ACTION_CARRY, /* transmit the link itself (-l) */
|
||||
} LinkAction;
|
||||
|
||||
typedef struct {
|
||||
char* path;
|
||||
struct stat stats;
|
||||
bool is_directory;
|
||||
/* True when the entry should be carried through as a SYMLINK (is_symlink)
|
||||
rather than a dereferenced file/directory. When true, `link_target` holds
|
||||
the owned target string to transmit (sender-munged under --munge-links);
|
||||
ownership transfers to the File built from this entry. */
|
||||
bool is_symlink;
|
||||
char* link_target;
|
||||
/* True when the entry was pruned by a user selection rule (--filter/-C/per-dir
|
||||
rules or the --exclude/--include layer) rather than skipped for another
|
||||
reason (unreadable, symlink policy, not applicable). */
|
||||
bool excluded;
|
||||
/* True when the entry was skipped specifically by --max-size/--min-size.
|
||||
Size pruning protects the destination mirror even under --delete-excluded,
|
||||
so it is recorded into a separate sink from `excluded`. */
|
||||
bool size_excluded;
|
||||
/* True when a symlink selected for dereferencing (-L/--copy-links or an
|
||||
unsafe target under --copy-unsafe-links) had no usable referent (a broken
|
||||
link or a stat() failure). rsync still reports this as a partial transfer
|
||||
(exit 23) even though the entry is skipped, so the scanner records it as a
|
||||
non-fatal I/O error. */
|
||||
bool referent_error;
|
||||
} ScannerEntry;
|
||||
|
||||
typedef enum {
|
||||
SCANNER_SPECIAL_REGULAR, /* ordinary file: transfer content */
|
||||
SCANNER_SPECIAL_RECREATE, /* is_special node to recreate on the receiver */
|
||||
SCANNER_SPECIAL_SKIP, /* non-regular entry not requested: skip */
|
||||
} ScannerSpecial;
|
||||
|
||||
/* scanner_filter.c */
|
||||
void filter_node_destroy(void* item);
|
||||
FilterNode* filter_node_alloc(FilterNode* parent, FilterRuleList* own);
|
||||
void dir_entry_destroy(void* item);
|
||||
DirEntry* dir_entry_create(const char* path, int depth, FilterNode* context);
|
||||
LinkAction scanner_link_action(const ScannerOptions* options, const char* path,
|
||||
const char* link_rel, char* target, size_t target_size);
|
||||
File* scanner_build_dir_file(const char* path, const struct stat* stats,
|
||||
const ScannerOptions* options);
|
||||
char* child_rel_path(const char* parent_rel, const char* name);
|
||||
char* scanner_prefix_send_path(const char* prefix, const char* rel);
|
||||
bool entry_passes_selection(const FileListSet* file_list, const FilterRuleList* base,
|
||||
const FilterNode* node, const char* rel, const char* leaf, bool is_dir,
|
||||
bool per_dir_filters, bool exclude_filter_files, bool* protect_out);
|
||||
void scanner_capture_xattrs(const DirectoryScanner* scanner, File* file);
|
||||
void scanner_assign_hardlink(DirectoryScanner* scanner, HardLinkTable* table, File* file,
|
||||
const struct stat* stats);
|
||||
ScannerSpecial scanner_prepare_special(bool preserve_devices, bool preserve_specials,
|
||||
bool copy_devices, File* file, const struct stat* stats);
|
||||
bool excluded_sink_append(ArrayList* list, mtx_t* mtx, const char* rel);
|
||||
void scanner_note_nonreg(const ScannerOptions* options, const char* fs_path);
|
||||
void scanner_note_mount(const ScannerOptions* options, const char* fs_path);
|
||||
void scanner_note_filter(const ScannerOptions* options, const char* name);
|
||||
void scanner_dir_count_count(const ScannerOptions* options);
|
||||
void scanner_dir_count_uncount(const ScannerOptions* options);
|
||||
void scanner_record_excluded(DirectoryScanner* scanner, const char* fs_path);
|
||||
void scanner_record_size_skipped(DirectoryScanner* scanner, const char* fs_path);
|
||||
bool scanner_record_synced_dir(const ScannerOptions* options, const char* fs_path, const char* rel,
|
||||
bool relative_mode);
|
||||
FilterRuleList* read_dir_filters(const ScannerOptions* options, const char* dir_path,
|
||||
const char* rel, bool* any_exists, char* err, size_t err_size);
|
||||
int open_directory_filter_context(DirectoryScanner* scanner, const FilterNode* inherited);
|
||||
int scanner_inspect_entry(const ScannerOptions* options, const char* containing_dir,
|
||||
const char* link_rel, const char* name, ScannerEntry* entry);
|
||||
|
||||
/* scanner.c */
|
||||
bool scanner_capture_dir_time(ArrayList* dir_entries, mtx_t* mutex, const char* root_path,
|
||||
const char* fs_path, bool relative_mode, const char* relative_prefix,
|
||||
bool preserve_atimes, bool preserve_crtimes, bool preserve_xattrs,
|
||||
bool preserve_acls, bool no_implied_dirs,
|
||||
const FileListSet* file_list);
|
||||
|
||||
#endif
|
||||
@@ -0,0 +1,703 @@
|
||||
#include "log.h"
|
||||
#include "scanner.h"
|
||||
#include "scanner_internal.h"
|
||||
#include "array_list.h"
|
||||
#include "chunk.h"
|
||||
#include "file.h"
|
||||
#include "queue.h"
|
||||
#include "utils.h"
|
||||
#include <dirent.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <sys/stat.h>
|
||||
#include <sys/sysmacros.h>
|
||||
#include <threads.h>
|
||||
#include <unistd.h>
|
||||
#include <limits.h>
|
||||
|
||||
#include "xattr.h"
|
||||
|
||||
typedef struct {
|
||||
ParallelScanner* ps;
|
||||
char** dirs;
|
||||
int dir_count;
|
||||
char* root_dir; /* the transfer root, for relative-path computation */
|
||||
ScannerOptions options;
|
||||
ProtocolSession* allocation_session;
|
||||
} ParallelWorkerArg;
|
||||
|
||||
static int parallel_worker_thread(void* arg) {
|
||||
ParallelWorkerArg* wa = (ParallelWorkerArg*)arg;
|
||||
ProtocolSession* allocation_session = wa->allocation_session;
|
||||
if (allocation_session)
|
||||
protocol_session_bind(allocation_session);
|
||||
for (int i = 0; i < wa->dir_count; i++) {
|
||||
DirectoryScanner* ds = directory_scanner_create_with_options(wa->dirs[i], &wa->options);
|
||||
if (!ds) {
|
||||
mtx_lock(&wa->ps->result_mutex);
|
||||
wa->ps->failed = true;
|
||||
atomic_store(&wa->ps->cancelled, true);
|
||||
cnd_broadcast(&wa->ps->result_not_empty);
|
||||
cnd_broadcast(&wa->ps->result_not_full);
|
||||
mtx_unlock(&wa->ps->result_mutex);
|
||||
for (int j = i; j < wa->dir_count; j++)
|
||||
free(wa->dirs[j]);
|
||||
break;
|
||||
}
|
||||
/* Root .rsync-filter rules (parsed by the parallel scanner) apply to the
|
||||
* contents of every assigned subdirectory. Relative paths (used by the
|
||||
* allow-set and per-directory rules) are computed against the transfer
|
||||
* root, not the subdirectory the worker is seeded with. Exclusion
|
||||
* recording shares one caller-owned list across the workers. */
|
||||
free(ds->root_path);
|
||||
ds->root_path = str_dup(wa->root_dir);
|
||||
ds->seed_node = wa->ps->root_filter_node;
|
||||
ds->options.excluded_mutex = &wa->ps->result_mutex;
|
||||
Chunk* chunk;
|
||||
while ((chunk = directory_scanner_next(ds)) != NULL) {
|
||||
if (!queue_enqueue_multithreaded_cancel(wa->ps->result_queue, chunk, &wa->ps->result_mutex,
|
||||
&wa->ps->result_not_empty, &wa->ps->result_not_full,
|
||||
&wa->ps->cancelled)) {
|
||||
chunk_destroy(chunk);
|
||||
break;
|
||||
}
|
||||
}
|
||||
if (directory_scanner_failed(ds)) {
|
||||
mtx_lock(&wa->ps->result_mutex);
|
||||
wa->ps->failed = true;
|
||||
atomic_store(&wa->ps->cancelled, true);
|
||||
cnd_broadcast(&wa->ps->result_not_empty);
|
||||
cnd_broadcast(&wa->ps->result_not_full);
|
||||
mtx_unlock(&wa->ps->result_mutex);
|
||||
} else if (directory_scanner_had_io_error(ds)) {
|
||||
/* --ignore-errors path: an unreadable directory was skipped, not fatal. */
|
||||
mtx_lock(&wa->ps->result_mutex);
|
||||
wa->ps->io_error = true;
|
||||
mtx_unlock(&wa->ps->result_mutex);
|
||||
}
|
||||
directory_scanner_destroy(ds);
|
||||
free(wa->dirs[i]);
|
||||
}
|
||||
ParallelScanner* ps = wa->ps;
|
||||
free(wa->root_dir);
|
||||
free(wa->dirs);
|
||||
free(wa);
|
||||
mtx_lock(&ps->result_mutex);
|
||||
ps->completed++;
|
||||
if (ps->completed >= ps->expected_threads) {
|
||||
ps->done = true;
|
||||
cnd_signal(&ps->result_not_empty);
|
||||
}
|
||||
mtx_unlock(&ps->result_mutex);
|
||||
if (allocation_session)
|
||||
protocol_session_unbind();
|
||||
return thrd_success;
|
||||
}
|
||||
|
||||
static void parallel_scanner_creation_failed(ParallelScanner* ps) {
|
||||
mtx_lock(&ps->result_mutex);
|
||||
ps->failed = true;
|
||||
atomic_store(&ps->cancelled, true);
|
||||
ps->expected_threads = ps->created_threads;
|
||||
if (ps->completed >= ps->expected_threads)
|
||||
ps->done = true;
|
||||
cnd_broadcast(&ps->result_not_empty);
|
||||
cnd_broadcast(&ps->result_not_full);
|
||||
mtx_unlock(&ps->result_mutex);
|
||||
}
|
||||
|
||||
/* Initialize result queue and synchronization primitives. Returns true on success. */
|
||||
static bool parallel_scanner_init(ParallelScanner* ps) {
|
||||
ps->result_queue = queue_create(100, chunk_destroy);
|
||||
if (!ps->result_queue)
|
||||
return false;
|
||||
atomic_init(&ps->cancelled, false);
|
||||
int init = 0;
|
||||
bool ok = true;
|
||||
if (mtx_init(&ps->result_mutex, mtx_plain) != thrd_success)
|
||||
ok = false;
|
||||
if (ok) {
|
||||
init++;
|
||||
if (cnd_init(&ps->result_not_empty) != thrd_success)
|
||||
ok = false;
|
||||
}
|
||||
if (ok) {
|
||||
// cppcheck-suppress unreadVariable
|
||||
init++;
|
||||
if (cnd_init(&ps->result_not_full) != thrd_success)
|
||||
ok = false;
|
||||
}
|
||||
if (!ok) {
|
||||
if (init >= 3)
|
||||
cnd_destroy(&ps->result_not_full);
|
||||
if (init >= 2)
|
||||
cnd_destroy(&ps->result_not_empty);
|
||||
if (init >= 1)
|
||||
mtx_destroy(&ps->result_mutex);
|
||||
queue_destroy(ps->result_queue);
|
||||
ps->result_queue = NULL;
|
||||
return false;
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
/* Split files into chunks of roughly chunk_size bytes. Returns the first chunk (also stored
|
||||
* chunks beyond the first are enqueued on `queue`). Nulls out consumed entries in `files`.
|
||||
* Sets *failed on allocation/enqueue errors. */
|
||||
static Chunk* batch_files(ArrayList* files, unsigned long long chunk_size, Queue* queue,
|
||||
bool* failed) {
|
||||
Chunk* first = NULL;
|
||||
if (files->size <= 0)
|
||||
return NULL;
|
||||
ArrayList* batch = array_list_create(NULL);
|
||||
if (!batch) {
|
||||
*failed = true;
|
||||
return NULL;
|
||||
}
|
||||
unsigned long long batch_size = 0;
|
||||
for (int i = 0; i < files->size; i++) {
|
||||
File* f = (File*)files->items[i];
|
||||
if (!array_list_add(batch, f)) {
|
||||
*failed = true;
|
||||
break;
|
||||
}
|
||||
batch_size += f->data->size;
|
||||
if (batch_size >= chunk_size || i == files->size - 1) {
|
||||
void** items = array_list_to_array(batch);
|
||||
if (!items) {
|
||||
*failed = true;
|
||||
array_list_delete(batch);
|
||||
batch = NULL;
|
||||
break;
|
||||
}
|
||||
Chunk* c = chunk_create((File**)items, batch->size);
|
||||
free(items);
|
||||
if (!c) {
|
||||
*failed = true;
|
||||
array_list_delete(batch);
|
||||
batch = NULL;
|
||||
break;
|
||||
}
|
||||
int batch_start = i - batch->size + 1;
|
||||
for (int j = batch_start; j <= i; j++)
|
||||
files->items[j] = NULL;
|
||||
batch->item_destroyer = NULL;
|
||||
array_list_delete(batch);
|
||||
batch = NULL;
|
||||
if (!first) {
|
||||
first = c;
|
||||
} else {
|
||||
if (!queue_enqueue(queue, c)) {
|
||||
chunk_destroy(c);
|
||||
*failed = true;
|
||||
}
|
||||
}
|
||||
if (i < files->size - 1) {
|
||||
batch = array_list_create(NULL);
|
||||
if (!batch) {
|
||||
*failed = true;
|
||||
break;
|
||||
}
|
||||
batch_size = 0;
|
||||
}
|
||||
}
|
||||
}
|
||||
if (batch) {
|
||||
batch->item_destroyer = NULL;
|
||||
array_list_delete(batch);
|
||||
}
|
||||
return first;
|
||||
}
|
||||
|
||||
/* Scan one root-directory entry into either the subdirs or files list. */
|
||||
static void scan_root_entry(const ScannerOptions* options, const FilterNode* root_node,
|
||||
const char* root_directory, const struct dirent* entry,
|
||||
ArrayList* root_files, ArrayList* subdirs, dev_t root_dev,
|
||||
ParallelScanner* ps) {
|
||||
ScannerEntry inspected;
|
||||
int inspection =
|
||||
scanner_inspect_entry(options, root_directory, entry->d_name, entry->d_name, &inspected);
|
||||
if (inspection < 0) {
|
||||
ps->failed = true;
|
||||
return;
|
||||
}
|
||||
if (inspection == 0) {
|
||||
if (inspected.referent_error)
|
||||
ps->io_error = true;
|
||||
ArrayList* sink = NULL;
|
||||
if (inspected.excluded)
|
||||
sink = inspected.size_excluded ? options->size_skipped_paths : options->excluded_paths;
|
||||
if (sink) {
|
||||
/* A root-level prune protects the destination mirror of the entry's wire
|
||||
path: under -R + --files-from that is the bare relative name, otherwise
|
||||
it is the full source path with a leading '/' removed (matching the
|
||||
send_path/file_wire_path the scanner hands the sender). */
|
||||
if (options->relative && options->file_list != NULL) {
|
||||
if (!excluded_sink_append(sink, options->excluded_mutex, entry->d_name))
|
||||
ps->failed = true;
|
||||
} else if (options->relative_prefix) {
|
||||
char* wrel = scanner_prefix_send_path(options->relative_prefix, entry->d_name);
|
||||
if (!wrel) {
|
||||
ps->failed = true;
|
||||
} else {
|
||||
if (!excluded_sink_append(sink, options->excluded_mutex, wrel))
|
||||
ps->failed = true;
|
||||
free(wrel);
|
||||
}
|
||||
} else {
|
||||
char* abs_path = path_cat(root_directory, entry->d_name);
|
||||
if (!abs_path) {
|
||||
ps->failed = true;
|
||||
} else {
|
||||
const char* rel = *abs_path == '/' ? abs_path + 1 : abs_path;
|
||||
if (!excluded_sink_append(sink, options->excluded_mutex, rel))
|
||||
ps->failed = true;
|
||||
free(abs_path);
|
||||
}
|
||||
}
|
||||
}
|
||||
return;
|
||||
}
|
||||
char* cur_path = inspected.path;
|
||||
struct stat st = inspected.stats;
|
||||
bool is_dir = inspected.is_directory;
|
||||
char* rel = str_dup(entry->d_name);
|
||||
if (!rel) {
|
||||
free(cur_path);
|
||||
ps->failed = true;
|
||||
return;
|
||||
}
|
||||
bool protect = false;
|
||||
bool passes = entry_passes_selection(options->file_list, options->base_filters, root_node, rel,
|
||||
entry->d_name, is_dir, options->per_dir_filters,
|
||||
options->exclude_per_dir_filter_files, &protect);
|
||||
/* -R + --files-from: root-level files keep their bare relative send path. */
|
||||
bool use_rel = options->relative && options->file_list != NULL;
|
||||
if (!passes || protect) {
|
||||
/* --files-from subset pruning is not a filter exclusion; -R bare-wire-path
|
||||
exclusions are never recorded (see ScannerOptions.excluded_paths). */
|
||||
bool files_from_prune = options->file_list && !file_list_affects(options->file_list, rel);
|
||||
if ((!files_from_prune && !use_rel) || protect) {
|
||||
const char* rel_path;
|
||||
char* prefixed = NULL;
|
||||
if (use_rel) {
|
||||
/* -R + --files-from: the destination/wire path is the bare relative
|
||||
name, not the source path. */
|
||||
rel_path = rel;
|
||||
} else if (options->relative_prefix) {
|
||||
prefixed = scanner_prefix_send_path(options->relative_prefix, entry->d_name);
|
||||
if (!prefixed) {
|
||||
free(rel);
|
||||
free(cur_path);
|
||||
ps->failed = true;
|
||||
return;
|
||||
}
|
||||
rel_path = prefixed;
|
||||
} else {
|
||||
rel_path = *cur_path == '/' ? cur_path + 1 : cur_path;
|
||||
}
|
||||
if (options->excluded_paths &&
|
||||
!excluded_sink_append(options->excluded_paths, options->excluded_mutex, rel_path))
|
||||
ps->failed = true;
|
||||
free(prefixed);
|
||||
}
|
||||
if (!passes) {
|
||||
scanner_note_filter(options, entry->d_name);
|
||||
free(rel);
|
||||
free(cur_path);
|
||||
return;
|
||||
}
|
||||
}
|
||||
if (is_dir) {
|
||||
if (!scanner_same_filesystem(options->one_file_system, root_dev, st.st_dev)) {
|
||||
if (options->one_file_system > 1) {
|
||||
/* -xx: drop the mount-point directory entirely (rsync) and print the
|
||||
--info=mount line when enabled. */
|
||||
scanner_note_mount(options, cur_path);
|
||||
free(rel);
|
||||
free(cur_path);
|
||||
return;
|
||||
}
|
||||
/* -x/--one-file-system: emit the mount-point directory entry (empty) but
|
||||
do not descend into it (see the sequential scanner for the same rule). */
|
||||
File* mount = file_create(cur_path);
|
||||
free(cur_path);
|
||||
if (mount == NULL) {
|
||||
free(rel);
|
||||
ps->failed = true;
|
||||
return;
|
||||
}
|
||||
mount->is_dir = true;
|
||||
if (options->use_metadata) {
|
||||
mount->metadata = file_metadata_create(mount->path, &st, options->preserve_atimes,
|
||||
options->preserve_crtimes);
|
||||
if (!mount->metadata) {
|
||||
free(rel);
|
||||
file_destroy(mount);
|
||||
ps->failed = true;
|
||||
return;
|
||||
}
|
||||
}
|
||||
if (options->relative_prefix) {
|
||||
mount->send_path = scanner_prefix_send_path(options->relative_prefix, rel);
|
||||
if (!mount->send_path) {
|
||||
free(rel);
|
||||
file_destroy(mount);
|
||||
ps->failed = true;
|
||||
return;
|
||||
}
|
||||
}
|
||||
free(rel);
|
||||
if (!array_list_add(root_files, mount)) {
|
||||
file_destroy(mount);
|
||||
ps->failed = true;
|
||||
}
|
||||
return;
|
||||
}
|
||||
free(rel);
|
||||
if (!array_list_add(subdirs, cur_path)) {
|
||||
free(cur_path);
|
||||
ps->failed = true;
|
||||
}
|
||||
return;
|
||||
}
|
||||
File* file = file_create(cur_path);
|
||||
free(cur_path);
|
||||
if (!file) {
|
||||
free(rel);
|
||||
free(inspected.link_target);
|
||||
inspected.link_target = NULL;
|
||||
ps->failed = true;
|
||||
return;
|
||||
}
|
||||
if (inspected.is_symlink) {
|
||||
file->is_symlink = true;
|
||||
file->symlink_target = inspected.link_target;
|
||||
inspected.link_target = NULL;
|
||||
} else {
|
||||
file->data->size = st.st_size;
|
||||
}
|
||||
if (use_rel) {
|
||||
file->send_path = rel;
|
||||
rel = NULL;
|
||||
} else if (options->relative_prefix) {
|
||||
file->send_path = scanner_prefix_send_path(options->relative_prefix, rel);
|
||||
free(rel);
|
||||
rel = NULL;
|
||||
if (!file->send_path) {
|
||||
file_destroy(file);
|
||||
ps->failed = true;
|
||||
return;
|
||||
}
|
||||
}
|
||||
ScannerSpecial special = scanner_prepare_special(
|
||||
options->preserve_devices, options->preserve_specials, options->copy_devices, file, &st);
|
||||
if (special == SCANNER_SPECIAL_SKIP) {
|
||||
scanner_note_nonreg(ps->options, file->path);
|
||||
free(rel);
|
||||
file_destroy(file);
|
||||
return;
|
||||
}
|
||||
if (options->hardlinks && S_ISREG(st.st_mode)) {
|
||||
int gid;
|
||||
bool is_first;
|
||||
char* first_path = NULL;
|
||||
if (!hardlink_table_assign((HardLinkTable*)options->hardlinks, file_wire_path(file), st.st_dev,
|
||||
st.st_ino, &gid, &is_first, &first_path)) {
|
||||
ps->failed = true;
|
||||
} else {
|
||||
file->link_group = gid;
|
||||
file->link_first = is_first;
|
||||
if (!is_first) {
|
||||
file->hardlink_target = first_path;
|
||||
file->data->size = 0;
|
||||
} else {
|
||||
free(first_path);
|
||||
}
|
||||
}
|
||||
}
|
||||
if (options->use_metadata)
|
||||
file->metadata =
|
||||
file_metadata_create(file->path, &st, options->preserve_atimes, options->preserve_crtimes);
|
||||
if (options->use_metadata && !file->metadata) {
|
||||
free(rel);
|
||||
file_destroy(file);
|
||||
ps->failed = true;
|
||||
return;
|
||||
}
|
||||
if ((options->preserve_xattrs || options->preserve_acls) &&
|
||||
!(file->link_group != 0 && !file->link_first))
|
||||
file->xattrs = xattr_capture_path(file->path, options->preserve_acls);
|
||||
if (!array_list_add(root_files, file)) {
|
||||
free(rel);
|
||||
file_destroy(file);
|
||||
ps->failed = true;
|
||||
return;
|
||||
}
|
||||
free(rel);
|
||||
}
|
||||
|
||||
/* Scan the root directory itself, collecting root files and subdirectories.
|
||||
* Returns false if the root directory could not be opened. */
|
||||
static bool scan_root_directory(ParallelScanner* ps, const char* root_directory,
|
||||
const ScannerOptions* options, const FilterNode* root_node,
|
||||
dev_t root_dev, ArrayList* root_files, ArrayList* subdirs) {
|
||||
DIR* dir = opendir(root_directory);
|
||||
if (!dir) {
|
||||
log_perror("Could not open root directory for parallel scan");
|
||||
return false;
|
||||
}
|
||||
/* The parallel scanner opens the transfer root directly (not through
|
||||
open_next_directory), so record it as synchronized here. */
|
||||
if (!scanner_record_synced_dir(options, root_directory, "",
|
||||
options->relative && options->file_list != NULL)) {
|
||||
closedir(dir);
|
||||
ps->failed = true;
|
||||
return false;
|
||||
}
|
||||
log_debug_message(LOG_DEBUG_FLIST, "flist: scanning %s", root_directory);
|
||||
const struct dirent* entry;
|
||||
while ((entry = readdir(dir)) != NULL) {
|
||||
if (strcmp(entry->d_name, ".") == 0 || strcmp(entry->d_name, "..") == 0)
|
||||
continue;
|
||||
scan_root_entry(options, root_node, root_directory, entry, root_files, subdirs, root_dev, ps);
|
||||
}
|
||||
closedir(dir);
|
||||
return true;
|
||||
}
|
||||
|
||||
/* Spawn worker threads, one per group of subdirectories. */
|
||||
static void spawn_parallel_workers(ParallelScanner* ps, ArrayList* subdirs,
|
||||
const ScannerOptions* options, const char* root_directory,
|
||||
unsigned long long cs) {
|
||||
if (subdirs->size <= 0)
|
||||
return;
|
||||
int n = options->num_threads > 0 ? options->num_threads : 4;
|
||||
if (n > subdirs->size)
|
||||
n = subdirs->size;
|
||||
|
||||
ps->num_threads = n;
|
||||
ps->expected_threads = n;
|
||||
ps->threads = calloc(n, sizeof(thrd_t));
|
||||
if (!ps->threads) {
|
||||
ps->num_threads = 0;
|
||||
ps->expected_threads = 0;
|
||||
ps->failed = true;
|
||||
return;
|
||||
}
|
||||
int dirs_per_thread = subdirs->size / n;
|
||||
int remainder = subdirs->size % n;
|
||||
int start = 0;
|
||||
ps->num_threads = 0;
|
||||
for (int t = 0; t < n; t++) {
|
||||
int count = dirs_per_thread + (t < remainder ? 1 : 0);
|
||||
if (count == 0)
|
||||
break;
|
||||
ParallelWorkerArg* wa = calloc(1, sizeof(ParallelWorkerArg));
|
||||
if (!wa) {
|
||||
parallel_scanner_creation_failed(ps);
|
||||
break;
|
||||
}
|
||||
wa->ps = ps;
|
||||
wa->dirs = calloc(count, sizeof(char*));
|
||||
wa->root_dir = str_dup(root_directory);
|
||||
if (!wa->dirs || !wa->root_dir) {
|
||||
free(wa->root_dir);
|
||||
free(wa->dirs);
|
||||
free(wa);
|
||||
parallel_scanner_creation_failed(ps);
|
||||
break;
|
||||
}
|
||||
bool dup_ok = true;
|
||||
for (int j = 0; j < count; j++) {
|
||||
wa->dirs[j] = str_dup((char*)subdirs->items[start + j]);
|
||||
if (!wa->dirs[j])
|
||||
dup_ok = false;
|
||||
}
|
||||
if (!dup_ok) {
|
||||
for (int j = 0; j < count; j++)
|
||||
free(wa->dirs[j]);
|
||||
free(wa->root_dir);
|
||||
free(wa->dirs);
|
||||
free(wa);
|
||||
parallel_scanner_creation_failed(ps);
|
||||
break;
|
||||
}
|
||||
wa->dir_count = count;
|
||||
wa->options = *options;
|
||||
wa->options.chunk_size = cs;
|
||||
wa->allocation_session = ps->allocation_session;
|
||||
start += count;
|
||||
if (thrd_create(&ps->threads[t], parallel_worker_thread, wa) != thrd_success) {
|
||||
for (int j = 0; j < count; j++)
|
||||
free(wa->dirs[j]);
|
||||
free(wa->root_dir);
|
||||
free(wa->dirs);
|
||||
free(wa);
|
||||
parallel_scanner_creation_failed(ps);
|
||||
break;
|
||||
}
|
||||
ps->num_threads++;
|
||||
ps->created_threads++;
|
||||
}
|
||||
}
|
||||
|
||||
ParallelScanner* parallel_scanner_create_with_options(const char* root_directory,
|
||||
const ScannerOptions* options,
|
||||
ProtocolSession* allocation_session) {
|
||||
if (!root_directory || !options)
|
||||
return NULL;
|
||||
ParallelScanner* ps = calloc(1, sizeof(ParallelScanner));
|
||||
if (!ps)
|
||||
return NULL;
|
||||
if (!parallel_scanner_init(ps)) {
|
||||
free(ps);
|
||||
return NULL;
|
||||
}
|
||||
ps->allocation_session = allocation_session;
|
||||
ps->options = options;
|
||||
|
||||
ArrayList* root_files = array_list_create(file_destroy);
|
||||
ArrayList* subdirs = array_list_create(free);
|
||||
if (!root_files || !subdirs) {
|
||||
array_list_delete(root_files);
|
||||
array_list_delete(subdirs);
|
||||
parallel_scanner_destroy(ps);
|
||||
return NULL;
|
||||
}
|
||||
|
||||
dev_t root_dev = 0;
|
||||
if (options->one_file_system) {
|
||||
struct stat root_stats;
|
||||
if (stat(root_directory, &root_stats) != 0) {
|
||||
log_perror("Could not stat source directory");
|
||||
array_list_delete(root_files);
|
||||
array_list_delete(subdirs);
|
||||
parallel_scanner_destroy(ps);
|
||||
return NULL;
|
||||
}
|
||||
root_dev = root_stats.st_dev;
|
||||
}
|
||||
|
||||
/* Build the root directory's per-directory filter context once; workers seed
|
||||
* their scanners with it so per-dir rules behave identically to the sequential
|
||||
* scanner. */
|
||||
FilterNode* root_node = NULL;
|
||||
{
|
||||
char err[256];
|
||||
bool any_exists = false;
|
||||
FilterRuleList* own =
|
||||
read_dir_filters(options, root_directory, "", &any_exists, err, sizeof(err));
|
||||
if (!own) {
|
||||
/* A parse/allocation failure must fail the scan even when an earlier
|
||||
merge file in the same directory existed (see the sequential scanner). */
|
||||
if (err[0] != '\0') {
|
||||
log_message(LOG_LEVEL_ERROR, "invalid per-directory filter in %s: %s", root_directory, err);
|
||||
array_list_delete(root_files);
|
||||
array_list_delete(subdirs);
|
||||
parallel_scanner_destroy(ps);
|
||||
return NULL;
|
||||
}
|
||||
/* no files exist: leave root_node NULL */
|
||||
} else if (any_exists && (own->count > 0 || own->dir_merge_count > 0)) {
|
||||
root_node = filter_node_alloc(NULL, own);
|
||||
if (!root_node) {
|
||||
filter_rule_list_free(own);
|
||||
array_list_delete(root_files);
|
||||
array_list_delete(subdirs);
|
||||
parallel_scanner_destroy(ps);
|
||||
return NULL;
|
||||
}
|
||||
} else {
|
||||
filter_rule_list_free(own);
|
||||
}
|
||||
}
|
||||
ps->root_filter_node = root_node;
|
||||
|
||||
if (!scan_root_directory(ps, root_directory, options, root_node, root_dev, root_files, subdirs)) {
|
||||
array_list_delete(root_files);
|
||||
array_list_delete(subdirs);
|
||||
parallel_scanner_destroy(ps);
|
||||
return NULL;
|
||||
}
|
||||
/* The root itself is a traversed directory (rsync counts it in
|
||||
`Number of files`); the worker DirectoryScanners account for every
|
||||
subdirectory below it. */
|
||||
scanner_dir_count_count(options);
|
||||
/* P7 Wave D: the parallel scanner never runs a DirectoryScanner over the
|
||||
transfer root itself (it hands the root's immediate subdirectories to
|
||||
workers), so capture the root's directory time here. */
|
||||
if (options->capture_dir_times &&
|
||||
!scanner_capture_dir_time(
|
||||
options->dir_entries, options->dir_entries_mutex, root_directory, root_directory,
|
||||
options->relative && options->file_list != NULL, options->relative_prefix,
|
||||
options->preserve_atimes, options->preserve_crtimes, options->preserve_xattrs,
|
||||
options->preserve_acls, options->no_implied_dirs, options->file_list)) {
|
||||
array_list_delete(root_files);
|
||||
array_list_delete(subdirs);
|
||||
parallel_scanner_destroy(ps);
|
||||
return NULL;
|
||||
}
|
||||
|
||||
unsigned long long cs = options->chunk_size > 0 ? options->chunk_size : DESIRED_CHUNK_SIZE;
|
||||
ps->initial_chunk = batch_files(root_files, cs, ps->result_queue, &ps->failed);
|
||||
array_list_delete(root_files);
|
||||
|
||||
spawn_parallel_workers(ps, subdirs, options, root_directory, cs);
|
||||
array_list_delete(subdirs);
|
||||
return ps;
|
||||
}
|
||||
|
||||
Chunk* parallel_scanner_next(ParallelScanner* ps) {
|
||||
if (ps->initial_chunk) {
|
||||
Chunk* c = ps->initial_chunk;
|
||||
ps->initial_chunk = NULL;
|
||||
return c;
|
||||
}
|
||||
if (ps->num_threads == 0) {
|
||||
mtx_lock(&ps->result_mutex);
|
||||
if (!queue_is_empty(ps->result_queue)) {
|
||||
Chunk* chunk = queue_dequeue(ps->result_queue);
|
||||
mtx_unlock(&ps->result_mutex);
|
||||
return chunk;
|
||||
}
|
||||
ps->done = true;
|
||||
mtx_unlock(&ps->result_mutex);
|
||||
return NULL;
|
||||
}
|
||||
Chunk* chunk = queue_dequeue_multithreaded(
|
||||
ps->result_queue, &ps->result_mutex, &ps->result_not_empty, &ps->result_not_full, &ps->done);
|
||||
return chunk;
|
||||
}
|
||||
|
||||
bool parallel_scanner_failed(const ParallelScanner* ps) {
|
||||
return ps == NULL || ps->failed;
|
||||
}
|
||||
|
||||
bool parallel_scanner_had_io_error(const ParallelScanner* ps) {
|
||||
return ps != NULL && ps->io_error;
|
||||
}
|
||||
|
||||
void parallel_scanner_destroy(ParallelScanner* ps) {
|
||||
if (!ps)
|
||||
return;
|
||||
mtx_lock(&ps->result_mutex);
|
||||
ps->done = true;
|
||||
atomic_store(&ps->cancelled, true);
|
||||
cnd_broadcast(&ps->result_not_empty);
|
||||
cnd_broadcast(&ps->result_not_full);
|
||||
mtx_unlock(&ps->result_mutex);
|
||||
for (int i = 0; i < ps->num_threads; i++)
|
||||
thrd_join(ps->threads[i], NULL);
|
||||
free(ps->threads);
|
||||
if (ps->root_filter_node)
|
||||
filter_node_destroy(ps->root_filter_node);
|
||||
if (ps->initial_chunk)
|
||||
chunk_destroy(ps->initial_chunk);
|
||||
queue_destroy(ps->result_queue);
|
||||
mtx_destroy(&ps->result_mutex);
|
||||
cnd_destroy(&ps->result_not_empty);
|
||||
cnd_destroy(&ps->result_not_full);
|
||||
free(ps);
|
||||
}
|
||||
+41
-21
@@ -19,7 +19,9 @@ void print_usage(void) {
|
||||
printf("\n");
|
||||
printf("Options:\n");
|
||||
printf(" -c, --checksum Verify content by checksum instead of size+mtime\n");
|
||||
printf(" -z, --compress [level] Enable compression (level 1-22, default 5)\n");
|
||||
printf(" -z, --compress [level] Enable compression. The default level is\n");
|
||||
printf(" per-codec: zstd 3 (range 1-22), zlib/zlibx 6, lz4\n");
|
||||
printf(" ignores the level\n");
|
||||
printf(" -a, --archive rsync archive mode (-rlptgoD): links, perms, times,\n");
|
||||
printf(" owner, group, devices and specials; not\n");
|
||||
printf(" compression/multithreading\n");
|
||||
@@ -36,8 +38,8 @@ void print_usage(void) {
|
||||
printf(" arguments, e.g. -e \"ssh -p 2222\"\n");
|
||||
printf(" --rsync-path <path> Alias for --fastsync-server-path (path to the\n");
|
||||
printf(" fastsync server binary on the remote side)\n");
|
||||
printf(" --blocking-io Leave the SSH transport socket without read/write\n");
|
||||
printf(" timeouts so it blocks naturally\n");
|
||||
printf(" --blocking-io SSH transport only: leave the socket without read/write\n");
|
||||
printf(" timeouts so it blocks naturally (no effect on TCP)\n");
|
||||
printf(" --outbuf=MODE stdout/stderr buffering: N (none/unbuffered),\n");
|
||||
printf(" L (line-buffered), or B (block-buffered, default)\n");
|
||||
printf(" --progress Show transfer progress\n");
|
||||
@@ -49,6 +51,7 @@ void print_usage(void) {
|
||||
printf(" converted before transmission and back on receipt; a\n");
|
||||
printf(" name that cannot be represented in the target charset\n");
|
||||
printf(" fails that transfer cleanly (rsync-compatible)\n");
|
||||
printf(" --no-iconv Disable --iconv charset conversion (same as --iconv=-)\n");
|
||||
printf(" --protocol=NUM Force the wire protocol version (must equal the current\n");
|
||||
printf(" PROTOCOL_VERSION; FastSync cannot speak older/virtual\n");
|
||||
printf(" wire formats)\n");
|
||||
@@ -62,8 +65,7 @@ void print_usage(void) {
|
||||
printf(" NOTE: the FastSync batch format is NOT interoperable with rsync's batch\n");
|
||||
printf(" files (different container format); do not mix the two tools.\n");
|
||||
printf(" --delete Delete files on receiver not in source\n");
|
||||
printf(" (default timing: delete only after the whole\n");
|
||||
printf(" transfer has succeeded)\n");
|
||||
printf(" (default timing: delete-during, like rsync --del)\n");
|
||||
printf(" --delete-before Delete extras before the transfer starts\n");
|
||||
printf(" (implies --delete)\n");
|
||||
printf(" --delete-during Delete a directory's extras as that directory is\n");
|
||||
@@ -72,7 +74,9 @@ void print_usage(void) {
|
||||
printf(" --delete-delay Record the extras during the scan but remove them\n");
|
||||
printf(" only after a successful transfer (implies --delete)\n");
|
||||
printf(" --delete-after Delete only after the whole transfer succeeded\n");
|
||||
printf(" (the default --delete timing; implies --delete)\n");
|
||||
printf(" (implies --delete)\n");
|
||||
printf(" --delete-commit FastSync-only: restore the late whole-tree commit\n");
|
||||
printf(" (identical to --delete-after; implies --delete)\n");
|
||||
printf(" --delete-excluded Also delete destination files that were excluded on\n");
|
||||
printf(" the source (default protects them, matching rsync)\n");
|
||||
printf(" --max-delete=NUM Delete at most NUM destination entries per run; if the\n");
|
||||
@@ -90,10 +94,11 @@ void print_usage(void) {
|
||||
printf(" entry's destination mirror receiver-side. Independent of\n");
|
||||
printf(" --delete (it does not imply --delete; a non-empty directory\n");
|
||||
printf(" mirror is removed only with --force or --delete)\n");
|
||||
printf(" -m, --prune-empty-dirs Do not transfer empty directory entries (--dirs mode);\n");
|
||||
printf(" recursive transfers never send empty dirs\n");
|
||||
printf(" -m, --prune-empty-dirs Do not create empty directories (a recursive transfer\n");
|
||||
printf(" otherwise recreates them, like rsync)\n");
|
||||
printf(" Note: each timing flag implies --delete. Combining a timing flag with\n");
|
||||
printf(" --no-delete (in either order) is rejected as a config error.\n");
|
||||
printf(" --no-delete (in either order) is rejected as a config error, as is more\n");
|
||||
printf(" than one timing flag.\n");
|
||||
printf(" --ignore-existing Skip files that already exist on receiver\n");
|
||||
printf(" --delay-updates Put updated files into place only at the end of transfer\n");
|
||||
printf(" --dirs, -d, --old-dirs, --old-d Transfer the named directory entries without\n");
|
||||
@@ -103,8 +108,9 @@ void print_usage(void) {
|
||||
printf(" -R, --relative With --files-from, preserve each listed entry's relative path\n");
|
||||
printf(" below the destination root instead of mirroring the full\n");
|
||||
printf(" source path (no effect without --files-from)\n");
|
||||
printf(" --no-implied-dirs With -R --files-from, refuse to place a listed file whose\n");
|
||||
printf(" parent directory is not itself listed\n");
|
||||
printf(" --no-implied-dirs With -R, do not apply the source metadata of a listed file's\n");
|
||||
printf(" implied parent directories (they are still created with\n");
|
||||
printf(" default attributes)\n");
|
||||
printf(" --mkpath Create the destination root directory on the server when it\n");
|
||||
printf(" does not exist yet\n");
|
||||
printf(" --exclude <pattern>, --exclude=<pattern> Exclude files matching pattern\n");
|
||||
@@ -137,6 +143,9 @@ void print_usage(void) {
|
||||
printf(" into the destination instead of transferring its data\n");
|
||||
printf(" --link-dest <dir> Like --copy-dest, but hard-links the unchanged file from DIR\n");
|
||||
printf(" into the destination (repeatable; earlier DIRs win)\n");
|
||||
printf(" --verify-basis FastSync-only: require a basis hit's content to match the\n");
|
||||
printf(" source by whole-file digest instead of trusting rsync's\n");
|
||||
printf(" size+mtime (or --size-only) quick-check\n");
|
||||
printf(" --checksum-choice, --cc <alg> Whole-file checksum algorithm for --incremental/\n");
|
||||
printf(" --checksum compares. Accepted: xxh128 (default), xxh3, xxh64\n");
|
||||
printf(" (aka xxhash), md5, md4, sha1, or none. A two-name\n");
|
||||
@@ -156,7 +165,7 @@ void print_usage(void) {
|
||||
printf(" --no-delta, or --no-incremental)\n");
|
||||
printf(" --no-fuzzy Disable --fuzzy\n");
|
||||
printf(" -B <n>, --block-size <n>, --delta-block <n>\n");
|
||||
printf(" Delta block size in bytes (default: %d)\n", DELTA_BLOCK_SIZE_DEFAULT);
|
||||
printf(" Delta block size in bytes (default: %u)\n", DELTA_BLOCK_SIZE_DEFAULT);
|
||||
printf(" --delta-max <n> Max file size for delta transfer (default: %llu)\n",
|
||||
DELTA_MAX_FILE_SIZE);
|
||||
printf(" -j, --threads[=N] Enable the multithreaded scanner/loader/sender\n");
|
||||
@@ -186,6 +195,10 @@ void print_usage(void) {
|
||||
printf(" -U, --atimes Preserve access times\n");
|
||||
printf(" -N, --crtimes Capture birth time; cannot be applied (documented\n");
|
||||
printf(" divergence)\n");
|
||||
printf(" -O, --omit-dir-times Do not apply modification times to directories\n");
|
||||
printf(" -J, --omit-link-times Do not apply times to symlinks\n");
|
||||
printf(" --open-noatime Open source files with O_NOATIME so reading for a\n");
|
||||
printf(" transfer does not update their access time\n");
|
||||
printf(" -X, --xattrs Preserve user extended attributes (user.* only;\n");
|
||||
printf(" privileged security.*/trusted.* namespaces are\n");
|
||||
printf(" never captured or applied)\n");
|
||||
@@ -243,7 +256,9 @@ void print_usage(void) {
|
||||
printf(" reusable digest is sent (keep the file mode 0600)\n");
|
||||
printf(" --no-motd Suppress display of the daemon's MOTD (the server\n");
|
||||
printf(" still sends it; the client just does not show it)\n");
|
||||
printf(" --bwlimit <KB/s> Bandwidth limit in kilobytes per second\n");
|
||||
printf(" --bwlimit=RATE Limit socket I/O bandwidth (default unit KiB/s,\n");
|
||||
printf(" rsync-style: 0 = no limit; K/M/G/T/P suffixes are\n");
|
||||
printf(" binary, KB/MB decimal, KiB/MiB binary; decimals allowed)\n");
|
||||
printf(" --tls Enable TLS encryption\n");
|
||||
printf(" --cert <path> TLS certificate file (PEM)\n");
|
||||
printf(" --key <path> TLS private key file (PEM)\n");
|
||||
@@ -279,8 +294,12 @@ void print_usage(void) {
|
||||
printf(" -x, --one-file-system Do not cross filesystem boundaries\n");
|
||||
printf(" --log-file <path>, --log-file=<path> Write log messages to file\n");
|
||||
printf(" --stderr=MODE Route logging to stderr: errors or all\n");
|
||||
printf(" --msgs2stderr Route all messages to stderr (deprecated spelling of\n");
|
||||
printf(" --stderr=all)\n");
|
||||
printf(" --no-msgs2stderr Select errors-only stderr (deprecated spelling; the\n");
|
||||
printf(" default)\n");
|
||||
printf(" --partial Keep partial files on interrupted transfer\n");
|
||||
printf(" --partial-dir <dir> Directory for partial files\n");
|
||||
printf(" --partial-dir <dir> Directory for partial files (implies --partial)\n");
|
||||
printf(" -T, --temp-dir <dir> Scratch dir for temp files before atomic install.\n");
|
||||
printf(" Confined to the receive root: a relative dir resolves below\n");
|
||||
printf(" it and an absolute/traversal dir is rejected. The dir must\n");
|
||||
@@ -329,7 +348,8 @@ void print_usage(void) {
|
||||
printf(" --append-verify Like --append, but verifies the retained prefix checksum\n");
|
||||
printf(" before appending (falls back to a full transfer on mismatch)\n");
|
||||
printf(" --fsync Fsync every written file before publication\n");
|
||||
printf(" --compress-level <n> Compression level (default: 5)\n");
|
||||
printf(" --compress-level <n> Compression level (per-codec default: zstd 3,\n");
|
||||
printf(" zlib/zlibx 6, lz4 ignores it)\n");
|
||||
printf(" --zl <n> Alias for --compress-level\n");
|
||||
printf(" --skip-compress=LIST Skip compression for suffixes in LIST (separated by\n");
|
||||
printf(" '/' as in rsync, or ','); a leading dot is optional. The\n");
|
||||
@@ -341,19 +361,19 @@ void print_usage(void) {
|
||||
}
|
||||
|
||||
void print_debug_usage(void) {
|
||||
printf("Emitting debug flags: IO,PROTO,PACK,UTIL,ALL,NONE\n");
|
||||
printf("Emitting debug flags: IO,PROTO,PACK,UTIL,FLIST,DEL,HASH,DELTASUM,\n");
|
||||
printf("RECV,FILTER,SEND,ALL,NONE\n");
|
||||
printf("Also accepted for rsync CLI parity (silent): ACL,BACKUP,BIND,CHDIR,\n");
|
||||
printf("CONNECT,CMD,DEL,DELTASUM,DUP,EXIT,FILTER,FLIST,FUZZY,GENR,HASH,HLINK,\n");
|
||||
printf("ICONV,NSTR,OWN,RECV,SEND,TIME.\n");
|
||||
printf("CONNECT,CMD,DUP,EXIT,FUZZY,GENR,HLINK,ICONV,NSTR,OWN,TIME.\n");
|
||||
printf("Flags may be comma-separated, for example: --debug=io,proto\n");
|
||||
printf("An optional level suffix is accepted (e.g. --debug=io2); level 0\n");
|
||||
printf("silences that item. Unknown names are rejected.\n");
|
||||
}
|
||||
|
||||
void print_info_usage(void) {
|
||||
printf("Emitting info flags: COPY,NAME,MISC,SKIP,STATS,ALL,NONE\n");
|
||||
printf("Also accepted for rsync CLI parity (silent): BACKUP,DEL,FLIST,MOUNT,\n");
|
||||
printf("NONREG,PROGRESS,REMOVE,SYMSAFE.\n");
|
||||
printf("Emitting info flags: COPY,MISC,SKIP,STATS,DEL,REMOVE,NAME,FLIST,\n");
|
||||
printf("NONREG,PROGRESS,MOUNT,ALL,NONE\n");
|
||||
printf("Also accepted for rsync CLI parity (silent): BACKUP,SYMS,SYMSAFE.\n");
|
||||
printf("Flags may be comma-separated, for example: --info=name,stats\n");
|
||||
printf("An optional level suffix is accepted (e.g. --info=stats2); level 0\n");
|
||||
printf("silences that item. Unknown names are rejected.\n");
|
||||
|
||||
+360
-193
@@ -61,13 +61,18 @@ bool receiver_send_final_success(int fd, const Config* config, const ReceiverOut
|
||||
}
|
||||
|
||||
bool receiver_send_stats_frame(int fd, const Config* config, const ReceiverStats* stats,
|
||||
const struct ArrayList* would_delete) {
|
||||
const struct ArrayList* would_delete,
|
||||
const struct ArrayList* deleted_paths) {
|
||||
if (!config->report_stats)
|
||||
return true;
|
||||
ReceiverStats local;
|
||||
memset(&local, 0, sizeof(local));
|
||||
const ReceiverStats* out = stats ? stats : &local;
|
||||
size_t count = would_delete ? (size_t)would_delete->size : 0;
|
||||
/* The path list carries the dry-run would-delete set for a -n run and the
|
||||
actually-removed set for a real --info=del run. */
|
||||
const struct ArrayList* paths =
|
||||
config->dry_run ? would_delete : (config->report_deletes ? deleted_paths : NULL);
|
||||
size_t count = paths ? (size_t)paths->size : 0;
|
||||
if (count > (size_t)MAX_MANIFEST_ENTRIES)
|
||||
count = MAX_MANIFEST_ENTRIES;
|
||||
ReceiverStats record = *out;
|
||||
@@ -76,7 +81,7 @@ bool receiver_send_stats_frame(int fd, const Config* config, const ReceiverStats
|
||||
!send_int(fd, (int)count))
|
||||
return false;
|
||||
for (size_t i = 0; i < count; i++) {
|
||||
const char* path = (const char*)would_delete->items[i];
|
||||
const char* path = (const char*)paths->items[i];
|
||||
if (!send_wire_str(fd, path ? path : ""))
|
||||
return false;
|
||||
}
|
||||
@@ -90,6 +95,25 @@ static void receiver_tally_deleted(const ReceiverSink* sink, size_t deleted) {
|
||||
sink->stats->deleted_files += deleted;
|
||||
}
|
||||
|
||||
/* Observer for --info=del: record each truly-removed destination-relative path
|
||||
in the ArrayList passed as the observer context, so the terminal STATUS_STATS
|
||||
frame can list it. A failed append is best-effort (the deletion already
|
||||
happened; output is cosmetic). Shared by the single-threaded receiver and
|
||||
the -m pipeline's deferred commit. */
|
||||
void receiver_record_deleted_path(void* context, const char* rel_path) {
|
||||
ArrayList* paths = context;
|
||||
if (!paths || !rel_path)
|
||||
return;
|
||||
/* Bound the retained list like the keep-set manifest: only MAX_MANIFEST_ENTRIES
|
||||
paths are ever transmitted in the terminal STATUS_STATS frame, so recording
|
||||
more only grows memory. A hostile/huge deletion set is therefore capped. */
|
||||
if ((size_t)paths->size >= (size_t)MAX_MANIFEST_ENTRIES)
|
||||
return;
|
||||
char* copy = str_dup(rel_path);
|
||||
if (copy && !array_list_add(paths, copy))
|
||||
free(copy);
|
||||
}
|
||||
|
||||
static bool receiver_process_chunk(Chunk* chunk, const ReceiverSink* sink) {
|
||||
if (!chunk || !sink || !sink->store_file)
|
||||
return false;
|
||||
@@ -279,16 +303,276 @@ int receiver_process(Config* config, int file_descriptor, const ReceiverSink* si
|
||||
return receiver_process_pending(config, file_descriptor, sink, NULL, NULL);
|
||||
}
|
||||
|
||||
/* Per-connection state threaded through the status handlers below. The parked
|
||||
keep-set / per-directory session live here so one teardown helper can release
|
||||
them on every exit path. */
|
||||
typedef struct {
|
||||
Config* config;
|
||||
int fd;
|
||||
const ReceiverSink* sink;
|
||||
DeleteManifest** pending_manifest;
|
||||
DeletePlanSession** pending_plans;
|
||||
/* Parked keep-set for the late/commit timing. Every exit path frees it
|
||||
exactly once; the only exception is the successful FINISHED handoff, which
|
||||
transfers ownership to *pending_manifest (used by the -m receiver). */
|
||||
DeleteManifest* deferred_manifest;
|
||||
/* Per-directory delete session for --delete-during/--delete-delay. During the
|
||||
loop it applies plans inline (during) or snapshots their extras (delay); on
|
||||
a successful FINISHED it is either committed here or handed to
|
||||
*pending_plans so the -m caller commits after its disk writer drained. */
|
||||
DeletePlanSession* plan_session;
|
||||
bool early_delete;
|
||||
bool per_dir_delete;
|
||||
bool delete_limit_noted;
|
||||
} ReceiverPendingState;
|
||||
|
||||
/* Outcome of one frame handler. NEXT reads the following status frame; FAIL
|
||||
tears the connection down without a peer STATUS_ERROR; ERROR tears it down
|
||||
and (when the sink owns error reporting) emits STATUS_ERROR. */
|
||||
typedef enum {
|
||||
RECEIVER_STEP_NEXT,
|
||||
RECEIVER_STEP_FAIL,
|
||||
RECEIVER_STEP_ERROR,
|
||||
} ReceiverStep;
|
||||
|
||||
static ReceiverStep receiver_handle_keepalive(ReceiverPendingState* state) {
|
||||
if (!send_status(state->fd, STATUS_KEEPALIVE))
|
||||
return RECEIVER_STEP_FAIL;
|
||||
return RECEIVER_STEP_NEXT;
|
||||
}
|
||||
|
||||
static ReceiverStep receiver_handle_abort(ReceiverPendingState* state) {
|
||||
(void)state;
|
||||
log_message(LOG_LEVEL_INFO, "Received abort from client, cleaning up");
|
||||
return RECEIVER_STEP_FAIL;
|
||||
}
|
||||
|
||||
static ReceiverStep receiver_handle_check(ReceiverPendingState* state) {
|
||||
bool skipped = false;
|
||||
bool would_transfer = false;
|
||||
File* file = receive_incremental_check_ex(state->fd, state->config, &skipped, &would_transfer);
|
||||
if (state->config->dry_run) {
|
||||
/* Server-contacting --dry-run: the reply has already been sent
|
||||
(STATUS_OK = up to date, STATUS_DRY_RUN_TRANSFER = would transfer) and
|
||||
nothing may be stored. Both flags false means a genuine protocol
|
||||
error (STATUS_ERROR already sent or sent by receive_error below). */
|
||||
if (!skipped && !would_transfer)
|
||||
return RECEIVER_STEP_ERROR;
|
||||
} else if (!skipped && (!file || !state->sink->store_file(file, state->sink->context))) {
|
||||
return RECEIVER_STEP_ERROR;
|
||||
}
|
||||
return RECEIVER_STEP_NEXT;
|
||||
}
|
||||
|
||||
static ReceiverStep receiver_handle_chunk(ReceiverPendingState* state) {
|
||||
Chunk* chunk = receive_chunk_data(state->fd, state->config);
|
||||
if (!chunk || !receiver_process_chunk(chunk, state->sink))
|
||||
return RECEIVER_STEP_ERROR;
|
||||
return RECEIVER_STEP_NEXT;
|
||||
}
|
||||
|
||||
static ReceiverStep receiver_handle_check_batch(ReceiverPendingState* state) {
|
||||
if (!receiver_process_batch(state->config, state->fd))
|
||||
return RECEIVER_STEP_FAIL;
|
||||
return RECEIVER_STEP_NEXT;
|
||||
}
|
||||
|
||||
static ReceiverStep receiver_handle_mkdir(ReceiverPendingState* state) {
|
||||
File* dir = file_receive_directory(state->fd, state->config);
|
||||
if (!dir || !state->sink->store_file(dir, state->sink->context))
|
||||
return RECEIVER_STEP_ERROR;
|
||||
return RECEIVER_STEP_NEXT;
|
||||
}
|
||||
|
||||
static ReceiverStep receiver_handle_dir_times(const ReceiverPendingState* state) {
|
||||
if (!receiver_process_dir_times(state->fd, state->config, state->sink))
|
||||
return RECEIVER_STEP_ERROR;
|
||||
return RECEIVER_STEP_NEXT;
|
||||
}
|
||||
|
||||
static ReceiverStep receiver_handle_hardlink(ReceiverPendingState* state) {
|
||||
File* file = file_receive_hardlink(state->fd);
|
||||
if (!file || !state->sink->store_file(file, state->sink->context))
|
||||
return RECEIVER_STEP_ERROR;
|
||||
return RECEIVER_STEP_NEXT;
|
||||
}
|
||||
|
||||
static ReceiverStep receiver_handle_symlink(ReceiverPendingState* state) {
|
||||
File* sym = file_receive_symlink(state->fd, state->config);
|
||||
if (!sym || !state->sink->store_file(sym, state->sink->context))
|
||||
return RECEIVER_STEP_ERROR;
|
||||
return RECEIVER_STEP_NEXT;
|
||||
}
|
||||
|
||||
static ReceiverStep receiver_handle_special(ReceiverPendingState* state) {
|
||||
File* file = file_receive_special(state->fd);
|
||||
if (!file || !state->sink->store_file(file, state->sink->context))
|
||||
return RECEIVER_STEP_ERROR;
|
||||
return RECEIVER_STEP_NEXT;
|
||||
}
|
||||
|
||||
static ReceiverStep receiver_handle_manifest(ReceiverPendingState* state) {
|
||||
Config* config = state->config;
|
||||
int fd = state->fd;
|
||||
const ReceiverSink* sink = state->sink;
|
||||
DeleteManifest* manifest = receive_manifest_entries(fd);
|
||||
if (!manifest)
|
||||
return RECEIVER_STEP_FAIL; /* receive_manifest_entries already sent STATUS_ERROR */
|
||||
if (config->dry_run) {
|
||||
/* Server-contacting --dry-run mutates nothing, so a keep-set manifest
|
||||
is consumed and discarded. The early-delete mode still needs its ACK
|
||||
so a sender blocked on the delete handshake is not left hanging.
|
||||
When would-delete reporting is armed, enumerate (read-only) the
|
||||
destination extras so the terminal STATUS_STATS frame can list them. */
|
||||
if (config->use_delete && sink->would_delete) {
|
||||
size_t count = 0;
|
||||
if (!manifest_would_delete_list(config, manifest, sink->would_delete, &count))
|
||||
log_message(LOG_LEVEL_WARNING, "dry-run: could not enumerate would-delete paths");
|
||||
}
|
||||
delete_manifest_free(manifest);
|
||||
if (state->early_delete && !send_status(fd, STATUS_OK))
|
||||
return RECEIVER_STEP_FAIL;
|
||||
return RECEIVER_STEP_NEXT;
|
||||
}
|
||||
if (state->early_delete) {
|
||||
/* --delete-before: the whole-tree manifest is authoritative the moment
|
||||
it arrives, before any file data. Delete now and acknowledge so the
|
||||
sender only starts streaming once the deletion committed (or failed).
|
||||
A later transfer failure does not restore these deletions. A
|
||||
--max-delete-capped commit still succeeds and the transfer proceeds;
|
||||
the terminal success frame reports the cap. */
|
||||
size_t deleted = 0;
|
||||
DeletePathObserver observer =
|
||||
(config->report_deletes && sink->deleted_paths) ? receiver_record_deleted_path : NULL;
|
||||
DeleteCommitResult deletion =
|
||||
(config->use_delete || config->delete_missing_args)
|
||||
? manifest_delete_all_observed(config, manifest, &deleted, observer,
|
||||
(void*)sink->deleted_paths)
|
||||
: DELETE_COMMIT_OK;
|
||||
receiver_tally_deleted(sink, deleted);
|
||||
delete_manifest_free(manifest);
|
||||
if (deletion == DELETE_COMMIT_ERROR) {
|
||||
send_status(fd, STATUS_ERROR);
|
||||
return RECEIVER_STEP_FAIL;
|
||||
}
|
||||
if (deletion == DELETE_COMMIT_LIMIT_REACHED && sink->note_delete_limit)
|
||||
sink->note_delete_limit(sink->context);
|
||||
if (!send_status(fd, STATUS_OK))
|
||||
return RECEIVER_STEP_FAIL;
|
||||
} else if (config->use_delete || config->delete_missing_args) {
|
||||
/* Plain --delete / --delete-after and the --delete-missing-args
|
||||
exact-path deletions: hold the manifest and commit it only after
|
||||
STATUS_FINISHED. The per-directory modes never send this frame. */
|
||||
if (state->deferred_manifest) {
|
||||
log_message(LOG_LEVEL_ERROR, "Received a second delete manifest");
|
||||
delete_manifest_free(state->deferred_manifest);
|
||||
state->deferred_manifest = NULL;
|
||||
delete_manifest_free(manifest);
|
||||
send_status(fd, STATUS_ERROR);
|
||||
return RECEIVER_STEP_FAIL;
|
||||
}
|
||||
state->deferred_manifest = manifest;
|
||||
} else {
|
||||
delete_manifest_free(manifest);
|
||||
}
|
||||
return RECEIVER_STEP_NEXT;
|
||||
}
|
||||
|
||||
static ReceiverStep receiver_handle_delete_plan(ReceiverPendingState* state) {
|
||||
Config* config = state->config;
|
||||
int fd = state->fd;
|
||||
const ReceiverSink* sink = state->sink;
|
||||
if (!state->per_dir_delete) {
|
||||
log_message(LOG_LEVEL_ERROR, "Received a per-directory delete plan without a per-dir "
|
||||
"delete timing");
|
||||
send_status(fd, STATUS_ERROR);
|
||||
return RECEIVER_STEP_FAIL;
|
||||
}
|
||||
if (!state->plan_session) {
|
||||
state->plan_session = delete_plan_session_create(config);
|
||||
if (state->plan_session && config->report_deletes && sink->deleted_paths)
|
||||
delete_plan_session_set_delete_observer(state->plan_session, receiver_record_deleted_path,
|
||||
(void*)sink->deleted_paths);
|
||||
}
|
||||
if (!state->plan_session || delete_plan_session_receive(state->plan_session, config, fd) != 0)
|
||||
return RECEIVER_STEP_FAIL;
|
||||
if (delete_plan_session_limit_reached(state->plan_session) && !state->delete_limit_noted &&
|
||||
sink->note_delete_limit) {
|
||||
sink->note_delete_limit(sink->context);
|
||||
state->delete_limit_noted = true;
|
||||
}
|
||||
return RECEIVER_STEP_NEXT;
|
||||
}
|
||||
|
||||
static ReceiverStep receiver_handle_file(ReceiverPendingState* state) {
|
||||
File* file = file_receive(state->config, state->fd);
|
||||
if (!file) {
|
||||
log_message(LOG_LEVEL_ERROR, "Failed to receive file");
|
||||
return RECEIVER_STEP_ERROR;
|
||||
}
|
||||
if (!state->sink->store_file(file, state->sink->context))
|
||||
return RECEIVER_STEP_ERROR;
|
||||
return RECEIVER_STEP_NEXT;
|
||||
}
|
||||
|
||||
/* One dispatch per admitted frame type; STATUS_NEXT (and any other
|
||||
data-bearing status) falls through to the regular file receiver. */
|
||||
static ReceiverStep receiver_dispatch_status(ReceiverPendingState* state, Status status) {
|
||||
switch (status) {
|
||||
case STATUS_KEEPALIVE:
|
||||
return receiver_handle_keepalive(state);
|
||||
case STATUS_ABORT:
|
||||
return receiver_handle_abort(state);
|
||||
case STATUS_CHECK:
|
||||
return receiver_handle_check(state);
|
||||
case STATUS_CHUNK:
|
||||
return receiver_handle_chunk(state);
|
||||
case STATUS_CHECK_BATCH:
|
||||
return receiver_handle_check_batch(state);
|
||||
case STATUS_MKDIR:
|
||||
return receiver_handle_mkdir(state);
|
||||
case STATUS_DIR_TIMES:
|
||||
return receiver_handle_dir_times(state);
|
||||
case STATUS_HARDLINK:
|
||||
return receiver_handle_hardlink(state);
|
||||
case STATUS_SYMLINK:
|
||||
return receiver_handle_symlink(state);
|
||||
case STATUS_SPECIAL:
|
||||
return receiver_handle_special(state);
|
||||
case STATUS_MANIFEST:
|
||||
return receiver_handle_manifest(state);
|
||||
case STATUS_DELETE_PLAN:
|
||||
return receiver_handle_delete_plan(state);
|
||||
default:
|
||||
return receiver_handle_file(state);
|
||||
}
|
||||
}
|
||||
|
||||
/* Release the parked keep-set / per-directory session exactly once on every
|
||||
failure exit. Never commit a deletion for a failed stream. */
|
||||
static void receiver_drop_pending(ReceiverPendingState* state) {
|
||||
if (state->deferred_manifest) {
|
||||
delete_manifest_free(state->deferred_manifest);
|
||||
state->deferred_manifest = NULL;
|
||||
}
|
||||
if (state->plan_session) {
|
||||
delete_plan_session_destroy(state->plan_session);
|
||||
state->plan_session = NULL;
|
||||
}
|
||||
}
|
||||
|
||||
/* Runs the whole receive loop. The delete manifest may legitimately arrive
|
||||
either FIRST (--delete-before / --delete-during: the sender transmits the
|
||||
validated keep-set before any file data) or LAST (plain --delete /
|
||||
--delete-after / --delete-delay: the manifest closes the data stream). In
|
||||
validated keep-set before any file data) or LAST (--delete-after /
|
||||
--delete-commit / --delete-delay: the manifest closes the data stream). In
|
||||
the early modes the receiver deletes as soon as the manifest has been read
|
||||
and acknowledges with STATUS_OK so the sender only starts streaming once the
|
||||
deletion has committed (or failed); in the late modes the manifest is held
|
||||
and the deletion is committed only after the terminal STATUS_FINISHED proves
|
||||
the whole transfer succeeded. See receiver_process_pending() for how the -m
|
||||
receiver defers that commit until its disk writer has drained. */
|
||||
the whole transfer succeeded. A plain --delete defaults to the per-directory
|
||||
delete-during plan mode (no manifest at all). See the per-frame handlers
|
||||
above for how the -m receiver defers that commit until its disk writer has
|
||||
drained. */
|
||||
int receiver_process_pending(Config* config, int file_descriptor, const ReceiverSink* sink,
|
||||
DeleteManifest** pending_manifest, DeletePlanSession** pending_plans) {
|
||||
Status status;
|
||||
@@ -303,158 +587,29 @@ int receiver_process_pending(Config* config, int file_descriptor, const Receiver
|
||||
last_progress = session_start;
|
||||
if (!receiver_note_status(&session_start, &last_progress, status, file_descriptor, sink))
|
||||
return -1;
|
||||
bool early_delete = config_delete_timing_early(config);
|
||||
bool per_dir_delete = config_delete_timing_per_dir(config);
|
||||
/* Parked keep-set for the late/commit timing. Every exit path below frees it
|
||||
exactly once; the only exception is the successful FINISHED handoff, which
|
||||
transfers ownership to *pending_manifest (used by the -m receiver). */
|
||||
DeleteManifest* deferred_manifest = NULL;
|
||||
/* Per-directory delete session for --delete-during/--delete-delay. During the
|
||||
loop it applies plans inline (during) or snapshots their extras (delay); on
|
||||
a successful FINISHED it is either committed here or handed to
|
||||
*pending_plans so the -m caller commits after its disk writer drained. */
|
||||
DeletePlanSession* plan_session = NULL;
|
||||
bool delete_limit_noted = false;
|
||||
ReceiverPendingState state = {
|
||||
.config = config,
|
||||
.fd = file_descriptor,
|
||||
.sink = sink,
|
||||
.pending_manifest = pending_manifest,
|
||||
.pending_plans = pending_plans,
|
||||
.deferred_manifest = NULL,
|
||||
.plan_session = NULL,
|
||||
.early_delete = config_delete_timing_early(config),
|
||||
.per_dir_delete = config_delete_timing_per_dir(config),
|
||||
.delete_limit_noted = false,
|
||||
};
|
||||
bool notify_peer = false;
|
||||
while (status == STATUS_NEXT || status == STATUS_CHUNK || status == STATUS_CHECK ||
|
||||
status == STATUS_KEEPALIVE || status == STATUS_ABORT || status == STATUS_CHECK_BATCH ||
|
||||
status == STATUS_MKDIR || status == STATUS_MANIFEST || status == STATUS_HARDLINK ||
|
||||
status == STATUS_SYMLINK || status == STATUS_SPECIAL || status == STATUS_DIR_TIMES ||
|
||||
status == STATUS_DELETE_PLAN) {
|
||||
if (status == STATUS_KEEPALIVE) {
|
||||
if (!send_status(file_descriptor, STATUS_KEEPALIVE))
|
||||
goto fail;
|
||||
goto next_status;
|
||||
}
|
||||
if (status == STATUS_ABORT) {
|
||||
log_message(LOG_LEVEL_INFO, "Received abort from client, cleaning up");
|
||||
ReceiverStep step = receiver_dispatch_status(&state, status);
|
||||
if (step == RECEIVER_STEP_FAIL)
|
||||
goto fail;
|
||||
}
|
||||
if (status == STATUS_CHECK) {
|
||||
bool skipped = false;
|
||||
bool would_transfer = false;
|
||||
File* file = receive_incremental_check_ex(file_descriptor, config, &skipped, &would_transfer);
|
||||
if (config->dry_run) {
|
||||
/* Server-contacting --dry-run: the reply has already been sent
|
||||
(STATUS_OK = up to date, STATUS_DRY_RUN_TRANSFER = would transfer) and
|
||||
nothing may be stored. Both flags false means a genuine protocol
|
||||
error (STATUS_ERROR already sent or sent by receive_error below). */
|
||||
if (!skipped && !would_transfer)
|
||||
goto receive_error;
|
||||
} else if (!skipped && (!file || !sink->store_file(file, sink->context))) {
|
||||
goto receive_error;
|
||||
}
|
||||
} else if (status == STATUS_CHUNK) {
|
||||
Chunk* chunk = receive_chunk_data(file_descriptor, config);
|
||||
if (!chunk || !receiver_process_chunk(chunk, sink))
|
||||
goto receive_error;
|
||||
} else if (status == STATUS_CHECK_BATCH) {
|
||||
if (!receiver_process_batch(config, file_descriptor))
|
||||
goto fail;
|
||||
goto next_status;
|
||||
} else if (status == STATUS_MKDIR) {
|
||||
File* dir = file_receive_directory(file_descriptor, config);
|
||||
if (!dir || !sink->store_file(dir, sink->context))
|
||||
goto receive_error;
|
||||
} else if (status == STATUS_DIR_TIMES) {
|
||||
if (!receiver_process_dir_times(file_descriptor, config, sink))
|
||||
goto receive_error;
|
||||
} else if (status == STATUS_HARDLINK) {
|
||||
File* file = file_receive_hardlink(file_descriptor);
|
||||
if (!file || !sink->store_file(file, sink->context))
|
||||
goto receive_error;
|
||||
} else if (status == STATUS_SYMLINK) {
|
||||
File* sym = file_receive_symlink(file_descriptor, config);
|
||||
if (!sym || !sink->store_file(sym, sink->context))
|
||||
goto receive_error;
|
||||
} else if (status == STATUS_SPECIAL) {
|
||||
File* file = file_receive_special(file_descriptor);
|
||||
if (!file || !sink->store_file(file, sink->context))
|
||||
goto receive_error;
|
||||
} else if (status == STATUS_MANIFEST) {
|
||||
DeleteManifest* manifest = receive_manifest_entries(file_descriptor);
|
||||
if (!manifest)
|
||||
goto fail; /* receive_manifest_entries already sent STATUS_ERROR */
|
||||
if (config->dry_run) {
|
||||
/* Server-contacting --dry-run mutates nothing, so a keep-set manifest
|
||||
is consumed and discarded. The early-delete mode still needs its ACK
|
||||
so a sender blocked on the delete handshake is not left hanging.
|
||||
When would-delete reporting is armed, enumerate (read-only) the
|
||||
destination extras so the terminal STATUS_STATS frame can list them. */
|
||||
if (config->use_delete && sink->would_delete) {
|
||||
size_t count = 0;
|
||||
if (!manifest_would_delete_list(config, manifest, sink->would_delete, &count))
|
||||
log_message(LOG_LEVEL_WARNING, "dry-run: could not enumerate would-delete paths");
|
||||
}
|
||||
delete_manifest_free(manifest);
|
||||
if (early_delete && !send_status(file_descriptor, STATUS_OK))
|
||||
goto fail;
|
||||
goto next_status;
|
||||
}
|
||||
if (early_delete) {
|
||||
/* --delete-before: the whole-tree manifest is authoritative the moment
|
||||
it arrives, before any file data. Delete now and acknowledge so the
|
||||
sender only starts streaming once the deletion committed (or failed).
|
||||
A later transfer failure does not restore these deletions. A
|
||||
--max-delete-capped commit still succeeds and the transfer proceeds;
|
||||
the terminal success frame reports the cap. */
|
||||
size_t deleted = 0;
|
||||
DeleteCommitResult deletion = (config->use_delete || config->delete_missing_args)
|
||||
? manifest_delete_all_counted(config, manifest, &deleted)
|
||||
: DELETE_COMMIT_OK;
|
||||
receiver_tally_deleted(sink, deleted);
|
||||
delete_manifest_free(manifest);
|
||||
if (deletion == DELETE_COMMIT_ERROR) {
|
||||
send_status(file_descriptor, STATUS_ERROR);
|
||||
goto fail;
|
||||
}
|
||||
if (deletion == DELETE_COMMIT_LIMIT_REACHED && sink->note_delete_limit)
|
||||
sink->note_delete_limit(sink->context);
|
||||
if (!send_status(file_descriptor, STATUS_OK))
|
||||
goto fail;
|
||||
} else if (config->use_delete || config->delete_missing_args) {
|
||||
/* Plain --delete / --delete-after and the --delete-missing-args
|
||||
exact-path deletions: hold the manifest and commit it only after
|
||||
STATUS_FINISHED. The per-directory modes never send this frame. */
|
||||
if (deferred_manifest) {
|
||||
log_message(LOG_LEVEL_ERROR, "Received a second delete manifest");
|
||||
delete_manifest_free(deferred_manifest);
|
||||
deferred_manifest = NULL;
|
||||
delete_manifest_free(manifest);
|
||||
send_status(file_descriptor, STATUS_ERROR);
|
||||
goto fail;
|
||||
}
|
||||
deferred_manifest = manifest;
|
||||
} else {
|
||||
delete_manifest_free(manifest);
|
||||
}
|
||||
goto next_status;
|
||||
} else if (status == STATUS_DELETE_PLAN) {
|
||||
if (!per_dir_delete) {
|
||||
log_message(LOG_LEVEL_ERROR, "Received a per-directory delete plan without a per-dir "
|
||||
"delete timing");
|
||||
send_status(file_descriptor, STATUS_ERROR);
|
||||
goto fail;
|
||||
}
|
||||
if (!plan_session)
|
||||
plan_session = delete_plan_session_create(config);
|
||||
if (!plan_session || delete_plan_session_receive(plan_session, config, file_descriptor) != 0)
|
||||
goto fail;
|
||||
if (delete_plan_session_limit_reached(plan_session) && !delete_limit_noted &&
|
||||
sink->note_delete_limit) {
|
||||
sink->note_delete_limit(sink->context);
|
||||
delete_limit_noted = true;
|
||||
}
|
||||
goto next_status;
|
||||
} else {
|
||||
File* file = file_receive(config, file_descriptor);
|
||||
if (!file) {
|
||||
log_message(LOG_LEVEL_ERROR, "Failed to receive file");
|
||||
goto receive_error;
|
||||
}
|
||||
if (!sink->store_file(file, sink->context))
|
||||
goto receive_error;
|
||||
}
|
||||
next_status:
|
||||
if (step == RECEIVER_STEP_ERROR)
|
||||
goto receive_error;
|
||||
if (!receive_status(file_descriptor, &status))
|
||||
goto receive_error;
|
||||
if (!receiver_note_status(&session_start, &last_progress, status, file_descriptor, sink))
|
||||
@@ -473,17 +628,19 @@ int receiver_process_pending(Config* config, int file_descriptor, const Receiver
|
||||
disk writer may still be draining; the caller commits after the writer has
|
||||
joined so no extra file is removed unless the transfer is known to have
|
||||
succeeded. */
|
||||
if (deferred_manifest) {
|
||||
if (pending_manifest) {
|
||||
*pending_manifest = deferred_manifest;
|
||||
deferred_manifest = NULL;
|
||||
if (state.deferred_manifest) {
|
||||
if (state.pending_manifest) {
|
||||
*state.pending_manifest = state.deferred_manifest;
|
||||
state.deferred_manifest = NULL;
|
||||
} else {
|
||||
size_t deleted = 0;
|
||||
DeleteCommitResult deletion =
|
||||
manifest_delete_all_counted(config, deferred_manifest, &deleted);
|
||||
DeletePathObserver observer =
|
||||
(config->report_deletes && sink->deleted_paths) ? receiver_record_deleted_path : NULL;
|
||||
DeleteCommitResult deletion = manifest_delete_all_observed(
|
||||
config, state.deferred_manifest, &deleted, observer, (void*)sink->deleted_paths);
|
||||
receiver_tally_deleted(sink, deleted);
|
||||
delete_manifest_free(deferred_manifest);
|
||||
deferred_manifest = NULL;
|
||||
delete_manifest_free(state.deferred_manifest);
|
||||
state.deferred_manifest = NULL;
|
||||
if (deletion == DELETE_COMMIT_ERROR) {
|
||||
send_status(file_descriptor, STATUS_ERROR);
|
||||
goto fail;
|
||||
@@ -497,25 +654,28 @@ int receiver_process_pending(Config* config, int file_descriptor, const Receiver
|
||||
nothing yet and applies its decompressed snapshot here. The -m receiver
|
||||
hands the session to its caller instead, which commits after the disk
|
||||
writer drained. */
|
||||
if (plan_session) {
|
||||
if (pending_plans) {
|
||||
*pending_plans = plan_session;
|
||||
plan_session = NULL;
|
||||
if (state.plan_session) {
|
||||
if (config->report_deletes && sink->deleted_paths)
|
||||
delete_plan_session_set_delete_observer(state.plan_session, receiver_record_deleted_path,
|
||||
(void*)sink->deleted_paths);
|
||||
if (state.pending_plans) {
|
||||
*state.pending_plans = state.plan_session;
|
||||
state.plan_session = NULL;
|
||||
} else if (config->dry_run) {
|
||||
/* Central dry-run no-op: never commit a deletion for a -n run. */
|
||||
delete_plan_session_destroy(plan_session);
|
||||
plan_session = NULL;
|
||||
delete_plan_session_destroy(state.plan_session);
|
||||
state.plan_session = NULL;
|
||||
} else {
|
||||
DeleteCommitResult deletion = delete_plan_session_commit(plan_session, config);
|
||||
bool limit = delete_plan_session_limit_reached(plan_session);
|
||||
receiver_tally_deleted(sink, delete_plan_session_deleted(plan_session));
|
||||
delete_plan_session_destroy(plan_session);
|
||||
plan_session = NULL;
|
||||
DeleteCommitResult deletion = delete_plan_session_commit(state.plan_session, config);
|
||||
bool limit = delete_plan_session_limit_reached(state.plan_session);
|
||||
receiver_tally_deleted(sink, delete_plan_session_deleted(state.plan_session));
|
||||
delete_plan_session_destroy(state.plan_session);
|
||||
state.plan_session = NULL;
|
||||
if (deletion == DELETE_COMMIT_ERROR) {
|
||||
send_status(file_descriptor, STATUS_ERROR);
|
||||
goto fail;
|
||||
}
|
||||
if (limit && !delete_limit_noted && sink->note_delete_limit)
|
||||
if (limit && !state.delete_limit_noted && sink->note_delete_limit)
|
||||
sink->note_delete_limit(sink->context);
|
||||
}
|
||||
}
|
||||
@@ -529,26 +689,14 @@ int receiver_process_pending(Config* config, int file_descriptor, const Receiver
|
||||
}
|
||||
return 0;
|
||||
|
||||
receive_error:
|
||||
notify_peer = true;
|
||||
fail:
|
||||
/* Failure exits that must not (or already did) report a STATUS_ERROR. The
|
||||
parked keep-set/session is dropped: never commit a deletion for a failed
|
||||
stream. */
|
||||
if (deferred_manifest) {
|
||||
delete_manifest_free(deferred_manifest);
|
||||
deferred_manifest = NULL;
|
||||
}
|
||||
if (plan_session)
|
||||
delete_plan_session_destroy(plan_session);
|
||||
return -1;
|
||||
|
||||
receive_error:
|
||||
if (deferred_manifest) {
|
||||
delete_manifest_free(deferred_manifest);
|
||||
deferred_manifest = NULL;
|
||||
}
|
||||
if (plan_session)
|
||||
delete_plan_session_destroy(plan_session);
|
||||
if (sink->send_error)
|
||||
receiver_drop_pending(&state);
|
||||
if (notify_peer && sink->send_error)
|
||||
send_status(file_descriptor, STATUS_ERROR);
|
||||
return -1;
|
||||
}
|
||||
@@ -569,11 +717,15 @@ typedef struct {
|
||||
--delete would-delete path list collected while processing the manifest. */
|
||||
ReceiverStats stats;
|
||||
ArrayList* would_delete;
|
||||
/* --info=del: actually-removed paths collected during the delete commit. */
|
||||
ArrayList* deleted_paths;
|
||||
} ReceiverSaveContext;
|
||||
|
||||
static bool receiver_save_file(File* file, void* context_pointer) {
|
||||
ReceiverSaveContext* context = context_pointer;
|
||||
FileSaveResult result = FILE_SAVE_ERROR;
|
||||
bool created = false;
|
||||
unsigned created_dirs = 0;
|
||||
if (context->config->dry_run) {
|
||||
/* Defense in depth: a dry-run receiver mutates nothing even if a data
|
||||
frame reaches the sink (the sender is not supposed to send one). */
|
||||
@@ -583,12 +735,17 @@ static bool receiver_save_file(File* file, void* context_pointer) {
|
||||
--remove-source-files sender keeps its source. */
|
||||
result = FILE_SAVE_SKIPPED;
|
||||
} else {
|
||||
result = file_save_to_disk_full(context->config->receive_root_directory, file, context->config);
|
||||
result = file_save_to_disk_full_ex(context->config->receive_root_directory, file,
|
||||
context->config, &created, &created_dirs);
|
||||
}
|
||||
/* Wire-stats tally: bytes reconstructed from the basis file (delta matches)
|
||||
count as matched data in the end-of-transfer report. */
|
||||
if (result != FILE_SAVE_ERROR && file->matched_bytes > 0)
|
||||
context->stats.matched_data += file->matched_bytes;
|
||||
/* Protocol 2.28.0: receiver-observed literal bytes and the created-entry
|
||||
breakdown (regular/dir/link/special) for the `--stats` report. */
|
||||
if (result == FILE_SAVE_WRITTEN)
|
||||
receiver_stats_note_saved(&context->stats, file, created, created_dirs);
|
||||
/* A directory's metadata is deferred, never applied inline: collect it now
|
||||
and apply it at the end. -O/--omit-dir-times and --preserve_perms/-times
|
||||
are honored by dir_metadata_list_apply's caller (see
|
||||
@@ -621,7 +778,8 @@ static void receiver_note_delete_limit(void* context_pointer) {
|
||||
static bool receiver_send_success_frame(int fd, void* context_pointer) {
|
||||
ReceiverSaveContext* context = context_pointer;
|
||||
Status final_status = context->delete_limit_reached ? STATUS_DELETE_LIMIT : STATUS_OK;
|
||||
if (!receiver_send_stats_frame(fd, context->config, &context->stats, context->would_delete))
|
||||
if (!receiver_send_stats_frame(fd, context->config, &context->stats, context->would_delete,
|
||||
context->deleted_paths))
|
||||
return false;
|
||||
/* Server-contacting --dry-run: nothing was staged or written, so there is
|
||||
nothing to publish and no directory times to stamp. */
|
||||
@@ -651,8 +809,15 @@ int receiver_receive_files(Config* config, int file_descriptor) {
|
||||
ReceiverSaveContext context = {.config = config, .outcomes = {0}};
|
||||
dir_time_list_init(&context.dir_times);
|
||||
context.would_delete = array_list_create(free);
|
||||
if (!context.would_delete)
|
||||
/* report_deletes (--info=del / -i / --out-format under --delete) is the only
|
||||
reason to retain the actually-removed paths; a plain --delete must not
|
||||
str_dup every removal. NULL is handled by every consumer. */
|
||||
context.deleted_paths = config->report_deletes ? array_list_create(free) : NULL;
|
||||
if (!context.would_delete || (config->report_deletes && !context.deleted_paths)) {
|
||||
array_list_delete(context.would_delete);
|
||||
array_list_delete(context.deleted_paths);
|
||||
return -1;
|
||||
}
|
||||
ReceiverSink sink = {receiver_save_file,
|
||||
&context,
|
||||
true,
|
||||
@@ -660,12 +825,14 @@ int receiver_receive_files(Config* config, int file_descriptor) {
|
||||
receiver_send_success_frame,
|
||||
receiver_note_delete_limit,
|
||||
&context.stats,
|
||||
context.would_delete};
|
||||
context.would_delete,
|
||||
context.deleted_paths};
|
||||
int ret = receiver_process(config, file_descriptor, &sink);
|
||||
if (ret != 0 && config->delay_updates && config->delay_context)
|
||||
delay_updates_cleanup(config->delay_context);
|
||||
receiver_outcomes_destroy(&context.outcomes);
|
||||
dir_time_list_free(&context.dir_times);
|
||||
array_list_delete(context.would_delete);
|
||||
array_list_delete(context.deleted_paths);
|
||||
return ret;
|
||||
}
|
||||
|
||||
+11
-1
@@ -46,11 +46,20 @@ typedef struct {
|
||||
carries the -n/--dry-run --delete path list. */
|
||||
ReceiverStats* stats;
|
||||
struct ArrayList* would_delete;
|
||||
/* When --info=del requested it, receiver-owned strings for every path the
|
||||
deletion commit ACTUALLY removed, sent in the terminal STATUS_STATS frame's
|
||||
path list so the sender can print rsync's `deleting PATH` lines. */
|
||||
struct ArrayList* deleted_paths;
|
||||
} ReceiverSink;
|
||||
|
||||
bool receiver_outcomes_append(ReceiverOutcomes* outcomes, unsigned char code);
|
||||
void receiver_outcomes_destroy(ReceiverOutcomes* outcomes);
|
||||
|
||||
/* DeletePathObserver implementation for --info=del: `context` is an ArrayList*
|
||||
that receives owned copies of every truly-removed destination-relative path.
|
||||
Shared by the single-threaded receiver and the -m pipeline's deferred commit. */
|
||||
void receiver_record_deleted_path(void* context, const char* rel_path);
|
||||
|
||||
/* Send the terminal success frame. `final_status` is usually STATUS_OK, or
|
||||
STATUS_DELETE_LIMIT when a --max-delete commit was capped. */
|
||||
bool receiver_send_final_success(int fd, const Config* config, const ReceiverOutcomes* outcomes,
|
||||
@@ -60,7 +69,8 @@ bool receiver_send_final_success(int fd, const Config* config, const ReceiverOut
|
||||
non-NULL, a count and that many wire strings) when the wire config requested
|
||||
report_stats. A no-op otherwise. */
|
||||
bool receiver_send_stats_frame(int fd, const Config* config, const ReceiverStats* stats,
|
||||
const struct ArrayList* would_delete);
|
||||
const struct ArrayList* would_delete,
|
||||
const struct ArrayList* deleted_paths);
|
||||
|
||||
int receiver_process(Config* config, int file_descriptor, const ReceiverSink* sink);
|
||||
/* receiver_process with an escape hatch for the commit-style (late) deletion:
|
||||
|
||||
@@ -31,6 +31,7 @@ PipelineContextReceiver* pipeline_context_receiver_create(Config* config, Queue*
|
||||
context->delete_limit_reached = false;
|
||||
memset(&context->stats, 0, sizeof(context->stats));
|
||||
context->would_delete = NULL;
|
||||
context->deleted_paths = NULL;
|
||||
atomic_init(&context->cancelled, false);
|
||||
int init = 0;
|
||||
if (mtx_init(&context->mutex, mtx_plain) != thrd_success)
|
||||
@@ -46,6 +47,15 @@ PipelineContextReceiver* pipeline_context_receiver_create(Config* config, Queue*
|
||||
context->would_delete = array_list_create(free);
|
||||
if (!context->would_delete)
|
||||
goto fail;
|
||||
/* The actually-removed path list is only needed to render rsync's
|
||||
`deleting PATH` lines, which the client requests via report_deletes
|
||||
(--info=del / -i / --out-format under --delete). A plain --delete run must
|
||||
not allocate it or observe every removal. */
|
||||
if (config->report_deletes) {
|
||||
context->deleted_paths = array_list_create(free);
|
||||
if (!context->deleted_paths)
|
||||
goto fail;
|
||||
}
|
||||
return context;
|
||||
|
||||
fail:
|
||||
@@ -56,6 +66,12 @@ fail:
|
||||
cnd_destroy(&context->condition_not_full);
|
||||
if (init >= 1)
|
||||
mtx_destroy(&context->mutex);
|
||||
/* Free every list that was already created before the failing allocation:
|
||||
`context` itself is freed below, so they would otherwise leak. */
|
||||
if (context->would_delete)
|
||||
array_list_delete(context->would_delete);
|
||||
if (context->deleted_paths)
|
||||
array_list_delete(context->deleted_paths);
|
||||
free(context);
|
||||
return NULL;
|
||||
}
|
||||
@@ -71,6 +87,8 @@ void pipeline_context_receiver_destroy(PipelineContextReceiver* context) {
|
||||
dir_time_list_free(&context->dir_times);
|
||||
if (context->would_delete)
|
||||
array_list_delete(context->would_delete);
|
||||
if (context->deleted_paths)
|
||||
array_list_delete(context->deleted_paths);
|
||||
mtx_destroy(&context->mutex);
|
||||
cnd_destroy(&context->condition_not_full);
|
||||
cnd_destroy(&context->condition_not_empty);
|
||||
@@ -184,7 +202,8 @@ int receive_thread(void* pipeline_context) {
|
||||
NULL,
|
||||
receiver_pipeline_note_delete_limit,
|
||||
&context->stats,
|
||||
context->would_delete};
|
||||
context->would_delete,
|
||||
context->deleted_paths};
|
||||
if (receiver_process_pending((Config*)config, file_descriptor, &sink, &context->deferred_manifest,
|
||||
&context->deferred_plans) != 0) {
|
||||
receiver_thread_fail(context);
|
||||
@@ -228,12 +247,23 @@ int write_thread(void* pipeline_context) {
|
||||
}
|
||||
size_t file_bytes = file->data ? file->data->size : 0;
|
||||
FileSaveResult result = FILE_SAVE_SKIPPED;
|
||||
bool created = false;
|
||||
unsigned created_dirs = 0;
|
||||
/* Server-contacting --dry-run: never write. The receiver thread does not
|
||||
enqueue anything on the dry-run path, but this keeps the writer thread
|
||||
provably mutation-free if a data frame ever reached it. */
|
||||
bool dry_run = context->config->dry_run;
|
||||
if (save_to_disk && !dry_run) {
|
||||
result = file_save_to_disk_full(root_directory, file, context->config);
|
||||
result =
|
||||
file_save_to_disk_full_ex(root_directory, file, context->config, &created, &created_dirs);
|
||||
if (result == FILE_SAVE_WRITTEN) {
|
||||
/* Protocol 2.28.0: fold the receiver-observed literal bytes and the
|
||||
created-entry type into the shared stats block under its mutex (the
|
||||
receive thread also writes stats.matched_data). */
|
||||
mtx_lock(&context->mutex);
|
||||
receiver_stats_note_saved(&context->stats, file, created, created_dirs);
|
||||
mtx_unlock(&context->mutex);
|
||||
}
|
||||
if (result == FILE_SAVE_ERROR) {
|
||||
file_destroy(file);
|
||||
pipeline_context_receiver_note_bytes_released(context, file_bytes);
|
||||
|
||||
@@ -61,6 +61,9 @@ typedef struct PipelineContextReceiver {
|
||||
/* -n/--dry-run --delete would-delete path list, collected by receive_thread
|
||||
and reported in the STATUS_STATS frame. */
|
||||
struct ArrayList* would_delete;
|
||||
/* --info=del actually-removed path list, collected by the deferred delete
|
||||
commit in server.c and reported in the STATUS_STATS frame. */
|
||||
struct ArrayList* deleted_paths;
|
||||
} PipelineContextReceiver;
|
||||
|
||||
PipelineContextReceiver* pipeline_context_receiver_create(Config* config, Queue* queue_receiver,
|
||||
|
||||
+313
-200
@@ -226,11 +226,6 @@ static void release_authorization(void) {
|
||||
close(root_fd);
|
||||
}
|
||||
|
||||
static bool path_is_within(const char* root, const char* path) {
|
||||
size_t n = strlen(root);
|
||||
return strncmp(root, path, n) == 0 && (path[n] == '\0' || path[n] == '/');
|
||||
}
|
||||
|
||||
/* --mkpath contract: when the client's destination root directory does not
|
||||
exist yet on the server side, --mkpath tells the server to create it (and
|
||||
any missing leading components) below the authorized root at connection
|
||||
@@ -710,69 +705,85 @@ static const char* server_module_gate(const Config* config, void* context) {
|
||||
return module_gate_install_root(config, module);
|
||||
}
|
||||
|
||||
void handler(int file_descriptor) {
|
||||
SSL* ssl = io_get_ssl();
|
||||
/* Per-connection state threaded through the handler phase helpers below. The
|
||||
* fields are a faithful split of the former handler() locals: the protocol
|
||||
* session, the config-frame gate context, the accepted config, the optional
|
||||
* multithreaded pipeline context and the teardown bookkeeping all live here so
|
||||
* the single `done` epilogue in handler() can release them exactly as before. */
|
||||
typedef struct ServerSession {
|
||||
int fd;
|
||||
SSL* ssl;
|
||||
ProtocolSession session;
|
||||
protocol_session_init(&session, file_descriptor, file_descriptor);
|
||||
protocol_session_set_ssl(&session, ssl);
|
||||
protocol_session_bind(&session);
|
||||
ModuleGateContext gate_ctx;
|
||||
gate_ctx.ssl = ssl;
|
||||
gate_ctx.fd = file_descriptor;
|
||||
gate_ctx.super_mode_override = -1;
|
||||
gate_ctx.has_peer_ip = false;
|
||||
gate_ctx.peer_ip[0] = '\0';
|
||||
gate_ctx.is_local = false;
|
||||
/* All teardown state starts empty so the single `done` epilogue is safe to
|
||||
* reach from any error path (including before the config frame arrives). */
|
||||
Config* config = NULL;
|
||||
PipelineContextReceiver* context = NULL;
|
||||
char* joined_destination = NULL;
|
||||
bool charset_ready = false;
|
||||
config = config_receive_with_validate(file_descriptor, server_module_gate, &gate_ctx);
|
||||
if (config == NULL) {
|
||||
Config* config;
|
||||
PipelineContextReceiver* context;
|
||||
char* joined_destination;
|
||||
bool charset_ready;
|
||||
} ServerSession;
|
||||
|
||||
/* Phase 1 -- config receipt + validation. Receives the client config frame
|
||||
* through the module gate, applies the super-mode override the gate recorded
|
||||
* exactly once, and installs the per-connection protocol/compression state.
|
||||
* Returns false when the config frame was refused (the gate has already
|
||||
* answered the client); the caller jumps to the shared `done` epilogue. */
|
||||
static bool server_accept_config(ServerSession* state) {
|
||||
state->config = config_receive_with_validate(state->fd, server_module_gate, &state->gate_ctx);
|
||||
if (state->config == NULL) {
|
||||
log_message(LOG_LEVEL_ERROR, "Failed to receive config");
|
||||
goto done;
|
||||
return false;
|
||||
}
|
||||
/* Apply the super-mode veto the gate decided on (operator --no-super, or a
|
||||
* daemon module without the `client owner = yes` opt-in) exactly once, so
|
||||
* every downstream gate (identity_apply_ownership via privilege_super_permitted,
|
||||
* device-node creation) sees SUPER_MODE_OFF. The gate never mutated the
|
||||
* received config. */
|
||||
if (gate_ctx.super_mode_override != -1)
|
||||
config->super_mode = (SuperMode)gate_ctx.super_mode_override;
|
||||
if (state->gate_ctx.super_mode_override != -1)
|
||||
state->config->super_mode = (SuperMode)state->gate_ctx.super_mode_override;
|
||||
/* Install the codec this connection negotiated before the receiver/writer
|
||||
* threads start (the server forks per connection, so the process-global
|
||||
* codec is private to this session). */
|
||||
compression_set_algo((CompressionAlgo)config->compression_algo);
|
||||
compression_set_algo((CompressionAlgo)state->config->compression_algo);
|
||||
/* If the client requested ownership but the effective super mode forbids it
|
||||
* (operator --no-super, a privileged standalone receiver's secure default, or
|
||||
* a daemon module without `client owner = yes`), say so ONCE per connection so
|
||||
* a successful -a/-o/-g transfer is not mistaken for preserved ownership. */
|
||||
if (config->super_mode == SUPER_MODE_OFF && identity_ownership_requested(config))
|
||||
if (state->config->super_mode == SUPER_MODE_OFF && identity_ownership_requested(state->config))
|
||||
log_message(LOG_LEVEL_WARNING,
|
||||
"requested ownership will NOT be applied: super-user activities are disabled "
|
||||
"for this connection (operator veto, or module without `client owner = yes`)");
|
||||
protocol_set_8_bit_output(config->eight_bit_output);
|
||||
protocol_set_8_bit_output(state->config->eight_bit_output);
|
||||
/* Server-side per-message protocol deadline for every frame from here on.
|
||||
* `timeout` is not serialized, so this is the server's own config (the server
|
||||
* has no --timeout CLI and defaults it to 0). A client's --timeout tightens
|
||||
* only that client's own protocol I/O; the server floors its own deadline at
|
||||
* SERVER_IO_TIMEOUT_SEC so a silent peer can never hold a session slot
|
||||
* forever (the socket layer gets the same floor at startup). */
|
||||
protocol_session_set_io_timeout(&session, protocol_server_io_timeout_sec(config->timeout));
|
||||
protocol_session_set_io_timeout(&state->session,
|
||||
protocol_server_io_timeout_sec(state->config->timeout));
|
||||
return true;
|
||||
}
|
||||
|
||||
/* Phase 2 -- security gates. The ORDER here is load-bearing and must not be
|
||||
* merged or reordered: transport/authentication (plaintext refusal, TLS
|
||||
* client-CN verification), then daemon-root confinement (absolute-destination
|
||||
* rejection, traversal + within-authorized-root), then delete/force
|
||||
* authorization -- exactly the sequence the former handler() used. Returns
|
||||
* false after logging the matching rejection; the caller jumps to the shared
|
||||
* `done` epilogue. */
|
||||
static bool server_apply_security_gates(ServerSession* state) {
|
||||
Config* config = state->config;
|
||||
const char* authorized_root = utils_get_authorized_root_path();
|
||||
if (!authorized_root) {
|
||||
log_message(LOG_LEVEL_ERROR, "No server-side destination root configured");
|
||||
goto done;
|
||||
return false;
|
||||
}
|
||||
if (!allow_unauthenticated && ssl == NULL) {
|
||||
if (!allow_unauthenticated && state->ssl == NULL) {
|
||||
log_message(LOG_LEVEL_ERROR, "Rejected unauthenticated plaintext connection");
|
||||
goto done;
|
||||
return false;
|
||||
}
|
||||
if (ssl && required_client_cn && !tls_client_identity_allowed(ssl)) {
|
||||
if (state->ssl && required_client_cn && !tls_client_identity_allowed(state->ssl)) {
|
||||
log_message(LOG_LEVEL_ERROR, "Rejected TLS client with unauthorized identity");
|
||||
goto done;
|
||||
return false;
|
||||
}
|
||||
/* Daemon mode: the module's root is the authorized root (installed by
|
||||
server_module_gate), and the client's destination is a MODULE-RELATIVE
|
||||
@@ -782,27 +793,27 @@ void handler(int file_descriptor) {
|
||||
if (g_daemon_conf && config->receive_root_directory && config->receive_root_directory[0] == '/') {
|
||||
log_message(LOG_LEVEL_ERROR, "Rejected absolute daemon destination (must be relative to the "
|
||||
"selected module root)");
|
||||
goto done;
|
||||
return false;
|
||||
}
|
||||
char* destination = config->receive_root_directory;
|
||||
if (destination && destination[0] != '/')
|
||||
joined_destination = path_cat(authorized_root, destination);
|
||||
if (joined_destination)
|
||||
destination = joined_destination;
|
||||
state->joined_destination = path_cat(authorized_root, destination);
|
||||
if (state->joined_destination)
|
||||
destination = state->joined_destination;
|
||||
if (!destination || has_path_traversal(destination) ||
|
||||
!path_is_within(authorized_root, destination)) {
|
||||
!path_is_within_root(authorized_root, destination)) {
|
||||
log_message(LOG_LEVEL_ERROR, "Rejected destination outside authorized root");
|
||||
free(joined_destination);
|
||||
joined_destination = NULL;
|
||||
goto done;
|
||||
free(state->joined_destination);
|
||||
state->joined_destination = NULL;
|
||||
return false;
|
||||
}
|
||||
if (joined_destination) {
|
||||
if (state->joined_destination) {
|
||||
free(config->receive_root_directory);
|
||||
config->receive_root_directory = joined_destination;
|
||||
joined_destination = NULL;
|
||||
config->receive_root_directory = state->joined_destination;
|
||||
state->joined_destination = NULL;
|
||||
}
|
||||
if (!config->receive_root_directory) {
|
||||
goto done;
|
||||
return false;
|
||||
}
|
||||
config->use_delete = config->use_delete && allow_delete;
|
||||
/* --force (receiver-side) is deletion authority too: it lets an incoming
|
||||
@@ -812,6 +823,18 @@ void handler(int file_descriptor) {
|
||||
* --delete-missing-args, so a client cannot use --force to bypass the delete
|
||||
* policy. */
|
||||
config->force_delete = config->force_delete && allow_delete;
|
||||
return true;
|
||||
}
|
||||
|
||||
/* Phase 3 -- session preparation. Installs the negotiated conversion, applies
|
||||
* the remaining deletion policy, materializes the destination root (--mkpath),
|
||||
* creates the --delay-updates staging tree, snapshots the identity policy, and
|
||||
* publishes the --keep-dirlinks/--trust-sender globals and the daemon MOTD.
|
||||
* All of it must happen before any receiver/writer thread is spawned. Returns
|
||||
* false after logging the matching failure; the caller jumps to the shared
|
||||
* `done` epilogue. */
|
||||
static bool server_prepare_session(ServerSession* state) {
|
||||
Config* config = state->config;
|
||||
/* --iconv (protocol 2.16.0): install the receiver-side wire->local conversion
|
||||
now that the client's full CONVERT_SPEC has been received and validated,
|
||||
before any received file name is decoded. The server's own --iconv (if
|
||||
@@ -822,14 +845,14 @@ void handler(int file_descriptor) {
|
||||
if (!charset_wire_init_receiver(config->iconv_spec, server_iconv_spec)) {
|
||||
log_message(LOG_LEVEL_ERROR,
|
||||
"--iconv: unsupported charset conversion requested (LOCAL[,REMOTE])");
|
||||
goto done;
|
||||
return false;
|
||||
}
|
||||
charset_ready = true;
|
||||
state->charset_ready = true;
|
||||
}
|
||||
/* --delete-missing-args deletes destination mirrors receiver-side, so it is
|
||||
deletion and stays gated by the same --allow-delete server policy. When
|
||||
the server policy is off the flag is inert (the missing entries are still
|
||||
skipped via its implied --ignore-missing-args, but nothing is deleted). */
|
||||
* deletion and stays gated by the same --allow-delete server policy. When
|
||||
* the server policy is off the flag is inert (the missing entries are still
|
||||
* skipped via its implied --ignore-missing-args, but nothing is deleted). */
|
||||
config->delete_missing_args = config->delete_missing_args && allow_delete;
|
||||
/* --mkpath: create the destination root (and its missing leading components)
|
||||
* before anything else; without it the root must pre-exist. The precondition
|
||||
@@ -844,7 +867,7 @@ void handler(int file_descriptor) {
|
||||
log_message(LOG_LEVEL_ERROR, "destination root is not available: %s",
|
||||
escaped_root ? escaped_root : "<allocation failed>");
|
||||
free(escaped_root);
|
||||
goto done;
|
||||
return false;
|
||||
}
|
||||
/* A --delay-updates transfer stages under a private 0700 directory inside
|
||||
the receive root. Create it up front (wiping leftovers of any previously
|
||||
@@ -854,7 +877,7 @@ void handler(int file_descriptor) {
|
||||
config->delay_context = delay_updates_context_create(config->receive_root_directory);
|
||||
if (!config->delay_context || !delay_updates_prepare(config->delay_context)) {
|
||||
log_message(LOG_LEVEL_ERROR, "Failed to initialize --delay-updates staging area");
|
||||
goto done;
|
||||
return false;
|
||||
}
|
||||
}
|
||||
/* Preserve the negotiated identity policy for the fd-relative ownership
|
||||
@@ -864,7 +887,7 @@ void handler(int file_descriptor) {
|
||||
rather than silently applying the wrong ownership policy. */
|
||||
if (!identity_set_active(config)) {
|
||||
log_message(LOG_LEVEL_ERROR, "Failed to activate identity policy");
|
||||
goto done;
|
||||
return false;
|
||||
}
|
||||
/* Persist the negotiated --keep-dirlinks policy once, here at config-accept,
|
||||
before any multithreaded receiver/writer threads are spawned, so the
|
||||
@@ -892,136 +915,195 @@ void handler(int file_descriptor) {
|
||||
Wave C note in config.h). */
|
||||
if (g_daemon_conf) {
|
||||
char* motd = motd_read_file(g_daemon_conf->global.motd_file);
|
||||
if (!motd_send(file_descriptor, motd ? motd : "")) {
|
||||
if (!motd_send(state->fd, motd ? motd : "")) {
|
||||
free(motd);
|
||||
log_message(LOG_LEVEL_ERROR, "Failed to send daemon MOTD");
|
||||
goto done;
|
||||
return false;
|
||||
}
|
||||
free(motd);
|
||||
}
|
||||
if (config->use_multithreading) {
|
||||
Queue* q = queue_create(100, file_destroy);
|
||||
if (q == NULL)
|
||||
goto done;
|
||||
context = pipeline_context_receiver_create(config, q, file_descriptor, ssl);
|
||||
if (context == NULL) {
|
||||
queue_destroy(q);
|
||||
goto done;
|
||||
}
|
||||
protocol_session_set_max_alloc(&context->session, config->max_alloc);
|
||||
protocol_session_set_io_timeout(&context->session,
|
||||
protocol_server_io_timeout_sec(config->timeout));
|
||||
atomic_store(&context->session.total_allocated_bytes,
|
||||
atomic_load(&session.total_allocated_bytes));
|
||||
pipeline_context_receiver_set_queue_byte_limit(context, RECEIVER_QUEUE_MAX_BYTES);
|
||||
thrd_t receiver = {0};
|
||||
thrd_t writer = {0};
|
||||
bool receiver_created = thrd_create(&receiver, receive_thread, context) == thrd_success;
|
||||
bool writer_created = false;
|
||||
if (receiver_created)
|
||||
writer_created = thrd_create(&writer, write_thread, context) == thrd_success;
|
||||
if (!receiver_created || !writer_created) {
|
||||
log_perror("Error creating Threads");
|
||||
if (receiver_created) {
|
||||
mtx_lock(&context->mutex);
|
||||
atomic_store(&context->cancelled, true);
|
||||
cnd_broadcast(&context->condition_not_full);
|
||||
cnd_broadcast(&context->condition_not_empty);
|
||||
mtx_unlock(&context->mutex);
|
||||
/* Unblock a worker parked in socket I/O without closing the fd (the
|
||||
* child owns the single close). shutdown() only affects sockets; for
|
||||
* the --stdio pipe the receiver's per-message poll timeout still
|
||||
* bounds the join, so do nothing there rather than close a descriptor
|
||||
* another thread may still be using. */
|
||||
struct stat fd_stat;
|
||||
if (fstat(file_descriptor, &fd_stat) == 0 && S_ISSOCK(fd_stat.st_mode))
|
||||
shutdown(file_descriptor, SHUT_RDWR);
|
||||
thrd_join(receiver, NULL);
|
||||
}
|
||||
if (writer_created)
|
||||
thrd_join(writer, NULL);
|
||||
goto done;
|
||||
}
|
||||
int receiver_result;
|
||||
int writer_result;
|
||||
thrd_join(receiver, &receiver_result);
|
||||
thrd_join(writer, &writer_result);
|
||||
bool transfer_ok = receiver_result == thrd_success && writer_result == thrd_success;
|
||||
if (transfer_ok && !config->dry_run) {
|
||||
/* Commit-style (late) deletion: receive_thread handed the keep-set
|
||||
manifest here instead of deleting while write_thread might still be
|
||||
draining, so by now every file is on disk and the whole transfer is
|
||||
known to have succeeded. Remove the extras before publishing a
|
||||
--delay-updates run; the walker skips the staging directory. A
|
||||
server-contacting --dry-run deletes nothing (no manifest is sent). */
|
||||
if (context->deferred_manifest) {
|
||||
size_t deleted = 0;
|
||||
DeleteCommitResult deletion =
|
||||
manifest_delete_all_counted(config, context->deferred_manifest, &deleted);
|
||||
context->stats.deleted_files += deleted;
|
||||
if (deletion == DELETE_COMMIT_ERROR) {
|
||||
transfer_ok = false;
|
||||
} else if (deletion == DELETE_COMMIT_LIMIT_REACHED) {
|
||||
/* The transfer still succeeds; the terminal frame reports the capped
|
||||
deletion so the sender exits 25 like rsync. */
|
||||
context->delete_limit_reached = true;
|
||||
}
|
||||
delete_manifest_free(context->deferred_manifest);
|
||||
context->deferred_manifest = NULL;
|
||||
}
|
||||
/* --delete-delay: receive_thread snapshotted each plan's extras as it
|
||||
arrived; with the disk writer drained, commit the deferred removals.
|
||||
--delete-during already applied its plans on the receive thread. */
|
||||
if (context->deferred_plans) {
|
||||
/* Defence in depth (the enclosing block already excludes dry-run): a
|
||||
-n run never commits a deletion. */
|
||||
DeleteCommitResult deletion =
|
||||
config->dry_run ? DELETE_COMMIT_OK
|
||||
: delete_plan_session_commit(context->deferred_plans, config);
|
||||
context->stats.deleted_files += delete_plan_session_deleted(context->deferred_plans);
|
||||
if (deletion == DELETE_COMMIT_ERROR) {
|
||||
transfer_ok = false;
|
||||
} else if (deletion == DELETE_COMMIT_LIMIT_REACHED) {
|
||||
context->delete_limit_reached = true;
|
||||
}
|
||||
delete_plan_session_destroy(context->deferred_plans);
|
||||
context->deferred_plans = NULL;
|
||||
}
|
||||
}
|
||||
if (transfer_ok && !config->dry_run) {
|
||||
/* --delay-updates: receive_thread has finished the whole protocol stream
|
||||
(including manifest/delete handling) and write_thread has drained its
|
||||
queue, so every staged file is complete. Publish atomically before the
|
||||
success/outcome frame so a --remove-source-files sender only learns of
|
||||
files that were actually installed. */
|
||||
if (config->delay_updates && config->delay_context &&
|
||||
!delay_updates_publish(config->delay_context, config)) {
|
||||
transfer_ok = false;
|
||||
}
|
||||
/* P7 Wave D: all writers have joined and the late deletion (and
|
||||
--delay-updates publication) has committed above, so it is finally safe
|
||||
to stamp directory times; a directory's mtime must not be clobbered by
|
||||
its children or by an extra removal. */
|
||||
if (transfer_ok)
|
||||
dir_metadata_list_apply(&context->dir_times, config->receive_root_directory, config);
|
||||
}
|
||||
if (transfer_ok) {
|
||||
Status final_status = context->delete_limit_reached ? STATUS_DELETE_LIMIT : STATUS_OK;
|
||||
/* Emit the optional wire-stats record first (protocol 2.25.0), then the
|
||||
success/outcome frame, exactly like the single-threaded receiver. */
|
||||
if (!receiver_send_stats_frame(file_descriptor, config, &context->stats,
|
||||
context->would_delete) ||
|
||||
!receiver_send_final_success(file_descriptor, config, &context->outcomes, final_status))
|
||||
transfer_ok = false;
|
||||
} else {
|
||||
send_error_detail(file_descriptor, "transfer failed on receiver");
|
||||
}
|
||||
if (!transfer_ok)
|
||||
log_message(LOG_LEVEL_ERROR, "Transfer failed");
|
||||
} else {
|
||||
if (receiver_receive_files(config, file_descriptor) != 0)
|
||||
log_message(LOG_LEVEL_ERROR, "Transfer failed");
|
||||
return true;
|
||||
}
|
||||
|
||||
/* Phase 4a -- transfer via the multithreaded receiver. Spawns the receive/write
|
||||
* thread pair, joins them, then commits the late deletion, --delay-updates
|
||||
* publication and directory times before emitting the terminal stats/success
|
||||
* frame. On any failure the helper just returns; the caller's `done` epilogue
|
||||
* releases the pipeline context (which owns the config and queue) exactly as the
|
||||
* former inline code did. */
|
||||
static void server_run_mt_receiver(ServerSession* state) {
|
||||
Config* config = state->config;
|
||||
Queue* q = queue_create(100, file_destroy);
|
||||
if (q == NULL)
|
||||
return;
|
||||
state->context = pipeline_context_receiver_create(config, q, state->fd, state->ssl);
|
||||
if (state->context == NULL) {
|
||||
queue_destroy(q);
|
||||
return;
|
||||
}
|
||||
protocol_session_set_max_alloc(&state->context->session, config->max_alloc);
|
||||
protocol_session_set_io_timeout(&state->context->session,
|
||||
protocol_server_io_timeout_sec(config->timeout));
|
||||
atomic_store(&state->context->session.total_allocated_bytes,
|
||||
atomic_load(&state->session.total_allocated_bytes));
|
||||
pipeline_context_receiver_set_queue_byte_limit(state->context, RECEIVER_QUEUE_MAX_BYTES);
|
||||
thrd_t receiver = {0};
|
||||
thrd_t writer = {0};
|
||||
bool receiver_created = thrd_create(&receiver, receive_thread, state->context) == thrd_success;
|
||||
bool writer_created = false;
|
||||
if (receiver_created)
|
||||
writer_created = thrd_create(&writer, write_thread, state->context) == thrd_success;
|
||||
if (!receiver_created || !writer_created) {
|
||||
log_perror("Error creating Threads");
|
||||
if (receiver_created) {
|
||||
mtx_lock(&state->context->mutex);
|
||||
atomic_store(&state->context->cancelled, true);
|
||||
cnd_broadcast(&state->context->condition_not_full);
|
||||
cnd_broadcast(&state->context->condition_not_empty);
|
||||
mtx_unlock(&state->context->mutex);
|
||||
/* Unblock a worker parked in socket I/O without closing the fd (the
|
||||
* child owns the single close). shutdown() only affects sockets; for
|
||||
* the --stdio pipe the receiver's per-message poll timeout still
|
||||
* bounds the join, so do nothing there rather than close a descriptor
|
||||
* another thread may still be using. */
|
||||
struct stat fd_stat;
|
||||
if (fstat(state->fd, &fd_stat) == 0 && S_ISSOCK(fd_stat.st_mode))
|
||||
shutdown(state->fd, SHUT_RDWR);
|
||||
thrd_join(receiver, NULL);
|
||||
}
|
||||
if (writer_created)
|
||||
thrd_join(writer, NULL);
|
||||
return;
|
||||
}
|
||||
int receiver_result;
|
||||
int writer_result;
|
||||
thrd_join(receiver, &receiver_result);
|
||||
thrd_join(writer, &writer_result);
|
||||
bool transfer_ok = receiver_result == thrd_success && writer_result == thrd_success;
|
||||
PipelineContextReceiver* context = state->context;
|
||||
if (transfer_ok && !config->dry_run) {
|
||||
/* Commit-style (late) deletion: receive_thread handed the keep-set
|
||||
manifest here instead of deleting while write_thread might still be
|
||||
draining, so by now every file is on disk and the whole transfer is
|
||||
known to have succeeded. Remove the extras before publishing a
|
||||
--delay-updates run; the walker skips the staging directory. A
|
||||
server-contacting --dry-run deletes nothing (no manifest is sent). */
|
||||
if (context->deferred_manifest) {
|
||||
size_t deleted = 0;
|
||||
DeletePathObserver observer = config->report_deletes ? receiver_record_deleted_path : NULL;
|
||||
DeleteCommitResult deletion = manifest_delete_all_observed(
|
||||
config, context->deferred_manifest, &deleted, observer, (void*)context->deleted_paths);
|
||||
context->stats.deleted_files += deleted;
|
||||
if (deletion == DELETE_COMMIT_ERROR) {
|
||||
transfer_ok = false;
|
||||
} else if (deletion == DELETE_COMMIT_LIMIT_REACHED) {
|
||||
/* The transfer still succeeds; the terminal frame reports the capped
|
||||
deletion so the sender exits 25 like rsync. */
|
||||
context->delete_limit_reached = true;
|
||||
}
|
||||
delete_manifest_free(context->deferred_manifest);
|
||||
context->deferred_manifest = NULL;
|
||||
}
|
||||
/* --delete-delay: receive_thread snapshotted each plan's extras as it
|
||||
arrived; with the disk writer drained, commit the deferred removals.
|
||||
--delete-during already applied its plans on the receive thread. */
|
||||
if (context->deferred_plans) {
|
||||
/* Defence in depth (the enclosing block already excludes dry-run): a
|
||||
-n run never commits a deletion. */
|
||||
if (config->report_deletes)
|
||||
delete_plan_session_set_delete_observer(
|
||||
context->deferred_plans, receiver_record_deleted_path, (void*)context->deleted_paths);
|
||||
DeleteCommitResult deletion =
|
||||
config->dry_run ? DELETE_COMMIT_OK
|
||||
: delete_plan_session_commit(context->deferred_plans, config);
|
||||
context->stats.deleted_files += delete_plan_session_deleted(context->deferred_plans);
|
||||
if (deletion == DELETE_COMMIT_ERROR) {
|
||||
transfer_ok = false;
|
||||
} else if (deletion == DELETE_COMMIT_LIMIT_REACHED) {
|
||||
context->delete_limit_reached = true;
|
||||
}
|
||||
delete_plan_session_destroy(context->deferred_plans);
|
||||
context->deferred_plans = NULL;
|
||||
}
|
||||
}
|
||||
if (transfer_ok && !config->dry_run) {
|
||||
/* --delay-updates: receive_thread has finished the whole protocol stream
|
||||
(including manifest/delete handling) and write_thread has drained its
|
||||
queue, so every staged file is complete. Publish atomically before the
|
||||
success/outcome frame so a --remove-source-files sender only learns of
|
||||
files that were actually installed. */
|
||||
if (config->delay_updates && config->delay_context &&
|
||||
!delay_updates_publish(config->delay_context, config)) {
|
||||
transfer_ok = false;
|
||||
}
|
||||
/* P7 Wave D: all writers have joined and the late deletion (and
|
||||
--delay-updates publication) has committed above, so it is finally safe
|
||||
to stamp directory times; a directory's mtime must not be clobbered by
|
||||
its children or by an extra removal. */
|
||||
if (transfer_ok)
|
||||
dir_metadata_list_apply(&context->dir_times, config->receive_root_directory, config);
|
||||
}
|
||||
if (transfer_ok) {
|
||||
Status final_status = context->delete_limit_reached ? STATUS_DELETE_LIMIT : STATUS_OK;
|
||||
/* Emit the optional wire-stats record first (protocol 2.25.0), then the
|
||||
success/outcome frame, exactly like the single-threaded receiver. */
|
||||
if (!receiver_send_stats_frame(state->fd, config, &context->stats, context->would_delete,
|
||||
context->deleted_paths) ||
|
||||
!receiver_send_final_success(state->fd, config, &context->outcomes, final_status))
|
||||
transfer_ok = false;
|
||||
} else {
|
||||
send_error_detail(state->fd, "transfer failed on receiver");
|
||||
}
|
||||
if (!transfer_ok)
|
||||
log_message(LOG_LEVEL_ERROR, "Transfer failed");
|
||||
}
|
||||
|
||||
/* Phase 4b -- transfer via the single-threaded receiver. Failure is logged
|
||||
* exactly as before; the caller's `done` epilogue then releases the config. */
|
||||
static void server_run_st_receiver(ServerSession* state) {
|
||||
if (receiver_receive_files(state->config, state->fd) != 0)
|
||||
log_message(LOG_LEVEL_ERROR, "Transfer failed");
|
||||
}
|
||||
|
||||
/* Phase 4 dispatch -- choose the receiver implementation the config asks for.
|
||||
* Both helpers own their success/failure logging; the caller falls through to
|
||||
* the shared `done` epilogue either way. */
|
||||
static void server_run_transfer(ServerSession* state) {
|
||||
if (state->config->use_multithreading)
|
||||
server_run_mt_receiver(state);
|
||||
else
|
||||
server_run_st_receiver(state);
|
||||
}
|
||||
|
||||
void handler(int file_descriptor) {
|
||||
/* Single per-connection state; every phase helper below advances it and
|
||||
* returns false on a logged failure. All teardown state starts empty so the
|
||||
* single `done` epilogue is safe to reach from any error path (including
|
||||
* before the config frame arrives). */
|
||||
ServerSession state;
|
||||
state.fd = file_descriptor;
|
||||
state.ssl = io_get_ssl();
|
||||
protocol_session_init(&state.session, file_descriptor, file_descriptor);
|
||||
protocol_session_set_ssl(&state.session, state.ssl);
|
||||
protocol_session_bind(&state.session);
|
||||
state.gate_ctx.ssl = state.ssl;
|
||||
state.gate_ctx.fd = file_descriptor;
|
||||
state.gate_ctx.super_mode_override = -1;
|
||||
state.gate_ctx.has_peer_ip = false;
|
||||
state.gate_ctx.peer_ip[0] = '\0';
|
||||
state.gate_ctx.is_local = false;
|
||||
state.config = NULL;
|
||||
state.context = NULL;
|
||||
state.joined_destination = NULL;
|
||||
state.charset_ready = false;
|
||||
|
||||
if (!server_accept_config(&state))
|
||||
goto done;
|
||||
if (!server_apply_security_gates(&state))
|
||||
goto done;
|
||||
if (!server_prepare_session(&state))
|
||||
goto done;
|
||||
server_run_transfer(&state);
|
||||
|
||||
done:
|
||||
/* Single cleanup epilogue: every error path jumps here, so the iconv
|
||||
@@ -1030,38 +1112,63 @@ done:
|
||||
* connection fd is deliberately NOT closed here -- the child functions own
|
||||
* its single close (plain_child_fn / tls_child_fn), and the --stdio call
|
||||
* site must leave stdin/stdout open. */
|
||||
if (charset_ready)
|
||||
if (state.charset_ready)
|
||||
charset_wire_free();
|
||||
/* The delay-updates staging tree is released by config_delete (which the
|
||||
branch below always reaches), so it is cleaned exactly once. */
|
||||
identity_clear_active();
|
||||
protocol_session_unbind();
|
||||
if (context != NULL) {
|
||||
if (state.context != NULL) {
|
||||
/* context owns both the config and the queue it was created with. */
|
||||
pipeline_context_receiver_destroy(context);
|
||||
context = NULL;
|
||||
config = NULL;
|
||||
pipeline_context_receiver_destroy(state.context);
|
||||
state.context = NULL;
|
||||
state.config = NULL;
|
||||
} else {
|
||||
config_delete(config);
|
||||
config = NULL;
|
||||
config_delete(state.config);
|
||||
state.config = NULL;
|
||||
}
|
||||
free(joined_destination);
|
||||
free(state.joined_destination);
|
||||
}
|
||||
|
||||
#ifndef FASTSYNC_SERVER_AS_LIB
|
||||
static Server* g_server = NULL;
|
||||
|
||||
/* Signal handler for the foreground daemon/standalone listener.
|
||||
*
|
||||
* Async-signal-safety: _exit(2) is on the POSIX async-signal-safe list and is
|
||||
* the ONLY thing done here. The previous body called server_delete()
|
||||
* (close/free/SSL_CTX_free), daemon_conf_free() and credentials_free(); none of
|
||||
* those (free/malloc, and much of OpenSSL teardown) are async-signal-safe, so a
|
||||
* signal delivered while the main thread was inside malloc/free could deadlock
|
||||
* or corrupt the heap.
|
||||
*
|
||||
* Residual (documented, not hidden): the in-memory teardown is skipped on the
|
||||
* signal path. That is safe because the parent daemon owns no persistent
|
||||
* resource that survives process exit -- the listening socket is closed by the
|
||||
* kernel, the connection registry is an anonymous MAP_SHARED mapping with no
|
||||
* named backing object, and the daemon config/credential stores are plain heap
|
||||
* allocations. Connection children are separate processes and handle their own
|
||||
* temp files/locks. The normal (non-signal) shutdown path in main() still runs
|
||||
* the full teardown, so no cleanup is dropped on the common path. Wiring the
|
||||
* accept loop (transport_tcp.c, outside this change's scope) to a flag-based
|
||||
* self-pipe shutdown would let the frees run context-safely; it is deliberately
|
||||
* deferred rather than risk restructuring the daemon loop. */
|
||||
static void cleanup(int sig) {
|
||||
(void)sig;
|
||||
if (g_server)
|
||||
server_delete(&g_server);
|
||||
daemon_conf_free(g_daemon_conf);
|
||||
g_daemon_conf = NULL;
|
||||
credentials_free(g_credentials);
|
||||
g_credentials = NULL;
|
||||
_exit(0);
|
||||
}
|
||||
|
||||
/* Install a signal handler with sigaction(2) (the required async-signal-safe
|
||||
* install primitive; signal(3) is not specified to be async-signal-safe). */
|
||||
static void install_cleanup_handler(int signo) {
|
||||
struct sigaction action;
|
||||
memset(&action, 0, sizeof(action));
|
||||
action.sa_handler = cleanup;
|
||||
sigemptyset(&action.sa_mask);
|
||||
action.sa_flags = 0;
|
||||
sigaction(signo, &action, NULL);
|
||||
}
|
||||
|
||||
static void print_server_usage(void) {
|
||||
printf("FastSync Server\n");
|
||||
printf("Usage: fastsync-server [options]\n\n");
|
||||
@@ -1174,13 +1281,19 @@ static bool daemonize(void) {
|
||||
close(devnull);
|
||||
}
|
||||
/* Do not pin the launch CWD (module-relative 'path' entries would resolve
|
||||
* against an unstable working directory) and drop the restrictive host umask
|
||||
* so modules can create files/dirs with the modes the config requests. */
|
||||
* against an unstable working directory). Set a conservative daemon umask
|
||||
* of 022 (the conventional service default): rsync never forces umask 0 --
|
||||
* it reads and restores the inherited umask and creates new entries as
|
||||
* 0777 & ~umask / source & ~umask without -p. Forcing 0 here made every
|
||||
* implied parent directory world-writable (0777) whenever -p metadata was not
|
||||
* applied. 022 gives 0755 directories and source&~022 files, matching rsync
|
||||
* under a normal daemon umask; -p/-a still restore the exact source mode via
|
||||
* fchmod, which is unaffected by the umask. */
|
||||
if (chdir("/") != 0)
|
||||
log_message(LOG_LEVEL_WARNING, "daemon: chdir to / failed: %s", strerror(errno));
|
||||
umask(0);
|
||||
umask(022);
|
||||
/* Refresh the cached umask: main() captured the launch umask before this
|
||||
* (single-threaded) umask(0), and file_mode_base() must see the daemon's
|
||||
* (single-threaded) umask(022), and file_mode_base() must see the daemon's
|
||||
* actual umask. */
|
||||
file_umask_capture();
|
||||
return true;
|
||||
@@ -1249,8 +1362,8 @@ int main(int argc, char* argv[]) {
|
||||
* this process-global policy cannot be re-enabled by a future caller. */
|
||||
server_allow_super = opts.allow_super && !opts.stdio_mode;
|
||||
server_iconv_spec = opts.iconv_spec;
|
||||
signal(SIGINT, cleanup);
|
||||
signal(SIGTERM, cleanup);
|
||||
install_cleanup_handler(SIGINT);
|
||||
install_cleanup_handler(SIGTERM);
|
||||
/* Server-owned socket deadline floor: the client default --timeout=0 would
|
||||
* otherwise leave accepted sockets without SO_RCVTIMEO/SO_SNDTIMEO and let a
|
||||
* silent peer hold a connection (and its process slot) forever. */
|
||||
|
||||
@@ -3,20 +3,12 @@
|
||||
#include "credentials.h"
|
||||
#include "utils.h"
|
||||
#include <limits.h>
|
||||
#include <stdarg.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <sys/socket.h>
|
||||
|
||||
static void set_error(char* err, size_t err_size, const char* fmt, ...) {
|
||||
if (!err || err_size == 0)
|
||||
return;
|
||||
va_list args;
|
||||
va_start(args, fmt);
|
||||
vsnprintf(err, err_size, fmt, args);
|
||||
va_end(args);
|
||||
}
|
||||
#define set_error utils_set_error
|
||||
|
||||
void server_cli_options_default(ServerCliOptions* opts) {
|
||||
if (!opts)
|
||||
|
||||
@@ -9,7 +9,7 @@
|
||||
ArrayList* array_list_create(void (*item_destroyer)(void* item)) {
|
||||
ArrayList* list = (ArrayList*)protocol_alloc(sizeof(ArrayList));
|
||||
if (list == NULL) {
|
||||
log_perror("ERROR: Could not allocate memory for array list struct");
|
||||
log_message(LOG_LEVEL_ERROR, "%s", "ERROR: Could not allocate memory for array list struct");
|
||||
return NULL;
|
||||
}
|
||||
|
||||
@@ -47,7 +47,7 @@ static bool array_list_extend(ArrayList* array_list) {
|
||||
new_capacity = INITIAL_ARRAY_SIZE;
|
||||
void* new_items = protocol_realloc(array_list->items, new_capacity * sizeof(void*));
|
||||
if (new_items == NULL) {
|
||||
log_perror("ERROR: Could not reallocate memory for array list items");
|
||||
log_message(LOG_LEVEL_ERROR, "%s", "ERROR: Could not reallocate memory for array list items");
|
||||
return false;
|
||||
}
|
||||
array_list->items = new_items;
|
||||
@@ -73,7 +73,7 @@ void** array_list_to_array(const ArrayList* array_list) {
|
||||
}
|
||||
void** array = protocol_alloc(array_list->size * sizeof(void*));
|
||||
if (array == NULL) {
|
||||
log_perror("Could not malloc space for array from array list!");
|
||||
log_message(LOG_LEVEL_ERROR, "%s", "Could not malloc space for array from array list!");
|
||||
return NULL;
|
||||
}
|
||||
memcpy(array, array_list->items, array_list->size * sizeof(void*));
|
||||
|
||||
+15
-10
@@ -168,9 +168,14 @@ bool charset_spec_valid_direction(const char* from_charset, const char* to_chars
|
||||
return direction_probe_valid(from_charset, to_charset);
|
||||
}
|
||||
|
||||
/* The receiver's real conversion is wire(client REMOTE) -> server-local (the
|
||||
* server's own --iconv LOCAL half, or the client's LOCAL half when the server
|
||||
* has no --iconv). A dedicated pre-ack check so an impossible direction is
|
||||
/* The receiver's conversion is wire charset -> destination charset. rsync's
|
||||
* CONVERT_SPEC is LOCAL,REMOTE and "stays the same whether you're pushing or
|
||||
* pulling", so for a PUSH (FastSync's only direction) the destination end's
|
||||
* charset is the spec's REMOTE half: the client converts LOCAL -> REMOTE on the
|
||||
* sender and the receiver writes the wire bytes verbatim. Only a server that
|
||||
* declares its OWN --iconv (the daemon "charset" analog) has a different local
|
||||
* charset, and then it is that spec's LOCAL half and the receiver converts
|
||||
* wire -> server-local. A dedicated pre-ack check so an impossible direction is
|
||||
* rejected before the connection instead of refusing mid-transfer. */
|
||||
bool charset_wire_receiver_spec_valid(const char* spec, const char* server_spec) {
|
||||
if (!spec)
|
||||
@@ -180,7 +185,7 @@ bool charset_wire_receiver_spec_valid(const char* spec, const char* server_spec)
|
||||
if (charset_spec_parse(spec, &local, &remote) != 0)
|
||||
return false;
|
||||
const char* wire = remote;
|
||||
const char* target_local = local;
|
||||
const char* target_local = remote;
|
||||
char* server_local = NULL;
|
||||
char* server_remote = NULL;
|
||||
if (server_spec) {
|
||||
@@ -302,13 +307,13 @@ bool charset_wire_init_receiver(const char* spec, const char* server_spec) {
|
||||
char* remote;
|
||||
if (charset_spec_parse(spec, &local, &remote) != 0)
|
||||
return false;
|
||||
/* The wire charset is the client spec's REMOTE half; the local charset is
|
||||
* the client spec's LOCAL half unless the server was itself started with
|
||||
* --iconv naming a different local charset (the server halves above never
|
||||
* travel, so the server's own flag is the only way its local charset can
|
||||
* differ from what the client assumed). */
|
||||
/* The wire charset is the client spec's REMOTE half (rsync's LOCAL,REMOTE
|
||||
* spec stays the same push or pull, so on a push the destination end's
|
||||
* charset is REMOTE and the receiver writes the wire bytes verbatim). Only a
|
||||
* server started with its own --iconv declares a different local charset (the
|
||||
* server halves above never travel), and then it is that spec's LOCAL half. */
|
||||
const char* wire = remote;
|
||||
const char* target_local = local;
|
||||
const char* target_local = remote;
|
||||
char* server_local = NULL;
|
||||
char* server_remote = NULL;
|
||||
if (server_spec) {
|
||||
|
||||
@@ -57,8 +57,9 @@ void charset_conversion_close(void* conversion);
|
||||
/* Process-wide wire conversion. charset_wire_init_sender (client side) opens
|
||||
* LOCAL->REMOTE; charset_wire_init_receiver (server side) opens
|
||||
* wire(REMOTE)->server-local. server_spec is the server's own --iconv, whose
|
||||
* LOCAL half may override the local charset the client assumed; NULL reuses
|
||||
* the client spec's LOCAL half. Both return false on an unsupported spec.
|
||||
* LOCAL half overrides the destination charset; NULL means the destination
|
||||
* charset is the client spec's REMOTE half (rsync's push semantics: the wire
|
||||
* bytes are written verbatim). Both return false on an unsupported spec.
|
||||
* The state is freed with charset_wire_free. */
|
||||
bool charset_wire_init_sender(const char* spec);
|
||||
bool charset_wire_init_receiver(const char* spec, const char* server_spec);
|
||||
@@ -66,9 +67,9 @@ void charset_wire_free(void);
|
||||
bool charset_wire_active(void);
|
||||
|
||||
/* Pre-ack receiver-direction sanity (see charset_wire_init_receiver): true
|
||||
* when the exact wire->server-local conversion the receiver will use (client
|
||||
* spec's REMOTE half into the server's own LOCAL half, or the client's LOCAL
|
||||
* half when the server has no --iconv) opens and produces NUL-free output. */
|
||||
* when the exact wire->destination conversion the receiver will use (client
|
||||
* spec's REMOTE half into the server's own LOCAL half, or REMOTE->REMOTE when
|
||||
* the server has no --iconv) opens and produces NUL-free output. */
|
||||
bool charset_wire_receiver_spec_valid(const char* spec, const char* server_spec);
|
||||
|
||||
/* Convert a path across the wire in the process direction. Returns a malloc'd
|
||||
|
||||
+53
-12
@@ -1,4 +1,5 @@
|
||||
#include "checksum.h"
|
||||
#include "utils.h"
|
||||
#include <fcntl.h>
|
||||
#include <openssl/evp.h>
|
||||
#include <string.h>
|
||||
@@ -223,17 +224,33 @@ bool checksum_digest_file(ChecksumAlgo algo, uint64_t seed, const char* path, ui
|
||||
if (fd < 0)
|
||||
return false;
|
||||
|
||||
bool ok = checksum_digest_fd(algo, seed, fd, out, out_capacity, out_len);
|
||||
close(fd);
|
||||
return ok;
|
||||
}
|
||||
|
||||
bool checksum_digest_fd(ChecksumAlgo algo, uint64_t seed, int fd, uint8_t* out, size_t out_capacity,
|
||||
size_t* out_len) {
|
||||
if (fd < 0 || !out || !out_len || out_capacity < CHECKSUM_MAX_DIGEST_LEN)
|
||||
return false;
|
||||
|
||||
if (algo == CHECKSUM_ALGO_NONE) {
|
||||
/* No checksum requested: nothing to read; an empty digest succeeds. */
|
||||
*out_len = 0;
|
||||
return true;
|
||||
}
|
||||
|
||||
uint8_t buffer[64 * 1024];
|
||||
bool ok = false;
|
||||
lseek(fd, 0, SEEK_SET);
|
||||
|
||||
if (algo == CHECKSUM_ALGO_MD5) {
|
||||
if (algo == CHECKSUM_ALGO_MD5 || algo == CHECKSUM_ALGO_SHA1) {
|
||||
const EVP_MD* md = algo == CHECKSUM_ALGO_MD5 ? EVP_md5() : EVP_sha1();
|
||||
EVP_MD_CTX* ctx = EVP_MD_CTX_new();
|
||||
if (!ctx) {
|
||||
close(fd);
|
||||
if (!ctx)
|
||||
return false;
|
||||
}
|
||||
unsigned int digest_len = 0;
|
||||
if (EVP_DigestInit_ex(ctx, EVP_md5(), NULL) == 1) {
|
||||
if (EVP_DigestInit_ex(ctx, md, NULL) == 1) {
|
||||
ok = true;
|
||||
ssize_t got;
|
||||
while ((got = read(fd, buffer, sizeof(buffer))) > 0) {
|
||||
@@ -250,7 +267,22 @@ bool checksum_digest_file(ChecksumAlgo algo, uint64_t seed, const char* path, ui
|
||||
ok = false;
|
||||
}
|
||||
EVP_MD_CTX_free(ctx);
|
||||
close(fd);
|
||||
return ok;
|
||||
}
|
||||
|
||||
if (algo == CHECKSUM_ALGO_MD4) {
|
||||
Md4Ctx ctx;
|
||||
md4_init(&ctx);
|
||||
ok = true;
|
||||
ssize_t got;
|
||||
while ((got = read(fd, buffer, sizeof(buffer))) > 0)
|
||||
md4_update(&ctx, buffer, (size_t)got);
|
||||
if (got < 0)
|
||||
ok = false;
|
||||
if (ok) {
|
||||
md4_final(&ctx, out);
|
||||
*out_len = 16;
|
||||
}
|
||||
return ok;
|
||||
}
|
||||
|
||||
@@ -260,16 +292,13 @@ bool checksum_digest_file(ChecksumAlgo algo, uint64_t seed, const char* path, ui
|
||||
XXH64_reset(&xxh64, seed);
|
||||
} else if (algo == CHECKSUM_ALGO_XXH3 || algo == CHECKSUM_ALGO_XXH128) {
|
||||
xxh3 = XXH3_createState();
|
||||
if (!xxh3) {
|
||||
close(fd);
|
||||
if (!xxh3)
|
||||
return false;
|
||||
}
|
||||
if (algo == CHECKSUM_ALGO_XXH3)
|
||||
XXH3_64bits_reset_withSeed(xxh3, seed);
|
||||
else
|
||||
XXH3_128bits_reset_withSeed(xxh3, seed);
|
||||
} else {
|
||||
close(fd);
|
||||
return false;
|
||||
}
|
||||
|
||||
@@ -303,7 +332,6 @@ bool checksum_digest_file(ChecksumAlgo algo, uint64_t seed, const char* path, ui
|
||||
}
|
||||
if (xxh3)
|
||||
XXH3_freeState(xxh3);
|
||||
close(fd);
|
||||
return ok;
|
||||
}
|
||||
|
||||
@@ -371,7 +399,7 @@ uint8_t checksum_digest_len(ChecksumAlgo algo) {
|
||||
return 0;
|
||||
}
|
||||
|
||||
ChecksumAlgo checksum_negotiate_default(void) {
|
||||
static ChecksumAlgo compiled_checksum_preference_first(void) {
|
||||
/* rsync 3.4.1 default preference order; every entry is compiled in, so this
|
||||
* resolves to xxh128. */
|
||||
static const ChecksumAlgo preference[] = {
|
||||
@@ -384,3 +412,16 @@ ChecksumAlgo checksum_negotiate_default(void) {
|
||||
}
|
||||
return CHECKSUM_ALGO_XXH64;
|
||||
}
|
||||
|
||||
int checksum_choice_resolve(void) {
|
||||
bool specified = false;
|
||||
int env = env_choice_first("RSYNC_CHECKSUM_LIST", checksum_algo_from_name, &specified);
|
||||
if (specified)
|
||||
return env; /* -1 = the list named no supported checksum */
|
||||
return (int)compiled_checksum_preference_first();
|
||||
}
|
||||
|
||||
ChecksumAlgo checksum_negotiate_default(void) {
|
||||
int resolved = checksum_choice_resolve();
|
||||
return resolved >= 0 ? (ChecksumAlgo)resolved : compiled_checksum_preference_first();
|
||||
}
|
||||
|
||||
@@ -51,6 +51,13 @@ bool checksum_digest(ChecksumAlgo algo, uint64_t seed, const void* data, size_t
|
||||
bool checksum_digest_file(ChecksumAlgo algo, uint64_t seed, const char* path, uint8_t* out,
|
||||
size_t out_capacity, size_t* out_len);
|
||||
|
||||
/* Descriptor form of the streaming digest: rewinds `fd` to the start and hashes
|
||||
* to EOF without closing it. Used by the --verify-basis path to hash an
|
||||
* already-open, root-confined basis descriptor. Same contract as
|
||||
* checksum_digest_file. */
|
||||
bool checksum_digest_fd(ChecksumAlgo algo, uint64_t seed, int fd, uint8_t* out, size_t out_capacity,
|
||||
size_t* out_len);
|
||||
|
||||
/* Resolve a --checksum-choice string (case-insensitive) to an algorithm id.
|
||||
* Accepts "xxh64"/"xxhash", "xxh3", "xxh128", "md5", "md4", "sha1", "none".
|
||||
* "auto" is not an algorithm here; the caller resolves it to the negotiated
|
||||
@@ -72,4 +79,11 @@ uint8_t checksum_digest_len(ChecksumAlgo algo);
|
||||
* xxh128 xxh3 xxh64 md5 md4 sha1 none). Used to resolve "auto". */
|
||||
ChecksumAlgo checksum_negotiate_default(void);
|
||||
|
||||
/* Resolve "auto" the way rsync does: the first supported name in
|
||||
* RSYNC_CHECKSUM_LIST (whitespace-separated, client half ends at '&'), then the
|
||||
* compiled-in preference order when the variable is unset/blank. Returns -1
|
||||
* when the variable is set but names no supported checksum (rsync's failed
|
||||
* negotiation), otherwise a valid ChecksumAlgo id. */
|
||||
int checksum_choice_resolve(void);
|
||||
|
||||
#endif /* CHECKSUM_H */
|
||||
|
||||
+5
-4
@@ -17,8 +17,9 @@
|
||||
#include "protocol.h"
|
||||
#include "utils.h"
|
||||
|
||||
/* Maximum individual file data size within a chunk (64 MB) */
|
||||
#define MAX_FILE_DATA_SIZE (64ULL * 1024 * 1024)
|
||||
/* Maximum individual file data size within a chunk (64 MB). Distinct from the
|
||||
* receiver's whole-file MAX_FILE_DATA_SIZE (256 MB) in file_receive.c. */
|
||||
#define MAX_CHUNK_FILE_DATA_SIZE (64ULL * 1024 * 1024)
|
||||
#define MAX_FILES_PER_CHUNK 65536U
|
||||
|
||||
/* Reserve `charge` against `session`'s connection budget. This mirrors the
|
||||
@@ -382,9 +383,9 @@ Chunk* chunk_deserialize(Data* data, bool use_metadata) {
|
||||
}
|
||||
|
||||
// Reject individual file data larger than the maximum allowed size.
|
||||
if (file_data_size > MAX_FILE_DATA_SIZE) {
|
||||
if (file_data_size > MAX_CHUNK_FILE_DATA_SIZE) {
|
||||
log_message(LOG_LEVEL_ERROR, "File data size %zu exceeds maximum %llu", file_data_size,
|
||||
(unsigned long long)MAX_FILE_DATA_SIZE);
|
||||
(unsigned long long)MAX_CHUNK_FILE_DATA_SIZE);
|
||||
goto error;
|
||||
}
|
||||
|
||||
|
||||
@@ -2,6 +2,7 @@
|
||||
#include "data.h"
|
||||
#include "log.h"
|
||||
#include "protocol.h"
|
||||
#include "utils.h"
|
||||
#include <limits.h>
|
||||
#include <lz4.h>
|
||||
#include <stdatomic.h>
|
||||
@@ -15,7 +16,14 @@
|
||||
#include <zstd.h>
|
||||
|
||||
#define INITIAL_DECOMPRESS_BUF_SIZE (1024 * 1024)
|
||||
#define MAX_DECOMPRESSED_SIZE (100ULL * 1024 * 1024) /* 100 MB hard ceiling */
|
||||
|
||||
/* Hard ceiling for a single decompression. The sender compresses whole files
|
||||
* up to the protocol's whole-file receive bound, so the decompressor must
|
||||
* accept payloads that large; referencing the protocol constant keeps the two
|
||||
* bounds from drifting apart (they previously did: a 100 MB ceiling rejected
|
||||
* 100-256 MB files). This remains a real bomb guard -- every allocation in the
|
||||
* paths below is clamped to it -- so it must not exceed the protocol bound. */
|
||||
#define MAX_DECOMPRESSED_SIZE MAX_RECEIVE_WHOLE_FILE_SIZE
|
||||
|
||||
/* rsync 3.4.1's built-in skip-compress suffix list (the `--skip-compress`
|
||||
* defaults, in the man page's order). rsync stores it as space-separated
|
||||
@@ -119,7 +127,7 @@ bool compression_algo_enabled(CompressionAlgo algo) {
|
||||
return algo != COMPRESSION_ALGO_NONE;
|
||||
}
|
||||
|
||||
CompressionAlgo compression_negotiate_default(void) {
|
||||
static CompressionAlgo compiled_preference_first(void) {
|
||||
/* rsync 3.4.1 default preference order; every entry is compiled in, so this
|
||||
* resolves to zstd. */
|
||||
static const CompressionAlgo preference[] = {
|
||||
@@ -133,6 +141,57 @@ CompressionAlgo compression_negotiate_default(void) {
|
||||
return COMPRESSION_ALGO_ZSTD;
|
||||
}
|
||||
|
||||
int compression_choice_resolve(void) {
|
||||
bool specified = false;
|
||||
int env = env_choice_first("RSYNC_COMPRESS_LIST", compression_algo_from_name, &specified);
|
||||
if (specified)
|
||||
return env; /* -1 = the list named no supported codec */
|
||||
return (int)compiled_preference_first();
|
||||
}
|
||||
|
||||
CompressionAlgo compression_negotiate_default(void) {
|
||||
int resolved = compression_choice_resolve();
|
||||
return resolved >= 0 ? (CompressionAlgo)resolved : compiled_preference_first();
|
||||
}
|
||||
|
||||
int compression_default_level(CompressionAlgo algo) {
|
||||
switch (algo) {
|
||||
case COMPRESSION_ALGO_ZSTD:
|
||||
return ZSTD_CLEVEL_DEFAULT;
|
||||
case COMPRESSION_ALGO_ZLIB:
|
||||
case COMPRESSION_ALGO_ZLIBX:
|
||||
return 6; /* rsync resolves zlib's Z_DEFAULT_COMPRESSION (-1) to 6 */
|
||||
case COMPRESSION_ALGO_LZ4:
|
||||
return 1; /* rsync lz4 level is 0/ignored; positive keeps the gate on */
|
||||
case COMPRESSION_ALGO_NONE:
|
||||
return 0;
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
int compression_clamp_level(CompressionAlgo algo, int level) {
|
||||
switch (algo) {
|
||||
case COMPRESSION_ALGO_ZSTD:
|
||||
if (level < 1)
|
||||
return 1;
|
||||
if (level > 22)
|
||||
return 22;
|
||||
return level;
|
||||
case COMPRESSION_ALGO_ZLIB:
|
||||
case COMPRESSION_ALGO_ZLIBX:
|
||||
if (level < 1)
|
||||
return 1;
|
||||
if (level > 9)
|
||||
return 9;
|
||||
return level;
|
||||
case COMPRESSION_ALGO_LZ4:
|
||||
return 1; /* ignored by lz4_compress; keeps the "compress" gate on */
|
||||
case COMPRESSION_ALGO_NONE:
|
||||
return 0;
|
||||
}
|
||||
return level;
|
||||
}
|
||||
|
||||
void compression_set_algo(CompressionAlgo algo) {
|
||||
if (compression_algo_valid((int)algo))
|
||||
atomic_store(&g_compression_algo, (int)algo);
|
||||
@@ -585,8 +644,7 @@ static Data* zstd_decompress(Data* compressed_data, size_t maximum_size) {
|
||||
}
|
||||
if (ret > 0 && output.pos == output.size) {
|
||||
if (buf_size >= hard_limit || buf_size > SIZE_MAX / 2) {
|
||||
log_message(LOG_LEVEL_ERROR, "Decompressed data exceeds %llu bytes",
|
||||
(unsigned long long)MAX_DECOMPRESSED_SIZE);
|
||||
log_message(LOG_LEVEL_ERROR, "Decompressed data exceeds %llu bytes", hard_limit);
|
||||
data_destroy(uncompressed_data);
|
||||
uncompressed_data = NULL;
|
||||
goto cleanup;
|
||||
|
||||
@@ -35,6 +35,26 @@ bool compression_algo_valid(int algo);
|
||||
* "auto". */
|
||||
CompressionAlgo compression_negotiate_default(void);
|
||||
|
||||
/* Resolve "auto" the way rsync does: the first supported name in
|
||||
* RSYNC_COMPRESS_LIST (whitespace-separated, client half ends at '&'), then the
|
||||
* compiled-in preference order when the variable is unset/blank. Returns -1
|
||||
* when the variable is set but names no supported codec (rsync's failed
|
||||
* negotiation), otherwise a valid CompressionAlgo id. */
|
||||
int compression_choice_resolve(void);
|
||||
|
||||
/* rsync 3.4.1's per-codec default level, applied when the user did not pass
|
||||
* --compress-level/--zl. zstd uses ZSTD_CLEVEL_DEFAULT (3) and zlib/zlibx the
|
||||
* resolved Z_DEFAULT_COMPRESSION (6). lz4 has no tunable level in rsync
|
||||
* (always the default acceleration); FastSync returns a positive placeholder so
|
||||
* its "level > 0" compression gate stays engaged, and lz4_compress ignores the
|
||||
* value, so the output is identical to rsync's. none is 0. */
|
||||
int compression_default_level(CompressionAlgo algo);
|
||||
|
||||
/* Clamp an explicit --compress-level to the codec's accepted range the way
|
||||
* rsync's init_compression_level() does: zstd 1..22, zlib/zlibx 1..9, lz4
|
||||
* ignored (fixed positive placeholder), none 0. */
|
||||
int compression_clamp_level(CompressionAlgo algo, int level);
|
||||
|
||||
/* True when the algorithm actually compresses (i.e. is not NONE). */
|
||||
bool compression_algo_enabled(CompressionAlgo algo);
|
||||
|
||||
|
||||
+169
-16
@@ -20,9 +20,9 @@
|
||||
|
||||
static void config_set_defaults(Config* config) {
|
||||
config->scanner_threads = 0;
|
||||
config->metadata_explicitly_disabled = false;
|
||||
config->preserve_perms_explicit_off = false;
|
||||
config->preserve_times_explicit_off = false;
|
||||
config->cli.preserve_perms_explicit_off = false;
|
||||
config->cli.preserve_times_explicit_off = false;
|
||||
config->cli.metadata_explicitly_disabled = false;
|
||||
config->show_progress = false;
|
||||
config->compression_threads = 0;
|
||||
config->ssh_port = 22;
|
||||
@@ -44,8 +44,8 @@ static void config_set_defaults(Config* config) {
|
||||
config->tls_ca = NULL;
|
||||
config->server_host = str_dup("127.0.0.1");
|
||||
config->server_port = 8080;
|
||||
config->server_port_set = false;
|
||||
config->server_host_set = false;
|
||||
config->cli.server_port_set = false;
|
||||
config->cli.server_host_set = false;
|
||||
/* rsync defaults: --timeout=0 (I/O timeouts disabled) and --contimeout=60.
|
||||
* A value of 0 disables the client's own deadline on both the socket layer
|
||||
* (tcp_set_timeouts) and the protocol layer
|
||||
@@ -68,8 +68,10 @@ static void config_set_defaults(Config* config) {
|
||||
config->human_readable = false;
|
||||
config->ignore_errors = false;
|
||||
config->ignore_missing_args = false;
|
||||
config->checksum_transfer_algo = CHECKSUM_ALGO_DEFAULT;
|
||||
config->cli_exit_code = 0;
|
||||
config->cli.checksum_transfer_algo = CHECKSUM_ALGO_DEFAULT;
|
||||
config->cli.cli_exit_code = 0;
|
||||
config->cli.compression_level_set = false;
|
||||
config->cli.checksum_choice_set = false;
|
||||
config->filters = NULL;
|
||||
config->files_from = NULL;
|
||||
config->files_from_set = NULL;
|
||||
@@ -77,13 +79,12 @@ static void config_set_defaults(Config* config) {
|
||||
config->cvs_exclude = false;
|
||||
config->per_dir_filter = false;
|
||||
config->per_dir_filter_count = 0;
|
||||
config->one_file_system = false;
|
||||
config->one_file_system = 0;
|
||||
config->no_implied_dirs = false;
|
||||
config->dirs = false;
|
||||
config->rsh_command = NULL;
|
||||
config->blocking_io = false;
|
||||
config->outbuf = OUTBUF_BLOCK;
|
||||
config->old_args = false;
|
||||
config->remote_options = NULL;
|
||||
config->remote_option_count = 0;
|
||||
config->address = NULL;
|
||||
@@ -99,7 +100,7 @@ static void config_set_defaults(Config* config) {
|
||||
config->trust_sender = false;
|
||||
config->stop_after_mins = 0;
|
||||
config->stop_at = 0;
|
||||
config->stop_at_set = false;
|
||||
config->cli.stop_at_set = false;
|
||||
config->write_batch = NULL;
|
||||
config->only_write_batch = NULL;
|
||||
config->read_batch = NULL;
|
||||
@@ -206,7 +207,8 @@ static bool validate_received_config(const Config* config) {
|
||||
valid_wire_bool(config->preserve_perms) && valid_wire_bool(config->preserve_times) &&
|
||||
valid_wire_bool(config->preserve_owner) && valid_wire_bool(config->preserve_group) &&
|
||||
valid_wire_bool(config->munge_links) && valid_wire_bool(config->keep_dirlinks) &&
|
||||
valid_wire_bool(config->fake_super) &&
|
||||
valid_wire_bool(config->fake_super) && valid_wire_bool(config->report_dest_info) &&
|
||||
valid_wire_bool(config->report_stats) && valid_wire_bool(config->report_deletes) &&
|
||||
(!config->copy_as_set || (config->copy_as_uid >= 0 && config->copy_as_gid >= 0)) &&
|
||||
(!config->use_compression ||
|
||||
(config->compression_level >= 1 && config->compression_level <= 22)) &&
|
||||
@@ -227,6 +229,13 @@ Config* config_create(void) {
|
||||
if (!config)
|
||||
return NULL;
|
||||
config_set_defaults(config);
|
||||
/* config_set_defaults() dups the default server host; a failure there leaves
|
||||
* server_host NULL and would crash later consumers, so fail the whole create
|
||||
* (every caller already handles a NULL return). */
|
||||
if (!config->server_host) {
|
||||
config_delete(config);
|
||||
return NULL;
|
||||
}
|
||||
return config;
|
||||
}
|
||||
|
||||
@@ -332,7 +341,8 @@ bool config_derived_use_metadata(const Config* config) {
|
||||
config->chown_uid_set || config->chown_gid_set || config->usermap_count > 0 ||
|
||||
config->groupmap_count > 0 || config->update)
|
||||
return true;
|
||||
return (config->use_incremental || config->use_delta) && !config->metadata_explicitly_disabled;
|
||||
return (config->use_incremental || config->use_delta) &&
|
||||
!config->cli.metadata_explicitly_disabled;
|
||||
}
|
||||
|
||||
bool config_has_basis(const Config* config) {
|
||||
@@ -683,9 +693,15 @@ int config_parse_ssh_dest(Config* config) {
|
||||
return daemon_dest_parse_error("invalid remote destination user@host (must not be empty or "
|
||||
"start with '-')",
|
||||
dest);
|
||||
config->transport = TRANSPORT_SSH;
|
||||
config->ssh_destination = str_dup(dest);
|
||||
char* ssh_destination = str_dup(dest);
|
||||
char* path = str_dup(colon + 1);
|
||||
if (!ssh_destination || !path) {
|
||||
free(ssh_destination);
|
||||
free(path);
|
||||
return daemon_dest_parse_error("out of memory parsing remote destination", dest);
|
||||
}
|
||||
config->transport = TRANSPORT_SSH;
|
||||
config->ssh_destination = ssh_destination;
|
||||
free(config->receive_root_directory);
|
||||
config->receive_root_directory = path;
|
||||
return 0;
|
||||
@@ -788,6 +804,8 @@ void config_delete(Config* config) {
|
||||
if (config->filters) {
|
||||
array_list_delete(config->filters);
|
||||
}
|
||||
filter_rule_list_free(config->protect_rules);
|
||||
config->protect_rules = NULL;
|
||||
/* A --delay-updates staging tree is transient receiver state: remove any
|
||||
leftovers on every exit path (success already emptied it). */
|
||||
if (config->delay_context)
|
||||
@@ -1023,6 +1041,134 @@ static bool receive_basis_entries(int fd, Config* c, ConfigStringBudget* budget)
|
||||
return true;
|
||||
}
|
||||
|
||||
/* Receiver-side delete-protection rules (protocol 2.28.0). The sender compiles
|
||||
* its command-line selection rules exactly as the scanner does and streams the
|
||||
* result as one bounded, self-describing block (count + per-rule records); the
|
||||
* receiver reconstructs a FilterRuleList for the --delete extras walk. owner
|
||||
* and pattern are charged through the shared ConfigStringBudget and the block
|
||||
* additionally enforces MAX_FILTER_RULES / MAX_FILTER_BYTES. */
|
||||
static bool send_protect_entries(int fd, const Config* c) {
|
||||
int count = c->filters ? c->filters->size : 0;
|
||||
const char** texts = NULL;
|
||||
if (count > 0) {
|
||||
texts = malloc((size_t)count * sizeof(char*));
|
||||
if (!texts)
|
||||
return false;
|
||||
for (int i = 0; i < count; i++)
|
||||
texts[i] = (const char*)c->filters->items[i];
|
||||
}
|
||||
char err[160];
|
||||
FilterRuleList* rules =
|
||||
filter_base_build(texts, count, c->cvs_exclude, c->delete_excluded, err, sizeof(err));
|
||||
free(texts);
|
||||
if (!rules) {
|
||||
log_message(LOG_LEVEL_ERROR, "invalid filter rule: %s", err);
|
||||
return false;
|
||||
}
|
||||
/* The receiver rejects any block with more than MAX_FILTER_RULES entries as a
|
||||
* protocol error; refuse to emit such a frame at all. filter_base_build()
|
||||
* can expand the client rule set (cvs-exclude, merge files), so this is the
|
||||
* authoritative bound, not config->filters->size. */
|
||||
if (rules->count < 0 || rules->count > MAX_FILTER_RULES) {
|
||||
log_message(LOG_LEVEL_ERROR, "too many filter rules: %d (maximum %d)", rules->count,
|
||||
MAX_FILTER_RULES);
|
||||
filter_rule_list_free(rules);
|
||||
return false;
|
||||
}
|
||||
bool ok = send_int(fd, rules->count);
|
||||
for (int i = 0; ok && i < rules->count; i++) {
|
||||
const FilterRule* r = rules->items[i];
|
||||
/* Mirror the receiver's limit so the peer never receives a rule it will
|
||||
reject as a protocol error. */
|
||||
if (r->pattern && strlen(r->pattern) > MAX_PROTECT_PATTERN_LEN) {
|
||||
log_message(LOG_LEVEL_ERROR, "filter pattern exceeds %d bytes", MAX_PROTECT_PATTERN_LEN);
|
||||
filter_rule_list_free(rules);
|
||||
return false;
|
||||
}
|
||||
ok = send_int(fd, (int)r->action) && send_int(fd, (int)r->sides) &&
|
||||
send_int(fd, r->anchored ? 1 : 0) && send_int(fd, r->dir_only ? 1 : 0) &&
|
||||
send_int(fd, r->negate ? 1 : 0) && send_str(fd, r->owner ? r->owner : "") &&
|
||||
send_str(fd, r->pattern ? r->pattern : "");
|
||||
}
|
||||
filter_rule_list_free(rules);
|
||||
return ok;
|
||||
}
|
||||
|
||||
static bool receive_protect_entries(int fd, Config* c, ConfigStringBudget* budget) {
|
||||
int count;
|
||||
if (!receive_int(fd, &count))
|
||||
return false;
|
||||
if (count < 0 || count > MAX_FILTER_RULES)
|
||||
return false;
|
||||
if (count == 0)
|
||||
return true;
|
||||
FilterRuleList* list = filter_rule_list_create();
|
||||
if (!list)
|
||||
return false;
|
||||
size_t pattern_bytes = 0;
|
||||
for (int i = 0; i < count; i++) {
|
||||
int action;
|
||||
int sides;
|
||||
bool anchored;
|
||||
bool dir_only;
|
||||
bool negate;
|
||||
if (!receive_int(fd, &action) ||
|
||||
(action != FILTER_ACTION_EXCLUDE && action != FILTER_ACTION_INCLUDE) ||
|
||||
!receive_int(fd, &sides) || sides < (int)FILTER_SIDE_SENDER ||
|
||||
sides > (int)(FILTER_SIDE_SENDER | FILTER_SIDE_RECEIVER) ||
|
||||
!receive_wire_bool(fd, &anchored) || !receive_wire_bool(fd, &dir_only) ||
|
||||
!receive_wire_bool(fd, &negate))
|
||||
goto fail;
|
||||
char* owner = config_receive_str(fd, budget);
|
||||
if (!owner)
|
||||
goto fail;
|
||||
char* pattern = config_receive_str(fd, budget);
|
||||
if (!pattern || pattern[0] == '\0') {
|
||||
free(owner);
|
||||
free(pattern);
|
||||
goto fail;
|
||||
}
|
||||
/* A pattern too long to be evaluated by glob_match against a PATH_MAX path
|
||||
would silently fail to match and leave a protect rule inert (fail-open:
|
||||
the entry is then deleted). Reject it up front as a protocol error
|
||||
rather than accept a rule that can never shield anything. */
|
||||
if (strlen(pattern) > MAX_PROTECT_PATTERN_LEN) {
|
||||
free(owner);
|
||||
free(pattern);
|
||||
goto fail;
|
||||
}
|
||||
size_t bytes = strlen(owner) + strlen(pattern);
|
||||
if (bytes > MAX_FILTER_BYTES - pattern_bytes) {
|
||||
free(owner);
|
||||
free(pattern);
|
||||
goto fail;
|
||||
}
|
||||
pattern_bytes += bytes;
|
||||
FilterRule* rule = calloc(1, sizeof(FilterRule));
|
||||
if (!rule) {
|
||||
free(owner);
|
||||
free(pattern);
|
||||
goto fail;
|
||||
}
|
||||
rule->action = (FilterAction)action;
|
||||
rule->sides = (unsigned)sides;
|
||||
rule->anchored = anchored;
|
||||
rule->dir_only = dir_only;
|
||||
rule->negate = negate;
|
||||
rule->owner = owner;
|
||||
rule->pattern = pattern;
|
||||
if (!filter_rule_list_add(list, rule)) {
|
||||
filter_rule_free(rule);
|
||||
goto fail;
|
||||
}
|
||||
}
|
||||
c->protect_rules = list;
|
||||
return true;
|
||||
fail:
|
||||
filter_rule_list_free(list);
|
||||
return false;
|
||||
}
|
||||
|
||||
static bool send_identity_entries(int fd, const IdentityMap* map, int count) {
|
||||
for (int i = 0; i < count; i++) {
|
||||
if (!send_int(fd, map[i].from) || !send_int(fd, map[i].from_hi) || !send_int(fd, map[i].to) ||
|
||||
@@ -1150,6 +1296,9 @@ fail:
|
||||
#define CONFIG_RECV_BLOCK_IDMAP(name) \
|
||||
receive_identity_entries(fd, budget, c->name##_count, &c->name)
|
||||
|
||||
#define CONFIG_SEND_BLOCK_PROTECT_RULES(name) send_protect_entries(fd, c)
|
||||
#define CONFIG_RECV_BLOCK_PROTECT_RULES(name) receive_protect_entries(fd, c, budget)
|
||||
|
||||
/* One table entry, applied in sequence. XSEND/XRECV are statement macros so
|
||||
* consecutive entries read as a plain sequence of assignments. */
|
||||
#define XSEND(name, ctype, def, kind) ok = ok && (CONFIG_SEND_##kind(name));
|
||||
@@ -1187,6 +1336,7 @@ CONFIG_DEFINE_SEND(send_privilege_options, CONFIG_WIRE_PRIVILEGE_FIELDS)
|
||||
CONFIG_DEFINE_SEND(send_copy_as_options, CONFIG_WIRE_COPY_AS_FIELDS)
|
||||
CONFIG_DEFINE_SEND(send_output_options, CONFIG_WIRE_OUTPUT_FIELDS)
|
||||
CONFIG_DEFINE_SEND(send_codec_options, CONFIG_WIRE_CODEC_FIELDS)
|
||||
CONFIG_DEFINE_SEND(send_protect_options, CONFIG_WIRE_PROTECT_FIELDS)
|
||||
|
||||
CONFIG_DEFINE_RECV(receive_core_fields, CONFIG_WIRE_CORE_FIELDS)
|
||||
CONFIG_DEFINE_RECV(receive_delta_fields, CONFIG_WIRE_DELTA_FIELDS)
|
||||
@@ -1207,6 +1357,7 @@ CONFIG_DEFINE_RECV(receive_privilege_options, CONFIG_WIRE_PRIVILEGE_FIELDS)
|
||||
CONFIG_DEFINE_RECV(receive_copy_as_options, CONFIG_WIRE_COPY_AS_FIELDS)
|
||||
CONFIG_DEFINE_RECV(receive_output_options, CONFIG_WIRE_OUTPUT_FIELDS)
|
||||
CONFIG_DEFINE_RECV(receive_codec_options, CONFIG_WIRE_CODEC_FIELDS)
|
||||
CONFIG_DEFINE_RECV(receive_protect_options, CONFIG_WIRE_PROTECT_FIELDS)
|
||||
|
||||
#undef XSEND
|
||||
#undef XRECV
|
||||
@@ -1325,7 +1476,8 @@ bool config_send_wire_block(int file_descriptor, const Config* config) {
|
||||
send_privilege_options(file_descriptor, config) &&
|
||||
send_copy_as_options(file_descriptor, config) &&
|
||||
send_output_options(file_descriptor, config) &&
|
||||
send_codec_options(file_descriptor, config);
|
||||
send_codec_options(file_descriptor, config) &&
|
||||
send_protect_options(file_descriptor, config);
|
||||
}
|
||||
|
||||
bool config_send(int file_descriptor, const Config* config) {
|
||||
@@ -1397,7 +1549,8 @@ Config* config_receive_with_validate(int file_descriptor, ConfigValidateFunc val
|
||||
!receive_privilege_options(file_descriptor, config, &budget) ||
|
||||
!receive_copy_as_options(file_descriptor, config, &budget) ||
|
||||
!receive_output_options(file_descriptor, config, &budget) ||
|
||||
!receive_codec_options(file_descriptor, config, &budget))
|
||||
!receive_codec_options(file_descriptor, config, &budget) ||
|
||||
!receive_protect_options(file_descriptor, config, &budget))
|
||||
goto error;
|
||||
/* Validate/normalize the negotiated codec. compress_choice is the human
|
||||
* spelling (NULL or "" when -z was not given); compression_algo is the
|
||||
|
||||
+148
-49
@@ -4,6 +4,7 @@
|
||||
#include "array_list.h"
|
||||
#include "checksum.h"
|
||||
#include "compression.h"
|
||||
#include "filter.h"
|
||||
#include <stdbool.h>
|
||||
#include <stdint.h>
|
||||
#include <stdio.h>
|
||||
@@ -82,7 +83,7 @@ typedef struct {
|
||||
typedef enum SuperMode { SUPER_MODE_AUTO = 0, SUPER_MODE_ON = 1, SUPER_MODE_OFF = 2 } SuperMode;
|
||||
|
||||
/* ===========================================================================
|
||||
* Config wire-field table (single source of truth for protocol 2.26.0).
|
||||
* Config wire-field table (single source of truth for protocol 2.28.0).
|
||||
*
|
||||
* Every field below crosses the wire. The table is the ONLY place a
|
||||
* serialized field is named: config.h expands CONFIG_WIRE_FIELDS() to declare
|
||||
@@ -197,9 +198,18 @@ typedef enum SuperMode { SUPER_MODE_AUTO = 0, SUPER_MODE_ON = 1, SUPER_MODE_OFF
|
||||
X(skip_compress_count, int, 0, INT_SKIPCOUNT) \
|
||||
X(skip_compress_suffixes, char**, NULL, BLOCK_SKIP_SUFFIXES)
|
||||
|
||||
/* FastSync-only --verify-basis (protocol 2.28.0, no version bump by project
|
||||
* decision): restores the stricter content equality on a basis hit. By
|
||||
* default a basis hit is accepted on rsync's metadata quick-check alone (equal
|
||||
* size plus equal mtime, or size alone under --size-only); with this flag the
|
||||
* receiver ALSO requires the basis bytes' whole-file digest (the negotiated
|
||||
* --checksum-choice algorithm) to equal the sender's, exactly FastSync's
|
||||
* historical behavior. It is a receiver policy and crosses the wire so the
|
||||
* receiver knows whether to read and hash the basis content. */
|
||||
#define CONFIG_WIRE_BASIS_FIELDS(X) \
|
||||
X(basis_count, int, 0, INT_BASISCOUNT) \
|
||||
X(basis_dirs, BasisDest*, NULL, BLOCK_BASIS)
|
||||
X(basis_dirs, BasisDest*, NULL, BLOCK_BASIS) \
|
||||
X(verify_basis, bool, false, BOOL)
|
||||
|
||||
#define CONFIG_WIRE_FUZZY_FIELDS(X) X(fuzzy, bool, false, BOOL)
|
||||
|
||||
@@ -259,9 +269,18 @@ typedef enum SuperMode { SUPER_MODE_AUTO = 0, SUPER_MODE_ON = 1, SUPER_MODE_OFF
|
||||
* for -n/--dry-run --delete, the destination-relative paths it WOULD have
|
||||
* deleted. It is set by the client only when --stats, --progress/-P, an
|
||||
* --out-format token needs a wire counter (%b/%c), or a dry-run carries
|
||||
* --delete; the transfer decision itself is unchanged. */
|
||||
* --delete; the transfer decision itself is unchanged.
|
||||
*
|
||||
* --info wave (protocol 2.27.0). report_deletes tells the receiver to include
|
||||
* the destination-relative paths it ACTUALLY removed in its terminal
|
||||
* STATUS_STATS record (the same path-list field the dry-run would-delete report
|
||||
* uses), so the sender can print rsync's `deleting PATH`/`*deleting` lines for a
|
||||
* real (non-dry-run) deletion. It is set when --delete is active and any of
|
||||
* --info=del, -i/--itemize-changes or --out-format requests per-file change
|
||||
* output; the transfer decision itself is unchanged. */
|
||||
#define CONFIG_WIRE_OUTPUT_FIELDS(X) \
|
||||
X(report_dest_info, bool, false, BOOL) X(report_stats, bool, false, BOOL)
|
||||
X(report_dest_info, bool, false, BOOL) \
|
||||
X(report_stats, bool, false, BOOL) X(report_deletes, bool, false, BOOL)
|
||||
|
||||
/* Codec-negotiation wave (protocol 2.26.0). compression_algo is the concrete
|
||||
* codec the client selected for this transfer (a CompressionAlgo id) and is the
|
||||
@@ -284,6 +303,20 @@ typedef enum SuperMode { SUPER_MODE_AUTO = 0, SUPER_MODE_ON = 1, SUPER_MODE_OFF
|
||||
#define CONFIG_WIRE_CODEC_FIELDS(X) \
|
||||
X(compression_algo, int, COMPRESSION_ALGO_ZSTD, INT_COMPRESSION_ALGO)
|
||||
|
||||
/* Receiver-side delete-protection filter rules (protocol 2.28.0). The sender
|
||||
* compiles its root-level selection rules exactly as the scanner does
|
||||
* (filter_base_build over --filter/-f/--exclude/--include/-C) and streams them
|
||||
* as one self-describing, bounded block (count followed by per-rule records).
|
||||
* The receiver reconstructs `protect_rules` and evaluates them against
|
||||
* DESTINATION-ONLY entries during the --delete extras walk, so a
|
||||
* `protect`/`P` rule protects an extra that never appeared on the sender
|
||||
* (rsync re-derives deletion protection from the filter list; FastSync
|
||||
* historically derived it only from the source scan). `protect_rules` is NULL
|
||||
* on the sender and is owned/freed by the receiver Config. Bounded by
|
||||
* MAX_FILTER_RULES and MAX_FILTER_BYTES; an unknown action/sides is a protocol
|
||||
* error. */
|
||||
#define CONFIG_WIRE_PROTECT_FIELDS(X) X(protect_rules, FilterRuleList*, NULL, BLOCK_PROTECT_RULES)
|
||||
|
||||
/* All serialized fields, in exact wire order. Concatenating the per-segment
|
||||
* lists here is what keeps the declaration order = the wire order. */
|
||||
#define CONFIG_WIRE_FIELDS(X) \
|
||||
@@ -306,7 +339,54 @@ typedef enum SuperMode { SUPER_MODE_AUTO = 0, SUPER_MODE_ON = 1, SUPER_MODE_OFF
|
||||
CONFIG_WIRE_PRIVILEGE_FIELDS(X) \
|
||||
CONFIG_WIRE_COPY_AS_FIELDS(X) \
|
||||
CONFIG_WIRE_OUTPUT_FIELDS(X) \
|
||||
CONFIG_WIRE_CODEC_FIELDS(X)
|
||||
CONFIG_WIRE_CODEC_FIELDS(X) \
|
||||
CONFIG_WIRE_PROTECT_FIELDS(X)
|
||||
|
||||
/* Client-only, CLI-parse bookkeeping (never serialized). These members exist
|
||||
* only so the client command-line parser can record HOW an option was
|
||||
* specified (explicitly set, explicitly negated, or a parser-requested exit
|
||||
* code); no other module and no wire peer ever needs them. Grouping them in
|
||||
* one nested member keeps the public Config free of client-CLI-only state. */
|
||||
typedef struct {
|
||||
/* Set when the user explicitly turned an attribute off with --no-perms /
|
||||
* --no-times (long or short form). --incremental/--delta historically
|
||||
* auto-enabled mode and mtime preservation; these flags let
|
||||
* cli_finalize_config restore that behavior while still honoring the
|
||||
* explicit per-attribute negation. A later -p/-t re-enables the attribute
|
||||
* directly, so the flag only prevents the incremental/delta implication,
|
||||
* never a POSITIVE request. */
|
||||
bool preserve_perms_explicit_off;
|
||||
bool preserve_times_explicit_off;
|
||||
/* Set by --no-preserve, the explicit opt-out of the whole preservation
|
||||
* bundle, so the --incremental/--delta auto-preserve implication stays off. */
|
||||
bool metadata_explicitly_disabled;
|
||||
/* True when --server-port/--port was explicitly given. --dry-run uses it to
|
||||
* decide whether a real server handshake was requested, so a plain local
|
||||
* destination (no explicit port) keeps the existing client-side dry-run
|
||||
* behavior instead of dialing the default 127.0.0.1:8080. */
|
||||
bool server_port_set;
|
||||
/* True when --server-host was explicitly given, and distinct from the
|
||||
* "127.0.0.1" default: --dry-run uses it to route an explicit remote target
|
||||
* to the server so it reports receiver state exactly like a real run,
|
||||
* instead of silently running the client-side manifest. */
|
||||
bool server_host_set;
|
||||
/* Codec-negotiation CLI state. The effective pre-transfer checksum is
|
||||
* Config->checksum_algo (serialized); checksum_transfer_algo is the rsync
|
||||
* "transfer" half of a two-name --checksum-choice form (validated and used
|
||||
* only to mirror rsync's whole-file forcing, since FastSync's per-block
|
||||
* strong hash is fixed). cli_exit_code carries a parser-requested process
|
||||
* exit status (rsync uses 4 for an unsupported checksum/compress algorithm)
|
||||
* so main() can mirror it. */
|
||||
int checksum_transfer_algo;
|
||||
int cli_exit_code;
|
||||
/* "The user explicitly chose" bits. They let the per-codec default level /
|
||||
* checksum list be applied only when the corresponding rsync option was
|
||||
* omitted (an explicit --compress-level / --checksum-choice always wins). */
|
||||
bool compression_level_set;
|
||||
bool checksum_choice_set;
|
||||
/* True when --stop-at was given. */
|
||||
bool stop_at_set;
|
||||
} ConfigCliParse;
|
||||
|
||||
typedef struct Config {
|
||||
/* -j/--threads=N: number of parallel scanner worker threads for the -m
|
||||
@@ -314,16 +394,8 @@ typedef struct Config {
|
||||
* scanner's built-in default" (4). CLIENT-ONLY: it is a local scheduling
|
||||
* concern and is NEVER serialized into the wire config frame. */
|
||||
int scanner_threads;
|
||||
bool metadata_explicitly_disabled;
|
||||
/* CLIENT-ONLY (never serialized; not in CONFIG_WIRE_FIELDS). Set when the
|
||||
* user explicitly turned an attribute off with --no-perms / --no-times (long
|
||||
* or short form). --incremental/--delta historically auto-enabled mode and
|
||||
* mtime preservation; these flags let cli_finalize_config restore that
|
||||
* behavior while still honoring the explicit per-attribute negation. A
|
||||
* later -p/-t re-enables the attribute directly, so the flag only prevents
|
||||
* the incremental/delta implication, never a POSITIVE request. */
|
||||
bool preserve_perms_explicit_off;
|
||||
bool preserve_times_explicit_off;
|
||||
/* Client-only CLI-parse bookkeeping (never serialized). See ConfigCliParse. */
|
||||
ConfigCliParse cli;
|
||||
bool show_progress;
|
||||
int compression_threads;
|
||||
int ssh_port;
|
||||
@@ -344,18 +416,6 @@ typedef struct Config {
|
||||
bool use_tls;
|
||||
char* server_host;
|
||||
int server_port;
|
||||
/* True when --server-port/--port was explicitly given. CLIENT-ONLY (never
|
||||
* serialized): --dry-run uses it to decide whether a real server handshake
|
||||
* was requested, so a plain local destination (no explicit port) keeps the
|
||||
* existing client-side dry-run behavior instead of dialing the default
|
||||
* 127.0.0.1:8080. */
|
||||
bool server_port_set;
|
||||
/* True when --server-host was explicitly given. CLIENT-ONLY (never
|
||||
* serialized), and distinct from the "127.0.0.1" default: --dry-run uses it
|
||||
* to route an explicit remote target to the server so it reports receiver
|
||||
* state exactly like a real run, instead of silently running the client-side
|
||||
* manifest. */
|
||||
bool server_host_set;
|
||||
char* tls_cert;
|
||||
char* tls_key;
|
||||
char* tls_ca;
|
||||
@@ -401,16 +461,6 @@ typedef struct Config {
|
||||
* enters the keep-set. Implied by --delete-missing-args. */
|
||||
bool ignore_missing_args;
|
||||
|
||||
/* Codec-negotiation CLI state (all client-only, never serialized). The
|
||||
* effective pre-transfer checksum is Config->checksum_algo (serialized);
|
||||
* checksum_transfer_algo is the rsync "transfer" half of a two-name
|
||||
* --checksum-choice form (validated and used only to mirror rsync's
|
||||
* whole-file forcing, since FastSync's per-block strong hash is fixed).
|
||||
* cli_exit_code carries a parser-requested process exit status (rsync uses 4
|
||||
* for an unsupported checksum/compress algorithm) so main() can mirror it. */
|
||||
int checksum_transfer_algo;
|
||||
int cli_exit_code;
|
||||
|
||||
// Issue #129: Advanced file selection. These fields are CLIENT-ONLY: they are
|
||||
// never serialized to the wire (the receiver must not learn them).
|
||||
ArrayList* filters; /* --filter=RULE rule strings, in order */
|
||||
@@ -423,9 +473,13 @@ typedef struct Config {
|
||||
* /.rsync-filter' (the .rsync-filter files themselves are transferred); a
|
||||
* repeated -F adds --filter='- .rsync-filter' so they are excluded too. */
|
||||
int per_dir_filter_count;
|
||||
bool one_file_system; /* -x/--one-file-system: do not cross filesystem boundaries */
|
||||
/* --no-implied-dirs: client-only. With -R + --files-from, refuse to place a
|
||||
* listed file whose ancestor directory is not itself explicitly listed. */
|
||||
int one_file_system; /* -x/--one-file-system: do not cross filesystem boundaries.
|
||||
Repeated -x (rsync's -xx) drops the mount-point
|
||||
directory entirely instead of recreating it empty. */
|
||||
/* --no-implied-dirs: client-only. With -R, do not transfer the source
|
||||
* metadata of the parent directories implied by a listed path; an unlisted
|
||||
* implied parent is still created (with default attributes) so the listed
|
||||
* file can be placed, matching rsync. */
|
||||
bool no_implied_dirs;
|
||||
/* -d/--dirs: client-only. Transfer the directory entries named by the
|
||||
* source argument / --files-from list without recursing into contents. */
|
||||
@@ -442,7 +496,6 @@ typedef struct Config {
|
||||
/* --outbuf mode (OutbufMode): stdout/stderr buffering. Client-only launch
|
||||
* concern: NEVER crosses the wire. */
|
||||
int outbuf;
|
||||
bool old_args;
|
||||
/* --remote-option=OPT (Phase 5, long form only): one or more extra command-line
|
||||
* options to append to the REMOTE server invocation over SSH. CLIENT-ONLY:
|
||||
* they are composed into the remote command line by ssh_build_remote_command()
|
||||
@@ -512,7 +565,6 @@ typedef struct Config {
|
||||
* process and are NEVER serialized into the config frame. */
|
||||
int stop_after_mins; /* --stop-after=MINS minutes; 0 when unset */
|
||||
time_t stop_at; /* --stop-at=... absolute wall-clock deadline */
|
||||
bool stop_at_set; /* true when --stop-at was given */
|
||||
|
||||
/* Client-only residual-batch paths. A residual batch is a self-contained
|
||||
* single-file record of the whole source tree (full file images using the
|
||||
@@ -600,10 +652,14 @@ typedef struct Config {
|
||||
source directory is streamed in directory order, and the receiver removes
|
||||
each directory's extras when its plan arrives (during) or snapshots them
|
||||
and removes them only after a successful transfer (delay). delete_after
|
||||
(and plain --delete) keep the whole-tree commit mode: extras are removed
|
||||
from a fresh end-of-transfer destination scan only after the whole transfer
|
||||
succeeded. See config_delete_timing_early()/config_delete_timing_per_dir()
|
||||
below. */
|
||||
keeps the whole-tree commit mode: extras are removed from a fresh
|
||||
end-of-transfer destination scan only after the whole transfer succeeded.
|
||||
A plain --delete with no explicit timing flag defaults to delete_during on
|
||||
the client (cli_finalize_config), matching rsync's --del default; the old
|
||||
late-commit behavior is selected explicitly by --delete-after or the
|
||||
FastSync-only long spelling --delete-commit (an exact alias for
|
||||
--delete-after, mapped onto the same wire field). See
|
||||
config_delete_timing_early()/config_delete_timing_per_dir() below. */
|
||||
/* partial_dir */
|
||||
// PR #174: Partial transfer resumption
|
||||
/* suffix */
|
||||
@@ -988,11 +1044,51 @@ typedef struct Config {
|
||||
* boundary, and the strict same-version handshake (config_receive rejects a
|
||||
* mismatched version before parsing anything else) keeps mixed deployments from
|
||||
* ever reaching that state. */
|
||||
#define PROTOCOL_VERSION "2.26.0"
|
||||
/* (7) --info=del report (protocol 2.27.0): the config frame gains one trailing
|
||||
* bool, report_deletes, appended after report_stats. When set, the receiver
|
||||
* lists the paths it actually removed in the terminal STATUS_STATS path list
|
||||
* (the same count-delimited list the -n/--dry-run would-delete report uses), so
|
||||
* the sender can print rsync's `deleting PATH` lines for a real deletion. No
|
||||
* change to the fixed STATUS_STATS record itself; only a new trailing config
|
||||
* bool, which still requires the version bump for the strict lockstep. */
|
||||
/* (8) --stats receiver-observed counters (protocol 2.28.0): the fixed
|
||||
* STATUS_STATS record grows from three counters to eight. The receiver now
|
||||
* reports the bytes it literally stored (`literal_data`) and the count of
|
||||
* destination entries it newly CREATED, split by type
|
||||
* (reg/dir/link/special), so the sender can print rsync's exact
|
||||
* `Number of created files: N (reg: X, dir: Y, link: Z, special: W)` line and
|
||||
* an exact `Literal data` total even for delta transfers. The config-frame
|
||||
* LAYOUT is unchanged (no new config field), but the STATUS_STATS body grows,
|
||||
* so a 2.27 peer that does not consume the five new fixed-width counters would
|
||||
* desynchronize on the trailing would-delete path list; the strict
|
||||
* same-version handshake (config_receive rejects a mismatched version before
|
||||
* parsing anything else) keeps mixed deployments from ever reaching that
|
||||
* state. */
|
||||
/* (9) Receiver-side delete protection (still protocol 2.28.0): the config frame
|
||||
* gains one trailing self-describing block carrying the sender's compiled base
|
||||
* filter rules so the receiver can protect DESTINATION-ONLY entries from
|
||||
* --delete with `protect`/`risk` rules (rsync parity). The block appends after
|
||||
* compression_algo; see CONFIG_WIRE_PROTECT_FIELDS. */
|
||||
#define PROTOCOL_VERSION "2.28.0"
|
||||
#define DEFAULT_CHUNK_SIZE (10 * 1024 * 1024)
|
||||
/* Upper bound on total basis-dir entries (rsync caps --link-dest at 20). */
|
||||
#define MAX_BASIS_DIRS 64
|
||||
|
||||
/* Bounds on the received receiver-side delete-protection rule block. The rule
|
||||
* count and the aggregate pattern+owner bytes are each capped so a hostile
|
||||
* peer cannot pin unbounded pre-auth memory; both are validated strictly on
|
||||
* receive (alongside the per-string ConfigStringBudget). */
|
||||
/* A peer may supply protect rules; cap the list so a crafted config cannot make
|
||||
* the receiver's delete walk evaluate an unbounded number of glob patterns per
|
||||
* destination entry (glob_match is O(pattern x path)). 1024 is far above any
|
||||
* legitimate selection. */
|
||||
#define MAX_FILTER_RULES 1024
|
||||
#define MAX_FILTER_BYTES (256 * 1024)
|
||||
/* glob_match's DP is capped at 64 Mi work units; a pattern longer than this
|
||||
* could exceed the cap against a PATH_MAX path and silently stop matching,
|
||||
* leaving a protect rule inert. Reject such a rule at receive time. */
|
||||
#define MAX_PROTECT_PATTERN_LEN 8192
|
||||
|
||||
/* Upper bound on the number of --skip-compress suffixes accepted from the wire.
|
||||
* Each suffix is an independent wire string (up to MAX_STRING_SIZE = 64 KiB), so
|
||||
* without this a hostile pre-auth client could otherwise retain
|
||||
@@ -1094,8 +1190,11 @@ bool config_delete_timing_early(const Config* config);
|
||||
* commits them only after a fully-successful transfer (delay). */
|
||||
bool config_delete_timing_per_dir(const Config* config);
|
||||
/* Delete-timing sanity: with deletion enabled at most one timing flag may be
|
||||
* set (none = the default delete-after commit timing); without deletion no
|
||||
* timing flag may be set (each timing flag implies --delete). */
|
||||
* set; without deletion no timing flag may be set (each timing flag implies
|
||||
* --delete). A plain --delete is normalized to delete_during by
|
||||
* cli_finalize_config on the client, so a transmitted use_delete config always
|
||||
* carries exactly one timing; the zero-timing case remains valid only for a
|
||||
* config that has not been through the CLI. */
|
||||
bool config_has_valid_delete_timing(const Config* config);
|
||||
|
||||
/* Single source of truth for the cross-field ("combination") invariants a
|
||||
|
||||
+204
-18
@@ -8,12 +8,13 @@
|
||||
#include <openssl/evp.h>
|
||||
#include <openssl/params.h>
|
||||
#include <openssl/rand.h>
|
||||
#include <stdarg.h>
|
||||
#include <poll.h>
|
||||
#include <stdint.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <sys/stat.h>
|
||||
#include <time.h>
|
||||
#include <unistd.h>
|
||||
|
||||
/* One store entry: a username and its salted PBKDF2 verifier. The plaintext
|
||||
@@ -58,19 +59,37 @@ struct CredentialStore {
|
||||
static const uint8_t k_dummy_stored_key[CREDENTIAL_KEY_LEN] = {0};
|
||||
static const uint8_t k_dummy_server_key[CREDENTIAL_KEY_LEN] = {0};
|
||||
|
||||
static void set_error(char* err, size_t err_size, const char* fmt, ...) {
|
||||
if (!err || err_size == 0)
|
||||
return;
|
||||
va_list args;
|
||||
va_start(args, fmt);
|
||||
vsnprintf(err, err_size, fmt, args);
|
||||
va_end(args);
|
||||
}
|
||||
#define set_error utils_set_error
|
||||
|
||||
static bool is_comment_char(char c) {
|
||||
return c == '#' || c == ';';
|
||||
}
|
||||
|
||||
/* True for a literal fd-backed store path: exactly "/dev/fd/<digits>" or
|
||||
* "/proc/self/fd/<digits>", with no trailing component and no "..". These name
|
||||
* the calling process's own open descriptors (e.g. a bash process substitution
|
||||
* `<(...)`, which passes /dev/fd/N), and both prefixes are symlinks by
|
||||
* construction. */
|
||||
static bool is_fd_backed_path(const char* path) {
|
||||
static const char* const prefixes[] = {"/dev/fd/", "/proc/self/fd/"};
|
||||
if (!path)
|
||||
return false;
|
||||
for (size_t i = 0; i < sizeof(prefixes) / sizeof(prefixes[0]); i++) {
|
||||
const char* prefix = prefixes[i];
|
||||
size_t prefix_len = strlen(prefix);
|
||||
if (strncmp(path, prefix, prefix_len) != 0)
|
||||
continue;
|
||||
const char* digits = path + prefix_len;
|
||||
if (*digits < '0' || *digits > '9')
|
||||
return false;
|
||||
const char* p = digits;
|
||||
while (*p >= '0' && *p <= '9')
|
||||
p++;
|
||||
return *p == '\0';
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
/* Open a --password-file / --early-input after verifying the EXACT inode we
|
||||
* will read: it must be owned by the effective user and grant no group/other
|
||||
* permission bit (so 0600 and stricter modes such as 0400 are accepted),
|
||||
@@ -80,12 +99,31 @@ static bool is_comment_char(char c) {
|
||||
* path and then fstat the resulting fd (rather than stat()ing the path first
|
||||
* and reopening it), so the permission decision is made on the same inode that
|
||||
* is read and cannot be raced by swapping the path between check and open.
|
||||
* The path may be a process-substitution pipe (`<(...)` -> /dev/fd/N), so
|
||||
* regular files and FIFOs are accepted when the ownership/mode checks pass.
|
||||
* O_NOFOLLOW refuses a symlinked path outright (ELOOP fails closed) instead of
|
||||
* following it before the owner/mode gate can run. The one exception is a
|
||||
* literal fd-backed path (/dev/fd/N or /proc/self/fd/N, see
|
||||
* is_fd_backed_path): those entries are symlinks to the CALLING process's own
|
||||
* descriptors, so following them is not the untrusted-symlink hazard
|
||||
* O_NOFOLLOW guards against, and requiring O_NOFOLLOW would break the
|
||||
* documented process-substitution/FIFO usage. For them only, O_NOFOLLOW is
|
||||
* omitted; the same fstat owner/mode gate still applies to the resolved inode.
|
||||
* O_NONBLOCK keeps the OPEN itself from
|
||||
* blocking forever on a writer-less FIFO (a blocking O_RDONLY open would wait
|
||||
* for a writer). The fd is left nonblocking for FIFOs so a read never blocks
|
||||
* either; the read loop (secret_read_line) absorbs the resulting EAGAIN by
|
||||
* waiting, under a bounded deadline, for the writer -- this is what makes a
|
||||
* slow process substitution (`--password-file <(sleep 1; ...)`) work while a
|
||||
* writer-less FIFO still fails after the deadline instead of hanging. Only
|
||||
* regular files and FIFOs pass the ownership/mode checks; O_NONBLOCK is
|
||||
* cleared for regular files, where it is a no-op anyway and no EAGAIN can
|
||||
* occur, so their stdio read path is byte-for-byte unchanged.
|
||||
*
|
||||
* Returns a FILE* the caller must fclose, or NULL with `err` filled. */
|
||||
static FILE* secret_file_open(const char* path, char* err, size_t err_size) {
|
||||
int fd = open(path, O_RDONLY | O_CLOEXEC);
|
||||
int flags = O_RDONLY | O_NONBLOCK | O_CLOEXEC;
|
||||
if (!is_fd_backed_path(path))
|
||||
flags |= O_NOFOLLOW;
|
||||
int fd = open(path, flags);
|
||||
if (fd < 0) {
|
||||
set_error(err, err_size, "cannot open secret file '%s': %s", path, strerror(errno));
|
||||
return NULL;
|
||||
@@ -105,6 +143,15 @@ static FILE* secret_file_open(const char* path, char* err, size_t err_size) {
|
||||
close(fd);
|
||||
return NULL;
|
||||
}
|
||||
/* O_NONBLOCK is only meaningful for the FIFO allowance. Restore blocking
|
||||
* mode on a regular file so its read path is exactly as before; a no-op on
|
||||
* most systems, but explicit. Failures here are ignored: O_NONBLOCK on a
|
||||
* regular file does not affect reads either way. */
|
||||
if (S_ISREG(st.st_mode)) {
|
||||
int status_flags = fcntl(fd, F_GETFL);
|
||||
if (status_flags >= 0)
|
||||
(void)fcntl(fd, F_SETFL, status_flags & ~O_NONBLOCK);
|
||||
}
|
||||
FILE* fp = fdopen(fd, "r");
|
||||
if (!fp) {
|
||||
set_error(err, err_size, "cannot read secret file '%s': %s", path, strerror(errno));
|
||||
@@ -114,6 +161,123 @@ static FILE* secret_file_open(const char* path, char* err, size_t err_size) {
|
||||
return fp;
|
||||
}
|
||||
|
||||
/* Overall bound on how long the reader waits for a process-substitution/FIFO
|
||||
* writer to produce data before giving up. It must comfortably exceed a
|
||||
* producer's startup delay (e.g. `--password-file <(sleep 1; ...)`) while still
|
||||
* bounding a writer-less FIFO, so a stray or hostile FIFO cannot stall the
|
||||
* daemon or client indefinitely. */
|
||||
#define CREDENTIAL_FIFO_READ_TIMEOUT_MS 3000
|
||||
|
||||
/* Monotonic milliseconds, used only for the read deadline (wall-clock changes
|
||||
* must not extend or shorten the wait). */
|
||||
static int64_t credential_monotonic_ms(void) {
|
||||
struct timespec ts;
|
||||
if (clock_gettime(CLOCK_MONOTONIC, &ts) != 0)
|
||||
return 0;
|
||||
return (int64_t)ts.tv_sec * 1000 + (int64_t)(ts.tv_nsec / 1000000);
|
||||
}
|
||||
|
||||
/* Wait until `fd` is readable or the deadline passes. Returns true when it is
|
||||
* readable, false on timeout or a poll error (err filled). EINTR is retried
|
||||
* against the same deadline, so signals cannot extend the wait. */
|
||||
static bool credential_wait_readable(int fd, int64_t deadline, const char* label, const char* path,
|
||||
char* err, size_t err_size) {
|
||||
for (;;) {
|
||||
int64_t remaining = deadline - credential_monotonic_ms();
|
||||
if (remaining <= 0)
|
||||
break;
|
||||
if (remaining > INT_MAX)
|
||||
remaining = INT_MAX;
|
||||
struct pollfd pfd = {.fd = fd, .events = POLLIN, .revents = 0};
|
||||
int rc = poll(&pfd, 1, (int)remaining);
|
||||
if (rc > 0)
|
||||
return true;
|
||||
if (rc == 0)
|
||||
break;
|
||||
if (errno != EINTR) {
|
||||
set_error(err, err_size, "error waiting for %s '%s': %s", label, path, strerror(errno));
|
||||
return false;
|
||||
}
|
||||
}
|
||||
set_error(err, err_size, "timed out after %d ms waiting for %s '%s'",
|
||||
CREDENTIAL_FIFO_READ_TIMEOUT_MS, label, path);
|
||||
return false;
|
||||
}
|
||||
|
||||
typedef enum {
|
||||
SECRET_READ_LINE,
|
||||
SECRET_READ_EOF,
|
||||
SECRET_READ_ERROR,
|
||||
} SecretReadResult;
|
||||
|
||||
/* Read one complete line from `fp` into `line` (capacity `cap`), including the
|
||||
* trailing newline when present and always NUL-terminating. `*out_len`
|
||||
* receives strlen(line).
|
||||
*
|
||||
* A regular file is read exactly as before: secret_file_open leaves it
|
||||
* blocking, so fgets never sees EAGAIN. A FIFO stays nonblocking, so fgets
|
||||
* returns NULL (or a partial line) with EAGAIN while the writer is still
|
||||
* starting up; instead of treating that as a fatal error the loop clearerr()s
|
||||
* and polls for readability against one overall deadline. The `used`
|
||||
* accumulator reassembles a line that arrived in several write()s into a single
|
||||
* line, so a split write is not misparsed as two entries.
|
||||
*
|
||||
* Returns SECRET_READ_LINE, SECRET_READ_EOF, or SECRET_READ_ERROR (err filled)
|
||||
* on timeout or a genuine read error. */
|
||||
static SecretReadResult secret_read_line(char* line, size_t cap, FILE* fp, const char* label,
|
||||
const char* path, size_t* out_len, char* err,
|
||||
size_t err_size) {
|
||||
int fd = fileno(fp);
|
||||
int64_t deadline = credential_monotonic_ms() + CREDENTIAL_FIFO_READ_TIMEOUT_MS;
|
||||
size_t used = 0;
|
||||
line[0] = '\0';
|
||||
for (;;) {
|
||||
errno = 0;
|
||||
if (fgets(line + used, (int)(cap - used), fp)) {
|
||||
used += strlen(line + used);
|
||||
if (used > 0 && line[used - 1] == '\n') {
|
||||
*out_len = used;
|
||||
return SECRET_READ_LINE;
|
||||
}
|
||||
if (feof(fp)) {
|
||||
*out_len = used; /* final unterminated line */
|
||||
return SECRET_READ_LINE;
|
||||
}
|
||||
/* No newline and not EOF. A full buffer is the caller's over-long-line
|
||||
* case; otherwise the line is only partially available (a nonblocking
|
||||
* FIFO under a slow writer), so any genuine read error fails and anything
|
||||
* else waits for the rest. */
|
||||
if (used >= cap - 1) {
|
||||
*out_len = used;
|
||||
return SECRET_READ_LINE;
|
||||
}
|
||||
int e = ferror(fp) ? errno : 0;
|
||||
if (e != 0 && e != EAGAIN && e != EWOULDBLOCK) {
|
||||
set_error(err, err_size, "error reading %s '%s': %s", label, path, strerror(e));
|
||||
return SECRET_READ_ERROR;
|
||||
}
|
||||
clearerr(fp);
|
||||
if (!credential_wait_readable(fd, deadline, label, path, err, err_size))
|
||||
return SECRET_READ_ERROR;
|
||||
continue;
|
||||
}
|
||||
/* fgets returned NULL: EOF, a not-yet-readable FIFO, or a real error. */
|
||||
if (feof(fp)) {
|
||||
*out_len = used;
|
||||
return used > 0 ? SECRET_READ_LINE : SECRET_READ_EOF;
|
||||
}
|
||||
if (errno == EAGAIN || errno == EWOULDBLOCK) {
|
||||
clearerr(fp);
|
||||
if (!credential_wait_readable(fd, deadline, label, path, err, err_size))
|
||||
return SECRET_READ_ERROR;
|
||||
continue;
|
||||
}
|
||||
set_error(err, err_size, "error reading %s '%s': %s", label, path,
|
||||
errno != 0 ? strerror(errno) : "read failed");
|
||||
return SECRET_READ_ERROR;
|
||||
}
|
||||
}
|
||||
|
||||
/* Trim leading/trailing ASCII space and tab in place; returns the new start. */
|
||||
static char* trim_space(char* s) {
|
||||
while (*s == ' ' || *s == '\t')
|
||||
@@ -509,9 +673,17 @@ static CredentialStore* load_store_file(const char* path, char* err, size_t err_
|
||||
char line[CREDENTIAL_MAX_LINE + 2];
|
||||
bool ok = true;
|
||||
|
||||
while (fgets(line, sizeof(line), fp)) {
|
||||
for (;;) {
|
||||
size_t len = 0;
|
||||
SecretReadResult rr =
|
||||
secret_read_line(line, sizeof(line), fp, "credential file", path, &len, err, err_size);
|
||||
if (rr == SECRET_READ_EOF)
|
||||
break;
|
||||
if (rr == SECRET_READ_ERROR) {
|
||||
ok = false;
|
||||
break;
|
||||
}
|
||||
line_no++;
|
||||
size_t len = strlen(line);
|
||||
if (len == CREDENTIAL_MAX_LINE + 1 && line[len - 1] != '\n' && !feof(fp)) {
|
||||
set_error(err, err_size, "credential file '%s' line %d exceeds the %d-byte limit", path,
|
||||
line_no, CREDENTIAL_MAX_LINE);
|
||||
@@ -1148,9 +1320,17 @@ int credentials_hash_file(const char* path, uint32_t iters, FILE* out, char* err
|
||||
int line_no = 0;
|
||||
int result = 0;
|
||||
char line[CREDENTIAL_MAX_LINE + 2];
|
||||
while (fgets(line, sizeof(line), fp)) {
|
||||
for (;;) {
|
||||
size_t len = 0;
|
||||
SecretReadResult rr =
|
||||
secret_read_line(line, sizeof(line), fp, "plaintext file", path, &len, err, err_size);
|
||||
if (rr == SECRET_READ_EOF)
|
||||
break;
|
||||
if (rr == SECRET_READ_ERROR) {
|
||||
result = -1;
|
||||
break;
|
||||
}
|
||||
line_no++;
|
||||
size_t len = strlen(line);
|
||||
if (len == CREDENTIAL_MAX_LINE + 1 && line[len - 1] != '\n' && !feof(fp)) {
|
||||
set_error(err, err_size, "plaintext file '%s' line %d exceeds the %d-byte limit", path,
|
||||
line_no, CREDENTIAL_MAX_LINE);
|
||||
@@ -1228,9 +1408,15 @@ int credentials_read_secret_file(const char* path, char** user_out, char** passw
|
||||
char line[CREDENTIAL_MAX_LINE + 2];
|
||||
int result = -1;
|
||||
|
||||
while (fgets(line, sizeof(line), fp)) {
|
||||
for (;;) {
|
||||
size_t len = 0;
|
||||
SecretReadResult rr =
|
||||
secret_read_line(line, sizeof(line), fp, "password file", path, &len, err, err_size);
|
||||
if (rr == SECRET_READ_EOF)
|
||||
break;
|
||||
if (rr == SECRET_READ_ERROR)
|
||||
goto done;
|
||||
line_no++;
|
||||
size_t len = strlen(line);
|
||||
if (len == CREDENTIAL_MAX_LINE + 1 && line[len - 1] != '\n' && !feof(fp)) {
|
||||
set_error(err, err_size, "password file '%s' line %d exceeds the %d-byte limit", path,
|
||||
line_no, CREDENTIAL_MAX_LINE);
|
||||
|
||||
@@ -6,7 +6,6 @@
|
||||
#include <errno.h>
|
||||
#include <limits.h>
|
||||
#include <netinet/in.h>
|
||||
#include <stdarg.h>
|
||||
#include <stdint.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
@@ -17,14 +16,7 @@
|
||||
/* helpers */
|
||||
/* ------------------------------------------------------------------ */
|
||||
|
||||
static void set_error(char* err, size_t err_size, const char* fmt, ...) {
|
||||
if (!err || err_size == 0)
|
||||
return;
|
||||
va_list args;
|
||||
va_start(args, fmt);
|
||||
vsnprintf(err, err_size, fmt, args);
|
||||
va_end(args);
|
||||
}
|
||||
#define set_error utils_set_error
|
||||
|
||||
/* Trim leading and trailing ASCII space/tab in place; returns the new start. */
|
||||
static char* trim_ws(char* s) {
|
||||
|
||||
@@ -0,0 +1,656 @@
|
||||
#include "delete.h"
|
||||
|
||||
#include "delay_updates.h"
|
||||
#include "filter.h"
|
||||
#include "log.h"
|
||||
#include "utils.h"
|
||||
#include <dirent.h>
|
||||
#include <errno.h>
|
||||
#include <fcntl.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <sys/stat.h>
|
||||
#include <unistd.h>
|
||||
|
||||
/* Build the keep-set index from the exact manifest entries only. A lookup of
|
||||
`rel` succeeds iff `rel` is a kept entry, a kept directory, or an ancestor
|
||||
directory of kept content (the old is_dir_in_manifest predicate); the sorted
|
||||
view answers "is an ancestor of kept content" without materializing any
|
||||
per-component prefix copy, so the index is O(manifest size) memory. */
|
||||
static bool build_keep_index(const ArrayList* manifest, PathIndex* index) {
|
||||
if (!manifest || manifest->size <= 0)
|
||||
return path_index_build(index, NULL, 0);
|
||||
return path_index_build(index, (const char* const*)manifest->items, (size_t)manifest->size);
|
||||
}
|
||||
|
||||
static bool keep_is_dir(const PathIndex* index, const char* rel_path) {
|
||||
return path_index_contains(index, rel_path) || path_index_has_descendant(index, rel_path);
|
||||
}
|
||||
|
||||
static bool keep_is_file(const PathIndex* index, const char* rel_path) {
|
||||
return path_index_contains(index, rel_path);
|
||||
}
|
||||
|
||||
/* True when child_rel is, or lies below, a protected entry. A prefix "a"
|
||||
therefore protects "a" and "a/b/c" but not "ab". Entries with top_level_only
|
||||
set only protect DIRECT children of the receive root (at_root); nested
|
||||
directories that share such a name stay ordinary destination content. */
|
||||
bool path_under_skip_prefix(const char* child_rel, bool at_root, const DeleteSkipEntry* skips,
|
||||
int skip_count) {
|
||||
for (int i = 0; i < skip_count; i++) {
|
||||
if (skips[i].top_level_only && !at_root)
|
||||
continue;
|
||||
size_t prefix_len = strlen(skips[i].prefix);
|
||||
if (strncmp(child_rel, skips[i].prefix, prefix_len) == 0 &&
|
||||
(child_rel[prefix_len] == '\0' || child_rel[prefix_len] == '/'))
|
||||
return true;
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
/* Per-run deletion budget and tallies. `max_delete` is the cap on the number
|
||||
of entries the walker may remove (SIZE_MAX = unlimited); once it is reached
|
||||
the remaining extras are counted in `skipped` and left in place, matching
|
||||
rsync's partial --max-delete behavior. */
|
||||
typedef struct {
|
||||
size_t max_delete;
|
||||
size_t deleted;
|
||||
size_t skipped;
|
||||
bool limit_hit;
|
||||
} DeleteBudget;
|
||||
|
||||
/* True when direct children of the directory named by `rel` may be removed.
|
||||
With no synchronization info (dirs == NULL) the whole tree is deletable; when
|
||||
a dirs index is supplied only its exact entries are (the receive root is the
|
||||
"." sentinel). */
|
||||
static bool is_synced_dir(const PathIndex* dirs, const char* rel) {
|
||||
if (!dirs)
|
||||
return true;
|
||||
return path_index_contains(dirs, rel[0] == '\0' ? "." : rel);
|
||||
}
|
||||
|
||||
/* Unsigned byte-wise string compare, matching rsync's u_strcmp (a signed
|
||||
strcmp would order bytes >= 0x80 differently). */
|
||||
static int delete_name_cmp(const char* a, const char* b) {
|
||||
const unsigned char* pa = (const unsigned char*)a;
|
||||
const unsigned char* pb = (const unsigned char*)b;
|
||||
while (*pa != '\0' && *pa == *pb) {
|
||||
pa++;
|
||||
pb++;
|
||||
}
|
||||
return (int)*pa - (int)*pb;
|
||||
}
|
||||
|
||||
bool delete_dir_entries_collect(int dirfd, DeleteDirEntry** out, size_t* count,
|
||||
bool* operation_ok) {
|
||||
*out = NULL;
|
||||
*count = 0;
|
||||
if (operation_ok)
|
||||
*operation_ok = true;
|
||||
int scanfd = openat(dirfd, ".", O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
|
||||
if (scanfd < 0)
|
||||
return false;
|
||||
DIR* dir = fdopendir(scanfd);
|
||||
if (!dir) {
|
||||
close(scanfd);
|
||||
return false;
|
||||
}
|
||||
DeleteDirEntry* entries = NULL;
|
||||
size_t used = 0;
|
||||
size_t capacity = 0;
|
||||
bool ok = true;
|
||||
const struct dirent* entry;
|
||||
while ((entry = readdir(dir)) != NULL) {
|
||||
if (strcmp(entry->d_name, ".") == 0 || strcmp(entry->d_name, "..") == 0)
|
||||
continue;
|
||||
struct stat st;
|
||||
if (fstatat(dirfd, entry->d_name, &st, AT_SYMLINK_NOFOLLOW) != 0) {
|
||||
if (errno != ENOENT && operation_ok)
|
||||
*operation_ok = false;
|
||||
continue;
|
||||
}
|
||||
if (used == capacity) {
|
||||
size_t next = capacity == 0 ? 16 : capacity * 2;
|
||||
DeleteDirEntry* grown = realloc(entries, next * sizeof(*grown));
|
||||
if (!grown) {
|
||||
ok = false;
|
||||
break;
|
||||
}
|
||||
entries = grown;
|
||||
capacity = next;
|
||||
}
|
||||
entries[used].name = str_dup(entry->d_name);
|
||||
if (!entries[used].name) {
|
||||
ok = false;
|
||||
break;
|
||||
}
|
||||
entries[used].is_dir = S_ISDIR(st.st_mode);
|
||||
used++;
|
||||
}
|
||||
closedir(dir);
|
||||
if (!ok) {
|
||||
delete_dir_entries_free(entries, used);
|
||||
return false;
|
||||
}
|
||||
*out = entries;
|
||||
*count = used;
|
||||
return true;
|
||||
}
|
||||
|
||||
void delete_dir_entries_free(DeleteDirEntry* entries, size_t count) {
|
||||
if (!entries)
|
||||
return;
|
||||
for (size_t i = 0; i < count; i++)
|
||||
free(entries[i].name);
|
||||
free(entries);
|
||||
}
|
||||
|
||||
/* rsync's extraneous-entry order: subdirectories before files, each group in
|
||||
descending name order. */
|
||||
int delete_dir_entry_cmp_desc(const void* a, const void* b) {
|
||||
const DeleteDirEntry* ea = a;
|
||||
const DeleteDirEntry* eb = b;
|
||||
if (ea->is_dir != eb->is_dir)
|
||||
return ea->is_dir ? -1 : 1;
|
||||
return -delete_name_cmp(ea->name, eb->name);
|
||||
}
|
||||
|
||||
/* rsync's kept-subdirectory order: plain ascending name. */
|
||||
int delete_dir_entry_cmp_asc(const void* a, const void* b) {
|
||||
const DeleteDirEntry* ea = a;
|
||||
const DeleteDirEntry* eb = b;
|
||||
return delete_name_cmp(ea->name, eb->name);
|
||||
}
|
||||
|
||||
/* How the shared classification/descent walk disposes of an extra it has
|
||||
identified. LIST records the destination-relative path without touching disk
|
||||
(the -n/--dry-run would-delete enumeration); DELETE unlinks/rmdirs it, charges
|
||||
the shared --max-delete budget and notifies the observer. Both modes classify
|
||||
and traverse identically, so the dry-run enumeration and the real deletion
|
||||
cannot drift. */
|
||||
typedef enum { DELETE_WALK_MODE_DELETE, DELETE_WALK_MODE_LIST } DeleteWalkMode;
|
||||
|
||||
typedef struct {
|
||||
DeleteWalkMode mode;
|
||||
DeleteBudget* budget; /* DELETE mode */
|
||||
ArrayList* out; /* LIST mode: receives strdup'd relative paths */
|
||||
size_t* recorded; /* LIST mode */
|
||||
DeletePathObserver observer; /* DELETE mode */
|
||||
void* observer_context; /* DELETE mode */
|
||||
} DeleteWalkState;
|
||||
|
||||
/* Remove the extras directly inside the directory open on `dirfd` (DELETE mode)
|
||||
or record the paths that WOULD be removed (LIST mode), recursing into every
|
||||
child directory so kept content below a synchronized prefix is reached.
|
||||
`all_removed` reports whether every child entry was removed (so the caller may
|
||||
rmdir this directory). A child directory is never removed when it is itself a
|
||||
synchronized directory or holds kept content; with a dirs index supplied,
|
||||
direct children of a non-synchronized directory are never extras at all (they
|
||||
are left in place but still descended into). Symlinks are unlinked like any
|
||||
other non-directory extra (never followed).
|
||||
|
||||
Entries are processed in rsync's order (extraneous subdirectories in
|
||||
descending name order, then extraneous files, then kept subdirectories in
|
||||
ascending order) rather than readdir() order, so `--max-delete` leaves the
|
||||
same survivors and the `--info=del`/dry-run line order matches rsync. */
|
||||
static bool delete_walk_fd(int dirfd, const char* rel_path, const PathIndex* keep,
|
||||
const PathIndex* dirs, DeleteWalkState* state,
|
||||
const DeleteSkipEntry* skips, int skip_count,
|
||||
const FilterRuleList* protect_rules, bool parent_deletable,
|
||||
bool* all_removed) {
|
||||
DeleteDirEntry* entries = NULL;
|
||||
size_t count = 0;
|
||||
bool collect_ok = true;
|
||||
if (!delete_dir_entries_collect(dirfd, &entries, &count, &collect_ok))
|
||||
return false;
|
||||
bool operation_ok = collect_ok;
|
||||
bool local_survives = false;
|
||||
bool* shielded = calloc(count ? count : 1, sizeof(bool));
|
||||
bool* is_extra = calloc(count ? count : 1, sizeof(bool));
|
||||
if (!shielded || !is_extra) {
|
||||
free(shielded);
|
||||
free(is_extra);
|
||||
delete_dir_entries_free(entries, count);
|
||||
return false;
|
||||
}
|
||||
/* A directory is deletable when it or ANY ancestor is synchronized; the
|
||||
`parent_deletable` flag carries that down the recursion so dest-only
|
||||
directories below a synchronized root are removed wholesale. */
|
||||
bool deletable = parent_deletable || is_synced_dir(dirs, rel_path);
|
||||
bool at_root = rel_path[0] == '\0';
|
||||
|
||||
/* Reproduce rsync's traversal order: extraneous subdirectories in descending
|
||||
name order, then extraneous files in descending name order, and kept
|
||||
subdirectories only afterwards (ascending). Sorting up front also fixes the
|
||||
identity of the survivors under a partial --max-delete. */
|
||||
if (count > 1)
|
||||
qsort(entries, count, sizeof(*entries), delete_dir_entry_cmp_desc);
|
||||
size_t dir_count = 0;
|
||||
while (dir_count < count && entries[dir_count].is_dir)
|
||||
dir_count++;
|
||||
|
||||
/* Classify every entry up front (the verdict does not depend on processing
|
||||
order) so the ordered passes below can act on it. */
|
||||
for (size_t i = 0; i < count; i++) {
|
||||
char* child_rel = path_cat((char*)rel_path, entries[i].name);
|
||||
if (!child_rel) {
|
||||
operation_ok = false;
|
||||
continue;
|
||||
}
|
||||
/* A --delay-updates run keeps its staging directory as a direct child of
|
||||
the receive root, and basis-dir snapshots live below it too. Their
|
||||
contents are not manifest entries, so descending into them would delete
|
||||
every staged / basis file as an "extra". Only the staging name (a
|
||||
top-level-only prefix) and the basis prefixes are protected: a nested
|
||||
destination directory that happens to be called .fastsync-stage is
|
||||
ordinary content. */
|
||||
if (path_under_skip_prefix(child_rel, at_root, skips, skip_count)) {
|
||||
shielded[i] = true;
|
||||
local_survives = true;
|
||||
} else if (protect_rules &&
|
||||
filter_rules_apply_side(protect_rules, child_rel, entries[i].name, entries[i].is_dir,
|
||||
FILTER_SIDE_RECEIVER) == FILTER_ACTION_PROTECT) {
|
||||
/* A first-match protect rule shields the extra; for a directory the whole
|
||||
subtree is shielded (rsync prunes an excluded directory), so do not
|
||||
descend. */
|
||||
shielded[i] = true;
|
||||
local_survives = true;
|
||||
} else if (entries[i].is_dir) {
|
||||
bool child_synced = dirs && path_index_contains(dirs, child_rel);
|
||||
is_extra[i] = deletable && !child_synced && !keep_is_dir(keep, child_rel);
|
||||
if (!is_extra[i])
|
||||
local_survives = true;
|
||||
} else {
|
||||
is_extra[i] = deletable && !keep_is_file(keep, child_rel);
|
||||
if (!is_extra[i])
|
||||
local_survives = true;
|
||||
}
|
||||
free(child_rel);
|
||||
}
|
||||
|
||||
/* Pass 1: extraneous subdirectories, descending. */
|
||||
for (size_t i = 0; i < dir_count; i++) {
|
||||
if (!is_extra[i])
|
||||
continue;
|
||||
char* child_rel = path_cat((char*)rel_path, entries[i].name);
|
||||
if (!child_rel) {
|
||||
operation_ok = false;
|
||||
continue;
|
||||
}
|
||||
int childfd = openat(dirfd, entries[i].name, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
|
||||
bool child_all_removed = false;
|
||||
if (childfd >= 0) {
|
||||
if (!delete_walk_fd(childfd, child_rel, keep, dirs, state, skips, skip_count, protect_rules,
|
||||
deletable, &child_all_removed))
|
||||
operation_ok = false;
|
||||
close(childfd);
|
||||
} else if (errno != ENOENT) {
|
||||
operation_ok = false;
|
||||
}
|
||||
if (child_all_removed && deletable) {
|
||||
if (state->mode == DELETE_WALK_MODE_LIST) {
|
||||
/* Record the directory with rsync's trailing slash. */
|
||||
size_t len = strlen(child_rel);
|
||||
char* copy = malloc(len + 2);
|
||||
if (!copy) {
|
||||
operation_ok = false;
|
||||
} else {
|
||||
memcpy(copy, child_rel, len);
|
||||
copy[len] = '/';
|
||||
copy[len + 1] = '\0';
|
||||
if (!array_list_add(state->out, copy)) {
|
||||
free(copy);
|
||||
operation_ok = false;
|
||||
} else {
|
||||
(*state->recorded)++;
|
||||
}
|
||||
}
|
||||
} else if (state->budget->deleted >= state->budget->max_delete) {
|
||||
state->budget->limit_hit = true;
|
||||
state->budget->skipped++;
|
||||
local_survives = true;
|
||||
} else if (unlinkat(dirfd, entries[i].name, AT_REMOVEDIR) != 0) {
|
||||
/* ENOENT: already gone (fine). ENOTEMPTY/EEXIST: the directory still
|
||||
holds entries the walker leaves in place (a protected excluded
|
||||
prefix, a kept file the manifest protects, a symlink); rsync leaves
|
||||
such a directory behind, so this is not an error. Only genuine I/O
|
||||
failures abort the deletion. */
|
||||
if (errno != ENOENT && errno != ENOTEMPTY && errno != EEXIST)
|
||||
operation_ok = false;
|
||||
local_survives = true;
|
||||
} else {
|
||||
state->budget->deleted++;
|
||||
/* rsync reports a removed directory with a trailing slash. */
|
||||
if (state->observer) {
|
||||
size_t len = strlen(child_rel);
|
||||
char* with_slash = malloc(len + 2);
|
||||
if (with_slash) {
|
||||
memcpy(with_slash, child_rel, len);
|
||||
with_slash[len] = '/';
|
||||
with_slash[len + 1] = '\0';
|
||||
state->observer(state->observer_context, with_slash);
|
||||
free(with_slash);
|
||||
} else {
|
||||
state->observer(state->observer_context, child_rel);
|
||||
}
|
||||
}
|
||||
}
|
||||
} else {
|
||||
local_survives = true;
|
||||
}
|
||||
free(child_rel);
|
||||
}
|
||||
|
||||
/* Pass 2: extraneous files, descending. */
|
||||
for (size_t i = dir_count; i < count; i++) {
|
||||
if (!is_extra[i])
|
||||
continue;
|
||||
if (state->mode == DELETE_WALK_MODE_LIST) {
|
||||
char* child_rel = path_cat((char*)rel_path, entries[i].name);
|
||||
if (!child_rel) {
|
||||
operation_ok = false;
|
||||
continue;
|
||||
}
|
||||
char* copy = str_dup(child_rel);
|
||||
if (!copy || !array_list_add(state->out, copy)) {
|
||||
free(copy);
|
||||
operation_ok = false;
|
||||
} else {
|
||||
(*state->recorded)++;
|
||||
}
|
||||
free(child_rel);
|
||||
} else if (state->budget->deleted >= state->budget->max_delete) {
|
||||
state->budget->limit_hit = true;
|
||||
state->budget->skipped++;
|
||||
local_survives = true;
|
||||
} else if (unlinkat(dirfd, entries[i].name, 0) != 0) {
|
||||
if (errno != ENOENT)
|
||||
operation_ok = false;
|
||||
local_survives = true;
|
||||
} else {
|
||||
state->budget->deleted++;
|
||||
char* child_rel = path_cat((char*)rel_path, entries[i].name);
|
||||
if (child_rel) {
|
||||
if (state->observer)
|
||||
state->observer(state->observer_context, child_rel);
|
||||
char* escaped_path = output_escape(child_rel, log_get_8_bit_output());
|
||||
fprintf(stderr, " Deleted: %s\n", escaped_path ? escaped_path : "<allocation failed>");
|
||||
free(escaped_path);
|
||||
}
|
||||
free(child_rel);
|
||||
}
|
||||
}
|
||||
|
||||
/* Pass 3: kept subdirectories, ascending (rsync descends into these only
|
||||
after the parent's own extras have been handled). */
|
||||
for (size_t i = dir_count; i-- > 0;) {
|
||||
if (is_extra[i] || shielded[i])
|
||||
continue;
|
||||
char* child_rel = path_cat((char*)rel_path, entries[i].name);
|
||||
if (!child_rel) {
|
||||
operation_ok = false;
|
||||
continue;
|
||||
}
|
||||
int childfd = openat(dirfd, entries[i].name, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
|
||||
bool child_all_removed = false;
|
||||
if (childfd >= 0) {
|
||||
if (!delete_walk_fd(childfd, child_rel, keep, dirs, state, skips, skip_count, protect_rules,
|
||||
deletable, &child_all_removed))
|
||||
operation_ok = false;
|
||||
close(childfd);
|
||||
} else if (errno != ENOENT) {
|
||||
operation_ok = false;
|
||||
}
|
||||
/* A kept/synchronized directory is never removed. */
|
||||
local_survives = true;
|
||||
free(child_rel);
|
||||
}
|
||||
|
||||
free(shielded);
|
||||
free(is_extra);
|
||||
delete_dir_entries_free(entries, count);
|
||||
*all_removed = !local_survives;
|
||||
return operation_ok;
|
||||
}
|
||||
|
||||
/* Open the receive root following the same authorized-root confinement the
|
||||
walker uses, or dest_root directly when no authorized root is installed. */
|
||||
static int open_destination_root(const char* dest_root) {
|
||||
int root_fd = utils_get_authorized_root_fd();
|
||||
if (root_fd >= 0) {
|
||||
if (utils_get_authorized_root_path())
|
||||
return utils_open_authorized_destination(dest_root);
|
||||
if (dest_root == NULL)
|
||||
return dup(root_fd);
|
||||
return -1;
|
||||
}
|
||||
return open(dest_root, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
|
||||
}
|
||||
|
||||
bool delete_extras_list(const char* dest_root, const ArrayList* manifest,
|
||||
const ArrayList* synced_dirs, const DeleteSkipEntry* skips, int skip_count,
|
||||
const FilterRuleList* protect_rules, ArrayList* out, size_t* count_out) {
|
||||
if (count_out)
|
||||
*count_out = 0;
|
||||
if (!manifest || !out)
|
||||
return false;
|
||||
PathIndex keep;
|
||||
if (!build_keep_index(manifest, &keep))
|
||||
return false;
|
||||
PathIndex dirs;
|
||||
bool have_dirs = synced_dirs != NULL;
|
||||
if (have_dirs &&
|
||||
!path_index_build(&dirs, (const char* const*)synced_dirs->items, (size_t)synced_dirs->size)) {
|
||||
path_index_free(&keep);
|
||||
return false;
|
||||
}
|
||||
int rootfd = open_destination_root(dest_root);
|
||||
if (rootfd < 0) {
|
||||
path_index_free(&keep);
|
||||
if (have_dirs)
|
||||
path_index_free(&dirs);
|
||||
return false;
|
||||
}
|
||||
bool all_removed = false;
|
||||
size_t recorded = 0;
|
||||
DeleteWalkState state = {.mode = DELETE_WALK_MODE_LIST,
|
||||
.budget = NULL,
|
||||
.out = out,
|
||||
.recorded = &recorded,
|
||||
.observer = NULL,
|
||||
.observer_context = NULL};
|
||||
bool ok = delete_walk_fd(rootfd, "", &keep, have_dirs ? &dirs : NULL, &state, skips, skip_count,
|
||||
protect_rules, false, &all_removed);
|
||||
if (close(rootfd) != 0)
|
||||
ok = false;
|
||||
path_index_free(&keep);
|
||||
if (have_dirs)
|
||||
path_index_free(&dirs);
|
||||
if (count_out)
|
||||
*count_out = recorded;
|
||||
return ok;
|
||||
}
|
||||
|
||||
DeleteWalkResult delete_extras_limited_observed(const char* dest_root, const ArrayList* manifest,
|
||||
const ArrayList* synced_dirs, size_t max_delete,
|
||||
const DeleteSkipEntry* skips, int skip_count,
|
||||
const FilterRuleList* protect_rules,
|
||||
size_t* deleted_out, size_t* skipped_out,
|
||||
DeletePathObserver observer,
|
||||
void* observer_context) {
|
||||
if (deleted_out)
|
||||
*deleted_out = 0;
|
||||
if (skipped_out)
|
||||
*skipped_out = 0;
|
||||
if (!manifest)
|
||||
return DELETE_WALK_ERROR;
|
||||
/* Index the keep-set (and the synchronized-dir set, when supplied) once so
|
||||
membership is answered in O(path length) instead of scanning every entry
|
||||
for every destination entry. */
|
||||
PathIndex keep;
|
||||
if (!build_keep_index(manifest, &keep))
|
||||
return DELETE_WALK_ERROR;
|
||||
PathIndex dirs;
|
||||
bool have_dirs = synced_dirs != NULL;
|
||||
if (have_dirs &&
|
||||
!path_index_build(&dirs, (const char* const*)synced_dirs->items, (size_t)synced_dirs->size)) {
|
||||
path_index_free(&keep);
|
||||
return DELETE_WALK_ERROR;
|
||||
}
|
||||
int rootfd = open_destination_root(dest_root);
|
||||
if (rootfd < 0) {
|
||||
path_index_free(&keep);
|
||||
if (have_dirs)
|
||||
path_index_free(&dirs);
|
||||
return DELETE_WALK_ERROR;
|
||||
}
|
||||
DeleteBudget budget = {.max_delete = max_delete, .deleted = 0, .skipped = 0, .limit_hit = false};
|
||||
bool all_removed = false;
|
||||
DeleteWalkState state = {.mode = DELETE_WALK_MODE_DELETE,
|
||||
.budget = &budget,
|
||||
.out = NULL,
|
||||
.recorded = NULL,
|
||||
.observer = observer,
|
||||
.observer_context = observer_context};
|
||||
bool ok = delete_walk_fd(rootfd, "", &keep, have_dirs ? &dirs : NULL, &state, skips, skip_count,
|
||||
protect_rules, false, &all_removed);
|
||||
if (close(rootfd) != 0)
|
||||
ok = false;
|
||||
path_index_free(&keep);
|
||||
if (have_dirs)
|
||||
path_index_free(&dirs);
|
||||
if (deleted_out)
|
||||
*deleted_out = budget.deleted;
|
||||
if (skipped_out)
|
||||
*skipped_out = budget.skipped;
|
||||
if (!ok)
|
||||
return DELETE_WALK_ERROR;
|
||||
return budget.limit_hit ? DELETE_WALK_LIMIT_REACHED : DELETE_WALK_OK;
|
||||
}
|
||||
|
||||
DeleteWalkResult delete_extras_limited(const char* dest_root, const ArrayList* manifest,
|
||||
const ArrayList* synced_dirs, size_t max_delete,
|
||||
const DeleteSkipEntry* skips, int skip_count,
|
||||
const FilterRuleList* protect_rules, size_t* deleted_out,
|
||||
size_t* skipped_out) {
|
||||
return delete_extras_limited_observed(dest_root, manifest, synced_dirs, max_delete, skips,
|
||||
skip_count, protect_rules, deleted_out, skipped_out, NULL,
|
||||
NULL);
|
||||
}
|
||||
|
||||
bool delete_extras(const char* dest_root, const ArrayList* manifest) {
|
||||
return delete_extras_limited(dest_root, manifest, NULL, SIZE_MAX, NULL, 0, NULL, NULL, NULL) ==
|
||||
DELETE_WALK_OK;
|
||||
}
|
||||
|
||||
/* Build the delete-walk protection prefix for one basis directory. The walker
|
||||
compares paths relative to the receive root, so a relative entry is already
|
||||
in the right form; an absolute entry that lies below the root is converted to
|
||||
its root-relative form, and one outside the root returns NULL (the walk
|
||||
cannot reach it, and it is not protected data beneath the root). Exposed so
|
||||
tests can exercise the root-of-"/" child mapping directly. */
|
||||
char* delete_basis_relative(const Config* config, const char* path) {
|
||||
if (!path)
|
||||
return NULL;
|
||||
if (path[0] != '/')
|
||||
return str_dup(path);
|
||||
const char* root = config->receive_root_directory;
|
||||
if (!root || root[0] != '/')
|
||||
return NULL;
|
||||
size_t root_len = strlen(root);
|
||||
while (root_len > 1 && root[root_len - 1] == '/')
|
||||
root_len--;
|
||||
if (strncmp(path, root, root_len) != 0)
|
||||
return NULL;
|
||||
if (root_len == 1) {
|
||||
/* `root` is "/" (the only single-character absolute root): every absolute
|
||||
path is below it, and the child relative form is everything after the
|
||||
leading '/'. */
|
||||
if (path[1] == '\0')
|
||||
return NULL; /* identical to the root, not a child */
|
||||
return str_dup(path + 1);
|
||||
}
|
||||
if (path[root_len] != '/')
|
||||
return NULL; /* identical or a sibling sharing a name prefix */
|
||||
return str_dup(path + root_len + 1);
|
||||
}
|
||||
|
||||
bool delete_skips_build(const Config* config, const ArrayList* protected_paths,
|
||||
const ArrayList* size_skipped, bool basis_root_relative,
|
||||
DeleteSkipSet* out) {
|
||||
if (!out)
|
||||
return false;
|
||||
out->entries = NULL;
|
||||
out->owned_prefixes = NULL;
|
||||
out->count = 0;
|
||||
out->owned_count = 0;
|
||||
if (!config)
|
||||
return false;
|
||||
int protected_count = protected_paths ? protected_paths->size : 0;
|
||||
int size_skipped_count = size_skipped ? size_skipped->size : 0;
|
||||
int count =
|
||||
(config->delay_updates ? 1 : 0) + config->basis_count + protected_count + size_skipped_count;
|
||||
if (count == 0)
|
||||
return true;
|
||||
out->entries = calloc((size_t)count, sizeof(DeleteSkipEntry));
|
||||
if (!out->entries)
|
||||
return false;
|
||||
if (basis_root_relative && config->basis_count > 0) {
|
||||
out->owned_prefixes = calloc((size_t)config->basis_count, sizeof(char*));
|
||||
if (!out->owned_prefixes) {
|
||||
free(out->entries);
|
||||
out->entries = NULL;
|
||||
return false;
|
||||
}
|
||||
out->owned_count = config->basis_count;
|
||||
}
|
||||
int idx = 0;
|
||||
if (config->delay_updates) {
|
||||
out->entries[idx].prefix = DELAY_UPDATES_STAGING_DIR;
|
||||
out->entries[idx].top_level_only = true;
|
||||
idx++;
|
||||
}
|
||||
for (int i = 0; i < config->basis_count; i++) {
|
||||
const char* prefix = config->basis_dirs[i].path;
|
||||
if (basis_root_relative) {
|
||||
/* An absolute basis outside the receive root is unreachable by this walk,
|
||||
so it contributes no protection prefix (and no slot). */
|
||||
char* relative = delete_basis_relative(config, config->basis_dirs[i].path);
|
||||
if (!relative)
|
||||
continue;
|
||||
out->owned_prefixes[i] = relative;
|
||||
prefix = relative;
|
||||
}
|
||||
out->entries[idx].prefix = prefix;
|
||||
out->entries[idx].top_level_only = false;
|
||||
idx++;
|
||||
}
|
||||
for (int i = 0; i < protected_count; i++) {
|
||||
out->entries[idx].prefix = (const char*)protected_paths->items[i];
|
||||
out->entries[idx].top_level_only = false;
|
||||
idx++;
|
||||
}
|
||||
for (int i = 0; i < size_skipped_count; i++) {
|
||||
out->entries[idx].prefix = (const char*)size_skipped->items[i];
|
||||
out->entries[idx].top_level_only = false;
|
||||
idx++;
|
||||
}
|
||||
out->count = idx;
|
||||
return true;
|
||||
}
|
||||
|
||||
void delete_skips_free(DeleteSkipSet* set) {
|
||||
if (!set)
|
||||
return;
|
||||
if (set->owned_prefixes) {
|
||||
for (int i = 0; i < set->owned_count; i++)
|
||||
free(set->owned_prefixes[i]);
|
||||
}
|
||||
free(set->owned_prefixes);
|
||||
free(set->entries);
|
||||
set->entries = NULL;
|
||||
set->owned_prefixes = NULL;
|
||||
set->count = 0;
|
||||
set->owned_count = 0;
|
||||
}
|
||||
@@ -0,0 +1,151 @@
|
||||
#ifndef DELETE_H
|
||||
#define DELETE_H
|
||||
|
||||
#include "array_list.h"
|
||||
#include "config.h"
|
||||
#include <stdbool.h>
|
||||
#include <stddef.h>
|
||||
|
||||
/* Delete engine.
|
||||
*
|
||||
* This module owns destination-relative delete traversal: the ordered directory
|
||||
* walker that reproduces rsync's extraneous-entry order, the skip-prefix
|
||||
* protection set shared by every delete pass, and the read-only enumeration
|
||||
* that mirrors the walker for -n/--dry-run. The budgeted manifest commit
|
||||
* (delete_commit.c) and the per-directory delete plans (delete_plan.c) are
|
||||
* built on the primitives exported here. */
|
||||
|
||||
/* Result of a bounded extra-file deletion run. */
|
||||
typedef enum {
|
||||
/* Every extra entry was removed (or there were none). */
|
||||
DELETE_WALK_OK = 0,
|
||||
/* The numeric cap for this run was reached before every extra was removed.
|
||||
The walker removed exactly the entries the cap allowed and skipped (without
|
||||
removing) the rest, matching rsync's partial --max-delete behavior. */
|
||||
DELETE_WALK_LIMIT_REACHED,
|
||||
/* A traversal or unlink failure aborted the deletion (partial removal is
|
||||
possible, mirroring the delete pass). */
|
||||
DELETE_WALK_ERROR
|
||||
} DeleteWalkResult;
|
||||
|
||||
/* One protected entry for the delete walker. When top_level_only is true the
|
||||
prefix is skipped only as a DIRECT child of dest_root (the --delay-updates
|
||||
staging directory, which must not hide genuine extras inside a nested
|
||||
destination directory that happens to share the staging name); otherwise the
|
||||
prefix is skipped at any depth (the --compare-dest/--copy-dest/--link-dest
|
||||
basis trees, and the sender-side protected filter-excluded prefixes, which
|
||||
are never destination content). */
|
||||
typedef struct {
|
||||
const char* prefix;
|
||||
bool top_level_only;
|
||||
} DeleteSkipEntry;
|
||||
|
||||
/* A built skip-prefix set. `entries`/`count` are what path_under_skip_prefix()
|
||||
consumes. `owned_prefixes` holds any prefix strings the builder had to
|
||||
allocate (root-relative basis-dir conversions); it is NULL when every prefix
|
||||
is borrowed from the config or the caller's lists. Release with
|
||||
delete_skips_free(). */
|
||||
typedef struct {
|
||||
DeleteSkipEntry* entries;
|
||||
char** owned_prefixes;
|
||||
int count;
|
||||
int owned_count;
|
||||
} DeleteSkipSet;
|
||||
|
||||
/* True when child_rel is, or lies below, one of the protected entries (a prefix
|
||||
"a" protects "a" and "a/b/c" but not "ab"; top_level_only entries protect
|
||||
only DIRECT children of the destination root, i.e. child_rel has no '/'). */
|
||||
bool path_under_skip_prefix(const char* child_rel, bool at_root, const DeleteSkipEntry* skips,
|
||||
int skip_count);
|
||||
|
||||
/* One destination-directory entry collected up front so the delete walkers can
|
||||
reproduce rsync's traversal order instead of readdir() order. rsync processes
|
||||
a directory's extraneous subdirectories first (descending name, depth-first),
|
||||
then its extraneous files (descending name), and only afterwards descends into
|
||||
its kept subdirectories (ascending name). */
|
||||
typedef struct {
|
||||
char* name;
|
||||
bool is_dir;
|
||||
} DeleteDirEntry;
|
||||
/* Collect the entries of the directory open on `dirfd` (excluding "." and ".."),
|
||||
stat'ing each with AT_SYMLINK_NOFOLLOW. On success *out is a malloc'd array of
|
||||
*count entries whose names the caller frees with delete_dir_entries_free().
|
||||
Returns false on an allocation/readdir failure; a vanished entry (ENOENT) is
|
||||
skipped, any other stat failure is reported through *operation_ok while the
|
||||
walk continues. */
|
||||
bool delete_dir_entries_collect(int dirfd, DeleteDirEntry** out, size_t* count, bool* operation_ok);
|
||||
void delete_dir_entries_free(DeleteDirEntry* entries, size_t count);
|
||||
/* Sort comparators: `_desc` orders subdirectories before files and each group by
|
||||
descending name (rsync's extraneous-entry order); `_asc` orders plain ascending
|
||||
name (rsync's kept-subdirectory order). */
|
||||
int delete_dir_entry_cmp_desc(const void* a, const void* b);
|
||||
int delete_dir_entry_cmp_asc(const void* a, const void* b);
|
||||
|
||||
/* Remove files/dirs/symlinks under dest_root that are not listed in manifest
|
||||
without ever descending into a protected prefix (see DeleteSkipEntry). When
|
||||
`synced_dirs` is non-NULL, extras are only removed directly inside a directory
|
||||
whose destination-relative path is an exact entry in that list (the receive
|
||||
root is the "." sentinel); directories outside the synchronized set are still
|
||||
descended into so kept content below a listed directory is preserved, but
|
||||
nothing in them is removed. A NULL `synced_dirs` keeps the legacy behavior of
|
||||
treating the whole destination tree as deletable. `max_delete` caps the
|
||||
number of removed entries (SIZE_MAX = unlimited): the walker removes up to the
|
||||
cap and returns DELETE_WALK_LIMIT_REACHED when more extras remained.
|
||||
`deleted_out`/`skipped_out` optionally receive the number of entries removed
|
||||
and the number skipped because of the cap. */
|
||||
DeleteWalkResult delete_extras_limited(const char* dest_root, const ArrayList* manifest,
|
||||
const ArrayList* synced_dirs, size_t max_delete,
|
||||
const DeleteSkipEntry* skips, int skip_count,
|
||||
const FilterRuleList* protect_rules, size_t* deleted_out,
|
||||
size_t* skipped_out);
|
||||
|
||||
/* Optional per-deletion observer: called for each destination-relative path
|
||||
actually removed (a file, symlink, or directory), in removal order, so the
|
||||
receiver can stream rsync's `--info=del`/`--info=remove` lines. */
|
||||
typedef void (*DeletePathObserver)(void* context, const char* rel_path);
|
||||
|
||||
/* `delete_extras_limited_observed` is delete_extras_limited with an optional
|
||||
* observer; the observer is invoked only for entries truly removed. When
|
||||
* `protect_rules` is non-NULL its receiver-side verdict is evaluated for every
|
||||
* candidate extra: a first-match PROTECT leaves the entry (and, for a
|
||||
* directory, its whole subtree) in place, while RISK/NONE fall through to the
|
||||
* ordinary skip-prefix/keep-set logic. */
|
||||
DeleteWalkResult delete_extras_limited_observed(const char* dest_root, const ArrayList* manifest,
|
||||
const ArrayList* synced_dirs, size_t max_delete,
|
||||
const DeleteSkipEntry* skips, int skip_count,
|
||||
const FilterRuleList* protect_rules,
|
||||
size_t* deleted_out, size_t* skipped_out,
|
||||
DeletePathObserver observer,
|
||||
void* observer_context);
|
||||
/* Read-only companion to delete_extras_limited: walk the destination exactly as
|
||||
the delete pass would and APPEND (strdup'd) destination-relative paths that
|
||||
WOULD be removed, without touching disk. Used for -n/--dry-run --delete
|
||||
would-delete reporting. Returns true on a clean walk; the caller owns the
|
||||
strings appended to `out` and receives their count in *count_out. */
|
||||
bool delete_extras_list(const char* dest_root, const ArrayList* manifest,
|
||||
const ArrayList* synced_dirs, const DeleteSkipEntry* skips, int skip_count,
|
||||
const FilterRuleList* protect_rules, ArrayList* out, size_t* count_out);
|
||||
bool delete_extras(const char* dest_root, const ArrayList* manifest);
|
||||
|
||||
/* Build the delete walk's skip-prefix set from the config's --delay-updates
|
||||
staging directory, its --compare-dest/--copy-dest/--link-dest basis dirs, and
|
||||
the caller-supplied protection lists, in that order. `protected_paths` and
|
||||
`size_skipped` are borrowed (may be NULL); every entry in them is protected at
|
||||
any depth. The staging directory is protected only as a DIRECT child of the
|
||||
receive root. `basis_root_relative` selects how a basis path becomes a
|
||||
prefix: true converts an absolute path under the receive root to its
|
||||
root-relative form (the whole-tree commit walk; an unreachable path
|
||||
contributes no slot), false keeps the configured path verbatim (the
|
||||
per-directory plan walk). On success the caller releases `*out` with
|
||||
delete_skips_free(); returns false on allocation failure. */
|
||||
bool delete_skips_build(const Config* config, const ArrayList* protected_paths,
|
||||
const ArrayList* size_skipped, bool basis_root_relative,
|
||||
DeleteSkipSet* out);
|
||||
void delete_skips_free(DeleteSkipSet* set);
|
||||
|
||||
/* Convert one basis-directory path to the receive-root-relative protection
|
||||
prefix the delete walker uses (NULL when it lies outside the root). Exposed
|
||||
for unit tests of the root-of-"/" and normalization edge cases. */
|
||||
char* delete_basis_relative(const Config* config, const char* path);
|
||||
|
||||
#endif
|
||||
@@ -0,0 +1,488 @@
|
||||
#include <errno.h>
|
||||
#include <ctype.h>
|
||||
#include <dirent.h>
|
||||
#include <fcntl.h>
|
||||
#include <libgen.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <sys/stat.h>
|
||||
#include <sys/sysmacros.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#include "array_list.h"
|
||||
#include "charset.h"
|
||||
#include "chmod.h"
|
||||
#include "chunk.h"
|
||||
#include "compression.h"
|
||||
#include "config.h"
|
||||
#include "data.h"
|
||||
#include "delay_updates.h"
|
||||
#include "delete_commit.h"
|
||||
#include "delta.h"
|
||||
#include "file.h"
|
||||
#include "format.h"
|
||||
#include "identity.h"
|
||||
#include "log.h"
|
||||
#include "metadata.h"
|
||||
#include "protocol.h"
|
||||
#include "utils.h"
|
||||
#include "xattr.h"
|
||||
|
||||
#define MAX_SERVER_DELETE_COUNT 100000U
|
||||
/* Retained cost of one delete-manifest entry beyond its path bytes: the
|
||||
ArrayList pointer slot plus an approximate malloc header/rounding for the
|
||||
heap copy. Charged against MAX_MANIFEST_BYTES so a frame full of tiny paths
|
||||
cannot retain far more than the byte budget (B5). */
|
||||
#define MANIFEST_ENTRY_OVERHEAD (sizeof(char*) + 16)
|
||||
|
||||
/* Read a delete-manifest frame (the STATUS_MANIFEST leading code has already
|
||||
been consumed): a keep-set entry count followed by that many
|
||||
destination-relative paths, then a protected-prefix count followed by that
|
||||
many destination-relative prefixes, then a missing-args count followed by that
|
||||
many destination-relative delete paths, then (protocol 2.23.0) a
|
||||
synchronized-directory count followed by that many destination-relative
|
||||
directory paths (the receive root is the "." sentinel). The frame is
|
||||
self-delimiting (the counts are authoritative), so the caller decides what to
|
||||
do next and continues reading the following STATUS_* frame. Every section is
|
||||
validated identically: an entry must be non-empty, relative and traversal-free
|
||||
and the aggregate length across ALL sections is capped by MAX_MANIFEST_BYTES
|
||||
(so the missing-args deletion requests are confined like the rest of the
|
||||
manifest). Returns an owned DeleteManifest, or NULL after sending STATUS_ERROR
|
||||
when the frame is malformed (bad count, empty/absolute path, path traversal,
|
||||
or an aggregate size beyond MAX_MANIFEST_BYTES). */
|
||||
static bool receive_manifest_section(int fd, ArrayList* list, size_t* manifest_bytes,
|
||||
size_t* manifest_entries) {
|
||||
int count;
|
||||
if (!receive_int(fd, &count)) {
|
||||
send_status(fd, STATUS_ERROR);
|
||||
return false;
|
||||
}
|
||||
if (count < 0 || count > MAX_MANIFEST_ENTRIES ||
|
||||
(size_t)count > MAX_MANIFEST_ENTRIES - *manifest_entries) {
|
||||
send_status(fd, STATUS_ERROR);
|
||||
return false;
|
||||
}
|
||||
for (int i = 0; i < count; i++) {
|
||||
char* s = receive_wire_str(fd);
|
||||
size_t entry_size = s ? strlen(s) + MANIFEST_ENTRY_OVERHEAD : 0;
|
||||
if (!s || s[0] == '\0' || s[0] == '/' || has_path_traversal(s) ||
|
||||
entry_size > MAX_MANIFEST_BYTES - *manifest_bytes ||
|
||||
(*manifest_bytes += entry_size) > MAX_MANIFEST_BYTES || !array_list_add(list, s)) {
|
||||
free(s);
|
||||
send_status(fd, STATUS_ERROR);
|
||||
return false;
|
||||
}
|
||||
}
|
||||
*manifest_entries += (size_t)count;
|
||||
return true;
|
||||
}
|
||||
|
||||
DeleteManifest* receive_manifest_entries(int fd) {
|
||||
DeleteManifest* manifest = calloc(1, sizeof(DeleteManifest));
|
||||
if (!manifest) {
|
||||
send_status(fd, STATUS_ERROR);
|
||||
return NULL;
|
||||
}
|
||||
manifest->keeps = array_list_create(free);
|
||||
manifest->protected = array_list_create(free);
|
||||
manifest->missing = array_list_create(free);
|
||||
manifest->dirs = array_list_create(free);
|
||||
if (!manifest->keeps || !manifest->protected || !manifest->missing || !manifest->dirs) {
|
||||
delete_manifest_free(manifest);
|
||||
send_status(fd, STATUS_ERROR);
|
||||
return NULL;
|
||||
}
|
||||
size_t manifest_bytes = 0;
|
||||
size_t manifest_entries = 0;
|
||||
if (!receive_manifest_section(fd, manifest->keeps, &manifest_bytes, &manifest_entries) ||
|
||||
!receive_manifest_section(fd, manifest->protected, &manifest_bytes, &manifest_entries) ||
|
||||
!receive_manifest_section(fd, manifest->missing, &manifest_bytes, &manifest_entries) ||
|
||||
!receive_manifest_section(fd, manifest->dirs, &manifest_bytes, &manifest_entries)) {
|
||||
delete_manifest_free(manifest);
|
||||
return NULL;
|
||||
}
|
||||
return manifest;
|
||||
}
|
||||
|
||||
void delete_manifest_free(DeleteManifest* manifest) {
|
||||
if (!manifest)
|
||||
return;
|
||||
array_list_delete(manifest->keeps);
|
||||
array_list_delete(manifest->protected);
|
||||
array_list_delete(manifest->missing);
|
||||
array_list_delete(manifest->dirs);
|
||||
free(manifest);
|
||||
}
|
||||
|
||||
/* Shared --max-delete budget for one receiver-side deletion commit. Both the
|
||||
--delete-missing-args exact-path removals and the ordinary extras walk draw
|
||||
from the same tally, matching rsync (whose --max-delete counts every deleted
|
||||
file or directory). `max_delete` is SIZE_MAX for an unlimited budget. */
|
||||
typedef struct {
|
||||
size_t max_delete;
|
||||
size_t deleted;
|
||||
size_t skipped;
|
||||
bool limit_hit;
|
||||
} DeleteBudgetState;
|
||||
|
||||
/* Remove every destination entry under the receive root that is not in the
|
||||
keep-set, bounded by the shared budget (a smaller client --max-delete=NUM
|
||||
replaces the server hard bound; rsync deletes up to the bound and skips the
|
||||
rest). With --delay-updates the not-yet-published staging directory is a
|
||||
direct child of the receive root and must not be treated as a set of extras;
|
||||
the manifest's protected prefixes (paths excluded on the source), the
|
||||
size-pruned prefixes (--max-size/--min-size, always protected) and the
|
||||
alternate basis directories are never destination content and are skipped at
|
||||
any depth. Returns true unless a traversal/unlink error aborted the walk;
|
||||
the budget's limit_hit/skipped fields report a cap-stopped run. */
|
||||
static bool delete_extras_budgeted_observed(const Config* config, const DeleteManifest* manifest,
|
||||
DeleteBudgetState* budget, DeletePathObserver observer,
|
||||
void* observer_context) {
|
||||
if (!config || !manifest || !manifest->keeps)
|
||||
return false;
|
||||
fprintf(stderr, "Deleting files not in manifest...\n");
|
||||
/* Protected entries: the --delay-updates staging name (only as a DIRECT child
|
||||
of the receive root), the alternate basis directories and the sender-side
|
||||
protected prefixes (filter-excluded and size-pruned source mirrors), all at
|
||||
any depth. See delete_skips_build(). */
|
||||
DeleteSkipSet skips;
|
||||
if (!delete_skips_build(config, manifest->protected, NULL, true, &skips))
|
||||
return false;
|
||||
/* Clamp rather than subtract: an accounting bug where deleted already exceeds
|
||||
max_delete must never underflow into an effectively unlimited budget. */
|
||||
size_t remaining;
|
||||
if (budget->max_delete == SIZE_MAX)
|
||||
remaining = SIZE_MAX;
|
||||
else if (budget->deleted >= budget->max_delete)
|
||||
remaining = 0;
|
||||
else
|
||||
remaining = budget->max_delete - budget->deleted;
|
||||
size_t deleted = 0;
|
||||
size_t skipped = 0;
|
||||
DeleteWalkResult result = delete_extras_limited_observed(
|
||||
config->receive_root_directory, manifest->keeps, manifest->dirs, remaining, skips.entries,
|
||||
skips.count, config->protect_rules, &deleted, &skipped, observer, observer_context);
|
||||
delete_skips_free(&skips);
|
||||
budget->deleted += deleted;
|
||||
budget->skipped += skipped;
|
||||
if (result == DELETE_WALK_LIMIT_REACHED) {
|
||||
budget->limit_hit = true;
|
||||
return true;
|
||||
}
|
||||
if (result != DELETE_WALK_OK) {
|
||||
log_message(LOG_LEVEL_ERROR, "deletion failed while removing extraneous files");
|
||||
return false;
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
static bool delete_extras_budgeted(const Config* config, const DeleteManifest* manifest,
|
||||
DeleteBudgetState* budget) {
|
||||
return delete_extras_budgeted_observed(config, manifest, budget, NULL, NULL);
|
||||
}
|
||||
|
||||
/* Prefixes every observed path with a fixed subtree root, so a nested walk
|
||||
(a recursively removed missing-arg directory) reports receive-root-relative
|
||||
names like the rest of the delete output. */
|
||||
typedef struct {
|
||||
DeletePathObserver inner;
|
||||
void* inner_context;
|
||||
const char* prefix;
|
||||
} PrefixedDeleteObserver;
|
||||
|
||||
static void prefixed_delete_observer(void* context, const char* rel) {
|
||||
PrefixedDeleteObserver* prefixed = context;
|
||||
if (!prefixed->inner || !rel)
|
||||
return;
|
||||
char* joined = path_cat((char*)prefixed->prefix, rel);
|
||||
if (joined) {
|
||||
prefixed->inner(prefixed->inner_context, joined);
|
||||
free(joined);
|
||||
}
|
||||
}
|
||||
|
||||
/* --delete-missing-args exact-path deletions: each destination mirror in
|
||||
manifest->missing is an explicit user request, so it is removed even when the
|
||||
ordinary extras walk (with its protected prefixes) would leave it alone. The
|
||||
--delay-updates staging directory and basis snapshots are receiver artifacts
|
||||
and stay protected exactly as in the extras walker. A regular file or
|
||||
symlink is unlinked, an empty directory removed, and a NON-empty directory is
|
||||
removed recursively only when --delete or --force is in effect (rsync parity:
|
||||
the man page says a non-empty directory mirror is only deleted with --force
|
||||
or --delete); otherwise it is left with a warning and the run continues. A
|
||||
mirror that does not exist is a no-op. Each removal draws from the shared
|
||||
--max-delete budget: once it is exhausted the remaining requests are skipped
|
||||
and counted. Returns false only on a genuine error (a confinement failure on
|
||||
a validated path or an I/O error), which fails the run. */
|
||||
static bool delete_missing_args_budgeted_observed(const Config* config,
|
||||
const DeleteManifest* manifest,
|
||||
DeleteBudgetState* budget,
|
||||
DeletePathObserver observer,
|
||||
void* observer_context) {
|
||||
if (!config || !manifest)
|
||||
return false;
|
||||
if (!manifest->missing || manifest->missing->size == 0)
|
||||
return true;
|
||||
fprintf(stderr, "Deleting destination mirrors of missing source arguments...\n");
|
||||
/* The staging directory and basis snapshots stay protected exactly as in the
|
||||
extras walker (the missing-args path overrides the ordinary protected
|
||||
prefixes, so those are not passed here). */
|
||||
DeleteSkipSet skips;
|
||||
if (!delete_skips_build(config, NULL, NULL, true, &skips))
|
||||
return false;
|
||||
bool ok = true;
|
||||
for (int i = 0; i < manifest->missing->size; i++) {
|
||||
const char* rel = (const char*)manifest->missing->items[i];
|
||||
if (!rel || *rel == '\0' || *rel == '/' || has_path_traversal(rel)) {
|
||||
/* Defensive only: receive_manifest_entries already validated every
|
||||
section identically, so a controlled peer never reaches this branch. */
|
||||
log_message(LOG_LEVEL_ERROR, "invalid missing-args delete path");
|
||||
ok = false;
|
||||
continue;
|
||||
}
|
||||
bool at_root = strchr(rel, '/') == NULL;
|
||||
if (path_under_skip_prefix(rel, at_root, skips.entries, skips.count)) {
|
||||
char* escaped = output_escape(rel, log_get_8_bit_output());
|
||||
log_message(LOG_LEVEL_WARNING,
|
||||
"missing-args path '%s' is protected (staging directory or basis snapshot); "
|
||||
"not deleting",
|
||||
escaped ? escaped : "<allocation failed>");
|
||||
free(escaped);
|
||||
continue;
|
||||
}
|
||||
char* full = path_cat(config->receive_root_directory, rel);
|
||||
if (!full) {
|
||||
ok = false;
|
||||
continue;
|
||||
}
|
||||
char* leaf = NULL;
|
||||
int parent_fd = file_open_secure_parent(full, &leaf, false);
|
||||
if (parent_fd < 0) {
|
||||
/* The mirror's parent directory may itself not exist on the destination
|
||||
(a deeper missing entry whose leading directories were never created).
|
||||
That is a no-op -- there is nothing to delete -- matching
|
||||
file_remove_tree_secure's absent-path handling; only a genuine I/O
|
||||
error (EACCES, a symlink loop, ...) fails the run. */
|
||||
bool absent = errno == ENOENT || errno == ENOTDIR;
|
||||
free(full);
|
||||
free(leaf);
|
||||
if (!absent)
|
||||
ok = false;
|
||||
continue;
|
||||
}
|
||||
struct stat st;
|
||||
if (fstatat(parent_fd, leaf, &st, AT_SYMLINK_NOFOLLOW) != 0) {
|
||||
/* Already absent: nothing to delete (a no-op, not a deletion). */
|
||||
if (errno != ENOENT)
|
||||
ok = false;
|
||||
close(parent_fd);
|
||||
free(leaf);
|
||||
free(full);
|
||||
continue;
|
||||
}
|
||||
/* An entry that exists is one deletion: skip it (and count it) when the
|
||||
shared --max-delete budget is already exhausted. */
|
||||
if (budget->deleted >= budget->max_delete) {
|
||||
budget->limit_hit = true;
|
||||
budget->skipped++;
|
||||
close(parent_fd);
|
||||
free(leaf);
|
||||
free(full);
|
||||
continue;
|
||||
}
|
||||
bool removed = false;
|
||||
if (S_ISDIR(st.st_mode)) {
|
||||
if (unlinkat(parent_fd, leaf, AT_REMOVEDIR) == 0) {
|
||||
removed = true;
|
||||
} else if (errno == ENOTEMPTY || errno == EEXIST) {
|
||||
close(parent_fd);
|
||||
parent_fd = -1;
|
||||
free(leaf);
|
||||
leaf = NULL;
|
||||
if (config->use_delete || config->force_delete) {
|
||||
/* Remove the contents entry-by-entry through the budgeted extras
|
||||
walker so every deleted file/dir counts toward --max-delete (rsync
|
||||
parity); the now-empty directory itself costs one more. A run that
|
||||
hits the cap leaves the remaining entries in place. */
|
||||
ArrayList* no_keeps = array_list_create(free);
|
||||
/* Never let an accounting slip (deleted > max_delete) underflow the
|
||||
remaining budget into SIZE_MAX, which would grant unlimited
|
||||
deletions. */
|
||||
size_t remaining =
|
||||
budget->deleted >= budget->max_delete ? 0 : budget->max_delete - budget->deleted;
|
||||
size_t contents_deleted = 0;
|
||||
size_t contents_skipped = 0;
|
||||
PrefixedDeleteObserver nested = {observer, observer_context, rel};
|
||||
DeleteWalkResult walk =
|
||||
no_keeps ? delete_extras_limited_observed(full, no_keeps, NULL, remaining, NULL, 0,
|
||||
NULL, &contents_deleted, &contents_skipped,
|
||||
observer ? prefixed_delete_observer : NULL,
|
||||
observer ? &nested : NULL)
|
||||
: DELETE_WALK_ERROR;
|
||||
if (no_keeps)
|
||||
array_list_delete(no_keeps);
|
||||
budget->deleted += contents_deleted;
|
||||
budget->skipped += contents_skipped;
|
||||
if (walk == DELETE_WALK_LIMIT_REACHED) {
|
||||
budget->limit_hit = true;
|
||||
} else if (walk != DELETE_WALK_OK) {
|
||||
ok = false;
|
||||
} else if (budget->deleted >= budget->max_delete) {
|
||||
budget->limit_hit = true;
|
||||
budget->skipped++;
|
||||
} else if (file_remove_tree_secure(full)) {
|
||||
/* The shared `if (removed)` tail charges this directory exactly
|
||||
once; counting it here too would consume two budget units. */
|
||||
removed = true;
|
||||
} else {
|
||||
ok = false;
|
||||
}
|
||||
} else {
|
||||
char* escaped = output_escape(rel, log_get_8_bit_output());
|
||||
log_message(LOG_LEVEL_WARNING,
|
||||
"missing-args destination '%s' is a non-empty directory; use --force or "
|
||||
"--delete to remove it",
|
||||
escaped ? escaped : "<allocation failed>");
|
||||
free(escaped);
|
||||
}
|
||||
} else if (errno != ENOENT) {
|
||||
ok = false;
|
||||
}
|
||||
} else {
|
||||
if (unlinkat(parent_fd, leaf, 0) == 0) {
|
||||
removed = true;
|
||||
} else if (errno != ENOENT) {
|
||||
ok = false;
|
||||
}
|
||||
}
|
||||
if (removed) {
|
||||
budget->deleted++;
|
||||
if (observer)
|
||||
observer(observer_context, rel);
|
||||
char* escaped = output_escape(rel, log_get_8_bit_output());
|
||||
fprintf(stderr, " Deleted: %s\n", escaped ? escaped : "<allocation failed>");
|
||||
free(escaped);
|
||||
}
|
||||
if (parent_fd >= 0)
|
||||
close(parent_fd);
|
||||
free(leaf);
|
||||
free(full);
|
||||
if (!ok)
|
||||
break;
|
||||
}
|
||||
delete_skips_free(&skips);
|
||||
return ok;
|
||||
}
|
||||
|
||||
/* Public wrappers used outside the commit path (and by unit tests): no
|
||||
--max-delete budget. */
|
||||
bool manifest_would_delete_list(const Config* config, const DeleteManifest* manifest,
|
||||
ArrayList* out, size_t* count_out) {
|
||||
if (count_out)
|
||||
*count_out = 0;
|
||||
if (!config || !manifest || !manifest->keeps || !out)
|
||||
return false;
|
||||
DeleteSkipSet skips;
|
||||
if (!delete_skips_build(config, manifest->protected, NULL, true, &skips))
|
||||
return false;
|
||||
bool ok = delete_extras_list(config->receive_root_directory, manifest->keeps, manifest->dirs,
|
||||
skips.entries, skips.count, config->protect_rules, out, count_out);
|
||||
delete_skips_free(&skips);
|
||||
return ok;
|
||||
}
|
||||
|
||||
bool manifest_delete_extras(const Config* config, const DeleteManifest* manifest) {
|
||||
DeleteBudgetState budget = {
|
||||
.max_delete = SIZE_MAX, .deleted = 0, .skipped = 0, .limit_hit = false};
|
||||
return delete_extras_budgeted(config, manifest, &budget);
|
||||
}
|
||||
|
||||
bool manifest_delete_missing_args(const Config* config, const DeleteManifest* manifest) {
|
||||
DeleteBudgetState budget = {
|
||||
.max_delete = SIZE_MAX, .deleted = 0, .skipped = 0, .limit_hit = false};
|
||||
return delete_missing_args_budgeted_observed(config, manifest, &budget, NULL, NULL);
|
||||
}
|
||||
|
||||
bool manifest_delete_missing_args_limited(const Config* config, const DeleteManifest* manifest,
|
||||
size_t max_delete, size_t* deleted, size_t* skipped,
|
||||
bool* limit_hit) {
|
||||
return manifest_delete_missing_args_limited_observed(config, manifest, max_delete, deleted,
|
||||
skipped, limit_hit, NULL, NULL);
|
||||
}
|
||||
|
||||
bool manifest_delete_missing_args_limited_observed(
|
||||
const Config* config, const DeleteManifest* manifest, size_t max_delete, size_t* deleted,
|
||||
size_t* skipped, bool* limit_hit, DeletePathObserver observer, void* observer_context) {
|
||||
DeleteBudgetState budget = {
|
||||
.max_delete = max_delete, .deleted = 0, .skipped = 0, .limit_hit = false};
|
||||
bool ok =
|
||||
delete_missing_args_budgeted_observed(config, manifest, &budget, observer, observer_context);
|
||||
if (deleted)
|
||||
*deleted = budget.deleted;
|
||||
if (skipped)
|
||||
*skipped = budget.skipped;
|
||||
if (limit_hit)
|
||||
*limit_hit = budget.limit_hit;
|
||||
return ok;
|
||||
}
|
||||
|
||||
/* Commit every deletion family the manifest carries. The --delete-missing-args
|
||||
exact-path deletions run FIRST: they are explicit user requests and must not
|
||||
be blocked by the extras walker's filter-exclusion protection (a protected
|
||||
leftover inside a missing-argument directory must not make that user-requested
|
||||
removal fail). The ordinary extras walk then runs when --delete is active.
|
||||
Both draw from one --max-delete budget; the result reports a cap-stopped
|
||||
(partial) commit distinctly so the client can exit 25 like rsync. */
|
||||
DeleteCommitResult manifest_delete_all(const Config* config, const DeleteManifest* manifest) {
|
||||
return manifest_delete_all_counted(config, manifest, NULL);
|
||||
}
|
||||
|
||||
DeleteCommitResult manifest_delete_all_counted(const Config* config, const DeleteManifest* manifest,
|
||||
size_t* deleted) {
|
||||
return manifest_delete_all_observed(config, manifest, deleted, NULL, NULL);
|
||||
}
|
||||
|
||||
DeleteCommitResult manifest_delete_all_observed(const Config* config,
|
||||
const DeleteManifest* manifest, size_t* deleted,
|
||||
DeletePathObserver observer,
|
||||
void* observer_context) {
|
||||
if (deleted)
|
||||
*deleted = 0;
|
||||
if (!config || !manifest)
|
||||
return DELETE_COMMIT_ERROR;
|
||||
/* Central no-mutation guard: a dry-run never deletes. No manifest is sent on
|
||||
the dry-run path, but a hostile/buggy peer could; treat it as a no-op so
|
||||
the receiver can never remove anything. */
|
||||
if (config->dry_run)
|
||||
return DELETE_COMMIT_OK;
|
||||
/* A client --max-delete=NUM smaller than the server's hard bound replaces it
|
||||
for this run; both still bound the commit. */
|
||||
bool user_limited =
|
||||
config->max_delete >= 0 && (size_t)config->max_delete < MAX_SERVER_DELETE_COUNT;
|
||||
DeleteBudgetState budget = {.max_delete = user_limited ? (size_t)config->max_delete
|
||||
: MAX_SERVER_DELETE_COUNT,
|
||||
.deleted = 0,
|
||||
.skipped = 0,
|
||||
.limit_hit = false};
|
||||
if (config->delete_missing_args &&
|
||||
!delete_missing_args_budgeted_observed(config, manifest, &budget, observer, observer_context))
|
||||
return DELETE_COMMIT_ERROR;
|
||||
if (config->use_delete &&
|
||||
!delete_extras_budgeted_observed(config, manifest, &budget, observer, observer_context))
|
||||
return DELETE_COMMIT_ERROR;
|
||||
if (deleted)
|
||||
*deleted = budget.deleted;
|
||||
if (budget.limit_hit) {
|
||||
if (user_limited) {
|
||||
log_message(LOG_LEVEL_ERROR, "Deletions stopped due to --max-delete limit (%zu skipped)",
|
||||
budget.skipped);
|
||||
} else {
|
||||
log_message(LOG_LEVEL_ERROR,
|
||||
"Deletions stopped due to the server deletion limit of %u (%zu skipped)",
|
||||
(unsigned)MAX_SERVER_DELETE_COUNT, budget.skipped);
|
||||
}
|
||||
return DELETE_COMMIT_LIMIT_REACHED;
|
||||
}
|
||||
return DELETE_COMMIT_OK;
|
||||
}
|
||||
@@ -0,0 +1,105 @@
|
||||
#ifndef DELETE_COMMIT_H
|
||||
#define DELETE_COMMIT_H
|
||||
|
||||
#include "array_list.h"
|
||||
#include "config.h"
|
||||
#include "delete.h"
|
||||
#include <stdbool.h>
|
||||
|
||||
/* Delete-commit module: delete-manifest receive plus the budgeted extras and
|
||||
* --delete-missing-args walkers. These declarations are re-exported by the
|
||||
* file_receive.h facade. */
|
||||
|
||||
/* A received delete-manifest frame: the keep-set (`keeps`, destination-relative
|
||||
paths the sender transferred/keeps) plus `protected`, destination-relative
|
||||
prefixes the sender asks the receiver never to delete (paths excluded on the
|
||||
source, protected at any depth). When --delete-excluded is given the sender
|
||||
transmits an empty protected list so excluded destination mirrors are treated
|
||||
as ordinary extras. With --delete-missing-args a third section (`missing`)
|
||||
carries the destination mirrors of explicitly-listed source entries that do
|
||||
not exist: each is an exact deletion request, independent of the ordinary
|
||||
extras walk (never blocked by the protected prefixes) and processed when the
|
||||
manifest is committed. */
|
||||
typedef struct DeleteManifest {
|
||||
ArrayList* keeps;
|
||||
ArrayList* protected;
|
||||
ArrayList* missing;
|
||||
/* Destination-relative paths of the directories the sender synchronized for
|
||||
this run. The extras walker only removes entries directly inside one of
|
||||
these (the receive root is the "." sentinel); `--files-from` runs therefore
|
||||
leave untransmitted directories and the unlisted parts of listed ones
|
||||
alone, matching rsync's "delete only in synchronized directories". */
|
||||
ArrayList* dirs;
|
||||
} DeleteManifest;
|
||||
|
||||
void delete_manifest_free(DeleteManifest* manifest);
|
||||
/* Read a delete-manifest frame (protocol 2.23.0): keep count + keeps, then
|
||||
protected count + protected prefixes, then missing count + missing paths,
|
||||
then synchronized-directory count + directory paths (self-delimiting; the
|
||||
leading STATUS_MANIFEST code has been consumed). Returns an owned
|
||||
DeleteManifest, or NULL after signalling STATUS_ERROR on a malformed frame. */
|
||||
DeleteManifest* receive_manifest_entries(int fd);
|
||||
/* Remove destination entries under config->receive_root_directory that are not
|
||||
in `manifest` (bounded, all-or-nothing walk; staging-dir, basis-dir and
|
||||
protected-prefix skips). `--max-delete` and `--force` are honored here. The
|
||||
caller decides WHEN to run it based on the negotiated delete timing. Returns
|
||||
false (and the transfer fails) when the deletion cannot be committed. */
|
||||
bool manifest_delete_extras(const Config* config, const DeleteManifest* manifest);
|
||||
/* --delete-missing-args exact-path deletions: remove each destination mirror
|
||||
in `manifest->missing` (never blocked by the protected prefixes, staging dir
|
||||
and basis dirs excluded). A regular file/symlink is unlinked; an empty
|
||||
directory is removed; a NON-empty directory is removed recursively only when
|
||||
--delete or --force is in effect, otherwise it is left with a warning (rsync
|
||||
parity). A missing path is a no-op. Returns false only on a genuine
|
||||
confinement or I/O error (the run then fails); tolerated per-path cases are
|
||||
reported and skipped. */
|
||||
bool manifest_delete_missing_args(const Config* config, const DeleteManifest* manifest);
|
||||
/* Budgeted form of manifest_delete_missing_args for the per-directory delete
|
||||
session: each removed mirror draws from `max_delete` (SIZE_MAX = unlimited)
|
||||
and the tallies are accumulated into `*deleted`/`*skipped`. `*limit_hit` is set
|
||||
when the budget stopped the pass with entries left over. Returns false only
|
||||
on a genuine deletion error. */
|
||||
bool manifest_delete_missing_args_limited(const Config* config, const DeleteManifest* manifest,
|
||||
size_t max_delete, size_t* deleted, size_t* skipped,
|
||||
bool* limit_hit);
|
||||
/* Observer-aware form of manifest_delete_missing_args_limited: `observer` (may
|
||||
be NULL) is invoked for every destination-relative path truly removed. */
|
||||
bool manifest_delete_missing_args_limited_observed(
|
||||
const Config* config, const DeleteManifest* manifest, size_t max_delete, size_t* deleted,
|
||||
size_t* skipped, bool* limit_hit, DeletePathObserver observer, void* observer_context);
|
||||
/* Outcome of committing a delete manifest. LIMIT_REACHED reports rsync's
|
||||
partial --max-delete result: the budget allowed some deletions and the rest
|
||||
were skipped (the run still stores all file data but the client exits 25). */
|
||||
typedef enum {
|
||||
DELETE_COMMIT_OK = 0,
|
||||
DELETE_COMMIT_LIMIT_REACHED,
|
||||
DELETE_COMMIT_ERROR
|
||||
} DeleteCommitResult;
|
||||
|
||||
/* Run every deletion family the manifest carries: the --delete-missing-args
|
||||
exact-path deletions first (user requests are not blocked by exclusion
|
||||
protection), then the ordinary extras walk when --delete is active. Both
|
||||
share one --max-delete budget. Returns DELETE_COMMIT_OK when nothing was to
|
||||
do or everything committed, DELETE_COMMIT_LIMIT_REACHED when the budget
|
||||
stopped part of the work, or DELETE_COMMIT_ERROR on a genuine failure. */
|
||||
DeleteCommitResult manifest_delete_all(const Config* config, const DeleteManifest* manifest);
|
||||
/* Like manifest_delete_all, but reports how many destination entries the commit
|
||||
removed (for the end-of-transfer wire stats). `deleted` may be NULL. */
|
||||
DeleteCommitResult manifest_delete_all_counted(const Config* config, const DeleteManifest* manifest,
|
||||
size_t* deleted);
|
||||
/* Observer-aware form of manifest_delete_all_counted: `observer` (may be NULL)
|
||||
is invoked for every destination-relative path truly removed. */
|
||||
DeleteCommitResult manifest_delete_all_observed(const Config* config,
|
||||
const DeleteManifest* manifest, size_t* deleted,
|
||||
DeletePathObserver observer,
|
||||
void* observer_context);
|
||||
|
||||
/* -n/--dry-run --delete would-delete reporting: walk the destination exactly as
|
||||
the delete pass would and append (strdup'd) destination-relative paths that
|
||||
WOULD be removed to `out`, without touching disk. Uses the same staging-dir,
|
||||
basis-dir and protected-prefix skips as the real commit. Returns true on a
|
||||
clean walk; `*count_out` receives the number of paths appended. */
|
||||
bool manifest_would_delete_list(const Config* config, const DeleteManifest* manifest,
|
||||
ArrayList* out, size_t* count_out);
|
||||
|
||||
#endif
|
||||
+289
-108
@@ -2,6 +2,7 @@
|
||||
|
||||
#include "charset.h"
|
||||
#include "delay_updates.h"
|
||||
#include "delete.h"
|
||||
#include "file.h"
|
||||
#include "log.h"
|
||||
#include "utils.h"
|
||||
@@ -331,6 +332,8 @@ static int send_plan_node(int fd, DeletePlanSender* sender, PlanNode* node) {
|
||||
return -1;
|
||||
sender->config_sent = true;
|
||||
}
|
||||
if (!send_int(fd, 1)) /* apply = true */
|
||||
return -1;
|
||||
if (!send_wire_str(fd, node->dir))
|
||||
return -1;
|
||||
if (send_str_section(fd, node->dirs) != 0 || send_str_section(fd, node->files) != 0)
|
||||
@@ -339,6 +342,29 @@ static int send_plan_node(int fd, DeletePlanSender* sender, PlanNode* node) {
|
||||
return 0;
|
||||
}
|
||||
|
||||
/* Transmit the one-shot per-run config block (protected prefixes, size-pruned
|
||||
* mirrors, --delete-missing-args exact paths) on its own carrier frame, with
|
||||
* apply=false so the receiver consumes the config but walks nothing. This is
|
||||
* how the config still reaches the receiver when the scope allows no directory
|
||||
* plan at all (a --files-from list of bare files synchronizes no directory):
|
||||
* without it, the missing-args exact deletions would be lost. Idempotent. */
|
||||
static int send_config_only(int fd, DeletePlanSender* sender) {
|
||||
if (!sender || sender->config_sent)
|
||||
return 0;
|
||||
if (!send_status(fd, STATUS_DELETE_PLAN) || !send_int(fd, 1))
|
||||
return -1;
|
||||
if (send_str_section(fd, sender->protected_prefixes) != 0 ||
|
||||
send_str_section(fd, sender->size_skipped) != 0 ||
|
||||
send_str_section(fd, sender->missing_args) != 0)
|
||||
return -1;
|
||||
sender->config_sent = true;
|
||||
if (!send_int(fd, 0)) /* apply = false */
|
||||
return -1;
|
||||
if (!send_wire_str(fd, ".") || !send_int(fd, 0) || !send_int(fd, 0))
|
||||
return -1;
|
||||
return 0;
|
||||
}
|
||||
|
||||
static int send_prefix_plan(int fd, DeletePlanSender* sender, const char* dir) {
|
||||
PlanNode* node = plan_find(sender, dir);
|
||||
if (!node || node->sent)
|
||||
@@ -354,6 +380,10 @@ int delete_plan_send_root(int fd, DeletePlanSender* sender) {
|
||||
const char* root = sender->walk_root ? sender->walk_root : ".";
|
||||
if (!plan_ensure(sender, root))
|
||||
return -1;
|
||||
/* Put the config block on the wire first, on its own carrier frame, so the
|
||||
receiver always sees it even when the scope permits no directory plan. */
|
||||
if (send_config_only(fd, sender) != 0)
|
||||
return -1;
|
||||
return send_prefix_plan(fd, sender, root);
|
||||
}
|
||||
|
||||
@@ -403,6 +433,17 @@ int delete_plan_send_remaining(int fd, DeletePlanSender* sender, const ArrayList
|
||||
return 0;
|
||||
}
|
||||
|
||||
int delete_plan_send_all(int fd, DeletePlanSender* sender, const ArrayList* dirs) {
|
||||
if (!sender)
|
||||
return -1;
|
||||
/* Root first: this also transmits the one-shot per-run config block on its
|
||||
own carrier frame (see send_config_only), so it reaches the receiver even
|
||||
when the scope permits no directory plan at all. */
|
||||
if (delete_plan_send_root(fd, sender) != 0)
|
||||
return -1;
|
||||
return delete_plan_send_remaining(fd, sender, dirs);
|
||||
}
|
||||
|
||||
/* ------------------------------------------------------------------ */
|
||||
/* Receiver: delete session */
|
||||
/* ------------------------------------------------------------------ */
|
||||
@@ -412,6 +453,17 @@ struct DeletePlanSession {
|
||||
bool dry_run;
|
||||
size_t max_delete;
|
||||
size_t deleted;
|
||||
/* Removals charged against --max-delete. The budget is charged on ACTUAL
|
||||
removals (an unlink/rmdir that succeeded), matching rsync: a snapshotted
|
||||
entry that fails removal consumes nothing, so a later extra is still
|
||||
deleted. `planned` and `deleted` advance together for the inline paths and
|
||||
`apply_missing`; `deleted` is the reported count. */
|
||||
size_t planned;
|
||||
/* Hard bound on the deferred snapshot list. Because the budget is no longer
|
||||
charged at snapshot time, this independent cap keeps a huge destination
|
||||
from growing the list without limit (it matches the receiver's overall
|
||||
deletion bound). */
|
||||
size_t defer_cap;
|
||||
size_t skipped;
|
||||
bool limit_hit;
|
||||
bool limit_logged;
|
||||
@@ -421,8 +473,34 @@ struct DeletePlanSession {
|
||||
ArrayList* size_skipped;
|
||||
ArrayList* missing;
|
||||
ArrayList* deferred;
|
||||
DeletePathObserver observer;
|
||||
void* observer_context;
|
||||
};
|
||||
|
||||
/* Report one path the session truly removed (no-op without an observer). */
|
||||
static void notify_deleted(DeletePlanSession* session, const char* rel) {
|
||||
if (session && session->observer && rel)
|
||||
session->observer(session->observer_context, rel);
|
||||
}
|
||||
|
||||
/* A removed directory is reported with rsync's trailing slash (`deleting dir/`)
|
||||
while files keep their bare path. */
|
||||
static void notify_deleted_dir(DeletePlanSession* session, const char* rel) {
|
||||
if (!session || !session->observer || !rel)
|
||||
return;
|
||||
size_t len = strlen(rel);
|
||||
char* with_slash = malloc(len + 2);
|
||||
if (!with_slash) {
|
||||
session->observer(session->observer_context, rel);
|
||||
return;
|
||||
}
|
||||
memcpy(with_slash, rel, len);
|
||||
with_slash[len] = '/';
|
||||
with_slash[len + 1] = '\0';
|
||||
session->observer(session->observer_context, with_slash);
|
||||
free(with_slash);
|
||||
}
|
||||
|
||||
DeletePlanSession* delete_plan_session_create(const Config* config) {
|
||||
if (!config)
|
||||
return NULL;
|
||||
@@ -435,6 +513,7 @@ DeletePlanSession* delete_plan_session_create(const Config* config) {
|
||||
config->max_delete >= 0 && (size_t)config->max_delete < DELETE_PLAN_SERVER_LIMIT;
|
||||
session->max_delete =
|
||||
user_limited ? (size_t)config->max_delete : (size_t)DELETE_PLAN_SERVER_LIMIT;
|
||||
session->defer_cap = DELETE_PLAN_SERVER_LIMIT;
|
||||
session->protected_prefixes = array_list_create(free);
|
||||
session->size_skipped = array_list_create(free);
|
||||
session->missing = array_list_create(free);
|
||||
@@ -523,49 +602,27 @@ static int open_plan_dir(const Config* config, const char* dir) {
|
||||
return fd;
|
||||
}
|
||||
|
||||
typedef struct PlanSkips {
|
||||
DeleteSkipEntry* entries;
|
||||
int count;
|
||||
typedef struct {
|
||||
DeleteSkipSet set;
|
||||
/* Receiver-side delete-protection rules received on the config frame (NULL
|
||||
when the sender sent none). Evaluated per extra so a protect/risk rule is
|
||||
honored under --delete-during/--delete-delay exactly like the whole-tree
|
||||
commit walker. */
|
||||
const FilterRuleList* protect_rules;
|
||||
} PlanSkips;
|
||||
|
||||
static bool build_plan_skips(const Config* config, const DeletePlanSession* session,
|
||||
PlanSkips* out) {
|
||||
out->entries = NULL;
|
||||
out->count = 0;
|
||||
int count = (config->delay_updates ? 1 : 0) + config->basis_count +
|
||||
session->protected_prefixes->size + session->size_skipped->size;
|
||||
if (count == 0)
|
||||
return true;
|
||||
out->entries = calloc((size_t)count, sizeof(DeleteSkipEntry));
|
||||
if (!out->entries)
|
||||
return false;
|
||||
int idx = 0;
|
||||
if (config->delay_updates) {
|
||||
out->entries[idx].prefix = DELAY_UPDATES_STAGING_DIR;
|
||||
out->entries[idx].top_level_only = true;
|
||||
idx++;
|
||||
}
|
||||
for (int i = 0; i < config->basis_count; i++) {
|
||||
out->entries[idx].prefix = config->basis_dirs[i].path;
|
||||
out->entries[idx].top_level_only = false;
|
||||
idx++;
|
||||
}
|
||||
for (int i = 0; i < session->protected_prefixes->size; i++) {
|
||||
out->entries[idx].prefix = (const char*)session->protected_prefixes->items[i];
|
||||
out->entries[idx].top_level_only = false;
|
||||
idx++;
|
||||
}
|
||||
for (int i = 0; i < session->size_skipped->size; i++) {
|
||||
out->entries[idx].prefix = (const char*)session->size_skipped->items[i];
|
||||
out->entries[idx].top_level_only = false;
|
||||
idx++;
|
||||
}
|
||||
out->count = idx;
|
||||
return true;
|
||||
out->protect_rules = config->protect_rules;
|
||||
/* The per-directory plan walk keeps each basis path verbatim (it does not
|
||||
convert an absolute under-root path to its root-relative form, unlike the
|
||||
whole-tree commit walk). */
|
||||
return delete_skips_build(config, session->protected_prefixes, session->size_skipped, false,
|
||||
&out->set);
|
||||
}
|
||||
|
||||
static bool budget_available(const DeletePlanSession* session) {
|
||||
return session->deleted < session->max_delete;
|
||||
return session->planned < session->max_delete;
|
||||
}
|
||||
|
||||
static void note_skipped(DeletePlanSession* session) {
|
||||
@@ -579,8 +636,15 @@ static void log_deleted(const char* rel) {
|
||||
free(escaped);
|
||||
}
|
||||
|
||||
/* Append a snapshot path for --delete-delay. */
|
||||
/* Append a snapshot path for --delete-delay. The budget is NOT charged here:
|
||||
* the remover charges --max-delete only when a path is actually unlinked (see
|
||||
* apply_deferred_path), so a snapshotted entry that survives ENOTEMPTY cannot
|
||||
* deny budget to a later extra. The independent `defer_cap` bounds the list. */
|
||||
static bool defer_add(DeletePlanSession* session, const char* rel) {
|
||||
if ((size_t)session->deferred->size >= session->defer_cap) {
|
||||
note_skipped(session);
|
||||
return true;
|
||||
}
|
||||
char* copy = str_dup(rel);
|
||||
if (!copy)
|
||||
return false;
|
||||
@@ -588,7 +652,6 @@ static bool defer_add(DeletePlanSession* session, const char* rel) {
|
||||
free(copy);
|
||||
return false;
|
||||
}
|
||||
session->deleted++;
|
||||
return true;
|
||||
}
|
||||
|
||||
@@ -621,19 +684,21 @@ static bool process_extra_dir(int dirfd, const char* name, const char* child_rel
|
||||
return false;
|
||||
if (survives)
|
||||
return true;
|
||||
if (!budget_available(session)) {
|
||||
note_skipped(session);
|
||||
return true;
|
||||
}
|
||||
if (session->defer && !force_now) {
|
||||
if (!defer_add(session, child_rel))
|
||||
return false;
|
||||
*removed = true;
|
||||
return true;
|
||||
}
|
||||
if (!budget_available(session)) {
|
||||
note_skipped(session);
|
||||
return true;
|
||||
}
|
||||
if (unlinkat(dirfd, name, AT_REMOVEDIR) == 0) {
|
||||
session->deleted++;
|
||||
session->planned++;
|
||||
log_deleted(child_rel);
|
||||
notify_deleted_dir(session, child_rel);
|
||||
*removed = true;
|
||||
return true;
|
||||
}
|
||||
@@ -648,16 +713,18 @@ static bool process_extra_dir(int dirfd, const char* name, const char* child_rel
|
||||
|
||||
static bool process_extra_file(int dirfd, const char* name, const char* child_rel, bool force_now,
|
||||
DeletePlanSession* session) {
|
||||
if (session->defer && !force_now) {
|
||||
return defer_add(session, child_rel);
|
||||
}
|
||||
if (!budget_available(session)) {
|
||||
note_skipped(session);
|
||||
return true;
|
||||
}
|
||||
if (session->defer && !force_now) {
|
||||
return defer_add(session, child_rel);
|
||||
}
|
||||
if (unlinkat(dirfd, name, 0) == 0) {
|
||||
session->deleted++;
|
||||
session->planned++;
|
||||
log_deleted(child_rel);
|
||||
notify_deleted(session, child_rel);
|
||||
} else if (errno != ENOENT) {
|
||||
return false;
|
||||
}
|
||||
@@ -668,70 +735,108 @@ static bool process_children(int dirfd, const char* dir_rel, const ArrayList* ke
|
||||
const ArrayList* keep_files, bool at_root, bool force_now,
|
||||
const PlanSkips* skips, DeletePlanSession* session, bool* survives) {
|
||||
*survives = false;
|
||||
int scanfd = openat(dirfd, ".", O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
|
||||
if (scanfd < 0)
|
||||
DeleteDirEntry* entries = NULL;
|
||||
size_t count = 0;
|
||||
bool collect_ok = true;
|
||||
if (!delete_dir_entries_collect(dirfd, &entries, &count, &collect_ok))
|
||||
return false;
|
||||
DIR* dir = fdopendir(scanfd);
|
||||
if (!dir) {
|
||||
close(scanfd);
|
||||
bool operation_ok = collect_ok;
|
||||
bool local_survives = false;
|
||||
bool* shielded = calloc(count ? count : 1, sizeof(bool));
|
||||
bool* is_extra = calloc(count ? count : 1, sizeof(bool));
|
||||
bool* force = calloc(count ? count : 1, sizeof(bool));
|
||||
if (!shielded || !is_extra || !force) {
|
||||
free(shielded);
|
||||
free(is_extra);
|
||||
free(force);
|
||||
delete_dir_entries_free(entries, count);
|
||||
return false;
|
||||
}
|
||||
bool operation_ok = true;
|
||||
bool local_survives = false;
|
||||
const struct dirent* entry;
|
||||
while ((entry = readdir(dir)) != NULL) {
|
||||
if (strcmp(entry->d_name, ".") == 0 || strcmp(entry->d_name, "..") == 0)
|
||||
continue;
|
||||
|
||||
/* rsync's order: extraneous subdirectories in descending name order, then
|
||||
extraneous files in descending name order (kept entries survive and are not
|
||||
touched here — a kept subdirectory gets its own per-directory plan). */
|
||||
if (count > 1)
|
||||
qsort(entries, count, sizeof(*entries), delete_dir_entry_cmp_desc);
|
||||
size_t dir_count = 0;
|
||||
while (dir_count < count && entries[dir_count].is_dir)
|
||||
dir_count++;
|
||||
|
||||
for (size_t i = 0; i < count; i++) {
|
||||
char* child_rel =
|
||||
(strcmp(dir_rel, ".") == 0) ? str_dup(entry->d_name) : path_cat(dir_rel, entry->d_name);
|
||||
(strcmp(dir_rel, ".") == 0) ? str_dup(entries[i].name) : path_cat(dir_rel, entries[i].name);
|
||||
if (!child_rel) {
|
||||
operation_ok = false;
|
||||
continue;
|
||||
}
|
||||
if (path_under_skip_prefix(child_rel, at_root, skips->entries, skips->count)) {
|
||||
if (path_under_skip_prefix(child_rel, at_root, skips->set.entries, skips->set.count)) {
|
||||
shielded[i] = true;
|
||||
local_survives = true;
|
||||
free(child_rel);
|
||||
continue;
|
||||
}
|
||||
struct stat st;
|
||||
if (fstatat(dirfd, entry->d_name, &st, AT_SYMLINK_NOFOLLOW) != 0) {
|
||||
if (errno != ENOENT)
|
||||
operation_ok = false;
|
||||
free(child_rel);
|
||||
continue;
|
||||
}
|
||||
bool is_dir = S_ISDIR(st.st_mode);
|
||||
bool in_keep_dirs = is_dir && list_contains_str(keep_dirs, entry->d_name);
|
||||
bool in_keep_files = !is_dir && list_contains_str(keep_files, entry->d_name);
|
||||
if (in_keep_dirs) {
|
||||
bool is_dir = entries[i].is_dir;
|
||||
bool in_keep_dirs = is_dir && list_contains_str(keep_dirs, entries[i].name);
|
||||
bool in_keep_files = !is_dir && list_contains_str(keep_files, entries[i].name);
|
||||
bool rule_protected =
|
||||
skips->protect_rules &&
|
||||
filter_rules_apply_side(skips->protect_rules, child_rel, entries[i].name, is_dir,
|
||||
FILTER_SIDE_RECEIVER) == FILTER_ACTION_PROTECT;
|
||||
if (in_keep_dirs || in_keep_files || rule_protected) {
|
||||
shielded[i] = true;
|
||||
local_survives = true;
|
||||
} else if (keep_dirs && !is_dir && list_contains_str(keep_dirs, entry->d_name)) {
|
||||
/* Destination file blocks a source directory: clear it now, whatever the
|
||||
delete timing, so the directory can be created. */
|
||||
if (!process_extra_file(dirfd, entry->d_name, child_rel, true, session))
|
||||
operation_ok = false;
|
||||
} else if (in_keep_files) {
|
||||
local_survives = true;
|
||||
} else if (keep_files && is_dir && list_contains_str(keep_files, entry->d_name)) {
|
||||
/* Destination directory blocks a source file: remove it now. */
|
||||
bool removed = false;
|
||||
if (!process_extra_dir(dirfd, entry->d_name, child_rel, true, skips, session, &removed))
|
||||
operation_ok = false;
|
||||
else if (!removed)
|
||||
local_survives = true;
|
||||
} else if (is_dir) {
|
||||
bool removed = false;
|
||||
if (!process_extra_dir(dirfd, entry->d_name, child_rel, force_now, skips, session, &removed))
|
||||
operation_ok = false;
|
||||
else if (!removed)
|
||||
local_survives = true;
|
||||
/* A destination directory blocks a source file of the same name: remove
|
||||
it now, whatever the delete timing, so the file can be created. */
|
||||
is_extra[i] = true;
|
||||
force[i] = keep_files && list_contains_str(keep_files, entries[i].name);
|
||||
} else {
|
||||
if (!process_extra_file(dirfd, entry->d_name, child_rel, force_now, session))
|
||||
operation_ok = false;
|
||||
/* A destination file blocks a source directory of the same name: clear it
|
||||
now so the directory can be created. */
|
||||
is_extra[i] = true;
|
||||
force[i] = keep_dirs && list_contains_str(keep_dirs, entries[i].name);
|
||||
}
|
||||
free(child_rel);
|
||||
}
|
||||
closedir(dir);
|
||||
|
||||
/* Pass 1: extraneous subdirectories, descending. */
|
||||
for (size_t i = 0; i < dir_count; i++) {
|
||||
if (!is_extra[i])
|
||||
continue;
|
||||
char* child_rel =
|
||||
(strcmp(dir_rel, ".") == 0) ? str_dup(entries[i].name) : path_cat(dir_rel, entries[i].name);
|
||||
if (!child_rel) {
|
||||
operation_ok = false;
|
||||
continue;
|
||||
}
|
||||
bool removed = false;
|
||||
if (!process_extra_dir(dirfd, entries[i].name, child_rel, force[i] || force_now, skips, session,
|
||||
&removed))
|
||||
operation_ok = false;
|
||||
else if (!removed)
|
||||
local_survives = true;
|
||||
free(child_rel);
|
||||
}
|
||||
|
||||
/* Pass 2: extraneous files, descending. */
|
||||
for (size_t i = dir_count; i < count; i++) {
|
||||
if (!is_extra[i])
|
||||
continue;
|
||||
char* child_rel =
|
||||
(strcmp(dir_rel, ".") == 0) ? str_dup(entries[i].name) : path_cat(dir_rel, entries[i].name);
|
||||
if (!child_rel) {
|
||||
operation_ok = false;
|
||||
continue;
|
||||
}
|
||||
if (!process_extra_file(dirfd, entries[i].name, child_rel, force[i] || force_now, session))
|
||||
operation_ok = false;
|
||||
free(child_rel);
|
||||
}
|
||||
|
||||
free(shielded);
|
||||
free(is_extra);
|
||||
free(force);
|
||||
delete_dir_entries_free(entries, count);
|
||||
*survives = local_survives;
|
||||
return operation_ok;
|
||||
}
|
||||
@@ -751,7 +856,7 @@ static bool apply_plan_dir(DeletePlanSession* session, const Config* config, con
|
||||
bool survives = false;
|
||||
bool ok = process_children(dirfd, dir, dirs, files, strcmp(dir, ".") == 0, false, &skips, session,
|
||||
&survives);
|
||||
free(skips.entries);
|
||||
delete_skips_free(&skips.set);
|
||||
close(dirfd);
|
||||
if (!ok)
|
||||
log_message(LOG_LEVEL_ERROR, "deletion failed while removing extraneous files");
|
||||
@@ -768,13 +873,15 @@ static bool apply_missing(DeletePlanSession* session, const Config* config) {
|
||||
return true;
|
||||
DeleteManifest manifest = {
|
||||
.keeps = NULL, .protected = NULL, .missing = session->missing, .dirs = NULL};
|
||||
size_t remaining = budget_available(session) ? session->max_delete - session->deleted : 0;
|
||||
size_t remaining = budget_available(session) ? session->max_delete - session->planned : 0;
|
||||
size_t deleted = 0;
|
||||
size_t skipped = 0;
|
||||
bool limit = false;
|
||||
bool ok = manifest_delete_missing_args_limited(config, &manifest, remaining, &deleted, &skipped,
|
||||
&limit);
|
||||
bool ok = manifest_delete_missing_args_limited_observed(config, &manifest, remaining, &deleted,
|
||||
&skipped, &limit, session->observer,
|
||||
session->observer_context);
|
||||
session->deleted += deleted;
|
||||
session->planned += deleted;
|
||||
session->skipped += skipped;
|
||||
if (limit)
|
||||
session->limit_hit = true;
|
||||
@@ -801,6 +908,14 @@ int delete_plan_session_receive(DeletePlanSession* session, const Config* config
|
||||
}
|
||||
session->config_seen = true;
|
||||
}
|
||||
/* apply=false is the config-only carrier frame: the receiver consumes the
|
||||
config (and the missing-args exact deletions) but must not walk any
|
||||
directory. Every real plan carries apply=true. */
|
||||
int apply;
|
||||
if (!receive_int(fd, &apply) || (apply != 0 && apply != 1)) {
|
||||
send_status(fd, STATUS_ERROR);
|
||||
return -1;
|
||||
}
|
||||
char* dir = receive_wire_str(fd);
|
||||
ArrayList* dirs = array_list_create(free);
|
||||
ArrayList* files = array_list_create(free);
|
||||
@@ -818,7 +933,7 @@ int delete_plan_session_receive(DeletePlanSession* session, const Config* config
|
||||
if (!session->dry_run && enabled) {
|
||||
if (!session->defer && !apply_missing(session, config))
|
||||
ok = false;
|
||||
if (ok && !apply_plan_dir(session, config, dir, dirs, files))
|
||||
if (ok && apply && !apply_plan_dir(session, config, dir, dirs, files))
|
||||
ok = false;
|
||||
}
|
||||
free(dir);
|
||||
@@ -837,37 +952,103 @@ int delete_plan_session_receive(DeletePlanSession* session, const Config* config
|
||||
}
|
||||
|
||||
/* Apply one snapshotted --delete-delay path (post-order: children precede their
|
||||
* parent directory). */
|
||||
* parent directory). A directory that is still present is re-scanned so content
|
||||
* created after the plan is removed too; every actual removal charges
|
||||
* --max-delete. */
|
||||
static bool apply_deferred_path(DeletePlanSession* session, const Config* config, const char* rel) {
|
||||
(void)session;
|
||||
char* full = path_cat(config->receive_root_directory, rel);
|
||||
if (!full)
|
||||
return false;
|
||||
char* leaf = NULL;
|
||||
int parent_fd = file_open_secure_parent(full, &leaf, false);
|
||||
int open_errno = errno;
|
||||
free(full);
|
||||
if (parent_fd < 0) {
|
||||
free(leaf);
|
||||
return errno == ENOENT || errno == ENOTDIR;
|
||||
return open_errno == ENOENT || open_errno == ENOTDIR;
|
||||
}
|
||||
struct stat st;
|
||||
if (fstatat(parent_fd, leaf, &st, AT_SYMLINK_NOFOLLOW) != 0) {
|
||||
bool absent = errno == ENOENT;
|
||||
bool absent = errno == ENOENT || errno == ENOTDIR;
|
||||
close(parent_fd);
|
||||
free(leaf);
|
||||
return absent;
|
||||
}
|
||||
int rc;
|
||||
if (S_ISDIR(st.st_mode))
|
||||
rc = unlinkat(parent_fd, leaf, AT_REMOVEDIR);
|
||||
else
|
||||
rc = unlinkat(parent_fd, leaf, 0);
|
||||
bool ok = rc == 0 || errno == ENOENT || errno == ENOTEMPTY || errno == EEXIST;
|
||||
if (rc == 0)
|
||||
if (S_ISDIR(st.st_mode)) {
|
||||
if (!budget_available(session)) {
|
||||
note_skipped(session);
|
||||
close(parent_fd);
|
||||
free(leaf);
|
||||
return true;
|
||||
}
|
||||
int dirfd = openat(parent_fd, leaf, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
|
||||
if (dirfd < 0) {
|
||||
bool absent = errno == ENOENT || errno == ENOTDIR;
|
||||
close(parent_fd);
|
||||
free(leaf);
|
||||
return absent;
|
||||
}
|
||||
PlanSkips skips;
|
||||
if (!build_plan_skips(config, session, &skips)) {
|
||||
close(dirfd);
|
||||
close(parent_fd);
|
||||
free(leaf);
|
||||
return false;
|
||||
}
|
||||
bool survives = false;
|
||||
bool ok = process_children(dirfd, rel, NULL, NULL, false, true, &skips, session, &survives);
|
||||
delete_skips_free(&skips.set);
|
||||
close(dirfd);
|
||||
if (!ok) {
|
||||
close(parent_fd);
|
||||
free(leaf);
|
||||
return false;
|
||||
}
|
||||
if (!survives) {
|
||||
if (!budget_available(session)) {
|
||||
note_skipped(session);
|
||||
} else if (unlinkat(parent_fd, leaf, AT_REMOVEDIR) == 0) {
|
||||
session->deleted++;
|
||||
session->planned++;
|
||||
log_deleted(rel);
|
||||
notify_deleted_dir(session, rel);
|
||||
} else if (errno != ENOENT && errno != ENOTEMPTY && errno != EEXIST) {
|
||||
close(parent_fd);
|
||||
free(leaf);
|
||||
return false;
|
||||
}
|
||||
}
|
||||
close(parent_fd);
|
||||
free(leaf);
|
||||
return true;
|
||||
}
|
||||
if (!budget_available(session)) {
|
||||
note_skipped(session);
|
||||
close(parent_fd);
|
||||
free(leaf);
|
||||
return true;
|
||||
}
|
||||
if (unlinkat(parent_fd, leaf, 0) == 0) {
|
||||
session->deleted++;
|
||||
session->planned++;
|
||||
log_deleted(rel);
|
||||
notify_deleted(session, rel);
|
||||
} else if (errno != ENOENT) {
|
||||
close(parent_fd);
|
||||
free(leaf);
|
||||
return false;
|
||||
}
|
||||
close(parent_fd);
|
||||
free(leaf);
|
||||
return ok;
|
||||
return true;
|
||||
}
|
||||
|
||||
void delete_plan_session_set_delete_observer(DeletePlanSession* session,
|
||||
DeletePathObserver observer, void* context) {
|
||||
if (!session)
|
||||
return;
|
||||
session->observer = observer;
|
||||
session->observer_context = context;
|
||||
}
|
||||
|
||||
DeleteCommitResult delete_plan_session_commit(DeletePlanSession* session, const Config* config) {
|
||||
|
||||
@@ -3,8 +3,10 @@
|
||||
|
||||
#include "array_list.h"
|
||||
#include "config.h"
|
||||
#include "delete.h"
|
||||
#include "file_receive.h"
|
||||
#include "protocol.h"
|
||||
#include "utils.h"
|
||||
#include <stdbool.h>
|
||||
|
||||
/* Per-directory delete plans (protocol 2.24.0).
|
||||
@@ -47,19 +49,32 @@ void delete_plan_sender_finalize(DeletePlanSender* sender, const ArrayList* sync
|
||||
Directory keep entries do not count, so an I/O error that hid every file
|
||||
still refuses to delete. */
|
||||
bool delete_plan_sender_empty(const DeletePlanSender* sender);
|
||||
/* Attach the global config sections advertised on the first plan frame. */
|
||||
/* Attach the global config sections advertised on the first plan frame. The
|
||||
* block is always transmitted by delete_plan_send_root(), on a config-only
|
||||
* carrier frame when the scope allows no directory plan. */
|
||||
void delete_plan_sender_set_config(DeletePlanSender* sender, const ArrayList* protected_prefixes,
|
||||
const ArrayList* size_skipped, const ArrayList* missing_args);
|
||||
/* Send the root plan (even before any data, so root extras are handled like
|
||||
* rsync's first generator directory). Returns -1 on I/O error. */
|
||||
* rsync's first generator directory), after transmitting the per-run config
|
||||
* block on its own carrier frame. Returns -1 on I/O error. */
|
||||
int delete_plan_send_root(int fd, DeletePlanSender* sender);
|
||||
/* Send the plans for every ancestor of `path` (root-first) and, when is_dir,
|
||||
* for `path` itself; already-sent plans are skipped. */
|
||||
int delete_plan_send_for_path(int fd, DeletePlanSender* sender, const char* path, bool is_dir);
|
||||
/* Send the plan for every directory in `dirs` that has not been transmitted
|
||||
* yet. Called after the data stream so an empty source directory's plan still
|
||||
* clears its destination extras even though no file frame triggered it. */
|
||||
* yet. */
|
||||
int delete_plan_send_remaining(int fd, DeletePlanSender* sender, const ArrayList* dirs);
|
||||
/* Transmit the COMPLETE per-directory plan set in one pass, before any data
|
||||
* frame: the root plan (with the one-shot per-run config block on its carrier
|
||||
* frame) followed by every directory in `dirs`. Because the whole plan set is
|
||||
* known from the path-only pre-scan, sending it all up front means a
|
||||
* mid-transfer abort has already applied every planned removal, matching
|
||||
* rsync's generator (which runs ahead of its throttled sender). A completed
|
||||
* run is unaffected. `dirs` is the set of directories whose direct children
|
||||
* were enumerated (the scanner's plan_dirs sink), so a merely listed but
|
||||
* untraversed directory never gets a plan and its mirror is left intact.
|
||||
* Returns -1 on I/O error. */
|
||||
int delete_plan_send_all(int fd, DeletePlanSender* sender, const ArrayList* dirs);
|
||||
|
||||
/* ---- Receiver: delete session ---- */
|
||||
|
||||
@@ -76,8 +91,16 @@ int delete_plan_session_receive(DeletePlanSession* session, const Config* config
|
||||
DeleteCommitResult delete_plan_session_commit(DeletePlanSession* session, const Config* config);
|
||||
/* True once the shared --max-delete budget stopped part of a deletion. */
|
||||
bool delete_plan_session_limit_reached(const DeletePlanSession* session);
|
||||
/* Number of destination entries the session's plans removed (or, for
|
||||
--delete-delay, snapshotted for removal), for the end-of-transfer stats. */
|
||||
/* Number of destination entries the session actually removed, for the
|
||||
end-of-transfer stats. For --delete-delay this excludes a snapshotted entry
|
||||
that survived (e.g. a refilled directory that failed ENOTEMPTY), even though
|
||||
that entry already consumed --max-delete budget at snapshot time. */
|
||||
size_t delete_plan_session_deleted(const DeletePlanSession* session);
|
||||
/* Install an observer invoked for every destination-relative path the session
|
||||
truly removes (including the deferred --delete-delay commit), so the receiver
|
||||
can report rsync's `deleting PATH` lines through the terminal STATUS_STATS
|
||||
record. Pass NULL/0 to clear. */
|
||||
void delete_plan_session_set_delete_observer(DeletePlanSession* session,
|
||||
DeletePathObserver observer, void* context);
|
||||
|
||||
#endif
|
||||
|
||||
+420
-31
@@ -15,6 +15,7 @@
|
||||
#include <unistd.h>
|
||||
|
||||
#include "data.h"
|
||||
#include "checksum.h"
|
||||
#include "delta.h"
|
||||
#include "file.h"
|
||||
#include "file_store.h"
|
||||
@@ -39,6 +40,31 @@ static bool write_all(int fd, const void* data, unsigned long long size) {
|
||||
return true;
|
||||
}
|
||||
|
||||
/* Streaming copy of an open source descriptor into the just-created destination
|
||||
`fd` (already at offset 0). Used by the --copy-dest basis install so a basis
|
||||
larger than any in-memory whole-file bound still materializes without
|
||||
buffering the entire file. `expected_size` is the caller-verified basis
|
||||
size; the copy must produce exactly that many bytes (a short source is a hard
|
||||
error, never a silently truncated destination). The final ftruncate drops
|
||||
any residual tail a raced-in longer source might have left. */
|
||||
static bool copy_fd_all(int dst_fd, int src_fd, unsigned long long expected_size) {
|
||||
unsigned char buf[1 << 20];
|
||||
unsigned long long done = 0;
|
||||
while (done < expected_size) {
|
||||
unsigned long long remaining = expected_size - done;
|
||||
size_t want = remaining < sizeof(buf) ? (size_t)remaining : sizeof(buf);
|
||||
ssize_t n = read(src_fd, buf, want);
|
||||
if (n < 0 && errno == EINTR)
|
||||
continue;
|
||||
if (n <= 0)
|
||||
return false;
|
||||
if (!write_all(dst_fd, buf, (unsigned long long)n))
|
||||
return false;
|
||||
done += (unsigned long long)n;
|
||||
}
|
||||
return ftruncate(dst_fd, (off_t)expected_size) == 0;
|
||||
}
|
||||
|
||||
/* Preallocate `size` bytes on `fd` before any data is written (--preallocate).
|
||||
* fallocate(2) reserves real disk blocks, so an out-of-space condition
|
||||
* (ENOSPC/EDQUOT) surfaces up front instead of partway through a transfer;
|
||||
@@ -140,6 +166,12 @@ bool file_checksum(File* file, ChecksumAlgo algo, uint64_t seed, uint8_t* out, s
|
||||
if (file->data->size == 0) {
|
||||
return checksum_digest(algo, seed, "", 0, out, out_capacity, out_len);
|
||||
}
|
||||
/* A streamed source (data not loaded) may exceed any in-memory whole-file
|
||||
bound; hash it from the file path in bounded buffers instead of forcing a
|
||||
full load. This is the same digest the receiver recomputes on the basis. */
|
||||
if (!file->data->data && file->path && file->data->size > STREAM_THRESHOLD &&
|
||||
checksum_digest_file(algo, seed, file->path, out, out_capacity, out_len))
|
||||
return true;
|
||||
if (!file->data->data && !file_load_data(file))
|
||||
return false;
|
||||
return checksum_digest(algo, seed, file->data->data, file->data->size, out, out_capacity,
|
||||
@@ -176,6 +208,7 @@ File* file_create(const char* path) {
|
||||
file->is_dir = false;
|
||||
file->dir_time_only = false;
|
||||
file->basis_link = NULL;
|
||||
file->basis_copy = NULL;
|
||||
file->link_group = 0;
|
||||
file->link_first = false;
|
||||
file->hardlink_target = NULL;
|
||||
@@ -187,6 +220,7 @@ File* file_create(const char* path) {
|
||||
file->xattrs = NULL;
|
||||
file->dest_state = (OutputDestState){0};
|
||||
file->matched_bytes = 0;
|
||||
file->literal_bytes = 0;
|
||||
return file;
|
||||
}
|
||||
|
||||
@@ -204,6 +238,8 @@ void file_destroy(void* item) {
|
||||
file->send_path = NULL;
|
||||
free(file->basis_link);
|
||||
file->basis_link = NULL;
|
||||
free(file->basis_copy);
|
||||
file->basis_copy = NULL;
|
||||
free(file->hardlink_target);
|
||||
file->hardlink_target = NULL;
|
||||
free(file->symlink_target);
|
||||
@@ -614,6 +650,103 @@ static int open_dir_beneath_root(const char* resolved, const char* root) {
|
||||
}
|
||||
|
||||
int file_open_secure_parent(const char* path, char** leaf_out, bool create_dirs) {
|
||||
return file_open_secure_parent_counted(path, leaf_out, create_dirs, NULL, NULL);
|
||||
}
|
||||
|
||||
/* The logical transfer root expressed in the same coordinate as the secure
|
||||
* parent walk's `rel_buf` (relative to the authorized root, with a leading
|
||||
* '/'), used as the floor at or below which a created directory is a real
|
||||
* file-list entry. The on-disk transfer root is the receive root joined to the
|
||||
* wire path; the mirror scaffolding above it (the absolute source path below
|
||||
* the destination root) is not an rsync entry. Returns an allocated string or
|
||||
* NULL (count every created component). */
|
||||
static char* transfer_root_floor(const Config* config) {
|
||||
if (!config || !config->send_directory || config->send_directory[0] == '\0')
|
||||
return NULL;
|
||||
const char* spec = config->send_directory;
|
||||
const char* after = spec;
|
||||
if (spec[0] == '.' && spec[1] == '/') {
|
||||
after = spec + 2;
|
||||
} else {
|
||||
const char* cut = strstr(spec, "/./");
|
||||
if (cut)
|
||||
after = cut + 3;
|
||||
}
|
||||
while (*after == '/')
|
||||
after++;
|
||||
char* wire_root = str_dup(after);
|
||||
if (!wire_root)
|
||||
return NULL;
|
||||
size_t wlen = strlen(wire_root);
|
||||
while (wlen > 0 && wire_root[wlen - 1] == '/')
|
||||
wire_root[--wlen] = '\0';
|
||||
if (wlen == 0) {
|
||||
free(wire_root);
|
||||
return NULL;
|
||||
}
|
||||
char* disk_root = config->receive_root_directory
|
||||
? path_cat(config->receive_root_directory, wire_root)
|
||||
: str_dup(wire_root);
|
||||
free(wire_root);
|
||||
if (!disk_root)
|
||||
return NULL;
|
||||
const char* root_path = utils_get_authorized_root_path();
|
||||
const char* floor = disk_root;
|
||||
if (root_path && root_path[0] == '/') {
|
||||
size_t rl = strlen(root_path);
|
||||
while (rl > 0 && root_path[rl - 1] == '/')
|
||||
rl--;
|
||||
if (strncmp(disk_root, root_path, rl) == 0 && (disk_root[rl] == '/' || disk_root[rl] == '\0'))
|
||||
floor = disk_root + rl;
|
||||
}
|
||||
while (*floor == '/')
|
||||
floor++;
|
||||
char* out = str_dup(floor);
|
||||
free(disk_root);
|
||||
if (!out)
|
||||
return NULL;
|
||||
if (out[0] == '\0') {
|
||||
free(out);
|
||||
return NULL;
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
/* A created parent component counts toward `Number of created files` only when
|
||||
* its receive-root-relative path is at or below the logical transfer root
|
||||
* (`count_floor`). The transfer root itself corresponds to rsync's `.` entry
|
||||
* (created on a fresh destination, pre-existing otherwise); the mirror
|
||||
* scaffolding above it is FastSync's absolute-path layout, not an rsync entry. */
|
||||
static bool created_dir_counts(const char* count_floor, const char* rel_buf,
|
||||
const char* component) {
|
||||
if (!count_floor)
|
||||
return true;
|
||||
char candidate[PATH_MAX];
|
||||
int n = snprintf(candidate, sizeof(candidate), "%s/%s", rel_buf, component);
|
||||
if (n < 0 || (size_t)n >= sizeof(candidate))
|
||||
return false;
|
||||
const char* cand = candidate;
|
||||
while (*cand == '/')
|
||||
cand++;
|
||||
size_t fl = strlen(count_floor);
|
||||
if (strncmp(cand, count_floor, fl) != 0)
|
||||
return false;
|
||||
return cand[fl] == '\0' || cand[fl] == '/';
|
||||
}
|
||||
|
||||
/* Public wrapper for the receiver's created-directory accounting: the logical
|
||||
* transfer root expressed receive-root-relative, or NULL when the wire paths
|
||||
* carry no mirror scaffolding above it (--relative and --files-from, whose
|
||||
* paths are already relative to the transfer root). The caller frees a
|
||||
* non-NULL result. */
|
||||
char* file_transfer_root_floor(const Config* config) {
|
||||
if (!config || config->relative || config->files_from_set != NULL)
|
||||
return NULL;
|
||||
return transfer_root_floor(config);
|
||||
}
|
||||
|
||||
int file_open_secure_parent_counted(const char* path, char** leaf_out, bool create_dirs,
|
||||
unsigned* dirs_created, const char* count_floor) {
|
||||
char* copy = str_dup(path);
|
||||
if (!copy)
|
||||
return -1;
|
||||
@@ -674,6 +807,14 @@ int file_open_secure_parent(const char* path, char** leaf_out, bool create_dirs)
|
||||
if (next < 0 && create_dirs && errno == ENOENT) {
|
||||
bool created = mkdirat(fd, component, (mode_t)(0777 & ~(mode_t)file_process_umask())) == 0;
|
||||
if (created || errno == EEXIST) {
|
||||
/* Protocol 2.28.0: only directories the logical file list would
|
||||
create count toward `Number of created files`; the mirror
|
||||
scaffolding above the transfer root (e.g. the absolute source path
|
||||
under the destination root) is not an rsync entry. `count_floor`
|
||||
is a receive-root-relative prefix that must be reached before a
|
||||
created component is counted. */
|
||||
if (created && dirs_created && created_dir_counts(count_floor, rel_buf, component))
|
||||
(*dirs_created)++;
|
||||
/* P7 Wave E: --copy-as owns EVERY entry, including the intermediate
|
||||
directories this walk creates implicitly. Its target ids are a
|
||||
global policy, so they are available here without per-entry source
|
||||
@@ -807,6 +948,25 @@ bool file_ensure_directory_secure(const char* path) {
|
||||
} else if (errno == EEXIST) {
|
||||
dir_fd = openat(parent_fd, leaf, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
|
||||
}
|
||||
} else if (dir_fd < 0 && errno == ENOTDIR) {
|
||||
/* rsync replaces a destination non-directory (regular file) with an
|
||||
incoming directory. Confined to the already-opened secure parent fd:
|
||||
the leaf is unlinked by name (never followed) and only a non-directory
|
||||
is ever removed, so this cannot escape the authorized root or remove a
|
||||
pre-existing directory tree. A symlink is left alone (openat with
|
||||
O_NOFOLLOW reports ELOOP, which takes no branch here), since replacing
|
||||
it is not required for FastSync's transferred directories and keeps
|
||||
--keep-dirlinks semantics untouched. */
|
||||
struct stat leaf_st;
|
||||
if (fstatat(parent_fd, leaf, &leaf_st, AT_SYMLINK_NOFOLLOW) == 0 && !S_ISDIR(leaf_st.st_mode) &&
|
||||
!S_ISLNK(leaf_st.st_mode)) {
|
||||
if (unlinkat(parent_fd, leaf, 0) == 0) {
|
||||
if (mkdirat(parent_fd, leaf, (mode_t)(0777 & ~(mode_t)file_process_umask())) == 0)
|
||||
created = true;
|
||||
/* On failure dir_fd stays < 0 below, so the caller still sees it. */
|
||||
dir_fd = openat(parent_fd, leaf, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
|
||||
}
|
||||
}
|
||||
}
|
||||
bool ok = dir_fd >= 0;
|
||||
/* --copy-as owns a directory this call just created (the final component;
|
||||
@@ -969,17 +1129,51 @@ int file_open_private_dir(const char* dir_path) {
|
||||
return fd;
|
||||
}
|
||||
|
||||
/* Open a --temp-dir scratch directory exactly as rsync does: the directory must
|
||||
* already exist and is used as given (an absolute path is used verbatim, a
|
||||
* relative one was already resolved against the destination root by the
|
||||
* caller). Unlike file_open_private_dir this neither creates it nor confines
|
||||
* it below the receive root, because rsync accepts any temp dir -- including
|
||||
* one outside the destination tree or on another filesystem. Returns an
|
||||
* O_DIRECTORY|O_CLOEXEC fd, or -1 on error. */
|
||||
/* Open a --temp-dir scratch directory. The directory must already exist (rsync
|
||||
* never creates it); a relative path was already resolved against the
|
||||
* destination root by the caller. Unlike file_open_private_dir this neither
|
||||
* creates it nor requires it to be a direct child of the receive root, because
|
||||
* rsync permits a scratch dir that (via a symlink) lands on another filesystem
|
||||
* -- but it MUST resolve inside the authorized receive root. The directory is
|
||||
* opened following symlinks and then judged by the REAL path of the opened fd
|
||||
* (through /proc/self/fd), so a client-planted symlink under the receive root
|
||||
* can never redirect receiver scratch files outside the sandbox while an
|
||||
* in-root link to another filesystem (the EXDEV fallback case) still works.
|
||||
* Returns an O_DIRECTORY|O_CLOEXEC fd, or -1 on error (errno set; an escaping
|
||||
* target is reported as EACCES with a logged reason). */
|
||||
int file_open_temp_dir(const char* dir_path) {
|
||||
if (!dir_path)
|
||||
return -1;
|
||||
return open(dir_path, O_RDONLY | O_DIRECTORY | O_CLOEXEC);
|
||||
int fd = open(dir_path, O_RDONLY | O_DIRECTORY | O_CLOEXEC);
|
||||
if (fd < 0)
|
||||
return -1;
|
||||
const char* root = utils_get_authorized_root_path();
|
||||
if (!root) {
|
||||
/* No authorized root (e.g. a local batch apply): nothing to confine
|
||||
against, so preserve the historical open-as-given behavior. */
|
||||
return fd;
|
||||
}
|
||||
char fd_path[64];
|
||||
int fd_path_length = snprintf(fd_path, sizeof(fd_path), "/proc/self/fd/%d", fd);
|
||||
char resolved[PATH_MAX];
|
||||
if (fd_path_length < 0 || (size_t)fd_path_length >= sizeof(fd_path) ||
|
||||
!realpath(fd_path, resolved)) {
|
||||
int saved_errno = errno;
|
||||
close(fd);
|
||||
errno = saved_errno;
|
||||
return -1;
|
||||
}
|
||||
if (!path_is_within_root(root, resolved)) {
|
||||
char* escaped = output_escape(dir_path, log_get_8_bit_output());
|
||||
log_message(LOG_LEVEL_ERROR,
|
||||
"--temp-dir '%s' resolves outside the authorized receive root; refusing",
|
||||
escaped ? escaped : "<allocation failed>");
|
||||
free(escaped);
|
||||
close(fd);
|
||||
errno = EACCES;
|
||||
return -1;
|
||||
}
|
||||
return fd;
|
||||
}
|
||||
|
||||
/* After the content and mode/times are restored on the just-written file, apply
|
||||
@@ -1007,15 +1201,14 @@ static void restore_extra_fd(int fd, const FileMetadata* metadata, const FileXat
|
||||
}
|
||||
}
|
||||
|
||||
static bool file_to_disk_secure_impl(const char* path, const void* data,
|
||||
unsigned long long data_size, bool inplace, bool sparse,
|
||||
bool preallocate, const FileMetadata* metadata,
|
||||
FileAttrPolicy policy, bool update, bool no_replace,
|
||||
bool use_fsync, const char* temp_dir,
|
||||
const FileXattrList* xattrs, bool fake_super,
|
||||
bool keep_partial) {
|
||||
static bool
|
||||
file_to_disk_secure_impl(const char* path, const void* data, unsigned long long data_size,
|
||||
bool inplace, bool sparse, bool preallocate, const FileMetadata* metadata,
|
||||
FileAttrPolicy policy, bool update, bool no_replace, bool use_fsync,
|
||||
const char* temp_dir, const FileXattrList* xattrs, bool fake_super,
|
||||
bool keep_partial, unsigned* dirs_created, const char* count_floor) {
|
||||
char* leaf = NULL;
|
||||
int dirfd = file_open_secure_parent(path, &leaf, true);
|
||||
int dirfd = file_open_secure_parent_counted(path, &leaf, true, dirs_created, count_floor);
|
||||
if (dirfd < 0)
|
||||
return false;
|
||||
int fd = -1;
|
||||
@@ -1306,7 +1499,7 @@ static bool file_to_disk_secure_impl(const char* path, const void* data,
|
||||
"non-atomic copy into the destination directory");
|
||||
return file_to_disk_secure_impl(path, data, data_size, inplace, sparse, preallocate, metadata,
|
||||
policy, update, no_replace, use_fsync, NULL, xattrs, fake_super,
|
||||
keep_partial);
|
||||
keep_partial, dirs_created, count_floor);
|
||||
}
|
||||
return ok;
|
||||
}
|
||||
@@ -1315,7 +1508,8 @@ bool file_to_disk_secure(const char* path, const void* data, unsigned long long
|
||||
bool inplace, bool sparse, bool preallocate, const FileMetadata* metadata,
|
||||
FileAttrPolicy policy, const char* temp_dir) {
|
||||
return file_to_disk_secure_impl(path, data, data_size, inplace, sparse, preallocate, metadata,
|
||||
policy, false, false, false, temp_dir, NULL, false, false);
|
||||
policy, false, false, false, temp_dir, NULL, false, false, NULL,
|
||||
NULL);
|
||||
}
|
||||
|
||||
bool file_to_disk_secure_update(const char* path, const void* data, unsigned long long data_size,
|
||||
@@ -1323,7 +1517,8 @@ bool file_to_disk_secure_update(const char* path, const void* data, unsigned lon
|
||||
const FileMetadata* metadata, FileAttrPolicy policy,
|
||||
const char* temp_dir) {
|
||||
return file_to_disk_secure_impl(path, data, data_size, inplace, sparse, preallocate, metadata,
|
||||
policy, true, false, false, temp_dir, NULL, false, false);
|
||||
policy, true, false, false, temp_dir, NULL, false, false, NULL,
|
||||
NULL);
|
||||
}
|
||||
|
||||
bool file_to_disk_secure_with_fsync(const char* path, const void* data,
|
||||
@@ -1331,7 +1526,8 @@ bool file_to_disk_secure_with_fsync(const char* path, const void* data,
|
||||
bool preallocate, const FileMetadata* metadata,
|
||||
FileAttrPolicy policy, bool use_fsync, const char* temp_dir) {
|
||||
return file_to_disk_secure_impl(path, data, data_size, inplace, sparse, preallocate, metadata,
|
||||
policy, false, false, use_fsync, temp_dir, NULL, false, false);
|
||||
policy, false, false, use_fsync, temp_dir, NULL, false, false,
|
||||
NULL, NULL);
|
||||
}
|
||||
|
||||
bool file_to_disk_secure_no_replace(const char* path, const void* data,
|
||||
@@ -1339,7 +1535,8 @@ bool file_to_disk_secure_no_replace(const char* path, const void* data,
|
||||
const FileMetadata* metadata, FileAttrPolicy policy,
|
||||
const char* temp_dir) {
|
||||
return file_to_disk_secure_impl(path, data, data_size, false, sparse, preallocate, metadata,
|
||||
policy, false, true, false, temp_dir, NULL, false, false);
|
||||
policy, false, true, false, temp_dir, NULL, false, false, NULL,
|
||||
NULL);
|
||||
}
|
||||
|
||||
/* Receiver write-path variant that also applies the per-file xattrs (-X/-A)
|
||||
@@ -1352,9 +1549,21 @@ bool file_to_disk_secure_attrs(const char* path, const void* data, unsigned long
|
||||
const FileMetadata* metadata, FileAttrPolicy policy, bool update,
|
||||
bool no_replace, bool use_fsync, const FileXattrList* xattrs,
|
||||
bool fake_super, bool keep_partial, const char* temp_dir) {
|
||||
return file_to_disk_secure_attrs_counted(path, data, data_size, inplace, sparse, preallocate,
|
||||
metadata, policy, update, no_replace, use_fsync, xattrs,
|
||||
fake_super, keep_partial, temp_dir, NULL, NULL);
|
||||
}
|
||||
|
||||
bool file_to_disk_secure_attrs_counted(const char* path, const void* data,
|
||||
unsigned long long data_size, bool inplace, bool sparse,
|
||||
bool preallocate, const FileMetadata* metadata,
|
||||
FileAttrPolicy policy, bool update, bool no_replace,
|
||||
bool use_fsync, const FileXattrList* xattrs, bool fake_super,
|
||||
bool keep_partial, const char* temp_dir,
|
||||
unsigned* dirs_created, const char* count_floor) {
|
||||
return file_to_disk_secure_impl(path, data, data_size, inplace, sparse, preallocate, metadata,
|
||||
policy, update, no_replace, use_fsync, temp_dir, xattrs,
|
||||
fake_super, keep_partial);
|
||||
fake_super, keep_partial, dirs_created, count_floor);
|
||||
}
|
||||
|
||||
/* Atomic --link-dest install. The destination is replaced (via a temporary
|
||||
@@ -1372,16 +1581,176 @@ bool file_to_disk_secure_attrs(const char* path, const void* data, unsigned long
|
||||
* basis). Likewise `xattrs`/`fake_super` are applied only on the copy
|
||||
* fallback, so a fallback copy preserves the per-file attributes instead of
|
||||
* silently dropping them. */
|
||||
/* Streaming --copy-dest basis install: atomically materialize `path` from the
|
||||
* bytes of `basis_path` without holding the file in memory, so a basis larger
|
||||
* than any whole-file bound still works. Mirrors the ordinary secure store
|
||||
* path (confined parent walk, temp + rename, --update/--ignore-existing/
|
||||
* --preallocate/--temp-dir) but sources the data from the basis descriptor
|
||||
* rather than a caller buffer, and applies the SOURCE metadata (rsync copies
|
||||
* then fixes attributes). A hard-link install that falls back to a byte copy
|
||||
* also routes through here when the caller supplies the basis path. */
|
||||
static bool file_copy_basis_stream_impl(const char* path, const char* basis_path,
|
||||
unsigned long long expected_size, bool preallocate,
|
||||
const FileMetadata* metadata, FileAttrPolicy policy,
|
||||
bool update, bool no_replace, bool use_fsync,
|
||||
const FileXattrList* xattrs, bool fake_super,
|
||||
const char* temp_dir, unsigned* dirs_created,
|
||||
const char* count_floor) {
|
||||
if (!path || !basis_path)
|
||||
return false;
|
||||
char* leaf = NULL;
|
||||
int dirfd = file_open_secure_parent_counted(path, &leaf, true, dirs_created, count_floor);
|
||||
if (dirfd < 0)
|
||||
return false;
|
||||
|
||||
char* basis_leaf = NULL;
|
||||
int basis_dirfd = file_open_secure_parent(basis_path, &basis_leaf, false);
|
||||
int src_fd = -1;
|
||||
if (basis_dirfd >= 0 && basis_leaf != NULL) {
|
||||
/* O_NONBLOCK rejects a raced-in FIFO without blocking; the S_ISREG gate
|
||||
below is the real type check. */
|
||||
src_fd = openat(basis_dirfd, basis_leaf, O_RDONLY | O_CLOEXEC | O_NOFOLLOW | O_NONBLOCK);
|
||||
struct stat src_st;
|
||||
if (src_fd >= 0 && (fstat(src_fd, &src_st) != 0 || !S_ISREG(src_st.st_mode))) {
|
||||
close(src_fd);
|
||||
src_fd = -1;
|
||||
}
|
||||
}
|
||||
if (basis_dirfd >= 0)
|
||||
close(basis_dirfd);
|
||||
free(basis_leaf);
|
||||
if (src_fd < 0) {
|
||||
close(dirfd);
|
||||
free(leaf);
|
||||
return false;
|
||||
}
|
||||
|
||||
struct stat destination_stat;
|
||||
bool destination_is_regular = fstatat(dirfd, leaf, &destination_stat, AT_SYMLINK_NOFOLLOW) == 0 &&
|
||||
S_ISREG(destination_stat.st_mode);
|
||||
if (update && metadata && destination_is_regular && stat_is_newer(&destination_stat, metadata)) {
|
||||
close(src_fd);
|
||||
close(dirfd);
|
||||
free(leaf);
|
||||
return true;
|
||||
}
|
||||
if (no_replace && file_path_exists_secure(path)) {
|
||||
close(src_fd);
|
||||
close(dirfd);
|
||||
free(leaf);
|
||||
return true;
|
||||
}
|
||||
|
||||
int scratch_dirfd = -1;
|
||||
if (temp_dir) {
|
||||
scratch_dirfd = file_open_temp_dir(temp_dir);
|
||||
if (scratch_dirfd < 0) {
|
||||
int saved_errno = errno;
|
||||
log_message(LOG_LEVEL_ERROR,
|
||||
"--temp-dir '%s' could not be opened (rsync requires it to already exist): %s",
|
||||
temp_dir, strerror(saved_errno));
|
||||
close(src_fd);
|
||||
close(dirfd);
|
||||
free(leaf);
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
int target_dirfd = scratch_dirfd >= 0 ? scratch_dirfd : dirfd;
|
||||
int tmp_size = snprintf(NULL, 0, ".%s.tmp.%ld.%llu", leaf, (long)getpid(), ~0ULL);
|
||||
char* tmp = NULL;
|
||||
bool ok = false;
|
||||
if (tmp_size >= 0)
|
||||
tmp = malloc((size_t)tmp_size + 1);
|
||||
if (tmp) {
|
||||
for (unsigned int i = 0; i < 100 && !ok; ++i) {
|
||||
if (scratch_dirfd >= 0)
|
||||
snprintf(tmp, (size_t)tmp_size + 1, ".%s.tmp.%ld.%llu", leaf, (long)getpid(),
|
||||
next_temp_sequence());
|
||||
else
|
||||
snprintf(tmp, (size_t)tmp_size + 1, ".%s.tmp.%ld.%u", leaf, (long)getpid(), i);
|
||||
int fd =
|
||||
openat(target_dirfd, tmp, O_WRONLY | O_CREAT | O_EXCL | O_CLOEXEC | O_NOFOLLOW, 0600);
|
||||
if (fd < 0) {
|
||||
if (errno != EEXIST)
|
||||
break;
|
||||
continue;
|
||||
}
|
||||
bool wrote = true;
|
||||
if (preallocate && expected_size > 0 && preallocate_fd(fd, expected_size) != 0)
|
||||
wrote = false;
|
||||
if (wrote)
|
||||
wrote = copy_fd_all(fd, src_fd, expected_size);
|
||||
if (wrote && metadata) {
|
||||
if (!policy.perms &&
|
||||
fchmod(fd, file_mode_base(metadata, destination_is_regular,
|
||||
destination_is_regular ? destination_stat.st_mode & 0777
|
||||
: 0)) != 0)
|
||||
wrote = false;
|
||||
if (wrote)
|
||||
wrote = file_restore_metadata_fd(fd, metadata, policy);
|
||||
} else if (wrote && fchmod(fd, S_IRUSR | S_IWUSR | S_IRGRP | S_IROTH) != 0) {
|
||||
wrote = false;
|
||||
}
|
||||
if (wrote)
|
||||
restore_extra_fd(fd, metadata, xattrs, fake_super, policy);
|
||||
if (wrote && use_fsync)
|
||||
wrote = fsync(fd) == 0;
|
||||
if (close(fd) != 0)
|
||||
wrote = false;
|
||||
if (wrote && renameat(target_dirfd, tmp, dirfd, leaf) != 0)
|
||||
wrote = false;
|
||||
if (!wrote)
|
||||
unlinkat(target_dirfd, tmp, 0);
|
||||
ok = wrote;
|
||||
}
|
||||
free(tmp);
|
||||
}
|
||||
if (!ok && scratch_dirfd >= 0) {
|
||||
/* Retry once with no scratch dir (rsync's EXDEV fallback). */
|
||||
close(scratch_dirfd);
|
||||
close(src_fd);
|
||||
close(dirfd);
|
||||
free(leaf);
|
||||
return file_copy_basis_stream_impl(path, basis_path, expected_size, preallocate, metadata,
|
||||
policy, update, no_replace, use_fsync, xattrs, fake_super,
|
||||
NULL, dirs_created, count_floor);
|
||||
}
|
||||
if (scratch_dirfd >= 0)
|
||||
close(scratch_dirfd);
|
||||
close(src_fd);
|
||||
close(dirfd);
|
||||
free(leaf);
|
||||
return ok;
|
||||
}
|
||||
|
||||
/* --copy-dest basis install (streaming). Applies the source metadata and the
|
||||
per-file xattrs / --fake-super record. */
|
||||
bool file_copy_basis_stream_attrs(const char* path, const char* basis_path,
|
||||
unsigned long long expected_size, bool preallocate,
|
||||
const FileMetadata* metadata, FileAttrPolicy policy, bool update,
|
||||
bool use_fsync, const FileXattrList* xattrs, bool fake_super,
|
||||
const char* temp_dir) {
|
||||
return file_copy_basis_stream_impl(path, basis_path, expected_size, preallocate, metadata, policy,
|
||||
update, false, use_fsync, xattrs, fake_super, temp_dir, NULL,
|
||||
NULL);
|
||||
}
|
||||
|
||||
static bool file_to_disk_secure_link_impl(const char* path, const char* basis_path,
|
||||
const void* data, unsigned long long data_size,
|
||||
bool preallocate, const FileMetadata* metadata,
|
||||
FileAttrPolicy policy, bool use_fsync,
|
||||
const FileXattrList* xattrs, bool fake_super,
|
||||
const char* temp_dir) {
|
||||
const char* temp_dir, unsigned* dirs_created,
|
||||
const char* count_floor) {
|
||||
if (!path || !basis_path)
|
||||
return false;
|
||||
/* The caller-supplied buffer is no longer used: the copy fallback streams
|
||||
from the basis path (which may hold an over-limit file). Kept in the
|
||||
signature for the existing API. */
|
||||
(void)data;
|
||||
char* leaf = NULL;
|
||||
int dirfd = file_open_secure_parent(path, &leaf, true);
|
||||
int dirfd = file_open_secure_parent_counted(path, &leaf, true, dirs_created, count_floor);
|
||||
if (dirfd < 0)
|
||||
return false;
|
||||
|
||||
@@ -1463,10 +1832,17 @@ static bool file_to_disk_secure_link_impl(const char* path, const char* basis_pa
|
||||
close(dirfd);
|
||||
free(leaf);
|
||||
/* The basis file could not be linked in (missing, cross-device, refused
|
||||
by the filesystem). Write a byte-identical local copy instead. */
|
||||
return file_to_disk_secure_attrs(path, data, data_size, false, false, preallocate, metadata,
|
||||
policy, false, false, use_fsync, xattrs, fake_super, false,
|
||||
temp_dir);
|
||||
by the filesystem). Stream a byte-identical local copy from the basis
|
||||
itself (never the possibly-absent caller buffer) so an over-limit basis
|
||||
still materializes. When the basis path is not a readable regular file
|
||||
(e.g. a directory raced in), fall back to the caller-supplied bytes. */
|
||||
if (file_copy_basis_stream_impl(path, basis_path, data_size, preallocate, metadata, policy,
|
||||
false, false, use_fsync, xattrs, fake_super, temp_dir,
|
||||
dirs_created, count_floor))
|
||||
return true;
|
||||
return file_to_disk_secure_attrs_counted(
|
||||
path, data, data_size, false, false, preallocate, metadata, policy, false, false, use_fsync,
|
||||
xattrs, fake_super, false, temp_dir, dirs_created, count_floor);
|
||||
}
|
||||
|
||||
if (scratch_dirfd >= 0)
|
||||
@@ -1481,7 +1857,7 @@ bool file_to_disk_secure_link(const char* path, const char* basis_path, const vo
|
||||
const FileMetadata* metadata, FileAttrPolicy policy, bool use_fsync,
|
||||
const char* temp_dir) {
|
||||
return file_to_disk_secure_link_impl(path, basis_path, data, data_size, preallocate, metadata,
|
||||
policy, use_fsync, NULL, false, temp_dir);
|
||||
policy, use_fsync, NULL, false, temp_dir, NULL, NULL);
|
||||
}
|
||||
|
||||
bool file_to_disk_secure_link_attrs(const char* path, const char* basis_path, const void* data,
|
||||
@@ -1489,14 +1865,27 @@ bool file_to_disk_secure_link_attrs(const char* path, const char* basis_path, co
|
||||
const FileMetadata* metadata, FileAttrPolicy policy,
|
||||
bool use_fsync, const FileXattrList* xattrs, bool fake_super,
|
||||
const char* temp_dir) {
|
||||
return file_to_disk_secure_link_attrs_counted(path, basis_path, data, data_size, preallocate,
|
||||
metadata, policy, use_fsync, xattrs, fake_super,
|
||||
temp_dir, NULL, NULL);
|
||||
}
|
||||
|
||||
bool file_to_disk_secure_link_attrs_counted(const char* path, const char* basis_path,
|
||||
const void* data, unsigned long long data_size,
|
||||
bool preallocate, const FileMetadata* metadata,
|
||||
FileAttrPolicy policy, bool use_fsync,
|
||||
const FileXattrList* xattrs, bool fake_super,
|
||||
const char* temp_dir, unsigned* dirs_created,
|
||||
const char* count_floor) {
|
||||
return file_to_disk_secure_link_impl(path, basis_path, data, data_size, preallocate, metadata,
|
||||
policy, use_fsync, xattrs, fake_super, temp_dir);
|
||||
policy, use_fsync, xattrs, fake_super, temp_dir,
|
||||
dirs_created, count_floor);
|
||||
}
|
||||
|
||||
bool file_write_to_disk(const char* path, const void* data, unsigned long long data_size,
|
||||
bool inplace, bool sparse) {
|
||||
if (!path || (!data && data_size != 0) || has_path_traversal(path))
|
||||
return false;
|
||||
FileAttrPolicy policy = {false, false, false, false};
|
||||
FileAttrPolicy policy = {0};
|
||||
return file_to_disk_secure(path, data, data_size, inplace, sparse, false, NULL, policy, NULL);
|
||||
}
|
||||
|
||||
+43
-2
@@ -87,6 +87,11 @@ bool file_path_exists_secure(const char* path);
|
||||
bool file_stat_secure(const char* path, struct stat* st);
|
||||
bool file_destination_is_newer_secure(const char* path, const FileMetadata* metadata);
|
||||
int file_open_secure_parent(const char* path, char** leaf_out, bool create_dirs);
|
||||
/* Protocol 2.28.0 variant: also increments *dirs_created for every missing
|
||||
* parent directory this walk creates that lies strictly below `count_floor`
|
||||
* (a receive-root-relative path, or NULL to count all of them). */
|
||||
int file_open_secure_parent_counted(const char* path, char** leaf_out, bool create_dirs,
|
||||
unsigned* dirs_created, const char* count_floor);
|
||||
bool file_ensure_directory_secure(const char* path);
|
||||
bool file_directory_exists_secure(const char* path);
|
||||
bool file_rename_secure(const char* old_path, const char* new_path);
|
||||
@@ -98,8 +103,12 @@ bool file_remove_tree_secure(const char* path);
|
||||
the authorized root. Used for the --delay-updates staging directory. */
|
||||
int file_open_private_dir(const char* dir_path);
|
||||
|
||||
/* Open an existing --temp-dir scratch directory as-is (absolute or relative;
|
||||
no creation, no root confinement), matching rsync's --temp-dir handling. */
|
||||
/* Open an existing --temp-dir scratch directory (relative or absolute; no
|
||||
creation). When an authorized receive root is configured the directory's
|
||||
REAL path (symlinks resolved) must lie within it, so a client-planted
|
||||
symlink cannot redirect receiver scratch files outside the sandbox; an
|
||||
in-root symlink to another filesystem is still allowed for rsync's EXDEV
|
||||
fallback. */
|
||||
int file_open_temp_dir(const char* dir_path);
|
||||
|
||||
/* The file_to_disk_secure* variants write a temporary copy in the destination
|
||||
@@ -158,5 +167,37 @@ bool file_to_disk_secure_link_attrs(const char* path, const char* basis_path, co
|
||||
const FileMetadata* metadata, FileAttrPolicy policy,
|
||||
bool use_fsync, const FileXattrList* xattrs, bool fake_super,
|
||||
const char* temp_dir);
|
||||
/* Streaming --copy-dest install: atomically materialize `path` by copying the
|
||||
* bytes of `basis_path` through a bounded buffer (no whole-file buffering, so
|
||||
* an arbitrarily large basis works), applying the SOURCE metadata and the
|
||||
* per-file xattrs / --fake-super record. `update` honors a newer destination;
|
||||
* a --temp-dir scratch location falls back to a direct write on EXDEV. */
|
||||
bool file_copy_basis_stream_attrs(const char* path, const char* basis_path,
|
||||
unsigned long long expected_size, bool preallocate,
|
||||
const FileMetadata* metadata, FileAttrPolicy policy, bool update,
|
||||
bool use_fsync, const FileXattrList* xattrs, bool fake_super,
|
||||
const char* temp_dir);
|
||||
/* Protocol 2.28.0 receiver-stat variants: like the two above but additionally
|
||||
* report through `dirs_created` (when non-NULL) how many parent directories the
|
||||
* confined secure walk had to create that lie strictly below `count_floor` (a
|
||||
* receive-root-relative prefix, or NULL for all). Used to reproduce rsync's
|
||||
* `Number of created files` directory count on a fresh destination. */
|
||||
bool file_to_disk_secure_attrs_counted(const char* path, const void* data,
|
||||
unsigned long long data_size, bool inplace, bool sparse,
|
||||
bool preallocate, const FileMetadata* metadata,
|
||||
FileAttrPolicy policy, bool update, bool no_replace,
|
||||
bool use_fsync, const FileXattrList* xattrs, bool fake_super,
|
||||
bool keep_partial, const char* temp_dir,
|
||||
unsigned* dirs_created, const char* count_floor);
|
||||
bool file_to_disk_secure_link_attrs_counted(const char* path, const char* basis_path,
|
||||
const void* data, unsigned long long data_size,
|
||||
bool preallocate, const FileMetadata* metadata,
|
||||
FileAttrPolicy policy, bool use_fsync,
|
||||
const FileXattrList* xattrs, bool fake_super,
|
||||
const char* temp_dir, unsigned* dirs_created,
|
||||
const char* count_floor);
|
||||
/* The logical transfer root expressed receive-root-relative, or NULL when the
|
||||
* wire paths carry no mirror scaffolding above it. Caller frees non-NULL. */
|
||||
char* file_transfer_root_floor(const Config* config);
|
||||
|
||||
#endif
|
||||
|
||||
@@ -29,6 +29,12 @@ typedef struct FileAttrPolicy {
|
||||
bool times; /* config->preserve_times: apply the source mtime */
|
||||
bool atimes; /* config->preserve_atimes (-U): apply the source atime */
|
||||
bool executability; /* config->use_executability (-E): exec-bits-only mode */
|
||||
/* privilege_super_mode_permitted(): when false (SUPER_MODE_OFF / --no-super,
|
||||
or a daemon that did not grant `client owner = yes`), the setuid/setgid/
|
||||
sticky bits are stripped from every applied mode (source mode and any
|
||||
--chmod result) even under --perms. When true, rsync's exact semantics are
|
||||
preserved: -p copies the special bits and the kernel decides. */
|
||||
bool super_permitted;
|
||||
} FileAttrPolicy;
|
||||
|
||||
/* Build the per-attribute policy from a connection's Config. A NULL config
|
||||
|
||||
+7
-3066
File diff suppressed because it is too large
Load Diff
+10
-104
@@ -2,10 +2,19 @@
|
||||
#define FILE_RECEIVE_H
|
||||
|
||||
#include "config.h"
|
||||
#include "delete_commit.h"
|
||||
#include "file_save.h"
|
||||
#include "file_types.h"
|
||||
#include "incremental_check.h"
|
||||
#include "utils.h"
|
||||
#include <stdbool.h>
|
||||
|
||||
/* Server-side file receive/save path. */
|
||||
/* Server-side file receive/save path.
|
||||
*
|
||||
* This header is the public facade for the file_receive module family: the
|
||||
* wire receive dispatch (this file) plus the save-to-disk (file_save.h), the
|
||||
* incremental check (incremental_check.h) and the delete-commit
|
||||
* (delete_commit.h) modules. */
|
||||
|
||||
/* Cumulative caps for the deferred directory-time accumulator. The sender may
|
||||
* legitimately split a large tree across repeated STATUS_DIR_TIMES frames, so a
|
||||
@@ -22,15 +31,6 @@ File* file_receive_dir_time(int file_descriptor, const Config* config);
|
||||
File* file_receive_hardlink(int file_descriptor);
|
||||
File* file_receive_symlink(int file_descriptor, const Config* config);
|
||||
File* file_receive_special(int file_descriptor);
|
||||
bool file_special_rdev_valid(int32_t major, int32_t minor, mode_t mode);
|
||||
File* receive_incremental_check(int fd, const Config* config, bool* skipped);
|
||||
/* Extended variant used by the receiver. `would_transfer` (may be NULL) is set
|
||||
* true only on the server-contacting --dry-run path when the file is not up to
|
||||
* date: the receiver has already sent STATUS_DRY_RUN_TRANSFER and returns NULL
|
||||
* without storing anything. On that path `*skipped` is true for an up-to-date
|
||||
* (STATUS_OK) file and both flags are false for a genuine error. */
|
||||
File* receive_incremental_check_ex(int fd, const Config* config, bool* skipped,
|
||||
bool* would_transfer);
|
||||
|
||||
/* P7 Wave D directory-time accumulator. The receiver collects the metadata of
|
||||
* every directory it creates/receives (STATUS_MKDIR with metadata and/or the
|
||||
@@ -74,98 +74,4 @@ bool dir_time_list_add(DirTimeList* list, const char* wire_path, const FileMetad
|
||||
void dir_metadata_list_apply(const DirTimeList* list, const char* root_directory,
|
||||
const Config* config);
|
||||
|
||||
/* A received delete-manifest frame: the keep-set (`keeps`, destination-relative
|
||||
paths the sender transferred/keeps) plus `protected`, destination-relative
|
||||
prefixes the sender asks the receiver never to delete (paths excluded on the
|
||||
source, protected at any depth). When --delete-excluded is given the sender
|
||||
transmits an empty protected list so excluded destination mirrors are treated
|
||||
as ordinary extras. With --delete-missing-args a third section (`missing`)
|
||||
carries the destination mirrors of explicitly-listed source entries that do
|
||||
not exist: each is an exact deletion request, independent of the ordinary
|
||||
extras walk (never blocked by the protected prefixes) and processed when the
|
||||
manifest is committed. */
|
||||
typedef struct DeleteManifest {
|
||||
ArrayList* keeps;
|
||||
ArrayList* protected;
|
||||
ArrayList* missing;
|
||||
/* Destination-relative paths of the directories the sender synchronized for
|
||||
this run. The extras walker only removes entries directly inside one of
|
||||
these (the receive root is the "." sentinel); `--files-from` runs therefore
|
||||
leave untransmitted directories and the unlisted parts of listed ones
|
||||
alone, matching rsync's "delete only in synchronized directories". */
|
||||
ArrayList* dirs;
|
||||
} DeleteManifest;
|
||||
|
||||
void delete_manifest_free(DeleteManifest* manifest);
|
||||
/* Read a delete-manifest frame (protocol 2.23.0): keep count + keeps, then
|
||||
protected count + protected prefixes, then missing count + missing paths,
|
||||
then synchronized-directory count + directory paths (self-delimiting; the
|
||||
leading STATUS_MANIFEST code has been consumed). Returns an owned
|
||||
DeleteManifest, or NULL after signalling STATUS_ERROR on a malformed frame. */
|
||||
DeleteManifest* receive_manifest_entries(int fd);
|
||||
/* Remove destination entries under config->receive_root_directory that are not
|
||||
in `manifest` (bounded, all-or-nothing walk; staging-dir, basis-dir and
|
||||
protected-prefix skips). `--max-delete` and `--force` are honored here. The
|
||||
caller decides WHEN to run it based on the negotiated delete timing. Returns
|
||||
false (and the transfer fails) when the deletion cannot be committed. */
|
||||
bool manifest_delete_extras(const Config* config, DeleteManifest* manifest);
|
||||
/* --delete-missing-args exact-path deletions: remove each destination mirror
|
||||
in `manifest->missing` (never blocked by the protected prefixes, staging dir
|
||||
and basis dirs excluded). A regular file/symlink is unlinked; an empty
|
||||
directory is removed; a NON-empty directory is removed recursively only when
|
||||
--delete or --force is in effect, otherwise it is left with a warning (rsync
|
||||
parity). A missing path is a no-op. Returns false only on a genuine
|
||||
confinement or I/O error (the run then fails); tolerated per-path cases are
|
||||
reported and skipped. */
|
||||
bool manifest_delete_missing_args(const Config* config, DeleteManifest* manifest);
|
||||
/* Budgeted form of manifest_delete_missing_args for the per-directory delete
|
||||
session: each removed mirror draws from `max_delete` (SIZE_MAX = unlimited)
|
||||
and the tallies are accumulated into `*deleted`/`*skipped`. `*limit_hit` is set
|
||||
when the budget stopped the pass with entries left over. Returns false only
|
||||
on a genuine deletion error. */
|
||||
bool manifest_delete_missing_args_limited(const Config* config, DeleteManifest* manifest,
|
||||
size_t max_delete, size_t* deleted, size_t* skipped,
|
||||
bool* limit_hit);
|
||||
/* Outcome of committing a delete manifest. LIMIT_REACHED reports rsync's
|
||||
partial --max-delete result: the budget allowed some deletions and the rest
|
||||
were skipped (the run still stores all file data but the client exits 25). */
|
||||
typedef enum {
|
||||
DELETE_COMMIT_OK = 0,
|
||||
DELETE_COMMIT_LIMIT_REACHED,
|
||||
DELETE_COMMIT_ERROR
|
||||
} DeleteCommitResult;
|
||||
|
||||
/* Run every deletion family the manifest carries: the --delete-missing-args
|
||||
exact-path deletions first (user requests are not blocked by exclusion
|
||||
protection), then the ordinary extras walk when --delete is active. Both
|
||||
share one --max-delete budget. Returns DELETE_COMMIT_OK when nothing was to
|
||||
do or everything committed, DELETE_COMMIT_LIMIT_REACHED when the budget
|
||||
stopped part of the work, or DELETE_COMMIT_ERROR on a genuine failure. */
|
||||
DeleteCommitResult manifest_delete_all(const Config* config, DeleteManifest* manifest);
|
||||
/* Like manifest_delete_all, but reports how many destination entries the commit
|
||||
removed (for the end-of-transfer wire stats). `deleted` may be NULL. */
|
||||
DeleteCommitResult manifest_delete_all_counted(const Config* config, DeleteManifest* manifest,
|
||||
size_t* deleted);
|
||||
|
||||
/* -n/--dry-run --delete would-delete reporting: walk the destination exactly as
|
||||
the delete pass would and append (strdup'd) destination-relative paths that
|
||||
WOULD be removed to `out`, without touching disk. Uses the same staging-dir,
|
||||
basis-dir and protected-prefix skips as the real commit. Returns true on a
|
||||
clean walk; `*count_out` receives the number of paths appended. */
|
||||
bool manifest_would_delete_list(const Config* config, DeleteManifest* manifest, ArrayList* out,
|
||||
size_t* count_out);
|
||||
/* Convert one basis-directory path to the receive-root-relative protection
|
||||
prefix the delete walker uses (NULL when it lies outside the root). Exposed
|
||||
for unit tests of the root-of-"/" and normalization edge cases. */
|
||||
char* file_receive_basis_delete_relative(const Config* config, const char* path);
|
||||
|
||||
/* Outcome of a single file_save_to_disk operation. The receiver needs to
|
||||
distinguish "written" from "skipped" so --remove-source-files can be told
|
||||
which sources were actually stored. */
|
||||
typedef enum { FILE_SAVE_ERROR = 0, FILE_SAVE_WRITTEN = 1, FILE_SAVE_SKIPPED = 2 } FileSaveResult;
|
||||
|
||||
FileSaveResult file_save_to_disk_full(const char* root_directory, const File* file,
|
||||
const Config* config);
|
||||
bool file_save_to_disk(const char* root_directory, const File* file, const Config* config);
|
||||
|
||||
#endif
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,40 @@
|
||||
#ifndef FILE_SAVE_H
|
||||
#define FILE_SAVE_H
|
||||
|
||||
#include "config.h"
|
||||
#include "file_types.h"
|
||||
#include "format.h"
|
||||
#include <stdbool.h>
|
||||
|
||||
/* Save-to-disk module: regular-file/symlink/hardlink/special install, xattr
|
||||
* application, --fake-super and the --delay-updates staging path. These
|
||||
* declarations are re-exported by the file_receive.h facade. */
|
||||
|
||||
/* Outcome of a single file_save_to_disk operation. The receiver needs to
|
||||
distinguish "written" from "skipped" so --remove-source-files can be told
|
||||
which sources were actually stored. */
|
||||
typedef enum { FILE_SAVE_ERROR = 0, FILE_SAVE_WRITTEN = 1, FILE_SAVE_SKIPPED = 2 } FileSaveResult;
|
||||
|
||||
bool file_special_rdev_valid(int32_t major, int32_t minor, mode_t mode);
|
||||
|
||||
FileSaveResult file_save_to_disk_full(const char* root_directory, const File* file,
|
||||
const Config* config);
|
||||
/* Protocol 2.28.0 variant: also reports through `created` (when non-NULL)
|
||||
* whether the destination entry did not exist before this save, and through
|
||||
* `created_dirs` how many parent directories the confined walk created, so the
|
||||
* receiver can build rsync's `Number of created files` breakdown. The plain
|
||||
* file_save_to_disk_full() is this with both out-params NULL. */
|
||||
FileSaveResult file_save_to_disk_full_ex(const char* root_directory, const File* file,
|
||||
const Config* config, bool* created,
|
||||
unsigned* created_dirs);
|
||||
bool file_save_to_disk(const char* root_directory, const File* file, const Config* config);
|
||||
|
||||
/* Protocol 2.28.0 receiver counter accumulator: fold one successfully saved
|
||||
* entry into `stats`, adding its receiver-observed literal bytes and, when
|
||||
* `created`, the matching created-by-type counter (regular file / symlink /
|
||||
* special) plus `created_dirs` implicitly-created parent directories.
|
||||
* Non-first hardlink siblings contribute no literal bytes. */
|
||||
void receiver_stats_note_saved(ReceiverStats* stats, const File* file, bool created,
|
||||
unsigned created_dirs);
|
||||
|
||||
#endif
|
||||
@@ -166,6 +166,8 @@ bool file_send_sendfile_with_skip(File* file, int file_descriptor, bool use_meta
|
||||
}
|
||||
struct pollfd pfd = {.fd = file_descriptor, .events = POLLOUT};
|
||||
int polled = poll(&pfd, 1, timeout);
|
||||
if (polled < 0 && errno == EINTR)
|
||||
continue;
|
||||
if (polled <= 0 || (pfd.revents & (POLLERR | POLLHUP | POLLNVAL))) {
|
||||
close(fd);
|
||||
return false;
|
||||
@@ -183,6 +185,7 @@ bool file_send_sendfile_with_skip(File* file, int file_descriptor, bool use_meta
|
||||
return false;
|
||||
}
|
||||
protocol_note_bytes_written((unsigned long long)sent);
|
||||
protocol_throttle_bytes(file_descriptor, (size_t)sent);
|
||||
}
|
||||
|
||||
close(fd);
|
||||
|
||||
@@ -57,6 +57,11 @@ typedef struct {
|
||||
* equals the incoming file, and `data` is kept as the cross-filesystem
|
||||
* fallback (a local copy) if the hard link cannot be created. */
|
||||
char* basis_link;
|
||||
/* Receiver-only, --copy-dest: when set (and basis_link is NULL), stream the
|
||||
* basis file's bytes into the destination instead of `data`/`data->size`.
|
||||
* This lets a basis larger than any whole-file bound materialize without
|
||||
* buffering it; the source metadata on `metadata` is applied afterwards. */
|
||||
char* basis_copy;
|
||||
/* --hard-links (-H), sender + receiver wire state. link_group is a run-local
|
||||
* id shared by every member of one source inode (0 = not part of a group).
|
||||
* The FIRST member (link_first == true) carries its data on the wire and is
|
||||
@@ -97,6 +102,11 @@ typedef struct {
|
||||
* basis file (matched delta blocks) for this entry. 0 when the file was sent
|
||||
* whole. Accumulated into ReceiverStats.matched_data by the receiver sink. */
|
||||
unsigned long long matched_bytes;
|
||||
/* Receiver-only (protocol 2.28.0) wire-stats tally: the literal delta fragment
|
||||
* bytes this entry carried (DELTA_INSTR_LITERAL). 0 when the file was sent
|
||||
* whole; the sink then falls back to the whole payload size. Accumulated
|
||||
* into ReceiverStats.literal_bytes. */
|
||||
unsigned long long literal_bytes;
|
||||
} File;
|
||||
|
||||
/* The path that should be sent on the wire and used for the receiver-side
|
||||
|
||||
+95
-27
@@ -4,22 +4,12 @@
|
||||
#include <ctype.h>
|
||||
#include <errno.h>
|
||||
#include <limits.h>
|
||||
#include <stdarg.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
|
||||
/* Write a diagnostic message into the caller's optional buffer. A NULL `err`
|
||||
* (or a zero size) is a no-op, so a caller that only needs the boolean status
|
||||
* may pass NULL without the snprintf-on-NULL undefined behaviour. */
|
||||
static void filter_set_error(char* err, size_t err_size, const char* fmt, ...) {
|
||||
if (!err || err_size == 0)
|
||||
return;
|
||||
va_list ap;
|
||||
va_start(ap, fmt);
|
||||
vsnprintf(err, err_size, fmt, ap);
|
||||
va_end(ap);
|
||||
}
|
||||
/* Write a diagnostic message into the caller's optional buffer. */
|
||||
#define filter_set_error utils_set_error
|
||||
|
||||
/* ---- Ordered rule lists ---- */
|
||||
|
||||
@@ -172,14 +162,74 @@ static bool is_modifier_char(char c) {
|
||||
return c == 's' || c == 'r' || c == 'p' || c == 'x' || c == '/' || c == '!' || c == 'C';
|
||||
}
|
||||
|
||||
/* merge/dir-merge rules are the only rules rsync accepts the merge-file
|
||||
* modifiers on. */
|
||||
static bool is_merge_rule(RuleKind kind) {
|
||||
return kind == RULE_KIND_MERGE || kind == RULE_KIND_DIR_MERGE;
|
||||
}
|
||||
|
||||
/* Merge-file modifiers rsync defines but FastSync does not implement:
|
||||
* 'e' exclude the merge file itself, 'n' do not inherit the merge file, 'w'
|
||||
* word-split the merge file. They are recognized as part of a modifier run on
|
||||
* every rule (so a pure e/n/w token is rejected rather than folded into the
|
||||
* pattern), but are accepted (and ignored) only on merge/dir-merge rules. */
|
||||
static bool is_unsupported_modifier_char(char c) {
|
||||
return c == 'e' || c == 'n' || c == 'w';
|
||||
}
|
||||
|
||||
/* Merge-file modifiers rsync accepts on merge/dir-merge rules: 'e', 'n', 'w'
|
||||
* and '-' (do not transfer the merge file). */
|
||||
static bool is_merge_modifier_char(char c) {
|
||||
return c == 'e' || c == 'n' || c == 'w' || c == '-';
|
||||
}
|
||||
|
||||
/* Characters that count as part of a modifier run for `kind` when deciding
|
||||
* whether a token is a pure modifier run. e/n/w count on every rule so that a
|
||||
* pure e/n/w token is rejected on non-merge rules; '-' only on merge rules. */
|
||||
static bool is_modifier_scan_char(char c, RuleKind kind) {
|
||||
return is_modifier_char(c) || is_unsupported_modifier_char(c) ||
|
||||
(is_merge_rule(kind) && is_merge_modifier_char(c));
|
||||
}
|
||||
|
||||
/* Characters actually consumed as modifiers for `kind`. The merge-file
|
||||
* modifiers are consumed only on merge/dir-merge rules; elsewhere e/n/w fall
|
||||
* through to the pattern (so mixed tokens such as "H,!secret" keep their
|
||||
* historical "ecret" pattern). */
|
||||
static bool is_consumed_modifier_char(char c, RuleKind kind) {
|
||||
return is_modifier_char(c) || (is_merge_rule(kind) && is_merge_modifier_char(c));
|
||||
}
|
||||
|
||||
/* Inspect the token that follows a rule name (up to the first space/underscore
|
||||
* or the end). If the token is composed *solely* of modifier characters and
|
||||
* includes one that is invalid for `kind`, it is unambiguously a modifier run:
|
||||
* return that character so the caller can reject it. A token that contains any
|
||||
* non-modifier character is a pattern (e.g. "-newfile") and returns '\0', which
|
||||
* keeps the historical parsing of mixed tokens such as "H,!secret" intact. */
|
||||
static char unsupported_modifier_in_token(const char* tok, RuleKind kind) {
|
||||
if (*tok == '\0' || *tok == ' ' || *tok == '_')
|
||||
return '\0';
|
||||
char bad = '\0';
|
||||
for (const char* q = tok; *q != '\0' && *q != ' ' && *q != '_'; q++) {
|
||||
if (!is_modifier_scan_char(*q, kind))
|
||||
return '\0';
|
||||
if (!is_merge_rule(kind) && is_unsupported_modifier_char(*q))
|
||||
bad = *q;
|
||||
}
|
||||
return bad;
|
||||
}
|
||||
|
||||
/* Parse "RULE[,MODIFIERS] [PATTERN]". On success `kind`, `sides`,
|
||||
* `sides_explicit`, `negate`, `anchored_mod`, `perishable`, `xattr`,
|
||||
* `cvs_inject` and the pattern span (`pat_start`/`pat_len`, possibly 0 for
|
||||
* merge/clear) are filled. Returns true on success. */
|
||||
* merge/clear) are filled. Returns true on success.
|
||||
*
|
||||
* On failure `*bad_mod` is set to the offending modifier character when the
|
||||
* rule carried a modifier FastSync does not implement, and left '\0' for a
|
||||
* generic syntax error so callers can emit a precise diagnostic. */
|
||||
static bool parse_rule_syntax(const char* text, RuleKind* kind, unsigned* sides,
|
||||
bool* sides_explicit, bool* negate, bool* anchored_mod,
|
||||
bool* perishable, bool* xattr, bool* cvs_inject,
|
||||
const char** pat_start, size_t* pat_len) {
|
||||
const char** pat_start, size_t* pat_len, char* bad_mod) {
|
||||
const char* p = text;
|
||||
*sides = FILTER_SIDE_SENDER | FILTER_SIDE_RECEIVER;
|
||||
*sides_explicit = false;
|
||||
@@ -190,6 +240,7 @@ static bool parse_rule_syntax(const char* text, RuleKind* kind, unsigned* sides,
|
||||
*cvs_inject = false;
|
||||
*pat_start = NULL;
|
||||
*pat_len = 0;
|
||||
*bad_mod = '\0';
|
||||
|
||||
bool is_short = false;
|
||||
if (short_rule_char(*p, kind)) {
|
||||
@@ -210,17 +261,25 @@ static bool parse_rule_syntax(const char* text, RuleKind* kind, unsigned* sides,
|
||||
/* Modifiers: long names require a comma; short names may attach directly.
|
||||
Only commit a modifier run that terminates at a separator or the end, so a
|
||||
pattern such as "*.tmp" written as "-*.tmp" is not mistaken for modifiers. */
|
||||
if (*p == ',') {
|
||||
*bad_mod = unsupported_modifier_in_token(p + 1, *kind);
|
||||
} else if (is_short) {
|
||||
*bad_mod = unsupported_modifier_in_token(p, *kind);
|
||||
}
|
||||
if (*bad_mod != '\0')
|
||||
return false;
|
||||
|
||||
const char* mod_start = p;
|
||||
const char* mod_end = p;
|
||||
if (*p == ',') {
|
||||
p++;
|
||||
mod_start = p;
|
||||
while (is_modifier_char(*p))
|
||||
while (is_consumed_modifier_char(*p, *kind))
|
||||
p++;
|
||||
mod_end = p;
|
||||
} else if (is_short) {
|
||||
const char* scan = p;
|
||||
while (is_modifier_char(*scan))
|
||||
while (is_consumed_modifier_char(*scan, *kind))
|
||||
scan++;
|
||||
if (*scan == '\0' || *scan == ' ' || *scan == '_') {
|
||||
mod_start = p;
|
||||
@@ -290,9 +349,13 @@ FilterRule* filter_rule_parse(const char* line, const FilterParseOptions* opts,
|
||||
bool sides_explicit, negate, anchored_mod, perishable, xattr, cvs_inject;
|
||||
const char* pat;
|
||||
size_t pat_len;
|
||||
char bad_mod;
|
||||
if (!parse_rule_syntax(p, &kind, &sides, &sides_explicit, &negate, &anchored_mod, &perishable,
|
||||
&xattr, &cvs_inject, &pat, &pat_len)) {
|
||||
filter_set_error(err, err_size, "unrecognized filter rule syntax");
|
||||
&xattr, &cvs_inject, &pat, &pat_len, &bad_mod)) {
|
||||
if (bad_mod != '\0')
|
||||
filter_set_error(err, err_size, "unsupported filter modifier '%c'", bad_mod);
|
||||
else
|
||||
filter_set_error(err, err_size, "unrecognized filter rule syntax");
|
||||
return NULL;
|
||||
}
|
||||
if (cvs_inject) {
|
||||
@@ -401,7 +464,6 @@ FilterRule* filter_rule_parse(const char* line, const FilterParseOptions* opts,
|
||||
rule->dir_only = dir_only;
|
||||
rule->negate = negate;
|
||||
rule->perishable = perishable;
|
||||
(void)xattr; /* xattr-name rules never match file/dir names; accepted/ignored */
|
||||
return rule;
|
||||
}
|
||||
|
||||
@@ -530,16 +592,27 @@ static bool filter_list_parse_append_depth(FilterRuleList* list, const char* lin
|
||||
bool sides_explicit, negate, anchored_mod, perishable, xattr, cvs_inject;
|
||||
const char* pat;
|
||||
size_t pat_len;
|
||||
char bad_mod;
|
||||
if (!parse_rule_syntax(p, &kind, &sides, &sides_explicit, &negate, &anchored_mod, &perishable,
|
||||
&xattr, &cvs_inject, &pat, &pat_len)) {
|
||||
filter_set_error(err, err_size, "unrecognized filter rule syntax: %s", p);
|
||||
&xattr, &cvs_inject, &pat, &pat_len, &bad_mod)) {
|
||||
if (bad_mod != '\0')
|
||||
filter_set_error(err, err_size, "unsupported filter modifier '%c': %s", bad_mod, p);
|
||||
else
|
||||
filter_set_error(err, err_size, "unrecognized filter rule syntax: %s", p);
|
||||
return false;
|
||||
}
|
||||
(void)sides_explicit;
|
||||
(void)negate;
|
||||
(void)anchored_mod;
|
||||
(void)perishable;
|
||||
(void)xattr;
|
||||
|
||||
/* xattr-name rules are not implemented; reject them everywhere (including on
|
||||
* merge/dir-merge, where the flag would otherwise be silently dropped) with
|
||||
* the same diagnostic the standalone parser gives. */
|
||||
if (xattr) {
|
||||
filter_set_error(err, err_size, "xattr-name filter rules (the x modifier) are not supported");
|
||||
return false;
|
||||
}
|
||||
|
||||
if (cvs_inject) {
|
||||
/* "C" injects the CVS defaults in place; no pattern is expected. */
|
||||
@@ -811,8 +884,3 @@ FilterAction filter_rules_apply_side(const FilterRuleList* list, const char* rel
|
||||
}
|
||||
return FILTER_ACTION_NONE;
|
||||
}
|
||||
|
||||
FilterAction filter_rules_apply(const FilterRuleList* list, const char* rel_path, const char* leaf,
|
||||
bool is_dir) {
|
||||
return filter_rules_apply_side(list, rel_path, leaf, is_dir, FILTER_SIDE_SENDER);
|
||||
}
|
||||
|
||||
+8
-7
@@ -23,7 +23,13 @@
|
||||
* dir-merge/: per-directory merge file (registered for the scanner)
|
||||
* clear/! clear the current rule list (takes no argument)
|
||||
* Modifiers: '/' absolute anchor, '!' negate match, 'C' inject CVS defaults,
|
||||
* 's' sender side, 'r' receiver side, 'p' perishable, 'x' xattr name rule.
|
||||
* 's' sender side, 'r' receiver side, 'p' perishable. The rsync 'x'
|
||||
* (xattr-name) modifier is not implemented and is rejected explicitly
|
||||
* everywhere. The merge-file modifiers 'e' (exclude the merge file itself),
|
||||
* 'n' (do not inherit the merge file), 'w' (word-split the merge file) and '-'
|
||||
* (do not transfer the merge file) are accepted and consumed only on merge/
|
||||
* dir-merge rules (rejected on every other rule, matching rsync); their
|
||||
* semantics are not implemented and they are otherwise ignored.
|
||||
* A trailing '/' makes a pattern match directories only. A leading '/' anchors
|
||||
* the pattern to its owner directory.
|
||||
*/
|
||||
@@ -53,7 +59,7 @@ typedef struct {
|
||||
char* pattern; /* cleaned glob pattern (no leading '/', no trailing '/') */
|
||||
} FilterRule;
|
||||
|
||||
typedef struct {
|
||||
typedef struct FilterRuleList {
|
||||
FilterRule** items; /* owned array of rule pointers */
|
||||
int count;
|
||||
int capacity;
|
||||
@@ -129,9 +135,4 @@ FilterRuleList* filter_file_read(const char* dir_path, const char* owner_rel, bo
|
||||
FilterAction filter_rules_apply_side(const FilterRuleList* list, const char* rel_path,
|
||||
const char* leaf, bool is_dir, unsigned side);
|
||||
|
||||
/* Sender-side convenience wrapper (kept for callers/tests that only need the
|
||||
* transfer decision). */
|
||||
FilterAction filter_rules_apply(const FilterRuleList* list, const char* rel_path, const char* leaf,
|
||||
bool is_dir);
|
||||
|
||||
#endif
|
||||
|
||||
+15
-13
@@ -105,25 +105,27 @@ bool format_dest_state_receive(int fd, OutputDestState* state) {
|
||||
bool format_stats_send(int fd, const ReceiverStats* stats) {
|
||||
if (!stats)
|
||||
return false;
|
||||
unsigned long long matched = stats->matched_data;
|
||||
unsigned long long deleted = stats->deleted_files;
|
||||
unsigned long long would = stats->would_delete_count;
|
||||
return send_n_data(fd, &matched, sizeof(matched)) && send_n_data(fd, &deleted, sizeof(deleted)) &&
|
||||
send_n_data(fd, &would, sizeof(would));
|
||||
unsigned long long fields[8] = {
|
||||
stats->matched_data, stats->deleted_files, stats->would_delete_count, stats->literal_bytes,
|
||||
stats->created_reg, stats->created_dir, stats->created_link, stats->created_special,
|
||||
};
|
||||
return send_n_data(fd, fields, sizeof(fields));
|
||||
}
|
||||
|
||||
bool format_stats_receive(int fd, ReceiverStats* stats) {
|
||||
if (!stats)
|
||||
return false;
|
||||
unsigned long long matched = 0;
|
||||
unsigned long long deleted = 0;
|
||||
unsigned long long would = 0;
|
||||
if (!receive_n_data(fd, &matched, sizeof(matched)) ||
|
||||
!receive_n_data(fd, &deleted, sizeof(deleted)) || !receive_n_data(fd, &would, sizeof(would)))
|
||||
unsigned long long fields[8] = {0};
|
||||
if (!receive_n_data(fd, fields, sizeof(fields)))
|
||||
return false;
|
||||
memset(stats, 0, sizeof(*stats));
|
||||
stats->matched_data = matched;
|
||||
stats->deleted_files = deleted;
|
||||
stats->would_delete_count = would;
|
||||
stats->matched_data = fields[0];
|
||||
stats->deleted_files = fields[1];
|
||||
stats->would_delete_count = fields[2];
|
||||
stats->literal_bytes = fields[3];
|
||||
stats->created_reg = fields[4];
|
||||
stats->created_dir = fields[5];
|
||||
stats->created_link = fields[6];
|
||||
stats->created_special = fields[7];
|
||||
return true;
|
||||
}
|
||||
|
||||
+39
-4
@@ -57,14 +57,26 @@ bool format_dest_state_send(int fd, const OutputDestState* state);
|
||||
bool format_dest_state_receive(int fd, OutputDestState* state);
|
||||
|
||||
/* End-of-transfer receiver counters reported through STATUS_STATS (protocol
|
||||
* 2.25.0) when the wire config carries report_stats. `would_delete_count` is
|
||||
* the number of destination-relative paths the receiver would have deleted in a
|
||||
* -n/--dry-run --delete run; that many wire strings immediately follow the
|
||||
* fixed record (sent/read by the caller). */
|
||||
* 2.25.0, extended in 2.28.0) when the wire config carries report_stats.
|
||||
* `would_delete_count` is the number of destination-relative paths the receiver
|
||||
* would have deleted in a -n/--dry-run --delete run; that many wire strings
|
||||
* immediately follow the fixed record (sent/read by the caller).
|
||||
*
|
||||
* Protocol 2.28.0 adds the receiver-observed counters the sender cannot see:
|
||||
* `literal_bytes` is the file data the receiver actually stored literally
|
||||
* (whole files plus the literal fragments of a delta) and the four `created_*`
|
||||
* counters split the destination entries the receiver newly created by type,
|
||||
* reproducing rsync's `Number of created files` breakdown and an exact
|
||||
* `Literal data` for a delta run. */
|
||||
typedef struct {
|
||||
unsigned long long matched_data;
|
||||
unsigned long long deleted_files;
|
||||
unsigned long long would_delete_count;
|
||||
unsigned long long literal_bytes;
|
||||
unsigned long long created_reg;
|
||||
unsigned long long created_dir;
|
||||
unsigned long long created_link;
|
||||
unsigned long long created_special;
|
||||
} ReceiverStats;
|
||||
|
||||
/* Fixed-width STATUS_STATS counter record. The status frame and the optional
|
||||
@@ -73,4 +85,27 @@ typedef struct {
|
||||
bool format_stats_send(int fd, const ReceiverStats* stats);
|
||||
bool format_stats_receive(int fd, ReceiverStats* stats);
|
||||
|
||||
/* Sender-side file-list accounting for rsync's `--stats` block. Filled while
|
||||
* the scan/send loops walk each entry: the flist counters describe every
|
||||
* scanned source entry (transferred or skipped), while the transferred/literal
|
||||
* counters describe only the regular files the receiver actually stored. The
|
||||
* type split lets the client print rsync's `Number of files` breakdown; the
|
||||
* receiver-only counters (matched data, deleted, created) come from
|
||||
* STATUS_STATS. */
|
||||
typedef struct {
|
||||
unsigned long long flist_reg;
|
||||
unsigned long long flist_dir;
|
||||
unsigned long long flist_link;
|
||||
unsigned long long flist_special;
|
||||
unsigned long long total_file_size; /* sum of entry sizes (link target len) */
|
||||
unsigned long long transferred_regular; /* regular files actually stored */
|
||||
unsigned long long transferred_file_size; /* source size of those files */
|
||||
/* Whole-file accuracy: the `--stats` "Literal data" row. The sender counts
|
||||
* the source size of every stored file, so a whole-file transfer matches
|
||||
* rsync. A delta run actually ships only the literal fragments of the diff
|
||||
* (the rest is matched/copied), so here the value is an upper bound, not
|
||||
* rsync's literal-byte total; see RSYNC_COMPAT.md's `--stats` row. */
|
||||
unsigned long long literal_data;
|
||||
} TransferStats;
|
||||
|
||||
#endif
|
||||
|
||||
+158
-11
@@ -3,6 +3,7 @@
|
||||
#include "utils.h"
|
||||
#include <errno.h>
|
||||
#include <fcntl.h>
|
||||
#include <fnmatch.h>
|
||||
#include <grp.h>
|
||||
#include <limits.h>
|
||||
#include <pwd.h>
|
||||
@@ -12,6 +13,8 @@
|
||||
#include <sys/stat.h>
|
||||
#include <unistd.h>
|
||||
|
||||
static bool identity_id_fits_int32(unsigned long id);
|
||||
|
||||
/* The active identity snapshot lives in a per-process global. The TCP server
|
||||
* forks one child process per connection, so a connection never shares this
|
||||
* with another; within a connection the multithreaded receiver reads it without
|
||||
@@ -419,17 +422,10 @@ static int identity_parse_from(const char* token, bool is_group, int32_t* out_fr
|
||||
/* Not a numeric LOW-HIGH range: fall through and treat as a name (a
|
||||
* hyphenated account name like "wayne-smith" must still resolve). */
|
||||
}
|
||||
/* A sender-side name. A wildcard other than the bare '*' is matched by rsync
|
||||
* against the sender's names; because FastSync transmits numeric ids only, the
|
||||
* receiver cannot evaluate it, so reject rather than silently mis-match. */
|
||||
if (identity_token_has_glob(token)) {
|
||||
log_message(LOG_LEVEL_ERROR,
|
||||
"%smap FROM '%s': name wildcards other than '*' are not supported "
|
||||
"(FastSync transmits numeric ids, so sender names are unavailable on the "
|
||||
"receiver)",
|
||||
is_group ? "--group" : "--user", token);
|
||||
return -1;
|
||||
}
|
||||
/* A sender-side name. A FROM name wildcard other than the bare '*' is handled
|
||||
* by identity_expand_from_glob() in the caller (it expands against the
|
||||
* sender's account database at CLI-parse time), so this function only sees the
|
||||
* bare '*' or a literal name here. */
|
||||
int32_t id;
|
||||
if (identity_resolve_token(token, is_group, &id) != 0)
|
||||
return -1;
|
||||
@@ -486,6 +482,138 @@ static int identity_append_rule(IdentityMap** map, int* count, const IdentityMap
|
||||
return 0;
|
||||
}
|
||||
|
||||
/* True when `lo` and `hi` are adjacent ids (no overflow at INT32_MAX). */
|
||||
static bool identity_ids_adjacent(int32_t lo, int32_t hi) {
|
||||
return lo < INT32_MAX && hi == lo + 1;
|
||||
}
|
||||
|
||||
static int identity_id_cmp(const void* a, const void* b) {
|
||||
int32_t x = *(const int32_t*)a;
|
||||
int32_t y = *(const int32_t*)b;
|
||||
return (x > y) - (x < y);
|
||||
}
|
||||
|
||||
static bool identity_ids_push(int32_t** ids, size_t* count, size_t* cap, int32_t id) {
|
||||
if (*count == *cap) {
|
||||
size_t grown_cap = *cap ? *cap * 2 : 16;
|
||||
int32_t* grown = realloc(*ids, grown_cap * sizeof(int32_t));
|
||||
if (!grown)
|
||||
return false;
|
||||
*ids = grown;
|
||||
*cap = grown_cap;
|
||||
}
|
||||
(*ids)[(*count)++] = id;
|
||||
return true;
|
||||
}
|
||||
|
||||
/* Expand a FROM name wildcard (rsync's match against sender-side account names)
|
||||
* into one rule per contiguous run of matching numeric ids, all sharing the same
|
||||
* TO side. FastSync transmits numeric ids only, so the wildcard must be
|
||||
* resolved here -- at CLI-parse time -- against the SENDER's passwd/group
|
||||
* database; the receiver has no sender names to match. Contiguous matched ids
|
||||
* are collapsed into a single LOW-HIGH range (a range of adjacent ids contains
|
||||
* exactly the ids it spans, so this is semantically exact). Returns 0 on
|
||||
* success, -1 on an allocation failure, a wildcard that matches no sender
|
||||
* account, or an expansion that would push the map past MAX_IDENTITY_MAP. */
|
||||
static int identity_expand_from_glob(Config* config, const char* glob, bool is_group,
|
||||
const IdentityMap* to_rule) {
|
||||
const char* optname = is_group ? "--groupmap" : "--usermap";
|
||||
size_t cap = 0;
|
||||
size_t n = 0;
|
||||
int32_t* ids = NULL;
|
||||
bool alloc_failed = false;
|
||||
|
||||
if (is_group) {
|
||||
setgrent();
|
||||
struct group* gr;
|
||||
while ((gr = getgrent()) != NULL) {
|
||||
if (fnmatch(glob, gr->gr_name, 0) != 0)
|
||||
continue;
|
||||
if (!identity_id_fits_int32((unsigned long)gr->gr_gid))
|
||||
continue;
|
||||
if (!identity_ids_push(&ids, &n, &cap, (int32_t)gr->gr_gid)) {
|
||||
alloc_failed = true;
|
||||
break;
|
||||
}
|
||||
}
|
||||
endgrent();
|
||||
} else {
|
||||
setpwent();
|
||||
struct passwd* pw;
|
||||
while ((pw = getpwent()) != NULL) {
|
||||
if (fnmatch(glob, pw->pw_name, 0) != 0)
|
||||
continue;
|
||||
if (!identity_id_fits_int32((unsigned long)pw->pw_uid))
|
||||
continue;
|
||||
if (!identity_ids_push(&ids, &n, &cap, (int32_t)pw->pw_uid)) {
|
||||
alloc_failed = true;
|
||||
break;
|
||||
}
|
||||
}
|
||||
endpwent();
|
||||
}
|
||||
|
||||
if (alloc_failed) {
|
||||
free(ids);
|
||||
log_message(LOG_LEVEL_ERROR, "%s: memory allocation failed expanding FROM '%s'", optname, glob);
|
||||
return -1;
|
||||
}
|
||||
if (n == 0) {
|
||||
free(ids);
|
||||
log_message(LOG_LEVEL_ERROR, "%s FROM '%s': no source account name matches the wildcard",
|
||||
optname, glob);
|
||||
return -1;
|
||||
}
|
||||
|
||||
qsort(ids, n, sizeof(int32_t), identity_id_cmp);
|
||||
size_t unique = 0;
|
||||
for (size_t i = 0; i < n; i++) {
|
||||
if (unique == 0 || ids[unique - 1] != ids[i])
|
||||
ids[unique++] = ids[i];
|
||||
}
|
||||
n = unique;
|
||||
|
||||
int runs = 0;
|
||||
for (size_t i = 0; i < n; i++) {
|
||||
if (i == 0 || !identity_ids_adjacent(ids[i - 1], ids[i]))
|
||||
runs++;
|
||||
}
|
||||
|
||||
IdentityMap** map = is_group ? &config->groupmap : &config->usermap;
|
||||
int* count = is_group ? &config->groupmap_count : &config->usermap_count;
|
||||
if (*count > MAX_IDENTITY_MAP - runs) {
|
||||
log_message(LOG_LEVEL_ERROR,
|
||||
"%s FROM '%s': the name wildcard expands to %d rule(s), which would exceed "
|
||||
"the maximum of %d map rules",
|
||||
optname, glob, runs, MAX_IDENTITY_MAP);
|
||||
free(ids);
|
||||
return -1;
|
||||
}
|
||||
|
||||
for (size_t i = 0; i < n;) {
|
||||
size_t j = i;
|
||||
while (j + 1 < n && identity_ids_adjacent(ids[j], ids[j + 1]))
|
||||
j++;
|
||||
IdentityMap rule;
|
||||
rule.from = ids[i];
|
||||
rule.from_hi = ids[j];
|
||||
rule.to = to_rule->to;
|
||||
rule.to_name = to_rule->to_name ? str_dup(to_rule->to_name) : NULL;
|
||||
if (to_rule->to_name && !rule.to_name) {
|
||||
free(ids);
|
||||
return -1;
|
||||
}
|
||||
if (identity_append_rule(map, count, &rule) != 0) {
|
||||
free(rule.to_name);
|
||||
free(ids);
|
||||
return -1;
|
||||
}
|
||||
i = j + 1;
|
||||
}
|
||||
free(ids);
|
||||
return 0;
|
||||
}
|
||||
|
||||
int identity_parse_map(Config* config, const char* value, bool is_group) {
|
||||
if (!config || !value || *value == '\0') {
|
||||
log_message(LOG_LEVEL_ERROR, "%smap requires a value", is_group ? "--group" : "--user");
|
||||
@@ -509,6 +637,25 @@ int identity_parse_map(Config* config, const char* value, bool is_group) {
|
||||
char* to_token = colon + 1;
|
||||
IdentityMap parsed;
|
||||
memset(&parsed, 0, sizeof(parsed));
|
||||
/* A FROM name wildcard (anything with a glob metacharacter other than the
|
||||
* bare '*') is expanded against the sender's account database here, while
|
||||
* the sender's passwd/group DB is still available; the resulting numeric
|
||||
* rules travel on the wire like an explicit list. The TO side is parsed
|
||||
* first so every expanded rule shares it. */
|
||||
if (strcmp(from_token, "*") != 0 && identity_token_has_glob(from_token)) {
|
||||
if (identity_parse_to(to_token, is_group, &parsed.to, &parsed.to_name) != 0) {
|
||||
log_message(LOG_LEVEL_ERROR, "%s could not parse TO '%s' in '%s'", optname, to_token,
|
||||
value);
|
||||
free(list);
|
||||
return -1;
|
||||
}
|
||||
if (identity_expand_from_glob(config, from_token, is_group, &parsed) != 0) {
|
||||
free(parsed.to_name);
|
||||
free(list);
|
||||
return -1;
|
||||
}
|
||||
continue;
|
||||
}
|
||||
if (identity_parse_from(from_token, is_group, &parsed.from, &parsed.from_hi) != 0) {
|
||||
log_message(LOG_LEVEL_ERROR,
|
||||
"%s could not resolve FROM '%s' in '%s' (a name must exist on the "
|
||||
|
||||
@@ -25,8 +25,14 @@
|
||||
|
||||
/* Parse one --usermap= / --groupmap= value (comma-separated FROM:TO rules,
|
||||
* first match wins) into config->usermap / config->groupmap. is_group selects
|
||||
* the group tables and name databases. Returns 0 on success, -1 on a
|
||||
* malformed spec or an unresolvable name (never a silent no-op). */
|
||||
* the group tables and name databases. A FROM name wildcard (containing `*`,
|
||||
* `?` or `[...]`, but not the bare `*`) is expanded against the SENDER's
|
||||
* account database at parse time into one or more numeric id/range rules
|
||||
* (contiguous ids collapse to a range) sharing the same TO, because only
|
||||
* numeric ids cross the wire; the expansion is capped at MAX_IDENTITY_MAP and a
|
||||
* wildcard matching no account is an error. Returns 0 on success, -1 on a
|
||||
* malformed spec, an unresolvable name, an unmatched wildcard, or a map that
|
||||
* would exceed MAX_IDENTITY_MAP (never a silent no-op). */
|
||||
int identity_parse_map(Config* config, const char* value, bool is_group);
|
||||
|
||||
/* Parse --chown=USER:GROUP. Supports USER:GROUP, USER (owner only), :GROUP
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,41 @@
|
||||
#ifndef INCREMENTAL_CHECK_H
|
||||
#define INCREMENTAL_CHECK_H
|
||||
|
||||
#include "config.h"
|
||||
#include "file_types.h"
|
||||
#include "protocol.h"
|
||||
#include <stdbool.h>
|
||||
|
||||
/* Incremental-check module: the per-file STATUS_CHECK state machine, the
|
||||
* incremental delta / alternate-basis / fuzzy matching helpers and the shared
|
||||
* xattr receive helper. These declarations are re-exported by the
|
||||
* file_receive.h facade. */
|
||||
|
||||
/* Whole-file payload bound shared by the plain receive path and the
|
||||
* incremental check paths. */
|
||||
#define MAX_FILE_DATA_SIZE MAX_RECEIVE_WHOLE_FILE_SIZE
|
||||
|
||||
/* Receive a file's xattr block (when the config enables xattr transport) and
|
||||
* attach it to `file`. Returns false on a malformed/oversized frame. */
|
||||
bool receive_file_xattrs(File* file, int fd, const Config* config);
|
||||
|
||||
File* receive_incremental_check(int fd, const Config* config, bool* skipped);
|
||||
/* Extended variant used by the receiver. `would_transfer` (may be NULL) is set
|
||||
* true only on the server-contacting --dry-run path when the file is not up to
|
||||
* date: the receiver has already sent STATUS_DRY_RUN_TRANSFER and returns NULL
|
||||
* without storing anything. On that path `*skipped` is true for an up-to-date
|
||||
* (STATUS_OK) file and both flags are false for a genuine error. */
|
||||
File* receive_incremental_check_ex(int fd, const Config* config, bool* skipped,
|
||||
bool* would_transfer);
|
||||
|
||||
/* Testable basis quick-check / verification policy. file_basis_quick_match is
|
||||
* rsync's metadata quick-check for a basis candidate (equal size is required
|
||||
* separately by the caller; this adds the --size-only / mtime / --modify-window
|
||||
* leg). file_basis_content_required reports whether a hit must ALSO be
|
||||
* confirmed by a whole-file content digest (--verify-basis; false is the
|
||||
* default rsync-parity behavior). */
|
||||
bool file_basis_quick_match(const Config* config, const struct stat* st, time_t check_mtime,
|
||||
long check_mtime_nsec);
|
||||
bool file_basis_content_required(const Config* config);
|
||||
|
||||
#endif
|
||||
+49
-5
@@ -5,6 +5,15 @@
|
||||
#include <stdbool.h>
|
||||
#include <stdint.h>
|
||||
|
||||
/* Ask the compiler to type-check the printf-style arguments of the variadic
|
||||
* logging helpers. Only enabled for GNU-compatible compilers (gcc/clang). */
|
||||
#if defined(__GNUC__)
|
||||
#define LOG_PRINTF_ATTR(fmt_idx, first_vararg_idx) \
|
||||
__attribute__((format(printf, fmt_idx, first_vararg_idx)))
|
||||
#else
|
||||
#define LOG_PRINTF_ATTR(fmt_idx, first_vararg_idx)
|
||||
#endif
|
||||
|
||||
typedef enum { LOG_LEVEL_DEBUG, LOG_LEVEL_INFO, LOG_LEVEL_WARNING, LOG_LEVEL_ERROR } LogLevel;
|
||||
typedef enum { LOG_STDERR_ERRORS, LOG_STDERR_ALL } LogStderrMode;
|
||||
|
||||
@@ -13,7 +22,19 @@ typedef enum {
|
||||
LOG_DEBUG_PROTO = 1u << 1,
|
||||
LOG_DEBUG_PACK = 1u << 2,
|
||||
LOG_DEBUG_UTIL = 1u << 3,
|
||||
LOG_DEBUG_ALL = (1u << 4) - 1,
|
||||
/* rsync --debug categories that now map to a natural FastSync event:
|
||||
* flist (file-list scan progress), del (deletions), hash/deltasum
|
||||
* (whole-file hashing and delta-sum generation), recv (receiver
|
||||
* responses/signatures), filter (selection/exclusion decisions) and send
|
||||
* (files handed to the sender). Only emitted when the category is
|
||||
* explicitly enabled; a normal run stays silent. */
|
||||
LOG_DEBUG_FLIST = 1u << 4,
|
||||
LOG_DEBUG_DEL = 1u << 5,
|
||||
LOG_DEBUG_HASH = 1u << 6,
|
||||
LOG_DEBUG_RECV = 1u << 7,
|
||||
LOG_DEBUG_FILTER = 1u << 8,
|
||||
LOG_DEBUG_SEND = 1u << 9,
|
||||
LOG_DEBUG_ALL = (1u << 10) - 1,
|
||||
} LogDebugFlag;
|
||||
|
||||
typedef enum {
|
||||
@@ -21,10 +42,33 @@ typedef enum {
|
||||
LOG_INFO_MISC = 1u << 1,
|
||||
LOG_INFO_SKIP = 1u << 2,
|
||||
LOG_INFO_STATS = 1u << 3,
|
||||
LOG_INFO_ALL = LOG_INFO_COPY | LOG_INFO_MISC | LOG_INFO_SKIP | LOG_INFO_STATS,
|
||||
/* rsync categories that map to a FastSync event (emitted in rsync's line
|
||||
* format): del (deletions), remove (sender-side source removal), name
|
||||
* (transferred entry names), flist (file-list header), nonreg (skipped
|
||||
* non-regular files), progress (per-file progress). rsync's `backup`
|
||||
* category is accepted for CLI parity but stays silent: the receiver does the
|
||||
* backing-up and FastSync has no backup event to report from the sender. */
|
||||
LOG_INFO_DEL = 1u << 4,
|
||||
LOG_INFO_REMOVE = 1u << 5,
|
||||
LOG_INFO_NAME = 1u << 6,
|
||||
LOG_INFO_FLIST = 1u << 7,
|
||||
LOG_INFO_NONREG = 1u << 8,
|
||||
LOG_INFO_PROGRESS = 1u << 9,
|
||||
/* Marker for `--info=name2` and higher: also print rsync's
|
||||
"NAME is uptodate" line for entries the receiver already has. It rides in
|
||||
the info_level bitset (there is no separate Config field) and is never set
|
||||
by --info=all (which selects level 1). */
|
||||
LOG_INFO_NAME_UPTODATE = 1u << 10,
|
||||
/* --info=mount: print rsync's `[sender] skipping mount-point dir NAME` when
|
||||
* -xx/--one-file-system drops a mount-point directory (FastSync's client is
|
||||
* the sender). */
|
||||
LOG_INFO_MOUNT = 1u << 11,
|
||||
LOG_INFO_ALL = LOG_INFO_COPY | LOG_INFO_MISC | LOG_INFO_SKIP | LOG_INFO_STATS | LOG_INFO_DEL |
|
||||
LOG_INFO_REMOVE | LOG_INFO_NAME | LOG_INFO_FLIST | LOG_INFO_NONREG |
|
||||
LOG_INFO_PROGRESS | LOG_INFO_MOUNT,
|
||||
} LogInfoFlag;
|
||||
|
||||
void log_message(LogLevel log_level, const char* message, ...);
|
||||
void log_message(LogLevel log_level, const char* message, ...) LOG_PRINTF_ATTR(2, 3);
|
||||
void log_perror(const char* context);
|
||||
void set_log_level(LogLevel level);
|
||||
void set_log_debug_flags(uint32_t flags);
|
||||
@@ -33,10 +77,10 @@ uint32_t get_log_debug_flags(void);
|
||||
* the debug log level is enabled AND the flag is selected. Hot paths use this
|
||||
* to skip expensive message formatting/escaping when the line is filtered. */
|
||||
bool log_debug_enabled(LogDebugFlag flag);
|
||||
void log_debug_message(LogDebugFlag flag, const char* message, ...);
|
||||
void log_debug_message(LogDebugFlag flag, const char* message, ...) LOG_PRINTF_ATTR(2, 3);
|
||||
void set_log_info_flags(uint32_t flags);
|
||||
uint32_t get_log_info_flags(void);
|
||||
void log_info_message(LogInfoFlag flag, const char* message, ...);
|
||||
void log_info_message(LogInfoFlag flag, const char* message, ...) LOG_PRINTF_ATTR(2, 3);
|
||||
void log_set_file(FILE* fp);
|
||||
void log_set_8_bit_output(bool enabled);
|
||||
bool log_get_8_bit_output(void);
|
||||
|
||||
+20
-5
@@ -211,13 +211,22 @@ FileMetadata* metadata_receive(int file_descriptor, int* ok) {
|
||||
|
||||
bool metadata_mode_for_policy(mode_t source_mode, mode_t current_mode, FileAttrPolicy policy,
|
||||
mode_t* out_mode) {
|
||||
const mode_t special_bits = (mode_t)(S_ISUID | S_ISGID | S_ISVTX);
|
||||
const mode_t execute_bits = S_IXUSR | S_IXGRP | S_IXOTH;
|
||||
if (policy.perms) {
|
||||
/* rsync --perms copies the source's permission and special bits exactly,
|
||||
* including group/other write and setuid/setgid/sticky. The kernel may
|
||||
* still clear setgid when the receiver is not in the file's group; the
|
||||
* caller logs a failed chmod rather than silently masking the bits here. */
|
||||
*out_mode = source_mode & (mode_t)(S_ISUID | S_ISGID | S_ISVTX | 0777);
|
||||
* caller logs a failed chmod rather than silently masking the bits here.
|
||||
* Setuid/setgid/sticky are super-user activities: when the connection did
|
||||
* not permit them (SUPER_MODE_OFF / --no-super) they are stripped, so a
|
||||
* client can never install a privileged bit on a receiver that forbade
|
||||
* super-user activities. This also covers bits introduced by --chmod,
|
||||
* whose result is fed in as source_mode. */
|
||||
mode_t bits = source_mode & (mode_t)(special_bits | 0777);
|
||||
if (!policy.super_permitted)
|
||||
bits &= ~special_bits;
|
||||
*out_mode = bits;
|
||||
return true;
|
||||
}
|
||||
if (policy.executability) {
|
||||
@@ -227,8 +236,11 @@ bool metadata_mode_for_policy(mode_t source_mode, mode_t current_mode, FileAttrP
|
||||
* execute); otherwise clear every execute bit. This runs on the
|
||||
* destination-derived base (pre-existing dest mode, or source&~umask for a
|
||||
* new file), and leaves the special bits untouched. --perms wins when both
|
||||
* are set (handled above). */
|
||||
mode_t base = current_mode & (mode_t)(S_ISUID | S_ISGID | S_ISVTX | 0777);
|
||||
* are set (handled above). The destination's own special bits survive
|
||||
* unless super-user activities are forbidden. */
|
||||
mode_t base = current_mode & (mode_t)(special_bits | 0777);
|
||||
if (!policy.super_permitted)
|
||||
base &= ~special_bits;
|
||||
if (source_mode & 0111)
|
||||
*out_mode = base | ((base & 0444) >> 2);
|
||||
else
|
||||
@@ -240,12 +252,13 @@ bool metadata_mode_for_policy(mode_t source_mode, mode_t current_mode, FileAttrP
|
||||
}
|
||||
|
||||
FileAttrPolicy file_attr_policy_from_config(const Config* config) {
|
||||
FileAttrPolicy policy = {false, false, false, false};
|
||||
FileAttrPolicy policy = {0};
|
||||
if (config) {
|
||||
policy.perms = config->preserve_perms;
|
||||
policy.times = config->preserve_times;
|
||||
policy.atimes = config->preserve_atimes;
|
||||
policy.executability = config->use_executability;
|
||||
policy.super_permitted = privilege_super_mode_permitted(config->super_mode);
|
||||
}
|
||||
return policy;
|
||||
}
|
||||
@@ -314,6 +327,8 @@ bool file_restore_symlink_metadata(const char* path, const FileMetadata* metadat
|
||||
transfer never fails over it. */
|
||||
if (policy.perms) {
|
||||
mode_t link_mode = metadata->mode & (mode_t)(S_ISUID | S_ISGID | S_ISVTX | 0777);
|
||||
if (!policy.super_permitted)
|
||||
link_mode &= ~(mode_t)(S_ISUID | S_ISGID | S_ISVTX);
|
||||
if (fchmodat(parent_fd, leaf, link_mode, AT_SYMLINK_NOFOLLOW) != 0 && errno != EOPNOTSUPP &&
|
||||
errno != ENOTSUP && errno != ENOSYS) {
|
||||
log_message(LOG_LEVEL_DEBUG, "Could not set symlink mode on %s: %s", path, strerror(errno));
|
||||
|
||||
@@ -38,16 +38,19 @@ PipelineContextSender* pipeline_context_sender_create(Config* config, Queue* que
|
||||
context->remove_source_files = NULL;
|
||||
context->early_delete = false;
|
||||
context->delete_plans = NULL;
|
||||
context->delete_suppressed = false;
|
||||
context->scan_stopped_early = false;
|
||||
context->total_files = 0;
|
||||
context->progress_bytes = 0;
|
||||
context->total_bytes = 0;
|
||||
memset(&context->stats, 0, sizeof(context->stats));
|
||||
context->sender_done = false;
|
||||
atomic_init(&context->cancelled, false);
|
||||
protocol_session_init(&context->allocation_session, -1, -1);
|
||||
protocol_session_set_max_alloc(&context->allocation_session, config->max_alloc);
|
||||
context->dir_entries = NULL;
|
||||
context->dir_entries_mutex_init = false;
|
||||
atomic_init(&context->dir_count, 0);
|
||||
context->delete_limit = false;
|
||||
int init = 0;
|
||||
if (config->use_metadata) {
|
||||
@@ -83,11 +86,13 @@ PipelineContextSender* pipeline_context_sender_create(Config* config, Queue* que
|
||||
return context;
|
||||
|
||||
fail:
|
||||
log_perror("Error initializing synchronization objects");
|
||||
log_message(LOG_LEVEL_ERROR, "%s", "Error initializing synchronization objects");
|
||||
if (context->dir_entries_mutex_init)
|
||||
mtx_destroy(&context->dir_entries_mutex);
|
||||
if (context->dir_entries)
|
||||
array_list_delete(context->dir_entries);
|
||||
if (init >= 7)
|
||||
mtx_destroy(&context->mutex_progress);
|
||||
if (init >= 6)
|
||||
cnd_destroy(&context->condition_not_empty_loader);
|
||||
if (init >= 5)
|
||||
|
||||
@@ -9,6 +9,7 @@
|
||||
#include "config.h"
|
||||
#include "delete_plan.h"
|
||||
#include "file.h"
|
||||
#include "format.h"
|
||||
#include "protocol.h"
|
||||
#include "queue.h"
|
||||
#include "stop_condition.h"
|
||||
@@ -82,10 +83,22 @@ typedef struct {
|
||||
thread transmits the root plan before any data and the remaining plans
|
||||
alongside the chunks. Set once before the worker threads start. */
|
||||
DeletePlanSender* delete_plans;
|
||||
/* A scan I/O error without --ignore-errors suppressed deletion: the prebuilt
|
||||
keep-set/plans were dropped, and the streaming scanner must not build a
|
||||
fresh manifest or re-send the per-directory plans. Set once before the
|
||||
worker threads start. */
|
||||
bool delete_suppressed;
|
||||
mtx_t mutex_progress;
|
||||
int total_files;
|
||||
unsigned long long progress_bytes;
|
||||
unsigned long long total_bytes;
|
||||
/* Per-type flist / transferred accounting for the rsync --stats breakdown and
|
||||
the progress `to-chk` denominator. Owned by the sender thread: it is the
|
||||
only writer (the entry/transfer notes in send_chunks_multithreaded) and it
|
||||
reads the totals in its completion tail, so no lock is needed. This is NOT
|
||||
guarded by mutex_progress (which covers total_files/progress_bytes/
|
||||
total_bytes/sender_done). */
|
||||
TransferStats stats;
|
||||
bool sender_done;
|
||||
atomic_bool cancelled;
|
||||
ProtocolSession allocation_session;
|
||||
@@ -106,6 +119,10 @@ typedef struct {
|
||||
ArrayList* dir_entries;
|
||||
mtx_t dir_entries_mutex;
|
||||
bool dir_entries_mutex_init;
|
||||
/* --stats directory accounting for a `-r` scan (no directory metadata):
|
||||
shared by the parallel scanner workers, read by the sender thread once the
|
||||
scanner is done. See ScannerOptions.dir_count. */
|
||||
atomic_ullong dir_count;
|
||||
/* Set by the sender thread when the receiver reported a --max-delete-capped
|
||||
deletion (STATUS_DELETE_LIMIT): the transfer succeeded and the process must
|
||||
exit 25 like rsync. Read by the caller after the sender thread is joined. */
|
||||
|
||||
+69
-10
@@ -183,12 +183,28 @@ void io_set_bwlimit(unsigned long long bytes_per_sec) {
|
||||
mtx_unlock(&bw_mutex);
|
||||
}
|
||||
|
||||
unsigned long long io_get_bwlimit(void) {
|
||||
return global_bwlimit();
|
||||
}
|
||||
|
||||
/* rsync's throttle (io.c sleep_for_bwlimit) sleeps once its unslept debt
|
||||
* reaches ~100 ms of bandwidth, so its effective initial burst is about 0.1 s
|
||||
* worth of bytes, not a full second. FastSync models the same with a token
|
||||
* bucket whose capacity is bwlimit/10, so a throttled run paces like rsync
|
||||
* instead of sending a full second's worth up front. */
|
||||
static long long bw_burst_capacity(unsigned long long bwlimit) {
|
||||
if (bwlimit == 0)
|
||||
return 0;
|
||||
long long burst = (long long)(bwlimit / 10);
|
||||
return burst > 0 ? burst : 1;
|
||||
}
|
||||
|
||||
void protocol_session_set_bwlimit(ProtocolSession* session, unsigned long long bytes_per_sec) {
|
||||
if (!session)
|
||||
return;
|
||||
session->bwlimit =
|
||||
bytes_per_sec > (unsigned long long)LLONG_MAX ? (unsigned long long)LLONG_MAX : bytes_per_sec;
|
||||
session->bw_tokens = (long long)session->bwlimit;
|
||||
session->bw_tokens = bw_burst_capacity(session->bwlimit);
|
||||
struct timespec now;
|
||||
clock_gettime(CLOCK_MONOTONIC, &now);
|
||||
session->bw_last_refill_sec = now.tv_sec;
|
||||
@@ -222,8 +238,9 @@ static void bw_throttle_session(ProtocolSession* session, size_t bytes_written)
|
||||
|
||||
long long tokens_to_add = (long long)((double)session->bwlimit * elapsed_ns / 1000000000.0);
|
||||
session->bw_tokens += tokens_to_add;
|
||||
if (session->bw_tokens > (long long)session->bwlimit)
|
||||
session->bw_tokens = (long long)session->bwlimit;
|
||||
long long burst = bw_burst_capacity(session->bwlimit);
|
||||
if (session->bw_tokens > burst)
|
||||
session->bw_tokens = burst;
|
||||
|
||||
session->bw_tokens -= bytes_written;
|
||||
|
||||
@@ -234,7 +251,11 @@ static void bw_throttle_session(ProtocolSession* session, size_t bytes_written)
|
||||
poll(NULL, 0, (int)(deficit_us / 1000));
|
||||
else
|
||||
usleep((useconds_t)deficit_us);
|
||||
/* Reset the bucket AFTER the sleep: crediting the sleep duration as elapsed
|
||||
refill time would cancel half the throttle (the next call would see the
|
||||
whole sleep as refill and immediately grant a fresh burst). */
|
||||
session->bw_tokens = 0;
|
||||
clock_gettime(CLOCK_MONOTONIC, &now);
|
||||
session->bw_last_refill_sec = now.tv_sec;
|
||||
session->bw_last_refill_nsec = now.tv_nsec;
|
||||
}
|
||||
@@ -280,6 +301,17 @@ static ProtocolSession* legacy_session(int read_fd, int write_fd) {
|
||||
return &legacy_io_session;
|
||||
}
|
||||
|
||||
/* Pace an out-of-band write that bypassed protocol_send_n_data (the plaintext
|
||||
* sendfile fast path). The bound/legacy session is resolved exactly as the
|
||||
* preceding send_n_data(fd, ...) resolved it, so the same token-bucket state is
|
||||
* throttled and the TLS and plaintext transports share identical --bwlimit
|
||||
* semantics. Passing the wire fd (rather than -1) is essential: the sendfile
|
||||
* send left legacy_io_session.write_fd bound to it, so resolving with -1 would
|
||||
* mismatch, re-initialize the session and hand out a second first-call burst. */
|
||||
void protocol_throttle_bytes(int file_descriptor, size_t bytes) {
|
||||
bw_throttle_session(legacy_session(-1, file_descriptor), bytes);
|
||||
}
|
||||
|
||||
bool send_n_data(int file_descriptor, const void* data, size_t data_size) {
|
||||
return protocol_send_n_data(legacy_session(-1, file_descriptor), data, data_size);
|
||||
}
|
||||
@@ -361,7 +393,7 @@ bool protocol_send_n_data(ProtocolSession* session, const void* data, size_t dat
|
||||
if (session->ssl)
|
||||
wait_events = POLLOUT;
|
||||
}
|
||||
log_debug_message(LOG_DEBUG_IO, " Send n Data: %zu", total_bytes_send);
|
||||
log_debug_message(LOG_DEBUG_IO, " Send n Data: %zd", total_bytes_send);
|
||||
atomic_fetch_add(&io_bytes_written, (unsigned long long)total_bytes_send);
|
||||
return true;
|
||||
}
|
||||
@@ -418,12 +450,18 @@ static bool protocol_receive_n_data_until(ProtocolSession* session, void* data,
|
||||
}
|
||||
|
||||
ssize_t bytes_received;
|
||||
if (session->ssl)
|
||||
bytes_received = SSL_read(session->ssl, (char*)data + total_bytes_received,
|
||||
data_size - total_bytes_received);
|
||||
else
|
||||
if (session->ssl) {
|
||||
/* SSL_read takes an int length; clamp a >INT_MAX request into chunks
|
||||
* (mirrors the send path) so the size_t downcast can never truncate into
|
||||
* a negative/partial read. */
|
||||
size_t ssl_chunk = data_size - total_bytes_received > (size_t)INT_MAX
|
||||
? (size_t)INT_MAX
|
||||
: data_size - total_bytes_received;
|
||||
bytes_received = SSL_read(session->ssl, (char*)data + total_bytes_received, (int)ssl_chunk);
|
||||
} else {
|
||||
bytes_received =
|
||||
read(fd, (char*)data + total_bytes_received, data_size - total_bytes_received);
|
||||
}
|
||||
if (bytes_received <= 0) {
|
||||
if (session->ssl) {
|
||||
int ssl_err = SSL_get_error(session->ssl, (int)bytes_received);
|
||||
@@ -529,6 +567,15 @@ static const char* status_to_string(Status status) {
|
||||
}
|
||||
}
|
||||
|
||||
/* Reject a raw wire status outside the known enum range before it is handed to
|
||||
* callers, so an unknown/corrupt frame fails as a protocol error instead of
|
||||
* being silently interpreted as an unexpected-but-valid verdict. STATUS_OK is
|
||||
* the first enumerator and STATUS_STATS the last, so the range check accepts
|
||||
* every status the protocol defines. */
|
||||
static bool status_is_valid(Status status) {
|
||||
return status >= STATUS_OK && status <= STATUS_STATS;
|
||||
}
|
||||
|
||||
/* Shared string send/receive implementation. `redact` selects whether the
|
||||
* payload body is written to the LOG_DEBUG_PROTO debug log: daemon auth material
|
||||
* (the username and the proof/signature fields) sets it so a --verbose log never
|
||||
@@ -612,7 +659,7 @@ bool protocol_send_data(ProtocolSession* session, const Data* data) {
|
||||
return false;
|
||||
if (!protocol_send_n_data(session, data->data, data_size))
|
||||
return false;
|
||||
log_debug_message(LOG_DEBUG_PROTO, "Send %lld data", data_size);
|
||||
log_debug_message(LOG_DEBUG_PROTO, "Send %llu data", data_size);
|
||||
return true;
|
||||
}
|
||||
|
||||
@@ -646,7 +693,7 @@ Data* protocol_receive_data_limited(ProtocolSession* session, unsigned long long
|
||||
protocol_release_memory_for_session(session, allocation_size);
|
||||
return NULL;
|
||||
}
|
||||
log_debug_message(LOG_DEBUG_PROTO, "Received %lld data", size);
|
||||
log_debug_message(LOG_DEBUG_PROTO, "Received %llu data", size);
|
||||
Data* result = data_create(data, (size_t)size);
|
||||
if (!result) {
|
||||
protocol_release_memory_for_session(session, allocation_size);
|
||||
@@ -756,6 +803,10 @@ bool protocol_receive_status(ProtocolSession* session, Status* status) {
|
||||
}
|
||||
if (!protocol_receive_n_data_until(session, status, sizeof(Status), deadline_ptr))
|
||||
return false;
|
||||
if (!status_is_valid(*status)) {
|
||||
log_message(LOG_LEVEL_ERROR, "Received unknown protocol status %d", *status);
|
||||
return false;
|
||||
}
|
||||
if (!protocol_capture_error_detail(session, status, deadline_ptr, NULL))
|
||||
return false;
|
||||
log_debug_message(LOG_DEBUG_PROTO, "Received Status: %s", status_to_string(*status));
|
||||
@@ -777,6 +828,10 @@ bool protocol_receive_status_timed(ProtocolSession* session, Status* status, int
|
||||
deadline.tv_sec += timeout_sec;
|
||||
if (!protocol_receive_n_data_until(session, status, sizeof(Status), &deadline))
|
||||
return false;
|
||||
if (!status_is_valid(*status)) {
|
||||
log_message(LOG_LEVEL_ERROR, "Received unknown protocol status %d", *status);
|
||||
return false;
|
||||
}
|
||||
if (!protocol_capture_error_detail(session, status, &deadline, NULL))
|
||||
return false;
|
||||
log_debug_message(LOG_DEBUG_PROTO, "Received Status: %s", status_to_string(*status));
|
||||
@@ -890,6 +945,10 @@ bool protocol_receive_status_keepalive(ProtocolSession* session, Status* status,
|
||||
Status received;
|
||||
if (!protocol_read_status_until(session, &received, &deadline))
|
||||
return false;
|
||||
if (!status_is_valid(received)) {
|
||||
log_message(LOG_LEVEL_ERROR, "Received unknown protocol status %d", received);
|
||||
return false;
|
||||
}
|
||||
if (!protocol_capture_error_detail(session, &received, &deadline, abort_check))
|
||||
return false;
|
||||
if (received == STATUS_KEEPALIVE) {
|
||||
|
||||
+20
-1
@@ -28,6 +28,10 @@
|
||||
|
||||
/* Maximum chunk size (64 MB) — prevents unbounded allocation from the wire */
|
||||
#define MAX_CHUNK_SIZE (64ULL * 1024 * 1024)
|
||||
/* Files larger than this are not kept fully in memory while loading: the
|
||||
* loader skips them so the sender streams from the path, and file_checksum
|
||||
* hashes them from disk in bounded buffers instead of forcing a full load. */
|
||||
#define STREAM_THRESHOLD (64ULL * 1024 * 1024)
|
||||
#define MAX_MANIFEST_ENTRIES (1024 * 1024)
|
||||
/* Aggregate bytes retained by one received deletion manifest. */
|
||||
#define MAX_MANIFEST_BYTES (16ULL * 1024 * 1024)
|
||||
@@ -191,7 +195,9 @@ enum NET_STATUS {
|
||||
* (--delete-delay). Payload: an int32 has_config flag (1 on the first plan
|
||||
* of the run, 0 afterwards); when set, the three global config sections
|
||||
* (protected-prefix count+paths, size-skipped count+paths, missing-args
|
||||
* count+paths); then the destination-relative directory path wire string
|
||||
* count+paths); then an int32 apply flag (1 for a real plan, 0 for a
|
||||
* config-only carrier frame that must not walk a directory); then the
|
||||
* destination-relative directory path wire string
|
||||
* ("." for the receive root); then the child-directory count + names and the
|
||||
* child-file count + names that must be kept. Appended after
|
||||
* STATUS_DEST_INFO so no existing status is renumbered. */
|
||||
@@ -207,6 +213,7 @@ enum NET_STATUS {
|
||||
|
||||
void io_set_fds(int read_fd, int write_fd);
|
||||
void io_set_bwlimit(unsigned long long bytes_per_sec);
|
||||
unsigned long long io_get_bwlimit(void);
|
||||
void io_set_ssl(SSL* ssl);
|
||||
SSL* io_get_ssl(void);
|
||||
|
||||
@@ -217,6 +224,15 @@ SSL* io_get_ssl(void);
|
||||
unsigned long long protocol_bytes_written(void);
|
||||
unsigned long long protocol_bytes_read(void);
|
||||
void protocol_note_bytes_written(unsigned long long bytes);
|
||||
/* Apply --bwlimit pacing to bytes written outside protocol_send_n_data (the
|
||||
* plaintext zero-copy sendfile fast path). `file_descriptor` is the wire fd
|
||||
* the bytes were written to, so the legacy session is resolved exactly as the
|
||||
* preceding send_n_data call resolved it (the bound TLS session still wins when
|
||||
* set); resolving with the same fd avoids re-initializing the legacy session
|
||||
* and granting a second first-call burst. Runs the same token-bucket throttle,
|
||||
* so the sendfile transport is paced identically to the buffered/TLS paths. A
|
||||
* no-op when the effective session has no bandwidth limit. */
|
||||
void protocol_throttle_bytes(int file_descriptor, size_t bytes);
|
||||
|
||||
void protocol_session_init(ProtocolSession* session, int read_fd, int write_fd);
|
||||
/* Transitional bridge for helpers whose signatures still carry only an fd. */
|
||||
@@ -260,6 +276,9 @@ bool protocol_send_int(ProtocolSession* session, int data);
|
||||
bool protocol_receive_int(ProtocolSession* session, int* data);
|
||||
bool protocol_send_status(ProtocolSession* session, Status status);
|
||||
bool protocol_receive_status(ProtocolSession* session, Status* status);
|
||||
/* As protocol_receive_status, but with an explicit per-message deadline
|
||||
* (seconds) instead of the session's configured io_timeout_sec. */
|
||||
bool protocol_receive_status_timed(ProtocolSession* session, Status* status, int timeout_sec);
|
||||
bool send_n_data(int file_descriptor, const void* data, size_t data_size);
|
||||
bool receive_n_data(int file_descriptor, void* data, size_t data_size);
|
||||
|
||||
|
||||
+18
-1
@@ -128,7 +128,7 @@ bool queue_enqueue_multithreaded_cancel(Queue* queue, void* item, mtx_t* mutex,
|
||||
|
||||
void* queue_dequeue(Queue* queue) {
|
||||
if (queue == NULL || queue_is_empty(queue)) {
|
||||
log_perror("ERROR: Could not dequeue from null or empty queue.");
|
||||
log_message(LOG_LEVEL_ERROR, "%s", "ERROR: Could not dequeue from null or empty queue.");
|
||||
return NULL;
|
||||
}
|
||||
|
||||
@@ -139,6 +139,23 @@ void* queue_dequeue(Queue* queue) {
|
||||
return item;
|
||||
}
|
||||
|
||||
bool queue_push(Queue* queue, void* item) {
|
||||
return queue_enqueue(queue, item);
|
||||
}
|
||||
|
||||
void* queue_pop(Queue* queue) {
|
||||
if (queue == NULL || queue_is_empty(queue)) {
|
||||
log_message(LOG_LEVEL_ERROR, "%s", "ERROR: Could not pop from null or empty queue.");
|
||||
return NULL;
|
||||
}
|
||||
|
||||
queue->rear = (queue->rear - 1 + queue->capacity) % queue->capacity;
|
||||
void* item = queue->items[queue->rear];
|
||||
queue->items[queue->rear] = NULL;
|
||||
queue->size--;
|
||||
return item;
|
||||
}
|
||||
|
||||
void* queue_dequeue_multithreaded(Queue* queue, mtx_t* mutex, cnd_t* condition_not_empty,
|
||||
cnd_t* condition_not_full, const bool* other_thread_done) {
|
||||
mtx_lock(mutex);
|
||||
|
||||
@@ -28,4 +28,11 @@ void* queue_dequeue(Queue* queue);
|
||||
void* queue_dequeue_multithreaded(Queue* queue, mtx_t* mutex, cnd_t* condition_not_empty,
|
||||
cnd_t* condition_not_full, const bool* other_thread_done);
|
||||
|
||||
/* LIFO stack operations over the same ring buffer. queue_push() is the enqueue
|
||||
primitive; queue_pop() removes from the rear, so a sequence of pushes is
|
||||
returned in reverse order. Used by the sequential scanner's depth-first
|
||||
traversal. */
|
||||
bool queue_push(Queue* queue, void* item);
|
||||
void* queue_pop(Queue* queue);
|
||||
|
||||
#endif
|
||||
|
||||
@@ -87,7 +87,7 @@ static int parse_remote_dest(const char* dest, RemoteDest* r) {
|
||||
return 0;
|
||||
}
|
||||
|
||||
char* ssh_build_remote_command(const char* server_path, bool old_args, char* const* remote_options,
|
||||
char* ssh_build_remote_command(const char* server_path, char* const* remote_options,
|
||||
int remote_option_count) {
|
||||
const char* path = server_path ? server_path : "fastsync-server";
|
||||
const char* suffix = " --stdio";
|
||||
@@ -105,9 +105,7 @@ char* ssh_build_remote_command(const char* server_path, bool old_args, char* con
|
||||
shell word (remote options below reuse the same escaping), then
|
||||
" --stdio". Quoting the path is the only injection-safe construction: an
|
||||
unquoted path would carry shell metacharacters straight into the remote
|
||||
shell command. --old-args is kept for CLI/ABI compatibility but no longer
|
||||
disables that protection. */
|
||||
(void)old_args;
|
||||
shell command. (rsync's --old-args no longer disables that protection.) */
|
||||
size_t quote_count = 0;
|
||||
for (const char* p = path; *p; p++)
|
||||
if (*p == '\'')
|
||||
@@ -296,8 +294,8 @@ void ssh_free_client_argv(char** argv) {
|
||||
}
|
||||
|
||||
Client* client_connect_ssh(const char* destination, int port, const char* server_path,
|
||||
bool old_args, const char* rsh_command, bool blocking_io,
|
||||
char* const* remote_options, int remote_option_count) {
|
||||
const char* rsh_command, bool blocking_io, char* const* remote_options,
|
||||
int remote_option_count) {
|
||||
RemoteDest r;
|
||||
if (parse_remote_dest(destination, &r) != 0) {
|
||||
char* escaped = output_escape(destination, false);
|
||||
@@ -376,7 +374,7 @@ Client* client_connect_ssh(const char* destination, int port, const char* server
|
||||
snprintf(ssh_user, ssh_user_len, "%s", r.host);
|
||||
|
||||
char* remote_command =
|
||||
ssh_build_remote_command(server_path, old_args, remote_options, remote_option_count);
|
||||
ssh_build_remote_command(server_path, remote_options, remote_option_count);
|
||||
if (!remote_command)
|
||||
ssh_child_setup_failed(exec_pipe[1]);
|
||||
char** ssh_argv = ssh_build_client_argv(rsh_command, port, ssh_user, remote_command);
|
||||
|
||||
@@ -4,17 +4,17 @@
|
||||
#include "transport_tcp.h"
|
||||
|
||||
Client* client_connect_ssh(const char* destination, int port, const char* server_path,
|
||||
bool old_args, const char* rsh_command, bool blocking_io,
|
||||
char* const* remote_options, int remote_option_count);
|
||||
const char* rsh_command, bool blocking_io, char* const* remote_options,
|
||||
int remote_option_count);
|
||||
/* Build the escaped remote-shell command string (the server program path always
|
||||
* quoted as one remote-shell word, followed by ` --stdio` and each
|
||||
* --remote-option value appended as an individually single-quoted shell word).
|
||||
* `old_args` is accepted for CLI/ABI compatibility but no longer disables
|
||||
* quoting: the path is always escaped so a metacharacter-bearing
|
||||
* --rsync-path can never be interpreted by the remote shell. Every
|
||||
* --remote-option value is individually escaped with the '\'' sequence and
|
||||
* values with empty/control characters are rejected at the CLI parse layer. */
|
||||
char* ssh_build_remote_command(const char* server_path, bool old_args, char* const* remote_options,
|
||||
* The path is always escaped so a metacharacter-bearing --rsync-path can never
|
||||
* be interpreted by the remote shell (the --old-args no-op does not disable
|
||||
* quoting). Every --remote-option value is individually escaped with the '\''
|
||||
* sequence and values with empty/control characters are rejected at the CLI
|
||||
* parse layer. */
|
||||
char* ssh_build_remote_command(const char* server_path, char* const* remote_options,
|
||||
int remote_option_count);
|
||||
/* Build the NULL-terminated child argv for the remote-shell client (argv[0] is
|
||||
* the exec/execvp program). rsh_command is whitespace-split into leading argv
|
||||
|
||||
+77
-383
@@ -2,10 +2,12 @@
|
||||
#include "array_list.h"
|
||||
#include "log.h"
|
||||
#include <arpa/inet.h>
|
||||
#include <ctype.h>
|
||||
#include <dirent.h>
|
||||
#include <errno.h>
|
||||
#include <fcntl.h>
|
||||
#include <netinet/in.h>
|
||||
#include <stdarg.h>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
@@ -47,11 +49,45 @@ const char* utils_get_authorized_root_path(void) {
|
||||
return authorized_root_path;
|
||||
}
|
||||
|
||||
void utils_set_error(char* err, size_t err_size, const char* fmt, ...) {
|
||||
if (!err || err_size == 0)
|
||||
return;
|
||||
va_list args;
|
||||
va_start(args, fmt);
|
||||
vsnprintf(err, err_size, fmt, args);
|
||||
va_end(args);
|
||||
}
|
||||
|
||||
bool path_is_within_root(const char* root, const char* path) {
|
||||
size_t root_len = strlen(root);
|
||||
return strncmp(root, path, root_len) == 0 && (path[root_len] == '\0' || path[root_len] == '/');
|
||||
}
|
||||
|
||||
/* Borrowed transfer-relative view of `path`: strip any leading '/' and then a
|
||||
* `root` prefix (its own leading/trailing slashes tolerated), returning a
|
||||
* pointer into `path`. Non-allocating, so it is safe on the hot scan/print
|
||||
* paths. A NULL/empty root, or a path not under `root`, leaves only the
|
||||
* leading-slash strip. `path` must be NUL-terminated and live in the caller. */
|
||||
const char* utils_strip_transfer_root(const char* path, const char* root) {
|
||||
if (path == NULL)
|
||||
return NULL;
|
||||
const char* rel = path;
|
||||
while (*rel == '/')
|
||||
rel++;
|
||||
if (root == NULL)
|
||||
return rel;
|
||||
while (*root == '/')
|
||||
root++;
|
||||
size_t root_len = strlen(root);
|
||||
while (root_len > 0 && root[root_len - 1] == '/')
|
||||
root_len--;
|
||||
if (root_len == 0)
|
||||
return rel;
|
||||
if (strncmp(rel, root, root_len) == 0 && (rel[root_len] == '/' || rel[root_len] == '\0'))
|
||||
return rel + root_len + (rel[root_len] == '/' ? 1 : 0);
|
||||
return rel;
|
||||
}
|
||||
|
||||
/* Open the destination root directory itself, confined to the authorized root.
|
||||
* NOTE (do not merge with file_open_secure_parent): this walk opens dest_root
|
||||
* (a directory that must already exist) and returns its fd, whereas
|
||||
@@ -117,6 +153,47 @@ char* str_dup(const char* string) {
|
||||
return new_string;
|
||||
}
|
||||
|
||||
int env_choice_first(const char* env_name, int (*resolve)(const char*), bool* specified) {
|
||||
if (specified)
|
||||
*specified = false;
|
||||
if (!env_name || !resolve)
|
||||
return -1;
|
||||
const char* env = getenv(env_name);
|
||||
if (!env)
|
||||
return -1;
|
||||
|
||||
bool saw_nonblank = false;
|
||||
const char* p = env;
|
||||
while (*p) {
|
||||
if (*p == '&')
|
||||
break;
|
||||
if (isspace((unsigned char)*p)) {
|
||||
p++;
|
||||
continue;
|
||||
}
|
||||
saw_nonblank = true;
|
||||
char token[64];
|
||||
size_t len = 0;
|
||||
while (*p && *p != '&' && !isspace((unsigned char)*p)) {
|
||||
if (len < sizeof(token) - 1)
|
||||
token[len++] = *p;
|
||||
p++;
|
||||
}
|
||||
token[len] = '\0';
|
||||
if (len > 0) {
|
||||
int id = resolve(token);
|
||||
if (id >= 0) {
|
||||
if (specified)
|
||||
*specified = true;
|
||||
return id;
|
||||
}
|
||||
}
|
||||
}
|
||||
if (specified)
|
||||
*specified = saw_nonblank;
|
||||
return -1;
|
||||
}
|
||||
|
||||
#define STR_HASH_SET_MIN_CAPACITY 16
|
||||
|
||||
static size_t str_hash_set_hash(const char* key, size_t len) {
|
||||
@@ -548,389 +625,6 @@ bool format_human_bytes(unsigned long long bytes, char* buffer, size_t buffer_si
|
||||
return written >= 0 && (size_t)written < buffer_size;
|
||||
}
|
||||
|
||||
/* Build the keep-set index from the exact manifest entries only. A lookup of
|
||||
`rel` succeeds iff `rel` is a kept entry, a kept directory, or an ancestor
|
||||
directory of kept content (the old is_dir_in_manifest predicate); the sorted
|
||||
view answers "is an ancestor of kept content" without materializing any
|
||||
per-component prefix copy, so the index is O(manifest size) memory. */
|
||||
static bool build_keep_index(const ArrayList* manifest, PathIndex* index) {
|
||||
if (!manifest || manifest->size <= 0)
|
||||
return path_index_build(index, NULL, 0);
|
||||
return path_index_build(index, (const char* const*)manifest->items, (size_t)manifest->size);
|
||||
}
|
||||
|
||||
static bool keep_is_dir(const PathIndex* index, const char* rel_path) {
|
||||
return path_index_contains(index, rel_path) || path_index_has_descendant(index, rel_path);
|
||||
}
|
||||
|
||||
static bool keep_is_file(const PathIndex* index, const char* rel_path) {
|
||||
return path_index_contains(index, rel_path);
|
||||
}
|
||||
|
||||
/* True when child_rel is, or lies below, a protected entry. A prefix "a"
|
||||
therefore protects "a" and "a/b/c" but not "ab". Entries with top_level_only
|
||||
set only protect DIRECT children of the receive root (at_root); nested
|
||||
directories that share such a name stay ordinary destination content. */
|
||||
bool path_under_skip_prefix(const char* child_rel, bool at_root, const DeleteSkipEntry* skips,
|
||||
int skip_count) {
|
||||
for (int i = 0; i < skip_count; i++) {
|
||||
if (skips[i].top_level_only && !at_root)
|
||||
continue;
|
||||
size_t prefix_len = strlen(skips[i].prefix);
|
||||
if (strncmp(child_rel, skips[i].prefix, prefix_len) == 0 &&
|
||||
(child_rel[prefix_len] == '\0' || child_rel[prefix_len] == '/'))
|
||||
return true;
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
/* Per-run deletion budget and tallies. `max_delete` is the cap on the number
|
||||
of entries the walker may remove (SIZE_MAX = unlimited); once it is reached
|
||||
the remaining extras are counted in `skipped` and left in place, matching
|
||||
rsync's partial --max-delete behavior. */
|
||||
typedef struct {
|
||||
size_t max_delete;
|
||||
size_t deleted;
|
||||
size_t skipped;
|
||||
bool limit_hit;
|
||||
} DeleteBudget;
|
||||
|
||||
/* True when direct children of the directory named by `rel` may be removed.
|
||||
With no synchronization info (dirs == NULL) the whole tree is deletable; when
|
||||
a dirs index is supplied only its exact entries are (the receive root is the
|
||||
"." sentinel). */
|
||||
static bool is_synced_dir(const PathIndex* dirs, const char* rel) {
|
||||
if (!dirs)
|
||||
return true;
|
||||
return path_index_contains(dirs, rel[0] == '\0' ? "." : rel);
|
||||
}
|
||||
|
||||
/* Remove the extras directly inside the directory open on `dirfd`, recursing
|
||||
into every child directory so kept content below a synchronized prefix is
|
||||
reached. `all_removed` reports whether every child entry was removed (so the
|
||||
caller may rmdir this directory). A child directory is never removed when it
|
||||
is itself a synchronized directory or holds kept content; with a dirs index
|
||||
supplied, direct children of a non-synchronized directory are never extras at
|
||||
all (they are left in place but still descended into). Symlinks are unlinked
|
||||
like any other non-directory extra (never followed). */
|
||||
static bool delete_extras_fd(int dirfd, const char* rel_path, const PathIndex* keep,
|
||||
const PathIndex* dirs, DeleteBudget* budget,
|
||||
const DeleteSkipEntry* skips, int skip_count, bool parent_deletable,
|
||||
bool* all_removed) {
|
||||
/* openat(dirfd, ".") opens an independent file description: a dup() would
|
||||
share dirfd's file offset and a prior pass could leave the stream drained. */
|
||||
int scanfd = openat(dirfd, ".", O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
|
||||
if (scanfd < 0)
|
||||
return false;
|
||||
DIR* dir = fdopendir(scanfd);
|
||||
if (!dir) {
|
||||
close(scanfd);
|
||||
return false;
|
||||
}
|
||||
bool operation_ok = true;
|
||||
bool local_survives = false;
|
||||
/* A directory is deletable when it or ANY ancestor is synchronized; the
|
||||
`parent_deletable` flag carries that down the recursion so dest-only
|
||||
directories below a synchronized root are removed wholesale. */
|
||||
bool deletable = parent_deletable || is_synced_dir(dirs, rel_path);
|
||||
const struct dirent* entry;
|
||||
while ((entry = readdir(dir)) != NULL) {
|
||||
if (strcmp(entry->d_name, ".") == 0 || strcmp(entry->d_name, "..") == 0)
|
||||
continue;
|
||||
char* child_rel = path_cat((char*)rel_path, entry->d_name);
|
||||
if (!child_rel) {
|
||||
operation_ok = false;
|
||||
continue;
|
||||
}
|
||||
/* A --delay-updates run keeps its staging directory as a direct child of
|
||||
the receive root, and basis-dir snapshots live below it too. Their
|
||||
contents are not manifest entries, so descending into them would delete
|
||||
every staged / basis file as an "extra". Only the staging name (a
|
||||
top-level-only prefix) and the basis prefixes are protected: a nested
|
||||
destination directory that happens to be called .fastsync-stage is
|
||||
ordinary content. */
|
||||
if (path_under_skip_prefix(child_rel, rel_path[0] == '\0', skips, skip_count)) {
|
||||
local_survives = true;
|
||||
free(child_rel);
|
||||
continue;
|
||||
}
|
||||
struct stat st;
|
||||
if (fstatat(dirfd, entry->d_name, &st, AT_SYMLINK_NOFOLLOW) != 0) {
|
||||
if (errno != ENOENT)
|
||||
operation_ok = false;
|
||||
free(child_rel);
|
||||
continue;
|
||||
}
|
||||
if (S_ISDIR(st.st_mode)) {
|
||||
int childfd = openat(dirfd, entry->d_name, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
|
||||
bool child_all_removed = false;
|
||||
if (childfd >= 0) {
|
||||
if (!delete_extras_fd(childfd, child_rel, keep, dirs, budget, skips, skip_count, deletable,
|
||||
&child_all_removed))
|
||||
operation_ok = false;
|
||||
close(childfd);
|
||||
} else if (errno != ENOENT) {
|
||||
operation_ok = false;
|
||||
}
|
||||
bool child_synced = dirs && path_index_contains(dirs, child_rel);
|
||||
if (child_synced || keep_is_dir(keep, child_rel)) {
|
||||
/* A synchronized directory and a directory holding kept content are
|
||||
never removed. */
|
||||
local_survives = true;
|
||||
} else if (child_all_removed && deletable) {
|
||||
if (budget->deleted >= budget->max_delete) {
|
||||
budget->limit_hit = true;
|
||||
budget->skipped++;
|
||||
local_survives = true;
|
||||
} else if (unlinkat(dirfd, entry->d_name, AT_REMOVEDIR) != 0) {
|
||||
/* ENOENT: already gone (fine). ENOTEMPTY/EEXIST: the directory
|
||||
still holds entries the walker leaves in place (a protected
|
||||
excluded prefix, a kept file the manifest protects, a symlink);
|
||||
rsync leaves such a directory behind, so this is not an error.
|
||||
Only genuine I/O failures abort the deletion. */
|
||||
if (errno != ENOENT && errno != ENOTEMPTY && errno != EEXIST)
|
||||
operation_ok = false;
|
||||
local_survives = true;
|
||||
} else {
|
||||
budget->deleted++;
|
||||
}
|
||||
} else {
|
||||
local_survives = true;
|
||||
}
|
||||
} else {
|
||||
bool found = keep_is_file(keep, child_rel);
|
||||
if (found || !deletable) {
|
||||
/* Kept file, or a child of a directory that is not synchronized: never
|
||||
an extra for this run. */
|
||||
local_survives = true;
|
||||
} else if (budget->deleted >= budget->max_delete) {
|
||||
budget->limit_hit = true;
|
||||
budget->skipped++;
|
||||
local_survives = true;
|
||||
} else if (unlinkat(dirfd, entry->d_name, 0) != 0) {
|
||||
if (errno != ENOENT)
|
||||
operation_ok = false;
|
||||
local_survives = true;
|
||||
} else {
|
||||
budget->deleted++;
|
||||
char* escaped_path = output_escape(child_rel, log_get_8_bit_output());
|
||||
fprintf(stderr, " Deleted: %s\n", escaped_path ? escaped_path : "<allocation failed>");
|
||||
free(escaped_path);
|
||||
}
|
||||
}
|
||||
free(child_rel);
|
||||
}
|
||||
closedir(dir);
|
||||
*all_removed = !local_survives;
|
||||
return operation_ok;
|
||||
}
|
||||
|
||||
/* Read-only mirror of delete_extras_fd: records the paths that WOULD be removed
|
||||
without unlinking anything. A child directory is reported after its own
|
||||
reportable children (depth-first), matching the delete pass's ordering. */
|
||||
static bool list_extras_fd(int dirfd, const char* rel_path, const PathIndex* keep,
|
||||
const PathIndex* dirs, ArrayList* out, size_t* recorded,
|
||||
const DeleteSkipEntry* skips, int skip_count, bool parent_deletable,
|
||||
bool* all_removed) {
|
||||
int scanfd = openat(dirfd, ".", O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
|
||||
if (scanfd < 0)
|
||||
return false;
|
||||
DIR* dir = fdopendir(scanfd);
|
||||
if (!dir) {
|
||||
close(scanfd);
|
||||
return false;
|
||||
}
|
||||
bool operation_ok = true;
|
||||
bool local_survives = false;
|
||||
bool deletable = parent_deletable || is_synced_dir(dirs, rel_path);
|
||||
const struct dirent* entry;
|
||||
while ((entry = readdir(dir)) != NULL) {
|
||||
if (strcmp(entry->d_name, ".") == 0 || strcmp(entry->d_name, "..") == 0)
|
||||
continue;
|
||||
char* child_rel = path_cat((char*)rel_path, entry->d_name);
|
||||
if (!child_rel) {
|
||||
operation_ok = false;
|
||||
continue;
|
||||
}
|
||||
if (path_under_skip_prefix(child_rel, rel_path[0] == '\0', skips, skip_count)) {
|
||||
local_survives = true;
|
||||
free(child_rel);
|
||||
continue;
|
||||
}
|
||||
struct stat st;
|
||||
if (fstatat(dirfd, entry->d_name, &st, AT_SYMLINK_NOFOLLOW) != 0) {
|
||||
if (errno != ENOENT)
|
||||
operation_ok = false;
|
||||
free(child_rel);
|
||||
continue;
|
||||
}
|
||||
if (S_ISDIR(st.st_mode)) {
|
||||
int childfd = openat(dirfd, entry->d_name, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
|
||||
bool child_all_removed = false;
|
||||
if (childfd >= 0) {
|
||||
if (!list_extras_fd(childfd, child_rel, keep, dirs, out, recorded, skips, skip_count,
|
||||
deletable, &child_all_removed))
|
||||
operation_ok = false;
|
||||
close(childfd);
|
||||
} else if (errno != ENOENT) {
|
||||
operation_ok = false;
|
||||
}
|
||||
bool child_synced = dirs && path_index_contains(dirs, child_rel);
|
||||
if (child_synced || keep_is_dir(keep, child_rel)) {
|
||||
local_survives = true;
|
||||
} else if (child_all_removed && deletable) {
|
||||
size_t len = strlen(child_rel);
|
||||
char* copy = malloc(len + 2);
|
||||
if (!copy) {
|
||||
operation_ok = false;
|
||||
} else {
|
||||
memcpy(copy, child_rel, len);
|
||||
copy[len] = '/';
|
||||
copy[len + 1] = '\0';
|
||||
if (!array_list_add(out, copy)) {
|
||||
free(copy);
|
||||
operation_ok = false;
|
||||
} else {
|
||||
(*recorded)++;
|
||||
}
|
||||
}
|
||||
} else {
|
||||
local_survives = true;
|
||||
}
|
||||
} else {
|
||||
bool found = keep_is_file(keep, child_rel);
|
||||
if (found || !deletable) {
|
||||
local_survives = true;
|
||||
} else {
|
||||
char* copy = str_dup(child_rel);
|
||||
if (!copy || !array_list_add(out, copy)) {
|
||||
free(copy);
|
||||
operation_ok = false;
|
||||
} else {
|
||||
(*recorded)++;
|
||||
}
|
||||
}
|
||||
}
|
||||
free(child_rel);
|
||||
}
|
||||
closedir(dir);
|
||||
*all_removed = !local_survives;
|
||||
return operation_ok;
|
||||
}
|
||||
|
||||
bool delete_extras_list(const char* dest_root, const ArrayList* manifest,
|
||||
const ArrayList* synced_dirs, const DeleteSkipEntry* skips, int skip_count,
|
||||
ArrayList* out, size_t* count_out) {
|
||||
if (count_out)
|
||||
*count_out = 0;
|
||||
if (!manifest || !out)
|
||||
return false;
|
||||
PathIndex keep;
|
||||
if (!build_keep_index(manifest, &keep))
|
||||
return false;
|
||||
PathIndex dirs;
|
||||
bool have_dirs = synced_dirs != NULL;
|
||||
if (have_dirs &&
|
||||
!path_index_build(&dirs, (const char* const*)synced_dirs->items, (size_t)synced_dirs->size)) {
|
||||
path_index_free(&keep);
|
||||
return false;
|
||||
}
|
||||
int rootfd;
|
||||
int root_fd = utils_get_authorized_root_fd();
|
||||
if (root_fd >= 0) {
|
||||
if (utils_get_authorized_root_path())
|
||||
rootfd = utils_open_authorized_destination(dest_root);
|
||||
else if (dest_root == NULL)
|
||||
rootfd = dup(root_fd);
|
||||
else
|
||||
rootfd = -1;
|
||||
} else {
|
||||
rootfd = open(dest_root, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
|
||||
}
|
||||
if (rootfd < 0) {
|
||||
path_index_free(&keep);
|
||||
if (have_dirs)
|
||||
path_index_free(&dirs);
|
||||
return false;
|
||||
}
|
||||
bool all_removed = false;
|
||||
size_t recorded = 0;
|
||||
bool ok = list_extras_fd(rootfd, "", &keep, have_dirs ? &dirs : NULL, out, &recorded, skips,
|
||||
skip_count, false, &all_removed);
|
||||
if (close(rootfd) != 0)
|
||||
ok = false;
|
||||
path_index_free(&keep);
|
||||
if (have_dirs)
|
||||
path_index_free(&dirs);
|
||||
if (count_out)
|
||||
*count_out = recorded;
|
||||
return ok;
|
||||
}
|
||||
|
||||
DeleteWalkResult delete_extras_limited(const char* dest_root, const ArrayList* manifest,
|
||||
const ArrayList* synced_dirs, size_t max_delete,
|
||||
const DeleteSkipEntry* skips, int skip_count,
|
||||
size_t* deleted_out, size_t* skipped_out) {
|
||||
if (deleted_out)
|
||||
*deleted_out = 0;
|
||||
if (skipped_out)
|
||||
*skipped_out = 0;
|
||||
if (!manifest)
|
||||
return DELETE_WALK_ERROR;
|
||||
/* Index the keep-set (and the synchronized-dir set, when supplied) once so
|
||||
membership is answered in O(path length) instead of scanning every entry
|
||||
for every destination entry. */
|
||||
PathIndex keep;
|
||||
if (!build_keep_index(manifest, &keep))
|
||||
return DELETE_WALK_ERROR;
|
||||
PathIndex dirs;
|
||||
bool have_dirs = synced_dirs != NULL;
|
||||
if (have_dirs &&
|
||||
!path_index_build(&dirs, (const char* const*)synced_dirs->items, (size_t)synced_dirs->size)) {
|
||||
path_index_free(&keep);
|
||||
return DELETE_WALK_ERROR;
|
||||
}
|
||||
int rootfd;
|
||||
int root_fd = utils_get_authorized_root_fd();
|
||||
if (root_fd >= 0) {
|
||||
if (utils_get_authorized_root_path())
|
||||
rootfd = utils_open_authorized_destination(dest_root);
|
||||
else if (dest_root == NULL)
|
||||
rootfd = dup(root_fd);
|
||||
else
|
||||
rootfd = -1;
|
||||
} else {
|
||||
rootfd = open(dest_root, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
|
||||
}
|
||||
if (rootfd < 0) {
|
||||
path_index_free(&keep);
|
||||
if (have_dirs)
|
||||
path_index_free(&dirs);
|
||||
return DELETE_WALK_ERROR;
|
||||
}
|
||||
DeleteBudget budget = {.max_delete = max_delete, .deleted = 0, .skipped = 0, .limit_hit = false};
|
||||
bool all_removed = false;
|
||||
bool ok = delete_extras_fd(rootfd, "", &keep, have_dirs ? &dirs : NULL, &budget, skips,
|
||||
skip_count, false, &all_removed);
|
||||
if (close(rootfd) != 0)
|
||||
ok = false;
|
||||
path_index_free(&keep);
|
||||
if (have_dirs)
|
||||
path_index_free(&dirs);
|
||||
if (deleted_out)
|
||||
*deleted_out = budget.deleted;
|
||||
if (skipped_out)
|
||||
*skipped_out = budget.skipped;
|
||||
if (!ok)
|
||||
return DELETE_WALK_ERROR;
|
||||
return budget.limit_hit ? DELETE_WALK_LIMIT_REACHED : DELETE_WALK_OK;
|
||||
}
|
||||
|
||||
bool delete_extras(const char* dest_root, const ArrayList* manifest) {
|
||||
return delete_extras_limited(dest_root, manifest, NULL, SIZE_MAX, NULL, 0, NULL, NULL) ==
|
||||
DELETE_WALK_OK;
|
||||
}
|
||||
|
||||
bool has_path_traversal(const char* path) {
|
||||
if (!path)
|
||||
return true;
|
||||
|
||||
+23
-53
@@ -2,6 +2,7 @@
|
||||
#define UTILS_H
|
||||
|
||||
#include "array_list.h"
|
||||
#include "filter.h"
|
||||
#include <stddef.h>
|
||||
#include <stdbool.h>
|
||||
#include <stdio.h>
|
||||
@@ -81,6 +82,17 @@ bool path_index_has_descendant(const PathIndex* index, const char* path);
|
||||
|
||||
char* str_dup(const char* string);
|
||||
char* output_escape(const char* string, bool eight_bit_output);
|
||||
|
||||
/* Resolve the first supported name from a rsync algorithm-preference
|
||||
* environment variable (RSYNC_COMPRESS_LIST / RSYNC_CHECKSUM_LIST). `resolve`
|
||||
* maps a case-insensitive name to an algorithm id (>= 0) or -1 for an unknown
|
||||
* name. rsync's syntax is a whitespace-separated list (comma/colon are NOT
|
||||
* separators); the client-side half ends at '&'. Unknown entries are skipped
|
||||
* and the first resolvable one wins. *specified is set true when the variable
|
||||
* holds at least one non-blank character. Returns the first resolvable id, or
|
||||
* -1 when the variable is unset/blank or names no supported algorithm. */
|
||||
int env_choice_first(const char* env_name, int (*resolve)(const char*), bool* specified);
|
||||
|
||||
/* Upper bound on one line/token read from a local list file (--files-from,
|
||||
* --exclude-from/--include-from, .rsync-filter). Mirrors MAX_STRING_SIZE and
|
||||
* stops a hostile multi-gigabyte line from forcing unbounded allocation. */
|
||||
@@ -94,59 +106,7 @@ char* output_escape(const char* string, bool eight_bit_output);
|
||||
ssize_t utils_getdelim_bounded(FILE* stream, char** line, size_t* cap, int delim, size_t max_len);
|
||||
char* path_cat(const char* path1, const char* path2);
|
||||
bool glob_match(const char* pattern, const char* str);
|
||||
/* Result of a bounded extra-file deletion run. */
|
||||
typedef enum {
|
||||
/* Every extra entry was removed (or there were none). */
|
||||
DELETE_WALK_OK = 0,
|
||||
/* The numeric cap for this run was reached before every extra was removed.
|
||||
The walker removed exactly the entries the cap allowed and skipped (without
|
||||
removing) the rest, matching rsync's partial --max-delete behavior. */
|
||||
DELETE_WALK_LIMIT_REACHED,
|
||||
/* A traversal or unlink failure aborted the deletion (partial removal is
|
||||
possible, mirroring the delete pass). */
|
||||
DELETE_WALK_ERROR
|
||||
} DeleteWalkResult;
|
||||
/* One protected entry for the delete walker. When top_level_only is true the
|
||||
prefix is skipped only as a DIRECT child of dest_root (the --delay-updates
|
||||
staging directory, which must not hide genuine extras inside a nested
|
||||
destination directory that happens to share the staging name); otherwise the
|
||||
prefix is skipped at any depth (the --compare-dest/--copy-dest/--link-dest
|
||||
basis trees, and the sender-side protected filter-excluded prefixes, which
|
||||
are never destination content). */
|
||||
typedef struct {
|
||||
const char* prefix;
|
||||
bool top_level_only;
|
||||
} DeleteSkipEntry;
|
||||
/* True when child_rel is, or lies below, one of the protected entries (a prefix
|
||||
"a" protects "a" and "a/b/c" but not "ab"; top_level_only entries protect
|
||||
only DIRECT children of the destination root, i.e. child_rel has no '/'). */
|
||||
bool path_under_skip_prefix(const char* child_rel, bool at_root, const DeleteSkipEntry* skips,
|
||||
int skip_count);
|
||||
/* Remove files/dirs/symlinks under dest_root that are not listed in manifest
|
||||
without ever descending into a protected prefix (see DeleteSkipEntry). When
|
||||
`synced_dirs` is non-NULL, extras are only removed directly inside a directory
|
||||
whose destination-relative path is an exact entry in that list (the receive
|
||||
root is the "." sentinel); directories outside the synchronized set are still
|
||||
descended into so kept content below a listed directory is preserved, but
|
||||
nothing in them is removed. A NULL `synced_dirs` keeps the legacy behavior of
|
||||
treating the whole destination tree as deletable. `max_delete` caps the
|
||||
number of removed entries (SIZE_MAX = unlimited): the walker removes up to the
|
||||
cap and returns DELETE_WALK_LIMIT_REACHED when more extras remained.
|
||||
`deleted_out`/`skipped_out` optionally receive the number of entries removed
|
||||
and the number skipped because of the cap. */
|
||||
DeleteWalkResult delete_extras_limited(const char* dest_root, const ArrayList* manifest,
|
||||
const ArrayList* synced_dirs, size_t max_delete,
|
||||
const DeleteSkipEntry* skips, int skip_count,
|
||||
size_t* deleted_out, size_t* skipped_out);
|
||||
/* Read-only companion to delete_extras_limited: walk the destination exactly as
|
||||
the delete pass would and APPEND (strdup'd) destination-relative paths that
|
||||
WOULD be removed, without touching disk. Used for -n/--dry-run --delete
|
||||
would-delete reporting. Returns true on a clean walk; the caller owns the
|
||||
strings appended to `out` and receives their count in *count_out. */
|
||||
bool delete_extras_list(const char* dest_root, const ArrayList* manifest,
|
||||
const ArrayList* synced_dirs, const DeleteSkipEntry* skips, int skip_count,
|
||||
ArrayList* out, size_t* count_out);
|
||||
bool delete_extras(const char* dest_root, const ArrayList* manifest);
|
||||
|
||||
/* Open the existing destination directory at `dest_root`, confined to the
|
||||
authorized root with an O_NOFOLLOW component walk (the same confinement the
|
||||
deletion walker uses for its root). Returns a new fd the caller owns, or -1
|
||||
@@ -171,12 +131,22 @@ void utils_set_authorized_root_fd(int fd);
|
||||
* threads spawn; see utils.c). */
|
||||
int utils_get_authorized_root_fd(void);
|
||||
const char* utils_get_authorized_root_path(void);
|
||||
/* Write a diagnostic message into a caller-supplied buffer, mirroring
|
||||
* vsnprintf. A NULL `err` or a zero `err_size` is a no-op, so a caller that
|
||||
* only needs the boolean status may safely pass NULL. Returns nothing; the
|
||||
* buffer is always NUL-terminated by vsnprintf when err_size > 0. */
|
||||
void utils_set_error(char* err, size_t err_size, const char* fmt, ...);
|
||||
/* True when `path` is `root` itself or lies directly beneath it: a lexical
|
||||
* prefix test requiring the byte after `root` to be '\0' or '/'. Both `root`
|
||||
* and `path` must be absolute canonical paths free of "."/".." components (the
|
||||
* callers guarantee this); this is containment by string, not by resolved
|
||||
* symlinks. Shared by the utils and file secure-walk root confinement. */
|
||||
bool path_is_within_root(const char* root, const char* path);
|
||||
/* Non-allocating transfer-relative view of `path`: strip any leading '/' and
|
||||
* then a `root` prefix (leading/trailing slashes tolerated), returning a
|
||||
* borrowed pointer into `path`. A NULL/empty root, or a path not under
|
||||
* `root`, yields just the leading-slash strip. `path`/`root` must stay alive. */
|
||||
const char* utils_strip_transfer_root(const char* path, const char* root);
|
||||
/* True when `path` contains a ".." component. This is a purely lexical
|
||||
* dot-dot check: an absolute path is NOT rejected here, because default
|
||||
* (non-relative) transfers legitimately put the sender's absolute source path
|
||||
|
||||
@@ -119,6 +119,13 @@ static void build_canonical_frame(void) {
|
||||
cfg->usermap[0].to = MAP_TO;
|
||||
cfg->usermap[0].to_name = NULL;
|
||||
}
|
||||
/* Force a non-empty receiver delete-protection block so the fuzzer mutates
|
||||
* its rule count, action/sides codes and pattern strings. */
|
||||
cfg->filters = array_list_create(free);
|
||||
if (cfg->filters) {
|
||||
array_list_add(cfg->filters, str_dup("P *.log"));
|
||||
array_list_add(cfg->filters, str_dup("+r keep/**"));
|
||||
}
|
||||
if (!cfg->send_directory || !cfg->receive_root_directory || !cfg->usermap) {
|
||||
config_delete(cfg);
|
||||
return;
|
||||
|
||||
@@ -0,0 +1,80 @@
|
||||
# Integration tests
|
||||
|
||||
The integration suite drives the built `build/server` and `build/client`
|
||||
against local corpora. Unit tests live in `tests/` (the custom C framework);
|
||||
the Python suite here covers the full transfer pipeline, transports, features,
|
||||
and rsync parity.
|
||||
|
||||
## Running
|
||||
|
||||
```bash
|
||||
# Full suite (excludes privilege-dependent tests on CI runners)
|
||||
python3 -m pytest tests/integration/ -n 4 --dist=load -m "not setpriv"
|
||||
|
||||
# Fast PR subset only
|
||||
python3 -m pytest tests/integration/ -n 4 --dist=load -m ci
|
||||
```
|
||||
|
||||
The tests expect `build/server` and `build/client` (configure/build with CMake
|
||||
first); `common.py` derives `BUILD_DIR` from the repository root.
|
||||
|
||||
## Differential rsync-parity gate
|
||||
|
||||
`test_differential_parity.py` runs the **same** transfer with real
|
||||
`rsync 3.4.1` and with FastSync over separate destinations, then compares:
|
||||
|
||||
- the destination trees — relative paths, file content hashes, symlink
|
||||
targets, modes (where the case is about perms), and hard-link grouping;
|
||||
- the normalized stdout for output-oriented flags (`-i`,
|
||||
`--out-format=...`, `--stats`), after stripping volatile fields
|
||||
(timings, rates, wire byte counts) and directory-only itemize lines that
|
||||
FastSync's recursive scanner documents as absent.
|
||||
|
||||
FastSync mirrors the absolute source path under its receive root (see
|
||||
`get_dest_received_dir`); the harness normalizes that layout (and the
|
||||
`-R`/`--files-from` layouts) before comparing.
|
||||
|
||||
```bash
|
||||
# Fast subset that guards the ✅ surface on pull requests
|
||||
python3 -m pytest tests/integration/test_differential_parity.py -n 4 --dist=load -m parity_ci
|
||||
|
||||
# Full differential case table (`_CASES`): every row is marked `parity`, and a
|
||||
# case with an allowlisted residual in `parity_caveats.py` is included too.
|
||||
python3 -m pytest tests/integration/test_differential_parity.py -n 4 --dist=load -m parity
|
||||
```
|
||||
|
||||
`-m parity` selects only the `_CASES` table in this module. Differential
|
||||
coverage for options outside that table (`--temp-dir`, `--delay-updates`,
|
||||
`--dry-run`, `--fuzzy`, the basis-dir options, `-M` over daemon/TCP, and
|
||||
receiver filter-protect) lives in dedicated modules (`test_option_parity.py`,
|
||||
`test_parity_blockers.py`, `test_parity_quickwins.py`, ...) and is not part of
|
||||
this gate. The suite skips cleanly when `rsync` is not installed.
|
||||
|
||||
## Allowlist (`parity_caveats.py`)
|
||||
|
||||
`parity_caveats.py` is the single data-driven allowlist of known differences.
|
||||
Each entry maps a case id to the aspects that may differ (`tree`, `stdout`,
|
||||
`extra`, `rc`) and cites the governing row in `RSYNC_COMPAT.md`:
|
||||
|
||||
```python
|
||||
CAVEATS = {
|
||||
# no known residuals at present -- the burn-down reached zero
|
||||
# "some_case_id": {"tree": "documented residual ... ref: RSYNC_COMPAT.md ..."},
|
||||
}
|
||||
```
|
||||
|
||||
A differential mismatch in an aspect that is **not** listed fails the gate with
|
||||
a readable tree/stdout diff.
|
||||
|
||||
If a case is allowlisted but now matches rsync, the gate emits a loud warning
|
||||
naming the stale entry — that is the parity burn-down signal. Run with
|
||||
`FASTSYNC_PARITY_STRICT=1` to make stale entries fail instead (the full CI
|
||||
parity job sets this). To add a residual:
|
||||
|
||||
1. Reproduce it with `-m parity` and read the failure's tree/stdout diff.
|
||||
2. Confirm it is a documented `⚠️`/`❌` residual (or get the `✅` row
|
||||
reclassified) and cite the row.
|
||||
3. Add the case id and aspect(s) to `CAVEATS`, keeping the reason concise.
|
||||
|
||||
Do not allowlist an undocumented divergence from a `✅` row — fix it or get the
|
||||
row reclassified first.
|
||||
+55
-12
@@ -20,6 +20,12 @@ CLIENT_CMD = [os.path.join(BUILD_DIR, "client")]
|
||||
_WORKER = os.environ.get("PYTEST_XDIST_WORKER")
|
||||
TEST_DATA_DIR = os.path.join(PROJECT_ROOT, f"test_data-{_WORKER}" if _WORKER else "test_data")
|
||||
|
||||
# Default wall-clock budget for a short-lived client invocation. Every client
|
||||
# is expected to finish well within this; the bound exists so a hung client
|
||||
# fails the test instead of stalling the whole CI run indefinitely. Callers
|
||||
# that legitimately need longer can pass an explicit ``timeout``.
|
||||
CLIENT_TIMEOUT = 180
|
||||
|
||||
|
||||
class ServerManager:
|
||||
"""Manages a long-lived server process. Reuses across test cases."""
|
||||
@@ -137,7 +143,40 @@ class CountingProxy:
|
||||
return result
|
||||
|
||||
|
||||
def run_client(source_dir, dest_dir, flags=None, port=None, extra_args=None):
|
||||
def _run_client_cmd(cmd, timeout):
|
||||
"""Run one client command, returning ``(result, duration)``.
|
||||
|
||||
On timeout the client is killed and a result-like ``CompletedProcess`` with
|
||||
a non-zero returncode is returned instead of raising, so callers keep the
|
||||
established ``(result, duration)`` contract and the failure carries the
|
||||
command plus whatever output was captured for diagnosis.
|
||||
"""
|
||||
start = time.monotonic()
|
||||
try:
|
||||
result = subprocess.run(cmd, text=True, capture_output=True, timeout=timeout)
|
||||
except subprocess.TimeoutExpired as exc:
|
||||
duration = time.monotonic() - start
|
||||
stdout = exc.stdout or ""
|
||||
stderr = exc.stderr or ""
|
||||
if isinstance(stdout, bytes):
|
||||
stdout = stdout.decode(errors="replace")
|
||||
if isinstance(stderr, bytes):
|
||||
stderr = stderr.decode(errors="replace")
|
||||
diagnostic = (
|
||||
f"client timed out after {timeout}s\n"
|
||||
f"command: {cmd!r}\n"
|
||||
f"--- captured stdout ---\n{stdout}\n"
|
||||
f"--- captured stderr ---\n{stderr}"
|
||||
)
|
||||
result = subprocess.CompletedProcess(cmd, returncode=-1,
|
||||
stdout=stdout, stderr=diagnostic)
|
||||
return result, duration
|
||||
duration = time.monotonic() - start
|
||||
return result, duration
|
||||
|
||||
|
||||
def run_client(source_dir, dest_dir, flags=None, port=None, extra_args=None,
|
||||
timeout=CLIENT_TIMEOUT):
|
||||
"""Run the client and return (result, duration)."""
|
||||
cmd = CLIENT_CMD + ["--source-dir", source_dir, "--dest-dir", dest_dir, "--save-to-disk"]
|
||||
if port:
|
||||
@@ -146,23 +185,18 @@ def run_client(source_dir, dest_dir, flags=None, port=None, extra_args=None):
|
||||
cmd += flags
|
||||
if extra_args:
|
||||
cmd += extra_args
|
||||
start = time.monotonic()
|
||||
result = subprocess.run(cmd, text=True, capture_output=True)
|
||||
duration = time.monotonic() - start
|
||||
return result, duration
|
||||
return _run_client_cmd(cmd, timeout)
|
||||
|
||||
|
||||
def run_client_posix(source_dir, dest_dir, flags=None, port=None):
|
||||
def run_client_posix(source_dir, dest_dir, flags=None, port=None,
|
||||
timeout=CLIENT_TIMEOUT):
|
||||
"""Run the client with positional args (rsync-style)."""
|
||||
cmd = CLIENT_CMD + [source_dir, dest_dir, "--save-to-disk"]
|
||||
if port:
|
||||
cmd += ["--server-port", str(port)]
|
||||
if flags:
|
||||
cmd += flags
|
||||
start = time.monotonic()
|
||||
result = subprocess.run(cmd, text=True, capture_output=True)
|
||||
duration = time.monotonic() - start
|
||||
return result, duration
|
||||
return _run_client_cmd(cmd, timeout)
|
||||
|
||||
|
||||
def generate_test_files(source_dir, full=False):
|
||||
@@ -238,8 +272,17 @@ def make_result(name, success, duration=None, error=""):
|
||||
|
||||
|
||||
def get_dest_received_dir(dest_dir, source_dir):
|
||||
"""Get the path where received files land inside dest_dir."""
|
||||
return os.path.join(dest_dir, os.path.abspath(source_dir).lstrip(os.sep))
|
||||
"""Get the path where received files land inside dest_dir.
|
||||
|
||||
FastSync mirrors the absolute source path below the receive root with the
|
||||
leading root separator removed. Strip that separator explicitly rather
|
||||
than with ``str.lstrip(os.sep)``: ``lstrip`` removes a *set* of characters
|
||||
rather than a path prefix, which is not the same operation.
|
||||
"""
|
||||
abs_source = os.path.abspath(source_dir)
|
||||
if abs_source.startswith(os.sep):
|
||||
abs_source = abs_source[len(os.sep):]
|
||||
return os.path.join(dest_dir, abs_source)
|
||||
|
||||
|
||||
def _find_free_port():
|
||||
|
||||
@@ -0,0 +1,34 @@
|
||||
"""Data-driven allowlist for the differential rsync-parity gate.
|
||||
|
||||
Every entry maps a case id (see ``test_differential_parity.py``) to the aspects
|
||||
that are *known* to differ from ``rsync 3.4.1`` and the documented reason. A
|
||||
differential mismatch in an aspect that is **not** listed here fails the gate.
|
||||
|
||||
Aspect keys
|
||||
-----------
|
||||
``tree`` destination tree differs (paths, file hashes, symlink targets,
|
||||
modes, hardlink grouping)
|
||||
``stdout`` normalized output for ``-i`` / ``--stats`` / ``--out-format``
|
||||
``extra`` a case-specific assertion differs (basis/inode checks, ...)
|
||||
``rc`` exit status differs
|
||||
|
||||
Burn-down
|
||||
---------
|
||||
If a case is listed here but now matches rsync, the gate emits a loud
|
||||
``pytest`` warning naming the stale entry: delete the entry (and, when the
|
||||
underlying row in ``RSYNC_COMPAT.md`` is now parity, update that row). Set
|
||||
``FASTSYNC_PARITY_STRICT=1`` to turn stale entries into failures in CI.
|
||||
|
||||
Keep the values concise but cite the governing row so the entry can be
|
||||
re-triaged when the row moves.
|
||||
"""
|
||||
|
||||
# case id -> {aspect: "reason (ref: RSYNC_COMPAT.md ...)"}
|
||||
CAVEATS = {}
|
||||
|
||||
# Accepted aspect names (guards against typos in this file).
|
||||
ASPECTS = ("tree", "stdout", "extra", "rc")
|
||||
|
||||
|
||||
def caveat_for(case_id: str) -> dict:
|
||||
return CAVEATS.get(case_id, {})
|
||||
@@ -0,0 +1,584 @@
|
||||
"""Differential rsync-parity harness.
|
||||
|
||||
Runs the SAME transfer with real ``rsync`` and with FastSync over separate
|
||||
destinations and compares the resulting trees and (optionally) normalized
|
||||
stdout. ``test_differential_parity.py`` drives this module with a table of
|
||||
cases; ``parity_caveats.py`` is the data-driven allowlist of documented
|
||||
residuals.
|
||||
|
||||
Design notes
|
||||
------------
|
||||
FastSync mirrors the *absolute* source path below its receive root, while
|
||||
rsync copies the source contents directly into the destination. ``Case.layout``
|
||||
tells the harness which pair of directory roots to compare:
|
||||
|
||||
* ``MIRROR`` -- rsync ``DEST/`` vs FastSync ``DEST/<abs-src>/`` (the common
|
||||
case; matches ``common.get_dest_received_dir``).
|
||||
* ``MIRROR_ABS`` -- ``rsync -R`` without a cut lays the full absolute path
|
||||
under the destination, so rsync ``DEST/<abs-src>/`` is compared against the
|
||||
same FastSync mirror path.
|
||||
* ``RELATIVE`` -- ``rsync -R --files-from`` lays bare relative paths under the
|
||||
destination and FastSync does the same, so both destination roots compare
|
||||
directly.
|
||||
|
||||
Only ``tests/integration/common.py`` is used to reach the build products and the
|
||||
server manager; the harness never duplicates that plumbing.
|
||||
"""
|
||||
import difflib
|
||||
import hashlib
|
||||
import os
|
||||
import re
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
from dataclasses import dataclass
|
||||
from typing import Callable, Dict, List, Optional, Tuple
|
||||
|
||||
sys.path.insert(0, os.path.dirname(__file__))
|
||||
from common import ( # noqa: E402 (path bootstrap above)
|
||||
TEST_DATA_DIR,
|
||||
clean_dir,
|
||||
get_dest_received_dir,
|
||||
run_client,
|
||||
)
|
||||
|
||||
RSYNC = shutil.which("rsync")
|
||||
|
||||
# Comparison layouts (see module docstring).
|
||||
MIRROR = "mirror"
|
||||
MIRROR_ABS = "mirror_abs"
|
||||
RELATIVE = "relative"
|
||||
|
||||
# stdout comparators.
|
||||
STDOUT_NONE = None
|
||||
STDOUT_ITEMIZE = "itemize"
|
||||
STDOUT_OUTFMT = "outfmt"
|
||||
STDOUT_STATS = "stats"
|
||||
STDOUT_PROGRESS = "progress"
|
||||
|
||||
# rsync --stats lines that are protocol-independent and must match exactly.
|
||||
# `Number of files` and `Number of created files` carry rsync's per-type
|
||||
# breakdown; protocol 2.28.0 reports the receiver-created split over
|
||||
# STATUS_STATS. Deliberately excluded: Total bytes sent/received (protocol
|
||||
# framing differs, see the `--stats` row in RSYNC_COMPAT.md).
|
||||
STATS_KEYS = (
|
||||
"Number of files",
|
||||
"Number of created files",
|
||||
"Number of deleted files",
|
||||
"Number of regular files transferred",
|
||||
"Total file size",
|
||||
"Total transferred file size",
|
||||
"Literal data",
|
||||
"Matched data",
|
||||
"File list size",
|
||||
)
|
||||
|
||||
_ITEMIZE_RE = re.compile(r"^(<|>|c|h|\.|\*)[fdLDS][.+\-][.+\-][.+\-][.+\-]")
|
||||
_PROGRESS_TOTAL_RE = re.compile(r"to-chk=\d+/(\d+)")
|
||||
_PROGRESS_XFR_RE = re.compile(r"xfr#(\d+)")
|
||||
|
||||
|
||||
@dataclass
|
||||
class Case:
|
||||
"""One differential scenario: a corpus, a flag set, and how to compare."""
|
||||
|
||||
id: str
|
||||
corpus: str
|
||||
flags: List[str]
|
||||
fastsync_flags: Optional[List[str]] = None
|
||||
layout: str = MIRROR
|
||||
server_args: Tuple[str, ...] = ("--allow-super",)
|
||||
seed: Optional[Callable] = None
|
||||
stdout: Optional[str] = STDOUT_NONE
|
||||
compare_modes: bool = False
|
||||
compare_hardlinks: bool = False
|
||||
ignore_paths: Tuple[str, ...] = ()
|
||||
extra_check: Optional[Callable] = None
|
||||
files_from: Optional[Tuple[str, ...]] = None
|
||||
# rsync receives ``src + "/"``; FastSync mirrors the path it is given, so a
|
||||
# trailing-slash-sensitive case must hand FastSync the same form.
|
||||
fs_src_suffix: str = ""
|
||||
# Some cases have an unspecified result (e.g. which extras survive a
|
||||
# partial --max-delete abort): assert the case-specific invariants via
|
||||
# extra_check and skip the exact-tree comparison.
|
||||
compare_tree: bool = True
|
||||
ci: bool = False
|
||||
ref: str = ""
|
||||
|
||||
def fs_flags(self) -> List[str]:
|
||||
return list(self.flags if self.fastsync_flags is None else self.fastsync_flags)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Corpora
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
# Deterministic mtimes so quick-check decisions are reproducible.
|
||||
_SRC_MTIME = 1_600_000_000
|
||||
|
||||
|
||||
def _write(path: str, data: bytes, mode: Optional[int] = None) -> None:
|
||||
os.makedirs(os.path.dirname(path), exist_ok=True)
|
||||
with open(path, "wb") as fh:
|
||||
fh.write(data)
|
||||
os.utime(path, (_SRC_MTIME, _SRC_MTIME))
|
||||
if mode is not None:
|
||||
os.chmod(path, mode)
|
||||
|
||||
|
||||
def _set_mode(path: str, mode: int) -> None:
|
||||
os.chmod(path, mode)
|
||||
|
||||
|
||||
def corpus_basic(root: str) -> None:
|
||||
"""Regular files + nested dirs (dirs are implied by their files)."""
|
||||
clean_dir(root)
|
||||
_write(os.path.join(root, "a.txt"), b"hello world\n")
|
||||
_write(os.path.join(root, "sub", "b.bin"),
|
||||
bytes((i * 7) & 0xFF for i in range(5000)))
|
||||
_write(os.path.join(root, "sub", "deep", "c.txt"), "w\u00f6rld\n".encode())
|
||||
|
||||
|
||||
def corpus_unicode(root: str) -> None:
|
||||
clean_dir(root)
|
||||
_write(os.path.join(root, "uni \u00f1\u6587.txt"), b"unicode\n")
|
||||
_write(os.path.join(root, "sub", "sp ace \u00e9.dat"), b"spaced\n")
|
||||
_set_mode(os.path.join(root, "sub"), 0o750)
|
||||
|
||||
|
||||
def corpus_links(root: str) -> None:
|
||||
corpus_basic(root)
|
||||
os.symlink("a.txt", os.path.join(root, "rel_link"))
|
||||
os.symlink("/etc/hostname", os.path.join(root, "abs_link"))
|
||||
os.symlink("nowhere/target", os.path.join(root, "broken_link"))
|
||||
|
||||
|
||||
def corpus_hardlinks(root: str) -> None:
|
||||
clean_dir(root)
|
||||
_write(os.path.join(root, "h1.txt"), b"hardlinked payload\n")
|
||||
os.link(os.path.join(root, "h1.txt"), os.path.join(root, "h2.txt"))
|
||||
_write(os.path.join(root, "other.txt"), b"other\n")
|
||||
|
||||
|
||||
def corpus_sparse(root: str) -> None:
|
||||
clean_dir(root)
|
||||
_write(os.path.join(root, "small.txt"), b"small\n")
|
||||
sparse = os.path.join(root, "sparse.bin")
|
||||
with open(sparse, "wb") as fh:
|
||||
fh.seek(1024 * 1024 - 1)
|
||||
fh.write(b"\0")
|
||||
os.utime(sparse, (_SRC_MTIME, _SRC_MTIME))
|
||||
|
||||
|
||||
def corpus_filters(root: str) -> None:
|
||||
clean_dir(root)
|
||||
_write(os.path.join(root, "keep.txt"), b"keep\n")
|
||||
_write(os.path.join(root, "drop.log"), b"log\n")
|
||||
_write(os.path.join(root, "sub", "keep2.txt"), b"keep2\n")
|
||||
_write(os.path.join(root, "sub", "drop2.log"), b"log2\n")
|
||||
_write(os.path.join(root, "sub", "data.bin"), b"bin\n")
|
||||
|
||||
|
||||
def corpus_empty_dir(root: str) -> None:
|
||||
clean_dir(root)
|
||||
_write(os.path.join(root, "keep.txt"), b"keep\n")
|
||||
os.makedirs(os.path.join(root, "emptydir"), exist_ok=True)
|
||||
os.utime(os.path.join(root, "emptydir"), (_SRC_MTIME, _SRC_MTIME))
|
||||
_write(os.path.join(root, "nonempty", "f.txt"), b"f\n")
|
||||
|
||||
|
||||
def corpus_multidir(root: str) -> None:
|
||||
"""Multi-directory tree for the --progress file-list naming/denominator.
|
||||
|
||||
Nested files, a directory-only branch, an empty directory and a symlink
|
||||
exercise every file-list entry type rsync counts in `to-chk` but FastSync's
|
||||
streaming scanner never emits as a transfer entry.
|
||||
"""
|
||||
clean_dir(root)
|
||||
_write(os.path.join(root, "a.txt"), b"alpha\n")
|
||||
_write(os.path.join(root, "b.txt"), b"bravo\n")
|
||||
_write(os.path.join(root, "sub1", "c.txt"), b"charlie\n")
|
||||
_write(os.path.join(root, "sub1", "deep", "d.txt"), b"delta\n")
|
||||
_write(os.path.join(root, "sub2", "e.txt"), b"echo\n")
|
||||
os.symlink("a.txt", os.path.join(root, "link1"))
|
||||
os.makedirs(os.path.join(root, "emptydir"), exist_ok=True)
|
||||
os.utime(os.path.join(root, "emptydir"), (_SRC_MTIME, _SRC_MTIME))
|
||||
|
||||
|
||||
def corpus_relative(root: str) -> None:
|
||||
"""Tree for the -R/--files-from cases."""
|
||||
clean_dir(root)
|
||||
_write(os.path.join(root, "a.txt"), b"a\n")
|
||||
_write(os.path.join(root, "b.txt"), b"b\n")
|
||||
_write(os.path.join(root, "sub", "x.txt"), b"x\n")
|
||||
_write(os.path.join(root, "sub", "y.txt"), b"y\n")
|
||||
os.makedirs(os.path.join(root, "dir1"), exist_ok=True)
|
||||
os.utime(os.path.join(root, "dir1"), (_SRC_MTIME, _SRC_MTIME))
|
||||
_write(os.path.join(root, "dir1", "keep.txt"), b"keep\n")
|
||||
|
||||
|
||||
def corpus_iconv(root: str) -> None:
|
||||
"""Latin-1 (ISO-8859-1) encoded filenames, matching the --iconv direction."""
|
||||
clean_dir(root)
|
||||
for rel, data in ((b"caf\xe9.txt", b"caf\xe9\n"),
|
||||
(os.path.join(b"sub", b"\xfcber.txt"), b"\xfcber\n")):
|
||||
full = os.path.join(os.fsencode(root), rel)
|
||||
os.makedirs(os.path.dirname(full), exist_ok=True)
|
||||
with open(full, "wb") as fh:
|
||||
fh.write(data)
|
||||
os.utime(full, (_SRC_MTIME, _SRC_MTIME))
|
||||
|
||||
|
||||
# Payload for the --fuzzy basis corpus: large enough for the delta engine's
|
||||
# 16 KiB minimum and with repeated content so a coinciding basis yields a
|
||||
# non-zero (and identical) Matched data count in both tools.
|
||||
FUZZY_PAYLOAD = (b"the quick brown fox jumps over the lazy dog\n" * 2000)[:65536]
|
||||
|
||||
|
||||
def corpus_fuzzy(root: str) -> None:
|
||||
"""A named regular file; the `fuzzy` seed adds the similar-suffix sibling."""
|
||||
clean_dir(root)
|
||||
_write(os.path.join(root, "report_v2.txt"), FUZZY_PAYLOAD)
|
||||
|
||||
|
||||
CORPORA: Dict[str, Callable[[str], None]] = {
|
||||
"basic": corpus_basic,
|
||||
"unicode": corpus_unicode,
|
||||
"links": corpus_links,
|
||||
"hardlinks": corpus_hardlinks,
|
||||
"sparse": corpus_sparse,
|
||||
"filters": corpus_filters,
|
||||
"empty_dir": corpus_empty_dir,
|
||||
"multidir": corpus_multidir,
|
||||
"relative": corpus_relative,
|
||||
"iconv": corpus_iconv,
|
||||
"fuzzy": corpus_fuzzy,
|
||||
}
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Tree snapshotting / comparison
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
def snapshot(root: str, compare_modes: bool = False) -> Dict[str, tuple]:
|
||||
"""Map relative path -> descriptor for every entry below ``root``.
|
||||
|
||||
Files hash their contents with SHA-256 (structural comparison, so differing
|
||||
quick-check metadata cannot mask a payload difference). Symlinks record
|
||||
their target. Empty directories are included (as ``("dir", ...)``) so the
|
||||
recursive-empty-directory residual is observable.
|
||||
"""
|
||||
out: Dict[str, tuple] = {}
|
||||
if not os.path.isdir(root):
|
||||
return out
|
||||
|
||||
def describe(path: str) -> Optional[tuple]:
|
||||
st = os.lstat(path)
|
||||
if os.path.islink(path):
|
||||
return ("link", os.readlink(path))
|
||||
if os.path.isdir(path):
|
||||
mode = oct(st.st_mode & 0o7777) if compare_modes else None
|
||||
return ("dir", mode)
|
||||
h = hashlib.sha256()
|
||||
with open(path, "rb") as fh:
|
||||
for chunk in iter(lambda: fh.read(65536), b""):
|
||||
h.update(chunk)
|
||||
mode = oct(st.st_mode & 0o7777) if compare_modes else None
|
||||
return ("file", h.hexdigest()[:16], mode)
|
||||
|
||||
# The comparison root itself is not part of the tree diff: a no-transfer
|
||||
# result legitimately leaves FastSync's mirror directory absent while rsync
|
||||
# leaves an existing (empty) destination root.
|
||||
for dirpath, dirnames, filenames in os.walk(root, followlinks=False):
|
||||
dirnames.sort()
|
||||
for name in sorted(dirnames):
|
||||
p = os.path.join(dirpath, name)
|
||||
rel = os.path.relpath(p, root)
|
||||
if os.path.islink(p):
|
||||
out[rel] = ("link", os.readlink(p))
|
||||
dirnames.remove(name)
|
||||
else:
|
||||
out[rel] = describe(p)
|
||||
for name in sorted(filenames):
|
||||
p = os.path.join(dirpath, name)
|
||||
out[os.path.relpath(p, root)] = describe(p)
|
||||
return out
|
||||
|
||||
|
||||
def _hardlink_groups(root: str) -> Dict[str, str]:
|
||||
"""Assign a stable group letter to each inode shared by >1 regular file."""
|
||||
inodes: Dict[tuple, List[str]] = {}
|
||||
for dirpath, _dirs, filenames in os.walk(root, followlinks=False):
|
||||
for name in filenames:
|
||||
p = os.path.join(dirpath, name)
|
||||
if os.path.islink(p):
|
||||
continue
|
||||
st = os.lstat(p)
|
||||
if st.st_nlink > 1:
|
||||
inodes.setdefault((st.st_dev, st.st_ino), []).append(
|
||||
os.path.relpath(p, root))
|
||||
groups: Dict[str, str] = {}
|
||||
for i, (_key, members) in enumerate(sorted(inodes.items())):
|
||||
for rel in members:
|
||||
groups[rel] = chr(ord("A") + i)
|
||||
return groups
|
||||
|
||||
|
||||
def _drop_ignored(tree: Dict[str, tuple], ignore_paths) -> Dict[str, tuple]:
|
||||
if not ignore_paths:
|
||||
return tree
|
||||
out = {}
|
||||
for rel, desc in tree.items():
|
||||
if any(rel == ig or rel.startswith(ig.rstrip("/") + "/") for ig in ignore_paths):
|
||||
continue
|
||||
out[rel] = desc
|
||||
return out
|
||||
|
||||
|
||||
def tree_diff(rsync_root: str, fs_root: str, case: Case) -> List[str]:
|
||||
"""Return a list of human-readable differences (empty when identical)."""
|
||||
rtree = _drop_ignored(snapshot(rsync_root, case.compare_modes), case.ignore_paths)
|
||||
ftree = _drop_ignored(snapshot(fs_root, case.compare_modes), case.ignore_paths)
|
||||
if case.compare_hardlinks:
|
||||
rgroups = _hardlink_groups(rsync_root)
|
||||
fgroups = _hardlink_groups(fs_root)
|
||||
else:
|
||||
rgroups = fgroups = {}
|
||||
diffs: List[str] = []
|
||||
for rel in sorted(set(rtree) | set(ftree)):
|
||||
r = rtree.get(rel)
|
||||
f = ftree.get(rel)
|
||||
if r == f:
|
||||
continue
|
||||
if r is None:
|
||||
diffs.append(f"+ fastsync-only: {rel!r} {f}")
|
||||
elif f is None:
|
||||
diffs.append(f"- rsync-only: {rel!r} {r}")
|
||||
else:
|
||||
diffs.append(f"~ differs: {rel!r} rsync={r} fastsync={f}")
|
||||
if case.compare_hardlinks:
|
||||
for rel in sorted(set(rgroups) | set(fgroups)):
|
||||
if rgroups.get(rel) != fgroups.get(rel):
|
||||
diffs.append(
|
||||
f"~ hardlink group: {rel!r} rsync={rgroups.get(rel)} "
|
||||
f"fastsync={fgroups.get(rel)}")
|
||||
return diffs
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# stdout normalization
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
def _parse_bytes(text: str) -> str:
|
||||
m = re.match(r"([\d,]+)", text.strip())
|
||||
return m.group(1).replace(",", "") if m else text.strip()
|
||||
|
||||
|
||||
def normalize_stdout(text: str, mode: Optional[str]) -> object:
|
||||
if mode == STDOUT_ITEMIZE:
|
||||
lines = []
|
||||
for line in (text or "").splitlines():
|
||||
line = line.rstrip()
|
||||
if not line:
|
||||
continue
|
||||
if line.startswith("*deleting"):
|
||||
lines.append(line)
|
||||
continue
|
||||
if not _ITEMIZE_RE.match(line):
|
||||
continue
|
||||
# Directories are not transfer entries in FastSync's recursive
|
||||
# scanner, so rsync's `cd+++++++++ name/` lines have no counterpart
|
||||
# (documented recursive-empty-dir residual). Compare file/link
|
||||
# itemization only.
|
||||
if line.rsplit(" ", 1)[-1].endswith("/"):
|
||||
continue
|
||||
lines.append(line)
|
||||
return sorted(lines)
|
||||
if mode == STDOUT_OUTFMT:
|
||||
lines = []
|
||||
for line in (text or "").splitlines():
|
||||
line = line.rstrip()
|
||||
if not line:
|
||||
continue
|
||||
# Directory entries are emitted by rsync but not by FastSync's
|
||||
# recursive scanner (documented residual). Tokens are either
|
||||
# `%n %l` (path first) or `%i %n` (path last); drop a line when
|
||||
# either end-token is a directory path.
|
||||
first = line.split(" ", 1)[0]
|
||||
last = line.rsplit(" ", 1)[-1]
|
||||
if first.endswith("/") or last.endswith("/"):
|
||||
continue
|
||||
lines.append(line)
|
||||
return sorted(lines)
|
||||
if mode == STDOUT_STATS:
|
||||
found = {}
|
||||
for line in (text or "").splitlines():
|
||||
for key in STATS_KEYS:
|
||||
if line.startswith(key + ":"):
|
||||
found[key] = _parse_bytes(line.split(":", 1)[1])
|
||||
return found
|
||||
if mode == STDOUT_PROGRESS:
|
||||
# rsync prints the file-list entries in sorted depth-first order while
|
||||
# FastSync's streaming scan emits them in readdir/BFS order; only the
|
||||
# entry set and deterministic fields are compared. The transfer-root
|
||||
# `./` line's trigger condition is a separate documented residual, and
|
||||
# the per-frame rate/elapsed/xfr#/to-chk numerator are wall-clock- or
|
||||
# order-dependent, so only the `to-chk` denominator and the name set are
|
||||
# asserted.
|
||||
names = []
|
||||
totals = set()
|
||||
max_xfr = 0
|
||||
for line in (text or "").splitlines():
|
||||
line = line.rstrip()
|
||||
if not line:
|
||||
continue
|
||||
if "%" in line:
|
||||
m = _PROGRESS_TOTAL_RE.search(line)
|
||||
if m:
|
||||
totals.add(int(m.group(1)))
|
||||
mx = _PROGRESS_XFR_RE.search(line)
|
||||
if mx:
|
||||
max_xfr = max(max_xfr, int(mx.group(1)))
|
||||
continue
|
||||
if line == "sending incremental file list":
|
||||
continue
|
||||
if line.startswith("created directory "):
|
||||
continue
|
||||
if line == "./":
|
||||
continue
|
||||
names.append(line)
|
||||
return {"names": sorted(names), "total": sorted(totals), "xfr": max_xfr}
|
||||
# raw
|
||||
return sorted(l.rstrip() for l in (text or "").splitlines() if l.strip())
|
||||
|
||||
|
||||
def stdout_diff(rsync_out: str, fs_out: str, mode: Optional[str]) -> List[str]:
|
||||
r = normalize_stdout(rsync_out, mode)
|
||||
f = normalize_stdout(fs_out, mode)
|
||||
if r == f:
|
||||
return []
|
||||
if mode == STDOUT_STATS:
|
||||
return [f"stats rsync={r}", f"stats fastsync={f}"]
|
||||
if mode == STDOUT_PROGRESS:
|
||||
return [f"progress rsync={r}", f"progress fastsync={f}"]
|
||||
return list(difflib.unified_diff(
|
||||
[str(x) for x in r], [str(x) for x in f],
|
||||
fromfile="rsync", tofile="fastsync", lineterm=""))
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Running one case
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
def run_rsync(src: str, rdst: str, flags: List[str]) -> subprocess.CompletedProcess:
|
||||
args = [RSYNC] + list(flags) + [src + "/", rdst + "/"]
|
||||
return subprocess.run(
|
||||
args, capture_output=True, text=True,
|
||||
env=dict(os.environ, LC_ALL="C"), timeout=180)
|
||||
|
||||
|
||||
def run_fastsync(src: str, fdst: str, flags: List[str], port: int):
|
||||
return run_client(src, fdst, flags=list(flags), port=port)
|
||||
|
||||
|
||||
def run_differential( # noqa: PLR0913 (explicit scenario parameters)
|
||||
src: str,
|
||||
rdst: str,
|
||||
fdst: str,
|
||||
rs_flags: List[str],
|
||||
fs_flags: List[str],
|
||||
server,
|
||||
layout: str = MIRROR,
|
||||
seed: Optional[Callable] = None,
|
||||
stdout: Optional[str] = STDOUT_NONE,
|
||||
compare_modes: bool = False,
|
||||
compare_hardlinks: bool = False,
|
||||
ignore_paths: Tuple[str, ...] = (),
|
||||
extra_check: Optional[Callable] = None,
|
||||
files_from: Optional[Tuple[str, ...]] = None,
|
||||
fs_src_suffix: str = "",
|
||||
compare_tree: bool = True,
|
||||
) -> Dict[str, object]:
|
||||
"""Run one rsync/FastSync pair and return the diff aspects.
|
||||
|
||||
Returned dict keys: ``rsync_rc``, ``fastsync_rc``, ``rsync_stderr``,
|
||||
``fastsync_stderr``, ``tree``, ``stdout``, ``extra``.
|
||||
"""
|
||||
clean_dir(rdst)
|
||||
clean_dir(fdst)
|
||||
abs_src = os.path.abspath(src)
|
||||
rel = abs_src.lstrip(os.sep)
|
||||
if layout == RELATIVE:
|
||||
rroot, froot = rdst, fdst
|
||||
elif layout == MIRROR_ABS:
|
||||
rroot, froot = os.path.join(rdst, rel), get_dest_received_dir(fdst, src)
|
||||
else:
|
||||
rroot, froot = rdst, get_dest_received_dir(fdst, src)
|
||||
# rsync's destination root always exists (clean_dir created it). FastSync's
|
||||
# logical transfer root is the mirror path below the destination argument,
|
||||
# so pre-create it too: `Number of created files` counts the root only when
|
||||
# it is genuinely absent, and the two tools must start from the same state.
|
||||
os.makedirs(froot, exist_ok=True)
|
||||
if seed:
|
||||
seed(src, rroot, froot)
|
||||
|
||||
rs_flags = list(rs_flags)
|
||||
fs_flags = list(fs_flags)
|
||||
if files_from is not None:
|
||||
list_path = os.path.join(TEST_DATA_DIR, "parity_" +
|
||||
os.path.basename(src) + ".list")
|
||||
write_list(list_path, files_from)
|
||||
rs_flags.append(f"--files-from={list_path}")
|
||||
fs_flags.append(f"--files-from={list_path}")
|
||||
|
||||
rs = run_rsync(src, rdst, rs_flags)
|
||||
fs_result, _ = run_fastsync(src + fs_src_suffix, fdst, fs_flags, server.port)
|
||||
|
||||
class _View:
|
||||
"""Adapter so tree_diff/extra_check keep the Case-shaped interface."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self.compare_modes = compare_modes
|
||||
self.compare_hardlinks = compare_hardlinks
|
||||
self.ignore_paths = ignore_paths
|
||||
|
||||
result = {
|
||||
"rsync_rc": rs.returncode,
|
||||
"fastsync_rc": fs_result.returncode,
|
||||
"rsync_stderr": rs.stderr,
|
||||
"fastsync_stderr": fs_result.stderr or fs_result.stdout,
|
||||
"tree": tree_diff(rroot, froot, _View()) if compare_tree else [],
|
||||
"stdout": [],
|
||||
"extra": [],
|
||||
}
|
||||
if stdout is not None:
|
||||
result["stdout"] = stdout_diff(rs.stdout, fs_result.stdout, stdout)
|
||||
if extra_check:
|
||||
result["extra"] = list(extra_check(src, rroot, froot, rs, fs_result) or [])
|
||||
return result
|
||||
|
||||
|
||||
def execute_case(case: Case, server) -> Dict[str, object]:
|
||||
"""Run a table-driven case and return the diff aspects."""
|
||||
tag = case.id
|
||||
src = os.path.join(TEST_DATA_DIR, f"parity_{tag}_src")
|
||||
rdst = os.path.join(TEST_DATA_DIR, f"parity_{tag}_rdst")
|
||||
fdst = os.path.join(TEST_DATA_DIR, f"parity_{tag}_fdst")
|
||||
CORPORA[case.corpus](src)
|
||||
return run_differential(
|
||||
src, rdst, fdst,
|
||||
case.flags, case.fs_flags(), server,
|
||||
layout=case.layout, seed=case.seed, stdout=case.stdout,
|
||||
compare_modes=case.compare_modes, compare_hardlinks=case.compare_hardlinks,
|
||||
ignore_paths=case.ignore_paths, extra_check=case.extra_check,
|
||||
files_from=case.files_from, fs_src_suffix=case.fs_src_suffix,
|
||||
compare_tree=case.compare_tree,
|
||||
)
|
||||
|
||||
|
||||
def write_list(path: str, entries) -> str:
|
||||
os.makedirs(os.path.dirname(path), exist_ok=True)
|
||||
with open(path, "w", encoding="utf-8") as fh:
|
||||
for e in entries:
|
||||
fh.write(e + "\n")
|
||||
return path
|
||||
@@ -0,0 +1,213 @@
|
||||
"""Differential tests for the client CLI's codec defaults and env lists.
|
||||
|
||||
Track 3a of the rsync-parity plan pins two rsync 3.4.1 behaviors that are
|
||||
resolved entirely on the client:
|
||||
|
||||
* the per-codec default ``--compress-level`` (zstd 3, zlib/zlibx 6, lz4
|
||||
ignored) applied when the user omits ``--compress-level``/``--zl``, with an
|
||||
explicit level clamped to the codec's range; and
|
||||
* the ``RSYNC_COMPRESS_LIST`` / ``RSYNC_CHECKSUM_LIST`` preference lists that
|
||||
rsync's ``auto`` consults before its compiled-in order (whitespace-separated,
|
||||
unknown names skipped, first supported wins, all-unknown is exit 4).
|
||||
|
||||
The rsync side is observed through ``--debug=NSTR1``; FastSync publishes its
|
||||
resolved codec/level through ``--debug=util``. The checksum side is confirmed
|
||||
byte-for-byte through ``--out-format %C``. The rsync-based tests skip cleanly
|
||||
when rsync is not installed.
|
||||
"""
|
||||
import os
|
||||
import re
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
import pytest
|
||||
|
||||
sys.path.insert(0, os.path.dirname(__file__))
|
||||
from common import (
|
||||
TEST_DATA_DIR,
|
||||
run_client,
|
||||
clean_dir,
|
||||
get_dest_received_dir,
|
||||
)
|
||||
|
||||
RSYNC = shutil.which("rsync")
|
||||
requires_rsync = pytest.mark.skipif(RSYNC is None, reason="rsync 3.4.1 not installed")
|
||||
|
||||
CODEC_ROOT = os.path.join(TEST_DATA_DIR, "cli_differential")
|
||||
|
||||
_COMPRESS_RE = re.compile(r"compress(?:ion)?: (\w+) \(level (-?\d+)\)")
|
||||
|
||||
|
||||
def _rsync(args):
|
||||
env = dict(os.environ, LC_ALL="C")
|
||||
return subprocess.run([RSYNC] + args, capture_output=True, text=True, env=env, timeout=120)
|
||||
|
||||
|
||||
def _scratch(tag):
|
||||
path = os.path.join(CODEC_ROOT, tag)
|
||||
clean_dir(path)
|
||||
os.makedirs(path, exist_ok=True)
|
||||
return path
|
||||
|
||||
|
||||
def _make_corpus(root):
|
||||
clean_dir(root)
|
||||
os.makedirs(root, exist_ok=True)
|
||||
with open(os.path.join(root, "big.bin"), "wb") as fh:
|
||||
fh.write(b"FastSync codec payload " * 4096)
|
||||
with open(os.path.join(root, "small.txt"), "wb") as fh:
|
||||
fh.write(b"hello codec world\n" * 32)
|
||||
return root
|
||||
|
||||
|
||||
def _rsync_compress_level(choice, level):
|
||||
src = _make_corpus(_scratch(f"lvl_src_{choice}_{level}"))
|
||||
dst = _scratch(f"lvl_rsync_{choice}_{level}")
|
||||
args = ["-a", "-z", f"--zc={choice}"]
|
||||
if level is not None:
|
||||
args.append(f"--zl={level}")
|
||||
args += ["--debug=NSTR1", src + "/", dst + "/"]
|
||||
result = _rsync(args)
|
||||
assert result.returncode == 0, result.stderr
|
||||
match = _COMPRESS_RE.search(result.stdout + result.stderr)
|
||||
assert match, (result.stdout, result.stderr)
|
||||
return match.group(1), int(match.group(2))
|
||||
|
||||
|
||||
def _fastsync_compress_level(choice, level, shared_server):
|
||||
src = _make_corpus(_scratch(f"lvl_src_fs_{choice}_{level}"))
|
||||
dst = _scratch(f"lvl_fs_{choice}_{level}")
|
||||
args = ["-a", "-z", f"--zc={choice}"]
|
||||
if level is not None:
|
||||
args.append(f"--zl={level}")
|
||||
args += ["-v", "--debug=util"]
|
||||
result, _ = run_client(src, dst, flags=args, port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
match = _COMPRESS_RE.search(result.stdout)
|
||||
assert match, result.stdout[:500]
|
||||
return match.group(1), int(match.group(2))
|
||||
|
||||
|
||||
class TestPerCodecCompressionLevelDefaults:
|
||||
"""``--compress-level`` defaults and clamping match rsync per codec."""
|
||||
|
||||
# FastSync uses a positive lz4 placeholder because its "level > 0" gate
|
||||
# enables compression; lz4_compress ignores the value, so rsync's level 0
|
||||
# and FastSync's level 1 produce the same bytes.
|
||||
CASES = [
|
||||
("zstd", None, 3),
|
||||
("zlib", None, 6),
|
||||
("zlibx", None, 6),
|
||||
("lz4", None, 1),
|
||||
("zstd", 10, 10),
|
||||
("zlib", 15, 9),
|
||||
("zlib", 3, 3),
|
||||
("lz4", 15, 1),
|
||||
]
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
@pytest.mark.parametrize("choice,level,fs_level", CASES)
|
||||
def test_level_matches_rsync(self, choice, level, fs_level, shared_server):
|
||||
rsync_algo, rsync_level = _rsync_compress_level(choice, level)
|
||||
fs_algo, fs_level_actual = _fastsync_compress_level(choice, level, shared_server)
|
||||
assert rsync_algo == choice
|
||||
assert fs_algo == choice
|
||||
if choice == "lz4":
|
||||
assert rsync_level == 0 and fs_level_actual > 0
|
||||
else:
|
||||
assert rsync_level == fs_level
|
||||
assert fs_level_actual == fs_level
|
||||
|
||||
|
||||
class TestEnvPreferenceLists:
|
||||
"""``RSYNC_COMPRESS_LIST`` / ``RSYNC_CHECKSUM_LIST`` drive auto like rsync."""
|
||||
|
||||
# (env value, expected codec, rsync level, FastSync level)
|
||||
COMPRESS_CASES = [
|
||||
("zlib lz4", "zlib", 6, 6),
|
||||
("lz4 zstd", "lz4", 0, 1),
|
||||
("bogus zstd zlib", "zstd", 3, 3),
|
||||
(" ", "zstd", 3, 3),
|
||||
]
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
@pytest.mark.parametrize("env,algo,rsync_level,fs_level", COMPRESS_CASES)
|
||||
def test_compress_list_matches_rsync(self, env, algo, rsync_level, fs_level, shared_server,
|
||||
monkeypatch):
|
||||
monkeypatch.setenv("RSYNC_COMPRESS_LIST", env)
|
||||
src = _make_corpus(_scratch(f"envc_src_{algo}"))
|
||||
rdst = _scratch(f"envc_rsync_{algo}")
|
||||
rs = _rsync(["-a", "-z", "--debug=NSTR1", src + "/", rdst + "/"])
|
||||
assert rs.returncode == 0, rs.stderr
|
||||
rm = _COMPRESS_RE.search(rs.stdout + rs.stderr)
|
||||
assert rm, (rs.stdout, rs.stderr)
|
||||
assert rm.group(1) == algo
|
||||
assert int(rm.group(2)) == rsync_level
|
||||
|
||||
fdst = _scratch(f"envc_fs_{algo}")
|
||||
result, _ = run_client(src, fdst, flags=["-a", "-z", "-v", "--debug=util"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
fm = _COMPRESS_RE.search(result.stdout)
|
||||
assert fm, result.stdout[:500]
|
||||
assert fm.group(1) == algo
|
||||
assert int(fm.group(2)) == fs_level
|
||||
received = get_dest_received_dir(fdst, src)
|
||||
assert _tree_bytes(received) == _tree_bytes(src)
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
@pytest.mark.parametrize("env,algo", [("md5", "md5"), ("sha1", "sha1"), ("xxh3 md5", "xxh3")])
|
||||
def test_checksum_list_matches_rsync(self, env, algo, shared_server, monkeypatch):
|
||||
monkeypatch.setenv("RSYNC_CHECKSUM_LIST", env)
|
||||
src = _make_corpus(_scratch(f"envcc_src_{algo}"))
|
||||
rdst = _scratch(f"envcc_rsync_{algo}")
|
||||
rs = _rsync(["-a", "--checksum", "--out-format=%C %n", src + "/", rdst + "/"])
|
||||
assert rs.returncode == 0, rs.stderr
|
||||
rs_digests = _digests(rs.stdout)
|
||||
|
||||
fdst = _scratch(f"envcc_fs_{algo}")
|
||||
result, _ = run_client(src, fdst, flags=["-a", "--checksum", "--out-format=%C %n"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
assert _digests(result.stdout) == rs_digests
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_all_unknown_lists_fail_like_rsync(self, shared_server, monkeypatch):
|
||||
src = _make_corpus(_scratch("envbad_src"))
|
||||
monkeypatch.setenv("RSYNC_COMPRESS_LIST", "bogus")
|
||||
rs = _rsync(["-a", "-z", src + "/", _scratch("envbad_rsync_c") + "/"])
|
||||
assert rs.returncode == 4, rs.stderr
|
||||
result, _ = run_client(src, _scratch("envbad_fs_c"), flags=["-a", "-z"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 4, (result.stderr or result.stdout)[:200]
|
||||
|
||||
monkeypatch.setenv("RSYNC_CHECKSUM_LIST", "bogus")
|
||||
rs = _rsync(["-a", src + "/", _scratch("envbad_rsync_s") + "/"])
|
||||
assert rs.returncode == 4, rs.stderr
|
||||
result, _ = run_client(src, _scratch("envbad_fs_s"), flags=["-a"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 4, (result.stderr or result.stdout)[:200]
|
||||
|
||||
|
||||
def _tree_bytes(root):
|
||||
out = {}
|
||||
for dirpath, _dirs, files in os.walk(root):
|
||||
for name in files:
|
||||
path = os.path.join(dirpath, name)
|
||||
with open(path, "rb") as fh:
|
||||
out[os.path.relpath(path, root)] = fh.read()
|
||||
return out
|
||||
|
||||
|
||||
def _digests(output):
|
||||
out = {}
|
||||
for line in output.splitlines():
|
||||
parts = line.split()
|
||||
if len(parts) == 2 and parts[0]:
|
||||
out[parts[1]] = parts[0]
|
||||
return out
|
||||
@@ -9,6 +9,7 @@ directory when the shared test server is launched), so every scratch tree lives
|
||||
under ``TEST_DATA_DIR`` rather than pytest's ``tmp_path``.
|
||||
"""
|
||||
import os
|
||||
import random
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
@@ -68,6 +69,84 @@ def _tree_bytes(root):
|
||||
return out
|
||||
|
||||
|
||||
_CC_DELTA_T0 = 1_600_000_000
|
||||
_CC_DELTA_T1 = 1_600_000_100
|
||||
|
||||
_DELTA_STATS_KEYS = (
|
||||
"Number of created files",
|
||||
"Number of regular files transferred",
|
||||
"Total transferred file size",
|
||||
"Literal data",
|
||||
"Matched data",
|
||||
)
|
||||
|
||||
|
||||
def _pin_tree(root, mtime):
|
||||
for dirpath, dirnames, filenames in os.walk(root):
|
||||
for name in dirnames + filenames:
|
||||
path = os.path.join(dirpath, name)
|
||||
if not os.path.islink(path):
|
||||
os.utime(path, (mtime, mtime))
|
||||
os.utime(root, (mtime, mtime))
|
||||
|
||||
|
||||
def _make_delta_basis(src, size=512 * 1024):
|
||||
"""Build a source and a matching pre-modification basis tree.
|
||||
|
||||
The source's ``big.bin`` is then modified in a few disjoint places and given
|
||||
a newer mtime so both tools take the delta path. Returns the basis dir.
|
||||
"""
|
||||
clean_dir(src)
|
||||
original = random.Random(20240101).randbytes(size)
|
||||
with open(os.path.join(src, "big.bin"), "wb") as fh:
|
||||
fh.write(original)
|
||||
with open(os.path.join(src, "small.txt"), "wb") as fh:
|
||||
fh.write(b"hello world\n")
|
||||
_pin_tree(src, _CC_DELTA_T0)
|
||||
|
||||
basis = src.rstrip("/") + "_basis"
|
||||
clean_dir(basis)
|
||||
shutil.copy2(os.path.join(src, "big.bin"), os.path.join(basis, "big.bin"))
|
||||
shutil.copy2(os.path.join(src, "small.txt"), os.path.join(basis, "small.txt"))
|
||||
_pin_tree(basis, _CC_DELTA_T0)
|
||||
|
||||
modified = bytearray(original)
|
||||
for off in (0, size // 3, 2 * size // 3, size - 64):
|
||||
for i in range(32):
|
||||
modified[off + i] ^= 0x5A
|
||||
with open(os.path.join(src, "big.bin"), "wb") as fh:
|
||||
fh.write(bytes(modified))
|
||||
os.utime(os.path.join(src, "big.bin"), (_CC_DELTA_T1, _CC_DELTA_T1))
|
||||
return basis
|
||||
|
||||
|
||||
def _seed_from_basis(basis, target):
|
||||
clean_dir(target)
|
||||
for name in os.listdir(basis):
|
||||
shutil.copy2(os.path.join(basis, name), os.path.join(target, name))
|
||||
|
||||
|
||||
def _delta_stats(text):
|
||||
found = {}
|
||||
for line in text.splitlines():
|
||||
for key in _DELTA_STATS_KEYS:
|
||||
if line.startswith(key + ":"):
|
||||
found[key] = line.split(":", 1)[1].strip()
|
||||
return found
|
||||
|
||||
|
||||
def _big_bin_outfmt(text):
|
||||
"""The ``(c, C)`` pair from the ``big.bin`` out-format line (`%c|%C %n`)."""
|
||||
for line in text.splitlines():
|
||||
stripped = line.strip()
|
||||
if "|" not in stripped or not stripped.endswith("big.bin"):
|
||||
continue
|
||||
c_field, rest = stripped.split("|", 1)
|
||||
fields = rest.split()
|
||||
return c_field.strip(), (fields[0] if fields else "")
|
||||
return None, None
|
||||
|
||||
|
||||
class TestCodecChoiceMatrix:
|
||||
"""The CLI accept/reject set and exit codes must match rsync 3.4.1."""
|
||||
|
||||
@@ -191,6 +270,76 @@ class TestCodecTransferDifferential:
|
||||
assert _tree_bytes(received) == _tree_bytes(rsync_dst)
|
||||
|
||||
|
||||
class TestChecksumChoiceDeltaSurface:
|
||||
"""--checksum-choice does not move the delta-transfer parity surface.
|
||||
|
||||
FastSync's delta BLOCK strong checksum is a fixed xxHash32, so the
|
||||
negotiated algorithm only selects the whole-file comparison digest (and the
|
||||
``%C`` transfer digest). A pre-seeded delta transfer must therefore land
|
||||
byte-identical bytes and report the same counters for every choice, while
|
||||
``%C`` -- the one token that tracks the choice -- stays byte-identical to
|
||||
rsync. This pins the Track-3b reclassification in RSYNC_COMPAT.md.
|
||||
"""
|
||||
|
||||
CHOICES = ["xxh64", "xxh128", "xxh3", "md5", "md4", "sha1",
|
||||
"xxh64,sha1", "sha1,xxh64"]
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_delta_surface_invariant_to_checksum_choice(self, shared_server):
|
||||
src = _scratch("ccdelta_src")
|
||||
basis = _make_delta_basis(src)
|
||||
source_bytes = _tree_bytes(src)
|
||||
|
||||
rsync_c, fastsync_c, digests = {}, {}, {}
|
||||
rsync_stats, fastsync_stats = {}, {}
|
||||
for choice in self.CHOICES:
|
||||
tag = choice.replace(",", "_")
|
||||
rsync_dst = _scratch(f"ccdelta_rs_{tag}")
|
||||
fs_dst = _scratch(f"ccdelta_fs_{tag}")
|
||||
fs_root = get_dest_received_dir(fs_dst, src)
|
||||
_seed_from_basis(basis, rsync_dst)
|
||||
_seed_from_basis(basis, fs_root)
|
||||
|
||||
# Pin the block size on both ends so the literal/matched split is
|
||||
# comparable (rsync's adaptive default would otherwise differ from
|
||||
# FastSync's 8192-byte default).
|
||||
rsync_result = _rsync(["-a", "--no-whole-file", "-B8192", "--stats",
|
||||
"--out-format=%c|%C %n", f"--cc={choice}",
|
||||
src + "/", rsync_dst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
result, _ = run_client(
|
||||
src, fs_dst,
|
||||
flags=["-a", "--incremental", "--delta", "-B8192", "--stats",
|
||||
"--out-format=%c|%C %n", f"--cc={choice}"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
|
||||
assert _tree_bytes(rsync_dst) == source_bytes, choice
|
||||
assert _tree_bytes(fs_root) == source_bytes, choice
|
||||
assert _delta_stats(rsync_result.stdout) == _delta_stats(result.stdout), choice
|
||||
|
||||
rs_c, rs_C = _big_bin_outfmt(rsync_result.stdout)
|
||||
fs_c, fs_C = _big_bin_outfmt(result.stdout)
|
||||
assert rs_C == fs_C, f"{choice}: %C rsync={rs_C!r} fastsync={fs_C!r}"
|
||||
rsync_c[choice] = rs_c
|
||||
fastsync_c[choice] = fs_c
|
||||
digests[choice] = fs_C
|
||||
rsync_stats[choice] = _delta_stats(rsync_result.stdout)
|
||||
fastsync_stats[choice] = _delta_stats(result.stdout)
|
||||
|
||||
# The choice is only observable in %C, and it is effective (the digests
|
||||
# are not all the same algorithm's output).
|
||||
assert len(set(digests.values())) > 1, digests
|
||||
# The compared --stats counters are invariant across choices in each tool
|
||||
# (and were asserted equal cross-tool inside the loop).
|
||||
assert len({tuple(sorted(s.items())) for s in rsync_stats.values()}) == 1, rsync_stats
|
||||
assert len({tuple(sorted(s.items())) for s in fastsync_stats.values()}) == 1, fastsync_stats
|
||||
# The block-checksum token (%c) is invariant across choices in each tool.
|
||||
assert len(set(rsync_c.values())) == 1, rsync_c
|
||||
assert len(set(fastsync_c.values())) == 1, fastsync_c
|
||||
|
||||
|
||||
class TestCodecNegotiationFallback:
|
||||
"""FastSync's auto negotiation and deterministic fallback order."""
|
||||
|
||||
|
||||
@@ -63,6 +63,10 @@ DETACH_MODULE = os.path.join(MODULE_ROOT, "detach")
|
||||
DETACH_CONF = os.path.join(TEST_DATA_DIR, "fastsyncd_detach.conf")
|
||||
DETACH_PORT = None
|
||||
|
||||
# A dedicated config for the umask test: the daemon must be launched in the real
|
||||
# (double-fork) detach path, whose daemonize() applies umask(022).
|
||||
UMASK_CONF = os.path.join(TEST_DATA_DIR, "fastsyncd_umask.conf")
|
||||
|
||||
# Passwords are never sent as plaintext and never logged; these literals are
|
||||
# only hashed into the server credential file / client password file.
|
||||
ALICE_PASS = "alice-s3cret"
|
||||
@@ -332,6 +336,44 @@ class TestDaemonModuleSelection:
|
||||
assert not missing, f"missing: {missing[:5]}"
|
||||
assert not mismatches, f"mismatch: {mismatches[:5]}"
|
||||
|
||||
def test_daemon_new_dirs_not_world_writable(self):
|
||||
"""The daemon must not force umask 0: implied parent directories created
|
||||
without -p are the source default (0755 under the daemon's 022 umask),
|
||||
never world-writable 0777.
|
||||
|
||||
This drives the real double-fork detach path, where the umask(022) fix
|
||||
lives (daemonize()); the --no-detach path never calls it. The launcher
|
||||
is run with umask 0, so without the fix the daemon would inherit 0 and
|
||||
create a 0777 directory; with the fix the assertion below fails only if
|
||||
the fix regresses."""
|
||||
port = _find_free_port()
|
||||
with open(UMASK_CONF, "w") as f:
|
||||
f.write("port = %d\n\n[files]\npath = %s\n" % (port, FILES_MODULE))
|
||||
sub = os.path.join(FILES_MODULE, "umask_check")
|
||||
shutil.rmtree(sub, ignore_errors=True)
|
||||
os.makedirs(sub, exist_ok=True)
|
||||
log_path = os.path.join(TEST_DATA_DIR, "fastsyncd_umask.log")
|
||||
log = open(log_path, "w")
|
||||
cmd = SERVER_CMD + ["--daemon", "--config", UMASK_CONF, "--allow-unauthenticated"]
|
||||
proc = subprocess.Popen(cmd, stdout=log, stderr=log, stdin=subprocess.DEVNULL,
|
||||
preexec_fn=lambda: os.umask(0))
|
||||
try:
|
||||
_wait_for_port(port, timeout=15)
|
||||
result = _push("127.0.0.1::files/umask_check", port)
|
||||
assert result.returncode == 0, result.stderr or result.stdout
|
||||
received = get_dest_received_dir(sub, SOURCE_DIR)
|
||||
nested = os.path.join(received, "nested")
|
||||
assert os.path.isdir(nested), f"nested dir missing under {received}"
|
||||
mode = stat.S_IMODE(os.stat(nested).st_mode)
|
||||
assert (mode & 0o022) == 0, f"implied directory is group/other writable: {oct(mode)}"
|
||||
finally:
|
||||
_kill_by_cmdline_marker(UMASK_CONF)
|
||||
log.close()
|
||||
try:
|
||||
proc.wait(timeout=5)
|
||||
except subprocess.TimeoutExpired:
|
||||
proc.kill()
|
||||
|
||||
|
||||
class TestDaemonRejection:
|
||||
def _tree_files(self):
|
||||
|
||||
@@ -0,0 +1,291 @@
|
||||
"""Differential coverage for the delete-timing ABORT BOUNDARY (A9/A10).
|
||||
|
||||
rsync's generator runs ahead of its throttled sender, so on a mid-transfer abort
|
||||
it has already removed every extra it planned. FastSync now transmits the
|
||||
COMPLETE per-directory plan set before the first data frame, so an abort has the
|
||||
same effect. Before that change FastSync only removed the extras of the
|
||||
directories its (slower) data stream had reached, and ``-d/--dirs`` used an
|
||||
end-of-transfer commit that removed nothing on abort.
|
||||
|
||||
These tests abort both tools mid-transfer and assert the destination extras
|
||||
removed match real ``rsync 3.4.1``. The rsync side is driven locally with
|
||||
``--bwlimit`` and a small timing window (its generator's delete list is computed
|
||||
long before the throttled payload finishes); the FastSync side uses the
|
||||
byte-deterministic slicing proxy from ``test_delete_timing_parity``.
|
||||
"""
|
||||
import os
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
|
||||
import pytest
|
||||
|
||||
sys.path.insert(0, os.path.dirname(__file__))
|
||||
from common import ( # noqa: E402
|
||||
TEST_DATA_DIR,
|
||||
ServerManager,
|
||||
clean_dir,
|
||||
get_dest_received_dir,
|
||||
run_client,
|
||||
)
|
||||
from test_delete_timing_parity import _SlicingProxy # noqa: E402
|
||||
|
||||
RSYNC = shutil.which("rsync")
|
||||
requires_rsync = pytest.mark.skipif(RSYNC is None, reason="rsync 3.4.1 not installed")
|
||||
|
||||
# Exceeds the 10 MiB scanner chunk, so the next directory lands in a later chunk
|
||||
# (still unreached when the proxy cuts the stream).
|
||||
BIG_BYTES = 16 * 1024 * 1024
|
||||
# Cut well past the (small) config + delete-plan frames and into the big payload,
|
||||
# so the receiver has provably processed every plan before the abort.
|
||||
MID_TRANSFER_BYTES = 256 * 1024
|
||||
PROXY_THROTTLE = 0.001
|
||||
# Throttle rsync's sender so the generator has deleted long before the payload
|
||||
# finishes, then interrupt it mid-transfer.
|
||||
RSYNC_BWLIMIT = 512 # KiB/s -> ~32 s for 16 MiB
|
||||
RSYNC_ABORT_DELAY = 1.5
|
||||
|
||||
|
||||
def _write(path, content):
|
||||
os.makedirs(os.path.dirname(path), exist_ok=True)
|
||||
with open(path, "wb") as fh:
|
||||
fh.write(content)
|
||||
|
||||
|
||||
def _rsync_aborted(args, delay=RSYNC_ABORT_DELAY):
|
||||
"""Start rsync, let its generator run, then interrupt it mid-transfer."""
|
||||
env = dict(os.environ, LC_ALL="C")
|
||||
proc = subprocess.Popen([RSYNC] + args, stdout=subprocess.PIPE, stderr=subprocess.PIPE,
|
||||
text=True, env=env)
|
||||
time.sleep(delay)
|
||||
proc.terminate()
|
||||
try:
|
||||
proc.wait(timeout=10)
|
||||
except subprocess.TimeoutExpired:
|
||||
proc.kill()
|
||||
proc.wait(timeout=5)
|
||||
return proc
|
||||
|
||||
|
||||
class TestDeleteDuringAbortBoundary:
|
||||
"""A9: on an abort, every planned removal has already been applied."""
|
||||
|
||||
def _seed_recursive(self, tag):
|
||||
source = os.path.join(TEST_DATA_DIR, f"dab_{tag}_src")
|
||||
clean_dir(source)
|
||||
# ``a/keep.bin`` sorts first, so the client streams it (and the proxy
|
||||
# cuts) before the data pass ever reaches ``z/deep``.
|
||||
_write(os.path.join(source, "a", "keep.bin"), b"B" * BIG_BYTES)
|
||||
_write(os.path.join(source, "z", "deep", "keep.txt"), b"keep\n")
|
||||
return source
|
||||
|
||||
@requires_rsync
|
||||
def test_recursive_abort_removes_all_planned_extras(self):
|
||||
# ---- FastSync: abort mid ``a/keep.bin``; ``z/deep`` is never reached.
|
||||
source = self._seed_recursive("rec_fs")
|
||||
dest = os.path.join(TEST_DATA_DIR, "dab_rec_fs_dst")
|
||||
clean_dir(dest)
|
||||
received = get_dest_received_dir(dest, source)
|
||||
os.makedirs(os.path.join(received, "a"), exist_ok=True)
|
||||
_write(os.path.join(received, "a", "a_extra"), b"stale\n")
|
||||
os.makedirs(os.path.join(received, "z", "deep"), exist_ok=True)
|
||||
_write(os.path.join(received, "z", "deep", "old_extra"), b"stale\n")
|
||||
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
proxy = _SlicingProxy(server.port, forward_limit=MID_TRANSFER_BYTES,
|
||||
throttle=PROXY_THROTTLE)
|
||||
result, _ = run_client(source, dest, flags=["--delete-during"], port=proxy.port)
|
||||
proxy.finish()
|
||||
assert result.returncode != 0, "truncated transfer reported success"
|
||||
assert not os.path.exists(os.path.join(received, "a", "a_extra"))
|
||||
assert not os.path.exists(os.path.join(received, "z", "deep", "old_extra")), (
|
||||
"FastSync left an extra in a directory it never reached before the abort"
|
||||
)
|
||||
|
||||
# ---- rsync 3.4.1: same tree, same abort, same delete outcome.
|
||||
source = self._seed_recursive("rec_rs")
|
||||
rsync_dst = os.path.join(TEST_DATA_DIR, "dab_rec_rs_dst")
|
||||
clean_dir(rsync_dst)
|
||||
os.makedirs(os.path.join(rsync_dst, "a"), exist_ok=True)
|
||||
_write(os.path.join(rsync_dst, "a", "a_extra"), b"stale\n")
|
||||
os.makedirs(os.path.join(rsync_dst, "z", "deep"), exist_ok=True)
|
||||
_write(os.path.join(rsync_dst, "z", "deep", "old_extra"), b"stale\n")
|
||||
|
||||
proc = _rsync_aborted(["-a", "--delete-during", f"--bwlimit={RSYNC_BWLIMIT}",
|
||||
source + "/", rsync_dst + "/"])
|
||||
assert proc.returncode != 0, "rsync was not actually interrupted"
|
||||
assert not os.path.exists(os.path.join(rsync_dst, "a", "a_extra"))
|
||||
assert not os.path.exists(os.path.join(rsync_dst, "z", "deep", "old_extra")), (
|
||||
"rsync's generator did not delete ahead of its sender"
|
||||
)
|
||||
|
||||
|
||||
class TestDirsDeleteAbortBoundary:
|
||||
"""A10: ``-d/--dirs`` uses per-directory plans like rsync.
|
||||
|
||||
The listed directory's direct extras are removed by the up-front plan while
|
||||
a kept but untraversed subdirectory (and its destination content) is
|
||||
shielded.
|
||||
"""
|
||||
|
||||
def _seed_dirs(self, tag):
|
||||
source = os.path.join(TEST_DATA_DIR, f"ddb_{tag}_src")
|
||||
clean_dir(source)
|
||||
_write(os.path.join(source, "big.bin"), b"B" * BIG_BYTES)
|
||||
_write(os.path.join(source, "subdir", "keep.txt"), b"inner\n")
|
||||
return source
|
||||
|
||||
@pytest.mark.parametrize("fs_flag,rs_flag", [("--delete-during", "--delete-during"),
|
||||
("--delete", "--delete")])
|
||||
@requires_rsync
|
||||
def test_dirs_abort_removes_direct_extras_only(self, fs_flag, rs_flag):
|
||||
label = f"{fs_flag.lstrip('-')}_{rs_flag.lstrip('-')}"
|
||||
# ---- FastSync: ``-d`` lists the immediate children; big.bin streams and
|
||||
# the abort lands mid-payload.
|
||||
source = self._seed_dirs(f"dirs_{label}_fs")
|
||||
dest = os.path.join(TEST_DATA_DIR, f"ddb_{label}_fs_dst")
|
||||
clean_dir(dest)
|
||||
received = get_dest_received_dir(dest, source)
|
||||
_write(os.path.join(received, "old_extra"), b"stale\n")
|
||||
os.makedirs(os.path.join(received, "subdir"), exist_ok=True)
|
||||
_write(os.path.join(received, "subdir", "stale.txt"), b"stale inner\n")
|
||||
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
proxy = _SlicingProxy(server.port, forward_limit=MID_TRANSFER_BYTES,
|
||||
throttle=PROXY_THROTTLE)
|
||||
result, _ = run_client(source + "/", dest, flags=["-d", fs_flag], port=proxy.port)
|
||||
proxy.finish()
|
||||
assert result.returncode != 0, f"{fs_flag}: truncated transfer reported success"
|
||||
assert not os.path.exists(os.path.join(received, "old_extra")), (
|
||||
f"{fs_flag}: the listed directory's direct extra survived the abort"
|
||||
)
|
||||
assert os.path.exists(os.path.join(received, "subdir", "stale.txt")), (
|
||||
f"{fs_flag}: descended into a kept, untraversed subdirectory"
|
||||
)
|
||||
|
||||
# ---- rsync 3.4.1: same shape and same abort.
|
||||
source = self._seed_dirs(f"dirs_{label}_rs")
|
||||
rsync_dst = os.path.join(TEST_DATA_DIR, f"ddb_{label}_rs_dst")
|
||||
clean_dir(rsync_dst)
|
||||
_write(os.path.join(rsync_dst, "old_extra"), b"stale\n")
|
||||
os.makedirs(os.path.join(rsync_dst, "subdir"), exist_ok=True)
|
||||
_write(os.path.join(rsync_dst, "subdir", "stale.txt"), b"stale inner\n")
|
||||
|
||||
proc = _rsync_aborted(["-d", rs_flag, f"--bwlimit={RSYNC_BWLIMIT}",
|
||||
source + "/", rsync_dst + "/"])
|
||||
assert proc.returncode != 0, "rsync was not actually interrupted"
|
||||
assert not os.path.exists(os.path.join(rsync_dst, "old_extra")), (
|
||||
f"rsync {rs_flag}: the listed directory's direct extra survived the abort"
|
||||
)
|
||||
assert os.path.exists(os.path.join(rsync_dst, "subdir", "stale.txt")), (
|
||||
f"rsync {rs_flag}: descended into a kept, untraversed subdirectory"
|
||||
)
|
||||
|
||||
|
||||
def _tree(root):
|
||||
out = []
|
||||
for dirpath, dirs, files in os.walk(root):
|
||||
for name in dirs:
|
||||
out.append(os.path.relpath(os.path.join(dirpath, name), root))
|
||||
for name in files:
|
||||
out.append(os.path.relpath(os.path.join(dirpath, name), root))
|
||||
return sorted(out)
|
||||
|
||||
|
||||
class TestDirsDeleteFinalStateParity:
|
||||
"""A10 completed run: ``-d DIR/ --delete`` (during default) and
|
||||
``--delete-during`` match rsync's final tree, including a kept but
|
||||
untraversed subdirectory whose destination content survives."""
|
||||
|
||||
@pytest.mark.parametrize("flag", ["--delete", "--delete-during"])
|
||||
@requires_rsync
|
||||
def test_dirs_final_state_matches_rsync(self, flag):
|
||||
source = os.path.join(TEST_DATA_DIR, f"ddf_{flag.lstrip('-')}_src")
|
||||
clean_dir(source)
|
||||
_write(os.path.join(source, "keep.txt"), b"new keep\n")
|
||||
_write(os.path.join(source, "subdir", "inner.txt"), b"inner\n")
|
||||
|
||||
def seed_dest(root):
|
||||
clean_dir(root)
|
||||
_write(os.path.join(root, "keep.txt"), b"old keep\n")
|
||||
_write(os.path.join(root, "extra.txt"), b"extra\n")
|
||||
_write(os.path.join(root, "extrasub", "ex.txt"), b"extra sub\n")
|
||||
_write(os.path.join(root, "subdir", "stale.txt"), b"stale inner\n")
|
||||
|
||||
rsync_dst = os.path.join(TEST_DATA_DIR, f"ddf_{flag.lstrip('-')}_rs_dst")
|
||||
seed_dest(rsync_dst)
|
||||
env = dict(os.environ, LC_ALL="C")
|
||||
rsync_result = subprocess.run(
|
||||
[RSYNC, "-d", flag, source + "/", rsync_dst + "/"],
|
||||
capture_output=True, text=True, env=env, timeout=120)
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
rsync_tree = _tree(rsync_dst)
|
||||
|
||||
dest = os.path.join(TEST_DATA_DIR, f"ddf_{flag.lstrip('-')}_fs_dst")
|
||||
clean_dir(dest)
|
||||
received = get_dest_received_dir(dest, source)
|
||||
seed_dest(received)
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
result, _ = run_client(source + "/", dest, flags=["-d", flag], port=server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
fastsync_tree = _tree(received)
|
||||
assert fastsync_tree == rsync_tree, (
|
||||
f"-d {flag}: fastsync tree {fastsync_tree} != rsync tree {rsync_tree}")
|
||||
|
||||
|
||||
class TestOneFileSystemDeleteParity:
|
||||
"""A9 side effect: the per-directory plan is now emitted only for directories
|
||||
whose children were enumerated, so a ``-x`` mount-point directory that is
|
||||
emitted but never traversed is shielded -- its destination content survives,
|
||||
exactly as rsync keeps a non-descended mount point under ``--delete``."""
|
||||
|
||||
@requires_rsync
|
||||
def test_mountpoint_content_survives_delete(self):
|
||||
local = os.stat(".")
|
||||
shm = "/dev/shm"
|
||||
if not os.path.isdir(shm) or os.stat(shm).st_dev == local.st_dev:
|
||||
pytest.skip("no cross-device filesystem available")
|
||||
probe = os.path.join(shm, f"fastsync_dofs_{os.getpid()}")
|
||||
clean_dir(probe)
|
||||
_write(os.path.join(probe, "inside.txt"), b"cross\n")
|
||||
try:
|
||||
source = os.path.join(TEST_DATA_DIR, "dofs_src")
|
||||
clean_dir(source)
|
||||
_write(os.path.join(source, "keep.txt"), b"keep\n")
|
||||
os.symlink(probe, os.path.join(source, "nested_link"))
|
||||
|
||||
def seed_dest(root):
|
||||
clean_dir(root)
|
||||
_write(os.path.join(root, "keep.txt"), b"old\n")
|
||||
_write(os.path.join(root, "nested_link", "stale.txt"), b"stale\n")
|
||||
|
||||
rsync_dst = os.path.join(TEST_DATA_DIR, "dofs_rs_dst")
|
||||
seed_dest(rsync_dst)
|
||||
env = dict(os.environ, LC_ALL="C")
|
||||
rsync_result = subprocess.run(
|
||||
[RSYNC, "-a", "--copy-links", "-x", "--delete-during",
|
||||
source + "/", rsync_dst + "/"],
|
||||
capture_output=True, text=True, env=env, timeout=120)
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
assert os.path.exists(os.path.join(rsync_dst, "nested_link", "stale.txt")), (
|
||||
"rsync unexpectedly descended into the mount point")
|
||||
|
||||
dest = os.path.join(TEST_DATA_DIR, "dofs_fs_dst")
|
||||
clean_dir(dest)
|
||||
received = get_dest_received_dir(dest, source)
|
||||
seed_dest(received)
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
result, _ = run_client(source, dest,
|
||||
flags=["-a", "--copy-links", "-x", "--delete-during"],
|
||||
port=server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
assert os.path.exists(os.path.join(received, "nested_link", "stale.txt")), (
|
||||
"FastSync descended into a non-traversed mount point under --delete")
|
||||
finally:
|
||||
clean_dir(probe)
|
||||
|
||||
@@ -0,0 +1,185 @@
|
||||
"""Differential coverage for ``--delete-delay`` + ``--max-delete`` with a
|
||||
refilled deferred directory.
|
||||
|
||||
FastSync snapshots a directory's extras at plan time (``defer_add``) but charges
|
||||
``--max-delete`` only when a path is actually removed, and its deferred commit
|
||||
re-scans a queued directory and removes content created after the plan -- the
|
||||
same rules as rsync. These tests run both tools on the same fixture and assert
|
||||
both sides remove the late content (recursively) and bound the deletion with
|
||||
``--max-delete`` identically.
|
||||
|
||||
They are not part of the fast PR gate because the rsync side needs a wide
|
||||
real-time injection window (a throttled transfer), while the FastSync side uses
|
||||
the existing byte-deterministic slicing proxy.
|
||||
|
||||
The refilled directory sits at the transfer ROOT, whose delete plan is always
|
||||
processed before any subdirectory's, so the budget is deterministically charged
|
||||
to the refilled entry; the second extra lives under ``b`` and is skipped.
|
||||
"""
|
||||
import os
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
import threading
|
||||
import time
|
||||
|
||||
import pytest
|
||||
|
||||
sys.path.insert(0, os.path.dirname(__file__))
|
||||
from common import ( # noqa: E402
|
||||
TEST_DATA_DIR,
|
||||
ServerManager,
|
||||
clean_dir,
|
||||
get_dest_received_dir,
|
||||
run_client,
|
||||
)
|
||||
from test_delete_timing_parity import _SlicingProxy # noqa: E402
|
||||
|
||||
RSYNC = shutil.which("rsync")
|
||||
requires_rsync = pytest.mark.skipif(RSYNC is None, reason="rsync 3.4.1 not installed")
|
||||
|
||||
BIG_BYTES = 8 * 1024 * 1024
|
||||
MID_TRANSFER_BYTES = 256 * 1024
|
||||
PROXY_THROTTLE = 0.001
|
||||
# rsync is driven locally, so the refill is injected on a wall-clock delay while
|
||||
# a throttled ~8 s transfer is in flight. 1.5 s is safely after rsync's plan
|
||||
# scan (t=0) and well before the deferred commit at the end.
|
||||
RSYNC_BWLIMIT = 1024 # 1 MiB/s
|
||||
RSYNC_INJECT_DELAY = 1.5
|
||||
|
||||
|
||||
def _write(path, content):
|
||||
os.makedirs(os.path.dirname(path), exist_ok=True)
|
||||
with open(path, "wb") as fh:
|
||||
fh.write(content)
|
||||
|
||||
|
||||
def _seed_source(tag):
|
||||
source = os.path.join(TEST_DATA_DIR, f"ddb_{tag}_src")
|
||||
clean_dir(source)
|
||||
_write(os.path.join(source, "a", "keep.bin"), b"B" * BIG_BYTES)
|
||||
_write(os.path.join(source, "b", "keep.txt"), b"keep\n")
|
||||
return source
|
||||
|
||||
|
||||
def _seed_fastsync(tag):
|
||||
"""FastSync mirrors the absolute source path under its receive root, so the
|
||||
extras live below ``received``."""
|
||||
source = _seed_source(tag)
|
||||
dest = os.path.join(TEST_DATA_DIR, f"ddb_{tag}_dst")
|
||||
clean_dir(dest)
|
||||
received = get_dest_received_dir(dest, source)
|
||||
os.makedirs(os.path.join(received, "xdir"), exist_ok=True)
|
||||
os.makedirs(os.path.join(received, "b", "ydir"), exist_ok=True)
|
||||
return source, dest, received
|
||||
|
||||
|
||||
def _seed_rsync(tag):
|
||||
"""rsync mirrors the source contents directly into the destination, so the
|
||||
extras are flat under ``rsync_dst``."""
|
||||
source = _seed_source(tag)
|
||||
rsync_dst = os.path.join(TEST_DATA_DIR, f"ddb_{tag}_dst")
|
||||
clean_dir(rsync_dst)
|
||||
os.makedirs(os.path.join(rsync_dst, "xdir"), exist_ok=True)
|
||||
os.makedirs(os.path.join(rsync_dst, "b", "ydir"), exist_ok=True)
|
||||
return source, rsync_dst
|
||||
|
||||
|
||||
def _deleted_count(text):
|
||||
for line in text.splitlines():
|
||||
if line.startswith("Number of deleted files:"):
|
||||
return int(line.split(":", 1)[1].split()[0])
|
||||
return None
|
||||
|
||||
|
||||
def _rsync(args, timeout=120):
|
||||
env = dict(os.environ, LC_ALL="C")
|
||||
return subprocess.run([RSYNC] + args, capture_output=True, text=True, env=env, timeout=timeout)
|
||||
|
||||
|
||||
class TestDeleteDelayRefilledDirVsRsync:
|
||||
"""Both tools charge --max-delete on actual removals and recurse."""
|
||||
|
||||
def _fastsync_refilled(self, tag, max_delete=None):
|
||||
"""Run FastSync with the refill injected deterministically by the proxy
|
||||
hook (fired once the receiver has processed the plan frames)."""
|
||||
source, dest, received = _seed_fastsync(tag)
|
||||
late = os.path.join(received, "xdir", "new.txt")
|
||||
|
||||
def hook():
|
||||
_write(late, b"created mid-transfer\n")
|
||||
|
||||
flags = ["--delete-delay", "--incremental", "--ignore-times", "--stats"]
|
||||
if max_delete is not None:
|
||||
flags.append(f"--max-delete={max_delete}")
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
proxy = _SlicingProxy(server.port, hook=hook, hook_after=MID_TRANSFER_BYTES,
|
||||
throttle=PROXY_THROTTLE, wait_for_reply=True)
|
||||
result, _ = run_client(source, dest, flags=flags, port=proxy.port)
|
||||
proxy.finish()
|
||||
assert proxy.hook_called.is_set(), "refill hook never fired"
|
||||
return result, received, late
|
||||
|
||||
@requires_rsync
|
||||
def test_max_delete_budget_bound_matches(self):
|
||||
# --- FastSync: the one actual removal is the late file; dirs survive ---
|
||||
result, received, late = self._fastsync_refilled("budget_fs", max_delete=1)
|
||||
assert result.returncode == 25, (result.stderr or result.stdout)[:300]
|
||||
assert _deleted_count(result.stdout) == 1, result.stdout
|
||||
assert not os.path.exists(late), "FastSync kept the late content of a queued dir"
|
||||
assert os.path.isdir(os.path.join(received, "xdir"))
|
||||
assert os.path.isdir(os.path.join(received, "b", "ydir")), (
|
||||
"FastSync did not bound the deletion with --max-delete=1"
|
||||
)
|
||||
|
||||
# --- rsync: same budget rule and recursive removal ---
|
||||
source, rsync_dst = _seed_rsync("budget_rsync")
|
||||
|
||||
def inject():
|
||||
time.sleep(RSYNC_INJECT_DELAY)
|
||||
_write(os.path.join(rsync_dst, "xdir", "new.txt"), b"created mid-transfer\n")
|
||||
|
||||
t = threading.Thread(target=inject)
|
||||
t.start()
|
||||
rsync_result = _rsync(
|
||||
["-a", "--delete-delay", "--max-delete=1", "--stats",
|
||||
f"--bwlimit={RSYNC_BWLIMIT}", source + "/", rsync_dst + "/"]
|
||||
)
|
||||
t.join()
|
||||
assert rsync_result.returncode == 25, rsync_result.stderr
|
||||
assert _deleted_count(rsync_result.stdout) == _deleted_count(result.stdout)
|
||||
assert not os.path.exists(os.path.join(rsync_dst, "xdir", "new.txt")), (
|
||||
"rsync kept late content inside a queued directory"
|
||||
)
|
||||
assert os.path.isdir(os.path.join(rsync_dst, "b", "ydir")), (
|
||||
"rsync did not bound the deletion with --max-delete=1"
|
||||
)
|
||||
assert os.path.isdir(os.path.join(received, "b", "ydir"))
|
||||
|
||||
@requires_rsync
|
||||
def test_refilled_extra_dir_recursive_removal_matches(self):
|
||||
"""Without --max-delete both tools remove the refilled extra directory
|
||||
(and its late content)."""
|
||||
result, received, late = self._fastsync_refilled("recur_fs")
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
assert not os.path.exists(late), "FastSync kept the refilled directory's late content"
|
||||
assert not os.path.isdir(os.path.join(received, "xdir"))
|
||||
|
||||
source, rsync_dst = _seed_rsync("recur_rsync")
|
||||
|
||||
def inject():
|
||||
time.sleep(RSYNC_INJECT_DELAY)
|
||||
_write(os.path.join(rsync_dst, "xdir", "new.txt"), b"created mid-transfer\n")
|
||||
|
||||
t = threading.Thread(target=inject)
|
||||
t.start()
|
||||
rsync_result = _rsync(
|
||||
["-a", "--delete-delay", "--stats", f"--bwlimit={RSYNC_BWLIMIT}",
|
||||
source + "/", rsync_dst + "/"]
|
||||
)
|
||||
t.join()
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
assert not os.path.exists(os.path.join(rsync_dst, "xdir")), (
|
||||
"rsync did not recursively remove the refilled extra directory"
|
||||
)
|
||||
@@ -217,7 +217,21 @@ class _SlicingProxy:
|
||||
|
||||
|
||||
class TestDeleteTimingFinalStateParity:
|
||||
"""On a successful transfer the per-directory timings match rsync's result."""
|
||||
"""On a successful transfer the per-directory timings match rsync's result.
|
||||
|
||||
Plain ``--delete`` has no rsync-incompatible spelling: it defaults to
|
||||
delete-during on both tools, so it is compared against rsync's own default.
|
||||
``--delete-commit`` is FastSync-only and selects the late whole-tree commit,
|
||||
which is rsync's ``--delete-after`` timing.
|
||||
"""
|
||||
|
||||
# (fastsync flag, rsync flag)
|
||||
PAIRS = [
|
||||
("--delete", "--delete"),
|
||||
("--delete-during", "--delete-during"),
|
||||
("--delete-delay", "--delete-delay"),
|
||||
("--delete-commit", "--delete-after"),
|
||||
]
|
||||
|
||||
def _run_fastsync(self, tag, timing):
|
||||
source, dest, received = _seed_pair(tag)
|
||||
@@ -226,28 +240,32 @@ class TestDeleteTimingFinalStateParity:
|
||||
result, _ = run_client(source, dest, flags=[timing], port=server.port)
|
||||
return result, received
|
||||
|
||||
@pytest.mark.parametrize("timing", ["--delete-during", "--delete-delay"])
|
||||
@pytest.mark.parametrize("fs_timing,rs_timing", PAIRS)
|
||||
@requires_rsync
|
||||
def test_success_final_state_matches_rsync(self, timing):
|
||||
def test_success_final_state_matches_rsync(self, fs_timing, rs_timing):
|
||||
# Worker-safe names: xdist may run the parametrizations concurrently, so
|
||||
# the flags are part of every fixture path.
|
||||
label = f"{fs_timing.lstrip('-')}_vs_{rs_timing.lstrip('-')}"
|
||||
# Build the rsync fixture from the same seed so both sides start equal.
|
||||
source, dest, received = _seed_pair("parity_rsync")
|
||||
source, dest, received = _seed_pair(f"parity_rsync_{label}")
|
||||
source2 = source
|
||||
rsync_dst = os.path.join(TEST_DATA_DIR, "dtp_parity_rsync_dst")
|
||||
rsync_dst = os.path.join(TEST_DATA_DIR, f"dtp_rsync_{label}_dst")
|
||||
clean_dir(rsync_dst)
|
||||
# rsync mirrors src/ into dst/; seed the same extra.
|
||||
_write(os.path.join(rsync_dst, "d", "old_extra"), b"stale extra\n")
|
||||
|
||||
rsync_result = _rsync(["-a", timing, source2 + "/", rsync_dst + "/"])
|
||||
rsync_result = _rsync(["-a", rs_timing, source2 + "/", rsync_dst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
rsync_tree = _tree(rsync_dst)
|
||||
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
result, _ = run_client(source, dest, flags=[timing], port=server.port)
|
||||
result, _ = run_client(source, dest, flags=[fs_timing], port=server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
fastsync_tree = _tree(received)
|
||||
assert fastsync_tree == rsync_tree, (
|
||||
f"{timing}: fastsync tree {fastsync_tree} != rsync tree {rsync_tree}"
|
||||
f"{fs_timing} vs rsync {rs_timing}: fastsync tree {fastsync_tree} != "
|
||||
f"rsync tree {rsync_tree}"
|
||||
)
|
||||
|
||||
|
||||
@@ -258,7 +276,8 @@ class TestDeleteTimingTypeConflictParity:
|
||||
@pytest.mark.parametrize("timing", ["--delete-during", "--delete-delay"])
|
||||
@requires_rsync
|
||||
def test_type_conflicts_match_rsync(self, timing):
|
||||
source = os.path.join(TEST_DATA_DIR, "dtc_src")
|
||||
label = timing.lstrip("-")
|
||||
source = os.path.join(TEST_DATA_DIR, f"dtc_{label}_src")
|
||||
clean_dir(source)
|
||||
_write(os.path.join(source, "foo"), b"now a file\n")
|
||||
_write(os.path.join(source, "bar", "inner.txt"), b"now a dir\n")
|
||||
@@ -268,13 +287,13 @@ class TestDeleteTimingTypeConflictParity:
|
||||
_write(os.path.join(root, "foo", "inner.txt"), b"was a dir\n")
|
||||
_write(os.path.join(root, "bar"), b"was a file\n")
|
||||
|
||||
rsync_dst = os.path.join(TEST_DATA_DIR, "dtc_rsync_dst")
|
||||
rsync_dst = os.path.join(TEST_DATA_DIR, f"dtc_{label}_rsync_dst")
|
||||
seed_dest(rsync_dst)
|
||||
rsync_result = _rsync(["-a", timing, source + "/", rsync_dst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
rsync_tree = _tree(rsync_dst)
|
||||
|
||||
dest = os.path.join(TEST_DATA_DIR, "dtc_dst")
|
||||
dest = os.path.join(TEST_DATA_DIR, f"dtc_{label}_dst")
|
||||
clean_dir(dest)
|
||||
received = get_dest_received_dir(dest, source)
|
||||
seed_dest(received)
|
||||
@@ -288,17 +307,28 @@ class TestDeleteTimingTypeConflictParity:
|
||||
|
||||
|
||||
class TestDeleteTimingFailure:
|
||||
"""A mid-transfer failure distinguishes during from delay."""
|
||||
"""A mid-transfer failure distinguishes the during timings from the late
|
||||
commit timings.
|
||||
|
||||
Plain ``--delete`` must behave like ``--delete-during`` (the rsync default),
|
||||
removing the extras of the directories already reached; ``--delete-commit``
|
||||
must behave like ``--delete-after`` and remove nothing until the transfer
|
||||
has fully succeeded.
|
||||
"""
|
||||
|
||||
@pytest.mark.parametrize("mt", [False, True])
|
||||
def test_during_removes_delay_preserves_on_failure(self, mt):
|
||||
source, dest, received = _seed_pair("failure", big=True)
|
||||
source, dest, received = _seed_pair(f"failure_mt{int(mt)}", big=True)
|
||||
extra = os.path.join(received, "d", "old_extra")
|
||||
assert os.path.exists(extra)
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
for timing, expect_removed in (("--delete-during", True),
|
||||
("--delete-delay", False)):
|
||||
for timing, expect_removed in (
|
||||
("--delete-during", True),
|
||||
("--delete", True),
|
||||
("--delete-delay", False),
|
||||
("--delete-commit", False),
|
||||
("--delete-after", False)):
|
||||
# Re-seed the extra before each run.
|
||||
_write(extra, b"stale extra\n")
|
||||
proxy = _SlicingProxy(server.port, forward_limit=MID_TRANSFER_BYTES, throttle=PROXY_THROTTLE)
|
||||
@@ -313,13 +343,99 @@ class TestDeleteTimingFailure:
|
||||
)
|
||||
|
||||
|
||||
class TestDeleteDelayDeletedCount:
|
||||
"""The reported deleted count must reflect entries actually removed."""
|
||||
|
||||
def test_refilled_deferred_dir_is_recursively_removed_and_counted(self):
|
||||
"""A directory snapshotted into a --delete-delay plan that is refilled
|
||||
before the commit is re-scanned and removed recursively (rsync parity):
|
||||
the late file and the directory are both counted as deleted."""
|
||||
source = os.path.join(TEST_DATA_DIR, "ddc_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "ddc_dst")
|
||||
clean_dir(source)
|
||||
clean_dir(dest)
|
||||
_write(os.path.join(source, "d", "keep.txt"), b"kept payload\n")
|
||||
_write(os.path.join(source, "d", "big.bin"), b"B" * BIG_BYTES)
|
||||
received = get_dest_received_dir(dest, source)
|
||||
extra_dir = os.path.join(received, "d", "extradir")
|
||||
os.makedirs(extra_dir, exist_ok=True)
|
||||
|
||||
def hook():
|
||||
# Runs while big.bin is in flight, after d's delete plan was processed.
|
||||
_write(os.path.join(extra_dir, "new.txt"), b"created mid-transfer\n")
|
||||
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
proxy = _SlicingProxy(server.port, hook=hook, hook_after=MID_TRANSFER_BYTES,
|
||||
throttle=PROXY_THROTTLE, wait_for_reply=True)
|
||||
flags = ["--delete-delay", "--incremental", "--ignore-times", "--stats"]
|
||||
result, _ = run_client(source, dest, flags=flags, port=proxy.port)
|
||||
proxy.finish()
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:400]
|
||||
assert proxy.hook_called.is_set(), "hook never fired"
|
||||
assert not os.path.exists(os.path.join(extra_dir, "new.txt")), "late file survived"
|
||||
assert not os.path.isdir(extra_dir), "refilled extra dir survived"
|
||||
deleted = None
|
||||
for line in result.stdout.splitlines():
|
||||
if line.startswith("Number of deleted files:"):
|
||||
deleted = int(line.split(":", 1)[1].split()[0])
|
||||
assert deleted == 2, (deleted, result.stdout)
|
||||
|
||||
|
||||
class TestDeleteDelayMaxDeleteParity:
|
||||
"""--max-delete with --delete-delay: a partial deletion still reports the
|
||||
number of entries actually removed, matching rsync (the exact surviving set
|
||||
can differ; only the count is compared)."""
|
||||
|
||||
@requires_rsync
|
||||
def test_max_delete_count_matches_rsync(self):
|
||||
source = os.path.join(TEST_DATA_DIR, "ddm_src")
|
||||
rsync_dst = os.path.join(TEST_DATA_DIR, "ddm_rsync_dst")
|
||||
clean_dir(source)
|
||||
clean_dir(rsync_dst)
|
||||
_write(os.path.join(source, "d", "keep.txt"), b"keep\n")
|
||||
for i in range(1, 6):
|
||||
_write(os.path.join(rsync_dst, "d", f"e{i}.txt"), f"extra{i}\n".encode())
|
||||
|
||||
rsync_result = _rsync(["-a", "--delete-delay", "--max-delete=2", "--stats",
|
||||
source + "/", rsync_dst + "/"])
|
||||
# rsync exits 25 ("the --max-delete limit stopped deletions").
|
||||
assert rsync_result.returncode == 25, rsync_result.stderr
|
||||
rsync_count = _deleted_count(rsync_result.stdout)
|
||||
assert rsync_count == 2, rsync_result.stdout
|
||||
|
||||
dest = os.path.join(TEST_DATA_DIR, "ddm_dst")
|
||||
clean_dir(dest)
|
||||
received = get_dest_received_dir(dest, source)
|
||||
for i in range(1, 6):
|
||||
_write(os.path.join(received, "d", f"e{i}.txt"), f"extra{i}\n".encode())
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
result, _ = run_client(
|
||||
source, dest,
|
||||
flags=["--delete-delay", "--max-delete=2", "--stats"],
|
||||
port=server.port,
|
||||
)
|
||||
# A capped --max-delete commit is a successful transfer that both tools
|
||||
# report with exit 25.
|
||||
assert result.returncode == 25, (result.stderr or result.stdout)[:300]
|
||||
assert _deleted_count(result.stdout) == rsync_count, result.stdout
|
||||
|
||||
|
||||
def _deleted_count(text):
|
||||
for line in text.splitlines():
|
||||
if line.startswith("Number of deleted files:"):
|
||||
return int(line.split(":", 1)[1].split()[0])
|
||||
return None
|
||||
|
||||
|
||||
class TestDeleteDelayVsAfterSnapshot:
|
||||
"""A destination entry created after its directory's scan survives under
|
||||
--delete-delay but is removed by --delete-after's fresh end scan."""
|
||||
|
||||
@pytest.mark.parametrize("mt", [False, True])
|
||||
def test_late_created_extra_survives_delay_not_after(self, mt):
|
||||
source, dest, received = _seed_pair("latecreate", big=True)
|
||||
source, dest, received = _seed_pair(f"latecreate_mt{int(mt)}", big=True)
|
||||
old_extra = os.path.join(received, "d", "old_extra")
|
||||
new_extra = os.path.join(received, "d", "new_extra")
|
||||
with ServerManager() as server:
|
||||
@@ -356,3 +472,68 @@ class TestDeleteDelayVsAfterSnapshot:
|
||||
f"{timing} (mt={mt}): new_extra present="
|
||||
f"{os.path.exists(new_extra)}, expected survives={new_survives}"
|
||||
)
|
||||
|
||||
|
||||
class TestDeleteAfterThreadsKeepSet:
|
||||
"""Regression: -j/--threads must still transmit the delete keep-set in every
|
||||
timing. PipelineContextSender.delete_suppressed was left uninitialized, so a
|
||||
garbage true silently skipped the late keep-set manifest under --threads.
|
||||
Plain --delete now uses the per-directory plans, while --delete-commit /
|
||||
--delete-after keep exercising the late whole-tree manifest."""
|
||||
|
||||
@pytest.mark.parametrize("delete_flag", ["--delete", "--delete-commit", "--delete-after"])
|
||||
def test_threads_delete_after_sends_keep_set(self, delete_flag):
|
||||
source, dest, received = _seed_pair("mtkeep")
|
||||
extra = os.path.join(received, "d", "old_extra")
|
||||
assert os.path.exists(extra)
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
result, _ = run_client(source, dest, flags=["--threads", delete_flag],
|
||||
port=server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
assert not os.path.exists(extra), (
|
||||
f"{delete_flag} --threads did not remove an extra: delete keep-set was suppressed"
|
||||
)
|
||||
|
||||
class TestDeleteDelayMaxDeleteRefilledDir:
|
||||
"""--delete-delay charges the --max-delete budget on ACTUAL removals: the
|
||||
refilled directory's late content is removed first (consuming the one slot),
|
||||
so the directory itself and a later extra are skipped, matching rsync.
|
||||
|
||||
The refilled directory is at the destination ROOT (its plan is always sent
|
||||
first) and the skipped extra is under a separate source directory, so the
|
||||
ordering that decides the budget charge is deterministic -- not readdir
|
||||
order. The refill is injected through the byte-barrier proxy so it is
|
||||
causally after the plan frame."""
|
||||
|
||||
def test_budget_charged_on_actual_removal(self):
|
||||
source = os.path.join(TEST_DATA_DIR, "ddmb_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "ddmb_dst")
|
||||
clean_dir(source)
|
||||
clean_dir(dest)
|
||||
_write(os.path.join(source, "a", "keep.bin"), b"B" * BIG_BYTES)
|
||||
_write(os.path.join(source, "b", "keep.txt"), b"keep\n")
|
||||
received = get_dest_received_dir(dest, source)
|
||||
refilled_dir = os.path.join(received, "xdir")
|
||||
os.makedirs(refilled_dir, exist_ok=True)
|
||||
later_dir = os.path.join(received, "b", "ydir")
|
||||
os.makedirs(later_dir, exist_ok=True)
|
||||
|
||||
def hook():
|
||||
_write(os.path.join(refilled_dir, "new.txt"), b"created mid-transfer\n")
|
||||
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
proxy = _SlicingProxy(server.port, hook=hook, hook_after=MID_TRANSFER_BYTES,
|
||||
throttle=PROXY_THROTTLE, wait_for_reply=True)
|
||||
flags = ["--delete-delay", "--max-delete=1", "--incremental", "--ignore-times", "--stats"]
|
||||
result, _ = run_client(source, dest, flags=flags, port=proxy.port)
|
||||
proxy.finish()
|
||||
assert result.returncode == 25, (result.stderr or result.stdout)[:400]
|
||||
assert proxy.hook_called.is_set(), "hook never fired"
|
||||
# The late content consumes the single budget slot; the refilled
|
||||
# directory itself and the later extra are skipped.
|
||||
assert not os.path.exists(os.path.join(refilled_dir, "new.txt")), "late file survived"
|
||||
assert os.path.isdir(later_dir), "later extra was not skipped by the budget"
|
||||
# The one actual removal is reported.
|
||||
assert _deleted_count(result.stdout) == 1, result.stdout
|
||||
|
||||
@@ -0,0 +1,719 @@
|
||||
"""Differential rsync-parity gate.
|
||||
|
||||
Runs real ``rsync 3.4.1`` and FastSync over the same corpora and flags, then
|
||||
compares the destination trees and the normalized output of the
|
||||
output-oriented flags. This is the executable counterpart of
|
||||
``RSYNC_COMPAT.md``: the fast subset (``-m parity_ci``) guards the ✅ surface on
|
||||
every pull request, and the full set (``-m parity``) burns the documented
|
||||
⚠️/❌ residuals down.
|
||||
|
||||
Known, documented differences live in ``parity_caveats.py``; anything else
|
||||
fails with a readable tree/stdout diff. A stale allowlist entry is reported
|
||||
loudly (and fails when ``FASTSYNC_PARITY_STRICT=1``).
|
||||
|
||||
Run locally::
|
||||
|
||||
python3 -m pytest tests/integration/test_differential_parity.py -n 4 --dist=load -m parity_ci
|
||||
python3 -m pytest tests/integration/test_differential_parity.py -n 4 --dist=load -m parity
|
||||
"""
|
||||
import os
|
||||
import shutil
|
||||
import sys
|
||||
import warnings
|
||||
|
||||
import pytest
|
||||
|
||||
sys.path.insert(0, os.path.dirname(__file__))
|
||||
from common import ( # noqa: E402
|
||||
ServerManager,
|
||||
TEST_DATA_DIR,
|
||||
clean_dir,
|
||||
get_dest_received_dir,
|
||||
)
|
||||
from parity_caveats import ASPECTS, caveat_for # noqa: E402
|
||||
import parity_harness as H # noqa: E402
|
||||
|
||||
RSYNC = shutil.which("rsync")
|
||||
requires_rsync = pytest.mark.skipif(RSYNC is None, reason="rsync 3.4.1 not installed")
|
||||
|
||||
# `--allow-super` matches the rest of the integration suite; `--allow-delete`
|
||||
# is needed only by the delete cases.
|
||||
SUPER = ("--allow-super",)
|
||||
DELETE = ("--allow-super", "--allow-delete")
|
||||
_OLD_MTIME = 1_500_000_000
|
||||
|
||||
parity = pytest.mark.parity
|
||||
parity_ci = pytest.mark.parity_ci
|
||||
|
||||
|
||||
@pytest.fixture(scope="session")
|
||||
def parity_server_factory():
|
||||
"""Lazily start one server per distinct extra-argument set, per xdist worker."""
|
||||
servers = {}
|
||||
|
||||
def get(extra):
|
||||
key = tuple(extra)
|
||||
if key not in servers:
|
||||
s = ServerManager()
|
||||
s.start(extra_args=list(extra))
|
||||
servers[key] = s
|
||||
return servers[key]
|
||||
|
||||
yield get
|
||||
for s in servers.values():
|
||||
s.stop()
|
||||
|
||||
|
||||
def _pin(path, mtime):
|
||||
os.utime(path, (mtime, mtime))
|
||||
|
||||
|
||||
def _mk(path, data, mtime=None):
|
||||
os.makedirs(os.path.dirname(path), exist_ok=True)
|
||||
with open(path, "wb") as fh:
|
||||
fh.write(data)
|
||||
if mtime is not None:
|
||||
_pin(path, mtime)
|
||||
|
||||
|
||||
# --- destination seeds ------------------------------------------------------
|
||||
|
||||
def seed_extras(_src, rroot, froot):
|
||||
for root in (rroot, froot):
|
||||
_mk(os.path.join(root, "extra.txt"), b"extra\n")
|
||||
_mk(os.path.join(root, "extradir", "z.txt"), b"z\n")
|
||||
|
||||
|
||||
def seed_update(_src, rroot, froot):
|
||||
for root in (rroot, froot):
|
||||
p = os.path.join(root, "a.txt")
|
||||
_mk(p, b"destination is newer and longer\n", 2_000_000_000)
|
||||
|
||||
|
||||
def seed_ignore_existing(_src, rroot, froot):
|
||||
for root in (rroot, froot):
|
||||
_mk(os.path.join(root, "a.txt"), b"destination-kept\n", _OLD_MTIME)
|
||||
|
||||
|
||||
def seed_append(_src, rroot, froot):
|
||||
for root in (rroot, froot):
|
||||
_mk(os.path.join(root, "a.txt"), b"hello ", _OLD_MTIME)
|
||||
|
||||
|
||||
def seed_backup(_src, rroot, froot):
|
||||
for root in (rroot, froot):
|
||||
_mk(os.path.join(root, "a.txt"), b"OLD-CONTENT\n", _OLD_MTIME)
|
||||
|
||||
|
||||
def seed_size_only(_src, rroot, froot):
|
||||
for root in (rroot, froot):
|
||||
_mk(os.path.join(root, "a.txt"), b"XXXXXXXXXXX\n", _OLD_MTIME)
|
||||
|
||||
|
||||
def seed_delete_excluded(_src, rroot, froot):
|
||||
for root in (rroot, froot):
|
||||
_mk(os.path.join(root, "drop.log"), b"stale log\n", _OLD_MTIME)
|
||||
_mk(os.path.join(root, "extra.txt"), b"extra\n", _OLD_MTIME)
|
||||
_mk(os.path.join(root, "keep.txt"), b"keep\n", _OLD_MTIME)
|
||||
|
||||
|
||||
def seed_filter_protect(_src, rroot, froot):
|
||||
"""Destination-only entries, including nested ones, for the receiver-side
|
||||
`protect` rule: the `.log` extras must survive --delete, the rest go."""
|
||||
for root in (rroot, froot):
|
||||
_mk(os.path.join(root, "extra.log"), b"dest-only log\n", _OLD_MTIME)
|
||||
_mk(os.path.join(root, "other.txt"), b"dest-only other\n", _OLD_MTIME)
|
||||
_mk(os.path.join(root, "sub", "extra2.log"), b"nested dest-only log\n", _OLD_MTIME)
|
||||
_mk(os.path.join(root, "sub", "other2.txt"), b"nested dest-only other\n", _OLD_MTIME)
|
||||
|
||||
|
||||
def seed_max_delete(_src, rroot, froot):
|
||||
for root in (rroot, froot):
|
||||
_mk(os.path.join(root, "extra1.txt"), b"e1\n", _OLD_MTIME)
|
||||
_mk(os.path.join(root, "extra2.txt"), b"e2\n", _OLD_MTIME)
|
||||
|
||||
|
||||
def fuzzy_basis_seed(_src, rroot, froot):
|
||||
"""Seed a same-suffix sibling whose name is one edit from the source and
|
||||
whose content matches it, with a DIFFERENT mtime so rsync's exact
|
||||
size+mtime pass cannot fire: both tools must select it via the
|
||||
name-distance pass. Where the two tools' basis choices coincide the
|
||||
block-level results are identical when the block size is pinned."""
|
||||
for root in (rroot, froot):
|
||||
_mk(os.path.join(root, "report_v1.txt"), H.FUZZY_PAYLOAD, _OLD_MTIME)
|
||||
|
||||
|
||||
def max_delete_count_check(_src, rroot, froot, _rs, _fs):
|
||||
"""The exact survivor set is order-dependent; the count must still match."""
|
||||
r = H.snapshot(rroot)
|
||||
f = H.snapshot(froot)
|
||||
if len(r) != len(f):
|
||||
return [f"survivor count differs: rsync={len(r)} fastsync={len(f)}"]
|
||||
return []
|
||||
|
||||
|
||||
# --- case table -------------------------------------------------------------
|
||||
|
||||
_CASES = [
|
||||
# --- core archive / recursion -----------------------------------------
|
||||
H.Case("archive", "basic", ["-a"], ci=True, ref="-a/--archive"),
|
||||
H.Case("recursive", "basic", ["-r"], ci=True, ref="-r/--recursive"),
|
||||
H.Case("unicode_names", "unicode", ["-a"], ci=True, ref="-a unicode names"),
|
||||
H.Case("links_archive", "links", ["-a"], ci=True, ref="-l/--links"),
|
||||
H.Case("copy_links", "links", ["-aL"], ref="-L/--copy-links"),
|
||||
H.Case("hardlinks", "hardlinks", ["-a", "-H"], compare_hardlinks=True,
|
||||
ci=True, ref="-H/--hard-links"),
|
||||
H.Case("hardlinks_without_H", "hardlinks", ["-a"], compare_hardlinks=True,
|
||||
ref="hardlinks without -H"),
|
||||
H.Case("sparse", "sparse", ["-a", "-S"], ref="-S/--sparse"),
|
||||
|
||||
# --- compression / checksums ------------------------------------------
|
||||
H.Case("compress_zstd", "basic", ["-a", "-z"], ci=True, ref="-z/--compress"),
|
||||
H.Case("checksum", "basic", ["-a", "-c"], ref="-c/--checksum"),
|
||||
H.Case("checksum_choice_xxh64", "basic",
|
||||
["-a", "-c", "--checksum-choice=xxh64"], ref="--checksum-choice"),
|
||||
|
||||
# --- selection --------------------------------------------------------
|
||||
H.Case("exclude", "filters", ["-a", "--exclude=*.log"], ci=True,
|
||||
ref="--exclude"),
|
||||
H.Case("include_exclude", "filters",
|
||||
["-a", "--include=*.txt", "--exclude=*"], ci=True,
|
||||
ref="--include/--exclude ordering"),
|
||||
H.Case("filter_rules", "filters",
|
||||
["-a", "-f", "- *.log", "-f", "+ *.txt", "-f", "- *"],
|
||||
ref="--filter/-f grammar"),
|
||||
H.Case("max_size", "basic", ["-a", "--max-size=1000"], ref="--max-size"),
|
||||
H.Case("min_size", "basic", ["-a", "--min-size=1000"], ref="--min-size"),
|
||||
|
||||
# --- output-oriented --------------------------------------------------
|
||||
H.Case("stats", "basic", ["-a", "--stats"], stdout=H.STDOUT_STATS,
|
||||
ci=True, ref="--stats"),
|
||||
H.Case("itemize", "links", ["-a", "-i"], stdout=H.STDOUT_ITEMIZE,
|
||||
ci=True, ref="-i/--itemize-changes"),
|
||||
H.Case("out_format_n_l", "basic", ["-a", "--out-format=%n %l"],
|
||||
stdout=H.STDOUT_OUTFMT, ref="--out-format %n %l"),
|
||||
H.Case("out_format_i_n", "basic", ["-a", "--out-format=%i %n"],
|
||||
stdout=H.STDOUT_OUTFMT, ref="--out-format %i %n"),
|
||||
H.Case("progress", "multidir", ["-a", "--progress"], stdout=H.STDOUT_PROGRESS,
|
||||
ci=True, ref="--progress multi-directory file list"),
|
||||
H.Case("progress_threads", "multidir", ["-a", "--progress"],
|
||||
fastsync_flags=["-a", "--progress", "--threads"],
|
||||
stdout=H.STDOUT_PROGRESS, ci=True,
|
||||
ref="--progress multi-directory file list (--threads)"),
|
||||
|
||||
# --- transfer modifications -------------------------------------------
|
||||
H.Case("update", "basic", ["-a", "--update"], seed=seed_update,
|
||||
ref="-u/--update"),
|
||||
H.Case("ignore_existing", "basic", ["-a", "--ignore-existing"],
|
||||
seed=seed_ignore_existing, ci=True, ref="--ignore-existing"),
|
||||
H.Case("size_only", "basic",
|
||||
["-a", "--size-only"], fastsync_flags=["-a", "--incremental", "--size-only"],
|
||||
seed=seed_size_only, ref="--size-only"),
|
||||
H.Case("append", "basic", ["-a", "--append"], seed=seed_append,
|
||||
ref="--append"),
|
||||
H.Case("append_verify", "basic", ["-a", "--append-verify"], seed=seed_append,
|
||||
ref="--append-verify"),
|
||||
H.Case("backup", "basic", ["-a", "--backup"], seed=seed_backup,
|
||||
ref="--backup"),
|
||||
H.Case("chmod", "basic", ["-a", "--chmod=Fu+rwx"], compare_modes=True,
|
||||
ci=True, ref="--chmod"),
|
||||
|
||||
# --- delta / similar-file basis (--fuzzy) -----------------------------
|
||||
# Basis choices coincide here (same-suffix sibling, name distance one edit,
|
||||
# content identical); with the block size pinned both tools report the same
|
||||
# Matched/Literal/transferred counters. The residual (FastSync's narrower
|
||||
# delta size window) is covered by TestFuzzy in test_parity_quickwins.py.
|
||||
H.Case("fuzzy_basis", "fuzzy",
|
||||
["-a", "--no-whole-file", "--fuzzy", "--stats", "-B8192"],
|
||||
fastsync_flags=["-a", "--incremental", "--delta", "--fuzzy",
|
||||
"--stats", "--delta-block=8192"],
|
||||
seed=fuzzy_basis_seed, stdout=H.STDOUT_STATS, ci=True,
|
||||
ref="-y/--fuzzy similar-file basis"),
|
||||
|
||||
# --- deletion ---------------------------------------------------------
|
||||
H.Case("delete", "basic", ["-a", "--delete"], seed=seed_extras,
|
||||
server_args=DELETE, ci=True, ref="--delete"),
|
||||
H.Case("delete_before", "basic", ["-a", "--delete-before"], seed=seed_extras,
|
||||
server_args=DELETE, ref="--delete-before"),
|
||||
H.Case("delete_during", "basic", ["-a", "--delete-during"], seed=seed_extras,
|
||||
server_args=DELETE, ref="--delete-during"),
|
||||
H.Case("delete_delay", "basic", ["-a", "--delete-delay"], seed=seed_extras,
|
||||
server_args=DELETE, ref="--delete-delay"),
|
||||
H.Case("delete_after", "basic", ["-a", "--delete-after"], seed=seed_extras,
|
||||
server_args=DELETE, ref="--delete-after"),
|
||||
H.Case("delete_commit", "basic", ["-a", "--delete-after"], seed=seed_extras,
|
||||
fastsync_flags=["-a", "--delete-commit"], server_args=DELETE,
|
||||
ref="FastSync-only --delete-commit == rsync --delete-after"),
|
||||
H.Case("delete_excluded", "filters",
|
||||
["-a", "--delete", "--delete-excluded", "--exclude=*.log"],
|
||||
seed=seed_delete_excluded, server_args=DELETE, ref="--delete-excluded"),
|
||||
H.Case("exclude_protect_dest_only", "filters",
|
||||
["-a", "--delete", "--exclude=*.log"],
|
||||
seed=seed_delete_excluded, server_args=DELETE, ci=True,
|
||||
ref="--delete protects a destination-only excluded entry like rsync"),
|
||||
H.Case("max_delete", "basic", ["-a", "--delete", "--max-delete=1"],
|
||||
seed=seed_max_delete, server_args=DELETE,
|
||||
extra_check=max_delete_count_check, compare_tree=False,
|
||||
ref="--max-delete"),
|
||||
H.Case("filter_protect", "filters",
|
||||
["-a", "--delete", "--filter=P *.log"],
|
||||
seed=seed_filter_protect, server_args=DELETE, ci=True,
|
||||
ref="--filter P/--protect receiver-side delete protection (default during)"),
|
||||
H.Case("filter_protect_during", "filters",
|
||||
["-a", "--delete-during", "--filter=P *.log"],
|
||||
seed=seed_filter_protect, server_args=DELETE, ci=True,
|
||||
ref="--filter P/--protect under --delete-during"),
|
||||
H.Case("filter_protect_delay", "filters",
|
||||
["-a", "--delete-delay", "--filter=P *.log"],
|
||||
seed=seed_filter_protect, server_args=DELETE, ci=True,
|
||||
ref="--filter P/--protect under --delete-delay"),
|
||||
H.Case("filter_protect_after", "filters",
|
||||
["-a", "--delete-after", "--filter=P *.log"],
|
||||
seed=seed_filter_protect, server_args=DELETE, ci=True,
|
||||
ref="--filter P/--protect under the whole-tree --delete-after commit"),
|
||||
|
||||
# --- relative / dirs --------------------------------------------------
|
||||
H.Case("relative_general", "basic", ["-a", "-R"], layout=H.MIRROR_ABS,
|
||||
compare_modes=True, ref="-R/--relative"),
|
||||
H.Case("relative_no_implied_dirs", "basic",
|
||||
["-a", "-R", "--no-implied-dirs"], layout=H.MIRROR_ABS,
|
||||
compare_modes=True, ref="--no-implied-dirs"),
|
||||
H.Case("files_from", "relative", ["--dirs", "-R"],
|
||||
files_from=("dir1", "sub/x.txt"), layout=H.RELATIVE, ci=True,
|
||||
ref="-d/--dirs + --files-from"),
|
||||
H.Case("dirs_plain", "basic", ["-d"], fs_src_suffix="/",
|
||||
ref="-d/--dirs (plain)"),
|
||||
H.Case("empty_dirs_recursive", "empty_dir", ["-a"],
|
||||
ref="recursive empty-directory residual"),
|
||||
H.Case("empty_dirs_files_from", "empty_dir", ["--dirs", "-R"],
|
||||
files_from=("emptydir",), layout=H.RELATIVE, ci=True,
|
||||
ref="-d/--dirs explicit empty directory"),
|
||||
|
||||
# --- codecs -----------------------------------------------------------
|
||||
H.Case("iconv_identity", "basic", ["-a", "--iconv=UTF-8,UTF-8"],
|
||||
ref="--iconv identity"),
|
||||
H.Case("iconv_convert", "iconv",
|
||||
["-a", "--iconv=ISO-8859-1,UTF-8"],
|
||||
server_args=("--allow-super", "--iconv=UTF-8"),
|
||||
ref="--iconv conversion (receiver declares its own charset)"),
|
||||
# rsync's spec is LOCAL,REMOTE and the destination end's charset is REMOTE
|
||||
# on a push, so a default server writes the wire (UTF-8) names verbatim.
|
||||
H.Case("iconv_default_server", "iconv",
|
||||
["-a", "--iconv=ISO-8859-1,UTF-8"],
|
||||
ref="--iconv push direction (default receiver charset = REMOTE)"),
|
||||
|
||||
# --- partial ----------------------------------------------------------
|
||||
H.Case("partial_complete", "basic", ["-a", "--partial"], ref="--partial"),
|
||||
]
|
||||
|
||||
# Cases that must always be tolerated (documented ⚠️/❌ residuals) get an
|
||||
# allowlist entry; the table below stays the exact ✅ surface.
|
||||
ALL_CASES = _CASES
|
||||
|
||||
|
||||
def _params():
|
||||
out = []
|
||||
for case in ALL_CASES:
|
||||
marks = [parity]
|
||||
if case.ci:
|
||||
marks.append(parity_ci)
|
||||
out.append(pytest.param(case, id=case.id, marks=marks))
|
||||
return out
|
||||
|
||||
|
||||
def _aspects_to_check(result):
|
||||
return {
|
||||
"tree": result["tree"],
|
||||
"stdout": result["stdout"],
|
||||
"extra": result["extra"],
|
||||
}
|
||||
|
||||
|
||||
def _assert_no_unexpected(case_id, mismatches, caveat, ref=""):
|
||||
unexpected = {a: v for a, v in mismatches.items() if v and a not in caveat}
|
||||
if unexpected:
|
||||
lines = [f"differential parity mismatch for case {case_id!r}:"]
|
||||
lines.append(f" ref: {ref or 'see RSYNC_COMPAT.md'}")
|
||||
for aspect, detail in unexpected.items():
|
||||
lines.append(f" --- {aspect} ---")
|
||||
lines.extend(" " + str(d) for d in detail)
|
||||
lines.append("If this is a documented residual, add it to "
|
||||
"tests/integration/parity_caveats.py with a RSYNC_COMPAT.md "
|
||||
"reference. Do not allowlist an undocumented divergence.")
|
||||
pytest.fail("\n".join(lines))
|
||||
|
||||
stale = [a for a in ASPECTS
|
||||
if a in caveat and a != "rc" and not mismatches.get(a)]
|
||||
if stale:
|
||||
msg = (f"stale parity allowlist entry for case {case_id!r}, aspect(s) "
|
||||
f"{stale}: FastSync now matches rsync. Remove it from "
|
||||
f"parity_caveats.py (and update RSYNC_COMPAT.md if the row moved).")
|
||||
if os.environ.get("FASTSYNC_PARITY_STRICT") == "1":
|
||||
pytest.fail(msg)
|
||||
warnings.warn(msg, stacklevel=2)
|
||||
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.parametrize("case", _params())
|
||||
def test_differential_case(case, parity_server_factory):
|
||||
server = parity_server_factory(case.server_args)
|
||||
result = H.execute_case(case, server)
|
||||
caveat = caveat_for(case.id)
|
||||
|
||||
mismatches = _aspects_to_check(result)
|
||||
if result["rsync_rc"] != result["fastsync_rc"]:
|
||||
mismatches["rc"] = [
|
||||
f"rsync rc={result['rsync_rc']} fastsync rc={result['fastsync_rc']} "
|
||||
f"(rsync stderr: {result['rsync_stderr'][:200]!r}, "
|
||||
f"fastsync stderr: {result['fastsync_stderr'][:200]!r})"]
|
||||
_assert_no_unexpected(case.id, mismatches, caveat, ref=case.ref)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Multi-run and setup-heavy scenarios (kept as explicit tests)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
def _result_aspects(result):
|
||||
return _aspects_to_check(result)
|
||||
|
||||
|
||||
_STANDALONE_REFS = {
|
||||
"incremental_modified": "-i/--itemize-changes + incremental second run",
|
||||
"compare_dest": "--compare-dest",
|
||||
"copy_dest": "--copy-dest",
|
||||
"link_dest": "--link-dest",
|
||||
"link_dest_stats": "--link-dest + --stats",
|
||||
"verify_basis": "--verify-basis (FastSync-only)",
|
||||
"verify_basis_default": "--verify-basis (default quick-check vs rsync)",
|
||||
"added_and_deleted": "--delete across two runs",
|
||||
"added_and_deleted_seed": "--delete across two runs",
|
||||
"one_file_system": "-x/--one-file-system",
|
||||
}
|
||||
|
||||
|
||||
def _run_and_check(case_id, result, ref=""):
|
||||
mismatches = _result_aspects(result)
|
||||
if result["rsync_rc"] != result["fastsync_rc"]:
|
||||
mismatches["rc"] = [
|
||||
f"rsync rc={result['rsync_rc']} fastsync rc={result['fastsync_rc']} "
|
||||
f"(rsync stderr: {result['rsync_stderr'][:200]!r}, "
|
||||
f"fastsync stderr: {result['fastsync_stderr'][:200]!r})"]
|
||||
_assert_no_unexpected(case_id, mismatches, caveat_for(case_id),
|
||||
ref=ref or _STANDALONE_REFS.get(case_id, ""))
|
||||
|
||||
|
||||
@requires_rsync
|
||||
@parity
|
||||
def test_incremental_modified_file(parity_server_factory):
|
||||
"""A second run sends only the modified file; destinations stay identical."""
|
||||
case_id = "incremental_modified"
|
||||
src = os.path.join(TEST_DATA_DIR, "parity_inc_src")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "parity_inc_rdst")
|
||||
fdst = os.path.join(TEST_DATA_DIR, "parity_inc_fdst")
|
||||
H.CORPORA["basic"](src)
|
||||
server = parity_server_factory(SUPER)
|
||||
|
||||
# Seed both destinations with the initial content.
|
||||
H.run_differential(src, rdst, fdst, ["-a"], ["-a"], server,
|
||||
extra_check=lambda *a: [])
|
||||
with open(os.path.join(src, "a.txt"), "wb") as fh:
|
||||
fh.write(b"hello world, now modified and longer\n")
|
||||
_pin(os.path.join(src, "a.txt"), 1_650_000_000)
|
||||
|
||||
result = H.run_differential(
|
||||
src, rdst, fdst, ["-a", "-i"], ["-a", "-i", "--incremental"], server,
|
||||
stdout=H.STDOUT_ITEMIZE)
|
||||
_run_and_check(case_id, result)
|
||||
|
||||
|
||||
def _seed_basis(rel_entries):
|
||||
def seed(src, rroot, froot):
|
||||
for root in (rroot, froot):
|
||||
os.makedirs(root, exist_ok=True)
|
||||
for rel, data in rel_entries.items():
|
||||
_mk(os.path.join(root, rel), data)
|
||||
return seed
|
||||
|
||||
|
||||
@requires_rsync
|
||||
@parity
|
||||
def test_compare_dest_skips_basis(parity_server_factory):
|
||||
"""--compare-dest: a file present in the basis is not copied."""
|
||||
case_id = "compare_dest"
|
||||
src = os.path.join(TEST_DATA_DIR, "parity_cmpd_src")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "parity_cmpd_rdst")
|
||||
fdst = os.path.join(TEST_DATA_DIR, "parity_cmpd_fdst")
|
||||
clean_dir(src)
|
||||
_mk(os.path.join(src, "f.txt"), b"basis-content\n")
|
||||
_pin(os.path.join(src, "f.txt"), _OLD_MTIME)
|
||||
server = parity_server_factory(SUPER)
|
||||
rel = os.path.abspath(src).lstrip(os.sep)
|
||||
|
||||
# rsync resolves --compare-dest relative to the destination dir; FastSync
|
||||
# resolves it under the receive root and appends the mirrored source path.
|
||||
# Both rely on rsync's size+mtime quick-check, so the basis mtime is pinned
|
||||
# to the source's to keep the match deterministic across a second boundary.
|
||||
def seed(_src, rroot, froot):
|
||||
_mk(os.path.join(rroot, "basis", "f.txt"), b"basis-content\n", _OLD_MTIME)
|
||||
_mk(os.path.join(fdst, "basis", rel, "f.txt"), b"basis-content\n", _OLD_MTIME)
|
||||
|
||||
def extra(_src, rroot, froot, _rs, _fs):
|
||||
out = []
|
||||
for label, root in (("rsync", rroot), ("fastsync", froot)):
|
||||
if os.path.exists(os.path.join(root, "f.txt")):
|
||||
out.append(f"{label} copied a file that is present in the "
|
||||
f"compare basis")
|
||||
return out
|
||||
|
||||
result = H.run_differential(
|
||||
src, rdst, fdst,
|
||||
["-a", "--compare-dest=basis"],
|
||||
["-a", f"--compare-dest={os.path.join(fdst, 'basis')}", "--incremental"],
|
||||
server, seed=seed, ignore_paths=("basis",), extra_check=extra)
|
||||
_run_and_check(case_id, result)
|
||||
|
||||
|
||||
@requires_rsync
|
||||
@parity
|
||||
def test_link_dest_hardlinks_basis(parity_server_factory):
|
||||
"""--link-dest: an unchanged file is hard-linked to the basis, not copied."""
|
||||
case_id = "link_dest"
|
||||
src = os.path.join(TEST_DATA_DIR, "parity_linkd_src")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "parity_linkd_rdst")
|
||||
fdst = os.path.join(TEST_DATA_DIR, "parity_linkd_fdst")
|
||||
clean_dir(src)
|
||||
_mk(os.path.join(src, "f.txt"), b"link-basis-content\n")
|
||||
_pin(os.path.join(src, "f.txt"), _OLD_MTIME)
|
||||
server = parity_server_factory(SUPER)
|
||||
rel = os.path.abspath(src).lstrip(os.sep)
|
||||
|
||||
def seed(_src, rroot, froot):
|
||||
_mk(os.path.join(rroot, "basis", "f.txt"), b"link-basis-content\n", _OLD_MTIME)
|
||||
_mk(os.path.join(fdst, "basis", rel, "f.txt"), b"link-basis-content\n", _OLD_MTIME)
|
||||
|
||||
def extra(_src, rroot, froot, _rs, _fs):
|
||||
r_basis = os.stat(os.path.join(rroot, "basis", "f.txt")).st_ino
|
||||
f_basis = os.stat(os.path.join(fdst, "basis", rel, "f.txt")).st_ino
|
||||
out = []
|
||||
for label, root, basis in (("rsync", rroot, r_basis),
|
||||
("fastsync", froot, f_basis)):
|
||||
target = os.path.join(root, "f.txt")
|
||||
if not os.path.exists(target):
|
||||
out.append(f"{label}: f.txt missing")
|
||||
elif os.stat(target).st_ino != basis:
|
||||
out.append(f"{label}: f.txt is not hard-linked to the basis")
|
||||
return out
|
||||
|
||||
result = H.run_differential(
|
||||
src, rdst, fdst,
|
||||
["-a", "--link-dest=basis"],
|
||||
["-a", f"--link-dest={os.path.join(fdst, 'basis')}", "--incremental"],
|
||||
server, seed=seed, ignore_paths=("basis",), extra_check=extra)
|
||||
_run_and_check(case_id, result)
|
||||
|
||||
|
||||
@requires_rsync
|
||||
@parity
|
||||
def test_link_dest_stats_matches_rsync(parity_server_factory):
|
||||
"""A basis hit must not be counted as created or literal data: rsync reports
|
||||
zero for both, so FastSync's receiver tallies must too (regression for the
|
||||
basis materialization over-report)."""
|
||||
case_id = "link_dest_stats"
|
||||
src = os.path.join(TEST_DATA_DIR, "parity_linkds_src")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "parity_linkds_rdst")
|
||||
fdst = os.path.join(TEST_DATA_DIR, "parity_linkds_fdst")
|
||||
clean_dir(src)
|
||||
_mk(os.path.join(src, "f.txt"), b"link-basis-content\n")
|
||||
_pin(os.path.join(src, "f.txt"), _OLD_MTIME)
|
||||
server = parity_server_factory(SUPER)
|
||||
rel = os.path.abspath(src).lstrip(os.sep)
|
||||
|
||||
def seed(_src, rroot, froot):
|
||||
_mk(os.path.join(rroot, "basis", "f.txt"), b"link-basis-content\n", _OLD_MTIME)
|
||||
_mk(os.path.join(fdst, "basis", rel, "f.txt"), b"link-basis-content\n", _OLD_MTIME)
|
||||
|
||||
result = H.run_differential(
|
||||
src, rdst, fdst,
|
||||
["-a", "--link-dest=basis", "--stats"],
|
||||
["-a", f"--link-dest={os.path.join(fdst, 'basis')}", "--incremental", "--stats"],
|
||||
server, seed=seed, ignore_paths=("basis",), stdout=H.STDOUT_STATS)
|
||||
_run_and_check(case_id, result, ref="--link-dest + --stats")
|
||||
|
||||
|
||||
@requires_rsync
|
||||
@parity
|
||||
def test_copy_dest_copies_basis(parity_server_factory):
|
||||
"""--copy-dest: a basis match is materialized as an independent copy with the
|
||||
source's attributes, matching rsync (copy then fix attributes)."""
|
||||
case_id = "copy_dest"
|
||||
src = os.path.join(TEST_DATA_DIR, "parity_copyd_src")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "parity_copyd_rdst")
|
||||
fdst = os.path.join(TEST_DATA_DIR, "parity_copyd_fdst")
|
||||
clean_dir(src)
|
||||
_mk(os.path.join(src, "f.txt"), b"copy-basis-content\n")
|
||||
_pin(os.path.join(src, "f.txt"), 1_600_000_000)
|
||||
os.chmod(os.path.join(src, "f.txt"), 0o755)
|
||||
server = parity_server_factory(SUPER)
|
||||
rel = os.path.abspath(src).lstrip(os.sep)
|
||||
|
||||
def seed(_src, rroot, froot):
|
||||
# Basis content matches the source; give the basis a different mode so a
|
||||
# wrong "keep basis attributes" implementation is visible.
|
||||
_mk(os.path.join(rroot, "basis", "f.txt"), b"copy-basis-content\n",
|
||||
1_600_000_000)
|
||||
os.chmod(os.path.join(rroot, "basis", "f.txt"), 0o644)
|
||||
_mk(os.path.join(fdst, "basis", rel, "f.txt"), b"copy-basis-content\n",
|
||||
1_600_000_000)
|
||||
os.chmod(os.path.join(fdst, "basis", rel, "f.txt"), 0o644)
|
||||
|
||||
def extra(_src, rroot, froot, _rs, _fs):
|
||||
out = []
|
||||
bases = {"rsync": os.path.join(rroot, "basis", "f.txt"),
|
||||
"fastsync": os.path.join(fdst, "basis", rel, "f.txt")}
|
||||
for label, root in (("rsync", rroot), ("fastsync", froot)):
|
||||
target = os.path.join(root, "f.txt")
|
||||
if not os.path.exists(target):
|
||||
out.append(f"{label}: f.txt missing")
|
||||
continue
|
||||
if os.stat(target).st_ino == os.stat(bases[label]).st_ino:
|
||||
out.append(f"{label}: f.txt is hard-linked, not copied")
|
||||
if (os.stat(target).st_mode & 0o777) != 0o755:
|
||||
out.append(f"{label}: f.txt mode "
|
||||
f"{oct(os.stat(target).st_mode & 0o777)} != 0o755")
|
||||
return out
|
||||
|
||||
result = H.run_differential(
|
||||
src, rdst, fdst,
|
||||
["-a", "--copy-dest=basis"],
|
||||
["-a", f"--copy-dest={os.path.join(fdst, 'basis')}", "--incremental"],
|
||||
server, seed=seed, ignore_paths=("basis",), extra_check=extra,
|
||||
compare_modes=True)
|
||||
_run_and_check(case_id, result)
|
||||
|
||||
|
||||
@requires_rsync
|
||||
@parity
|
||||
def test_verify_basis_restores_strict_content(parity_server_factory):
|
||||
"""Default matches rsync's metadata quick-check; FastSync-only
|
||||
`--verify-basis` restores strict content equality and transfers the source
|
||||
when a same-size/different-content basis would otherwise be trusted."""
|
||||
case_id = "verify_basis"
|
||||
src = os.path.join(TEST_DATA_DIR, "parity_vbasis_src")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "parity_vbasis_rdst")
|
||||
fdst = os.path.join(TEST_DATA_DIR, "parity_vbasis_fdst")
|
||||
clean_dir(src)
|
||||
_mk(os.path.join(src, "f.txt"), b"AAAA\n")
|
||||
_pin(os.path.join(src, "f.txt"), _OLD_MTIME)
|
||||
server = parity_server_factory(SUPER)
|
||||
rel = os.path.abspath(src).lstrip(os.sep)
|
||||
|
||||
def seed(_src, rroot, froot):
|
||||
# Same size and mtime as the source, different bytes: a metadata
|
||||
# quick-check trusts it; --verify-basis must not.
|
||||
for root, basis_rel in ((rroot, os.path.join("basis", "f.txt")),
|
||||
(fdst, os.path.join("basis", rel, "f.txt"))):
|
||||
_mk(os.path.join(root, basis_rel), b"BBBB\n", _OLD_MTIME)
|
||||
|
||||
# Default: both tools trust the basis (rsync's quick check), so the
|
||||
# destination carries the basis bytes and the trees match.
|
||||
result = H.run_differential(
|
||||
src, rdst, fdst,
|
||||
["-a", "--link-dest=basis"],
|
||||
["-a", f"--link-dest={os.path.join(fdst, 'basis')}", "--incremental"],
|
||||
server, seed=seed, ignore_paths=("basis",))
|
||||
_run_and_check(case_id + "_default", result)
|
||||
|
||||
# --verify-basis (FastSync only): the digest mismatch rejects the basis and
|
||||
# the source is transferred, so the destination is the source bytes. rsync
|
||||
# has no such flag; assert the FastSync outcome directly against the source.
|
||||
fdst2 = os.path.join(TEST_DATA_DIR, "parity_vbasis_fdst2")
|
||||
clean_dir(fdst2)
|
||||
for root, basis_rel in ((fdst2, os.path.join("basis", rel, "f.txt")),):
|
||||
_mk(os.path.join(root, basis_rel), b"BBBB\n", _OLD_MTIME)
|
||||
result, _ = H.run_fastsync(src, fdst2,
|
||||
["-a", f"--link-dest={os.path.join(fdst2, 'basis')}",
|
||||
"--incremental", "--verify-basis"], server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
target = os.path.join(get_dest_received_dir(fdst2, src), "f.txt")
|
||||
with open(target, "rb") as fh:
|
||||
assert fh.read() == b"AAAA\n", \
|
||||
"--verify-basis must reject the same-size/different-content basis"
|
||||
|
||||
|
||||
@requires_rsync
|
||||
@parity
|
||||
def test_added_and_deleted_between_runs(parity_server_factory):
|
||||
"""A source deletion and addition sync correctly under --delete."""
|
||||
case_id = "added_and_deleted"
|
||||
src = os.path.join(TEST_DATA_DIR, "parity_addel_src")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "parity_addel_rdst")
|
||||
fdst = os.path.join(TEST_DATA_DIR, "parity_addel_fdst")
|
||||
server = parity_server_factory(DELETE)
|
||||
H.CORPORA["basic"](src)
|
||||
|
||||
seed = seed_extras
|
||||
result = H.run_differential(
|
||||
src, rdst, fdst, ["-a", "--delete"], ["-a", "--delete"], server,
|
||||
seed=seed)
|
||||
_run_and_check(case_id + "_seed", result)
|
||||
|
||||
os.remove(os.path.join(src, "a.txt"))
|
||||
_mk(os.path.join(src, "added.txt"), b"added between runs\n")
|
||||
result = H.run_differential(
|
||||
src, rdst, fdst, ["-a", "--delete", "-i"],
|
||||
["-a", "--delete", "-i", "--incremental"], server,
|
||||
stdout=H.STDOUT_ITEMIZE)
|
||||
_run_and_check(case_id, result)
|
||||
|
||||
|
||||
@requires_rsync
|
||||
@parity
|
||||
def test_one_file_system(parity_server_factory):
|
||||
"""-x emits the mount-point directory but not its contents."""
|
||||
case_id = "one_file_system"
|
||||
local = os.stat(".")
|
||||
shm = "/dev/shm"
|
||||
if not os.path.isdir(shm):
|
||||
pytest.skip("/dev/shm not available")
|
||||
if os.stat(shm).st_dev == local.st_dev:
|
||||
pytest.skip("no cross-device filesystem available")
|
||||
|
||||
src = os.path.join(TEST_DATA_DIR, "parity_ofs_src")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "parity_ofs_rdst")
|
||||
fdst = os.path.join(TEST_DATA_DIR, "parity_ofs_fdst")
|
||||
clean_dir(src)
|
||||
_mk(os.path.join(src, "keep.txt"), b"keep\n")
|
||||
probe = os.path.join(shm, f"fastsync_ofs_{os.getpid()}")
|
||||
shutil.rmtree(probe, ignore_errors=True)
|
||||
os.makedirs(probe)
|
||||
_mk(os.path.join(probe, "inside.txt"), b"cross\n")
|
||||
try:
|
||||
os.symlink(probe, os.path.join(src, "nested_link"))
|
||||
server = parity_server_factory(SUPER)
|
||||
result = H.run_differential(
|
||||
src, rdst, fdst,
|
||||
["-a", "--copy-links", "-x"],
|
||||
["-a", "--copy-links", "-x"], server)
|
||||
_run_and_check(case_id, result)
|
||||
finally:
|
||||
shutil.rmtree(probe, ignore_errors=True)
|
||||
|
||||
|
||||
@requires_rsync
|
||||
@parity
|
||||
def test_parity_caveats_reference_known_cases():
|
||||
"""Every allowlist entry must name a real case id and aspect."""
|
||||
from parity_caveats import CAVEATS
|
||||
known = {c.id for c in ALL_CASES} | {
|
||||
"incremental_modified", "compare_dest", "link_dest",
|
||||
"added_and_deleted", "added_and_deleted_seed", "one_file_system",
|
||||
}
|
||||
problems = []
|
||||
for case_id, entry in CAVEATS.items():
|
||||
if case_id not in known:
|
||||
problems.append(f"unknown case id in parity_caveats.py: {case_id!r}")
|
||||
for aspect in entry:
|
||||
if aspect not in ASPECTS:
|
||||
problems.append(
|
||||
f"{case_id!r}: unknown aspect {aspect!r} (expected {ASPECTS})")
|
||||
assert not problems, "\n".join(problems)
|
||||
@@ -36,7 +36,7 @@ from common import ( # noqa: E402
|
||||
verify_transfer,
|
||||
)
|
||||
|
||||
PROTOCOL_VERSION = b"2.26.0"
|
||||
PROTOCOL_VERSION = b"2.28.0"
|
||||
STATUS_MANIFEST = 5
|
||||
STATUS_OK = 0
|
||||
|
||||
|
||||
+401
-106
@@ -583,6 +583,57 @@ class TestRemoteDryRun:
|
||||
assert os.path.exists(extra), f"{flags} deleted an extra in dry-run"
|
||||
assert _snapshot_tree(received) == before, f"{flags} mutated the destination"
|
||||
|
||||
@pytest.mark.skipif(shutil.which("rsync") is None, reason="rsync not installed")
|
||||
def test_dry_run_delete_lines_match_rsync(self):
|
||||
"""`-n --delete` lists exactly the destination extras rsync would remove.
|
||||
|
||||
Covers the three cases that a real run protects: the file being updated
|
||||
(in the keep set), a filter-excluded source entry (protected prefix), and
|
||||
a --max-size-pruned source entry (always-protected prefix). Track 4a
|
||||
adds a fourth: a destination-only entry matching the exclude rule is
|
||||
re-derived on the receiver and also protected, so only the genuine
|
||||
destination-only `extra.txt` appears.
|
||||
"""
|
||||
source = os.path.join(TEST_DATA_DIR, "dryrep_src")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "dryrep_rdst")
|
||||
fdst = os.path.join(TEST_DATA_DIR, "dryrep_fdst")
|
||||
clean_dir(source)
|
||||
clean_dir(rdst)
|
||||
clean_dir(fdst)
|
||||
for name, data in (("a.txt", b"new content\n"), ("keep.log", b"log\n"),
|
||||
("big.bin", b"B" * 2000)):
|
||||
with open(os.path.join(source, name), "wb") as fh:
|
||||
fh.write(data)
|
||||
os.utime(os.path.join(source, "a.txt"), (1_700_000_000, 1_700_000_000))
|
||||
received = get_dest_received_dir(fdst, source)
|
||||
os.makedirs(received, exist_ok=True)
|
||||
for root in (rdst, received):
|
||||
for name, data in (("a.txt", b"old\n"), ("keep.log", b"log\n"),
|
||||
("big.bin", b"B" * 2000), ("extra.txt", b"extra\n"),
|
||||
("stray.log", b"dest only\n")):
|
||||
with open(os.path.join(root, name), "wb") as fh:
|
||||
fh.write(data)
|
||||
os.utime(os.path.join(root, name), (1_500_000_000, 1_500_000_000))
|
||||
|
||||
flags = ["-a", "-n", "-i", "--delete", "--exclude=*.log", "--max-size=1000"]
|
||||
r = subprocess.run(["rsync", "-an", "-i", "--delete", "--exclude=*.log",
|
||||
"--max-size=1000", source + "/", rdst + "/"],
|
||||
capture_output=True, text=True,
|
||||
env=dict(os.environ, LC_ALL="C"))
|
||||
assert r.returncode == 0, r.stderr
|
||||
rsync_del = sorted(l for l in r.stdout.splitlines() if l.startswith("*deleting"))
|
||||
assert rsync_del == ["*deleting extra.txt"], f"unexpected rsync set: {rsync_del}"
|
||||
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
result, _ = run_client(source, fdst, flags=flags, port=server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
fs_del = sorted(l for l in (result.stdout or "").splitlines()
|
||||
if l.startswith("*deleting"))
|
||||
assert fs_del == rsync_del, f"rsync={rsync_del}\nfastsync={fs_del}"
|
||||
assert os.path.exists(os.path.join(received, "stray.log")), \
|
||||
"destination-only exclude match must be protected in the dry-run report"
|
||||
|
||||
@pytest.mark.ci
|
||||
def test_remote_dry_run_quiet_is_silent(self, shared_server):
|
||||
source = os.path.join(TEST_DATA_DIR, "remote_dry_quiet_src")
|
||||
@@ -2224,6 +2275,36 @@ class TestPartialDir:
|
||||
partial = os.path.join(dest, ".partial", os.path.relpath(source_file, os.path.sep))
|
||||
assert not os.path.exists(partial)
|
||||
|
||||
def test_partial_dir_alone_implies_partial(self, shared_server):
|
||||
"""--partial-dir=DIR with no --partial implies --partial, like rsync.
|
||||
|
||||
rsync 3.4.1 retains the staged partial when --partial-dir is given by
|
||||
itself; before the implication was added FastSync discarded it. The
|
||||
transfer is made to fail deterministically by placing a non-empty
|
||||
directory at the destination path, so the final partial-dir ->
|
||||
destination rename fails and whatever was staged under the partial dir
|
||||
stays on disk."""
|
||||
source = os.path.join(TEST_DATA_DIR, "partial_dir_implied_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "partial_dir_implied_dst")
|
||||
clean_dir(source)
|
||||
clean_dir(dest)
|
||||
source_file = os.path.join(source, "f.bin")
|
||||
with open(source_file, "wb") as f:
|
||||
f.write(b"partial payload")
|
||||
|
||||
received = get_dest_received_dir(dest, source)
|
||||
os.makedirs(os.path.join(received, "f.bin"))
|
||||
with open(os.path.join(received, "f.bin", "keep"), "wb") as f:
|
||||
f.write(b"keep")
|
||||
|
||||
result, _ = run_client(source, dest, flags=["--partial-dir=.partial"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode != 0, "expected the blocked install to fail"
|
||||
|
||||
partial = os.path.join(dest, ".partial", os.path.relpath(source_file, os.path.sep))
|
||||
assert os.path.exists(partial), \
|
||||
"--partial-dir alone must imply --partial and retain the partial file"
|
||||
|
||||
|
||||
class TestLargeFile:
|
||||
def test_transfer_100mb_file(self, shared_server):
|
||||
@@ -2376,6 +2457,61 @@ class TestTempDir:
|
||||
assert not mismatches, f"Mismatch: {mismatches}"
|
||||
self._assert_clean_scratch(os.path.join(dest, "scratch"))
|
||||
|
||||
@pytest.mark.skipif(shutil.which("rsync") is None, reason="rsync not installed")
|
||||
def test_relative_temp_dir_matches_rsync_absolute_rejected(self):
|
||||
"""Differential: a relative --temp-dir is resolved under the destination
|
||||
by both (rsync 3.4.1 and FastSync), producing identical trees. An
|
||||
absolute --temp-dir is used verbatim by rsync standalone, but the
|
||||
receiver deliberately confines it to the receive root (security
|
||||
invariant), so FastSync rejects it without writing outside the root.
|
||||
"""
|
||||
source = self._make_source("tempdir_diff_src")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "tempdir_diff_rdst")
|
||||
fdst = os.path.join(TEST_DATA_DIR, "tempdir_diff_fdst")
|
||||
clean_dir(rdst)
|
||||
clean_dir(fdst)
|
||||
os.makedirs(os.path.join(rdst, "scratch"), exist_ok=True)
|
||||
os.makedirs(os.path.join(fdst, "scratch"), exist_ok=True)
|
||||
r = subprocess.run(["rsync", "-a", "--temp-dir=scratch", source + "/", rdst + "/"],
|
||||
capture_output=True, text=True,
|
||||
env=dict(os.environ, LC_ALL="C"))
|
||||
assert r.returncode == 0, r.stderr
|
||||
with ServerManager() as server:
|
||||
server.start()
|
||||
result, _ = run_client(source, fdst, flags=["--temp-dir=scratch"],
|
||||
port=server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
# rsync lays the source contents directly in rdst; FastSync mirrors the
|
||||
# absolute source path below fdst. Compare the mirrored content trees
|
||||
# (the scratch dir lives at each destination root).
|
||||
rtree = sorted(os.path.relpath(os.path.join(dp, n), rdst)
|
||||
for dp, dn, fn in os.walk(rdst)
|
||||
for n in dn + fn if os.path.join(dp, n) != os.path.join(rdst, "scratch"))
|
||||
mirror = get_dest_received_dir(fdst, source)
|
||||
ftree = sorted(os.path.relpath(os.path.join(dp, n), mirror)
|
||||
for dp, dn, fn in os.walk(mirror) for n in dn + fn)
|
||||
assert rtree == ftree, f"relative temp-dir tree mismatch: {rtree} != {ftree}"
|
||||
assert _walk_tmp_files(os.path.join(rdst, "scratch")) == []
|
||||
assert _walk_tmp_files(os.path.join(fdst, "scratch")) == []
|
||||
|
||||
# Absolute temp dir: rsync accepts it; FastSync rejects it safely.
|
||||
abs_scratch = os.path.join(TEST_DATA_DIR, "tempdir_diff_abs")
|
||||
clean_dir(abs_scratch)
|
||||
rdst2 = os.path.join(TEST_DATA_DIR, "tempdir_diff_rdst2")
|
||||
clean_dir(rdst2)
|
||||
r2 = subprocess.run(["rsync", "-a", "--temp-dir=" + abs_scratch, source + "/", rdst2 + "/"],
|
||||
capture_output=True, text=True,
|
||||
env=dict(os.environ, LC_ALL="C"))
|
||||
assert r2.returncode == 0, r2.stderr
|
||||
fdst2 = os.path.join(TEST_DATA_DIR, "tempdir_diff_fdst2")
|
||||
clean_dir(fdst2)
|
||||
with ServerManager() as server:
|
||||
server.start()
|
||||
result2, _ = run_client(source, fdst2, flags=["--temp-dir", abs_scratch],
|
||||
port=server.port)
|
||||
assert result2.returncode != 0, "an absolute --temp-dir must be rejected (confined)"
|
||||
assert os.listdir(abs_scratch) == [], "receiver wrote into an unconfined temp dir"
|
||||
|
||||
def test_default_behavior_has_no_scratch_dir(self, shared_server):
|
||||
source = self._make_source("tempdir_default_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "tempdir_default_dst")
|
||||
@@ -2500,35 +2636,30 @@ class TestTimeoutAndAllocLimits:
|
||||
mismatches, missing = verify_transfer(source, received)
|
||||
assert not missing and not mismatches
|
||||
|
||||
def test_temp_dir_cross_filesystem_fallback(self, shared_server):
|
||||
"""A confined relative --temp-dir that resolves (via a symlink under the
|
||||
destination root) to another filesystem must fall back to a non-atomic
|
||||
copy instead of aborting (rsync parity). Skipped when no second
|
||||
filesystem is available."""
|
||||
shm = "/dev/shm"
|
||||
if not os.path.isdir(shm):
|
||||
pytest.skip("/dev/shm not available")
|
||||
if os.stat(shm).st_dev == os.stat(TEST_DATA_DIR).st_dev:
|
||||
pytest.skip("/dev/shm is on the same filesystem as the test data")
|
||||
scratch = os.path.join(shm, f"fastsync_tmp_{os.getpid()}")
|
||||
shutil.rmtree(scratch, ignore_errors=True)
|
||||
os.makedirs(scratch)
|
||||
def test_temp_dir_symlink_escape_rejected(self, shared_server):
|
||||
"""A symlink planted inside the destination root pointing outside it
|
||||
must not redirect receiver scratch files: --temp-dir=<that link> is
|
||||
refused and nothing is written at the link target. An in-root symlink
|
||||
(e.g. to a mount point that stays inside the authorized root) is still
|
||||
accepted, preserving the engine's EXDEV cross-filesystem fallback."""
|
||||
source, dest = self._seed("tempdir_escape_src")
|
||||
outside = "/tmp/fastsync_tempdir_escape_%d" % os.getpid()
|
||||
shutil.rmtree(outside, ignore_errors=True)
|
||||
os.makedirs(outside)
|
||||
link = os.path.join(dest, "escape_scratch")
|
||||
if os.path.lexists(link):
|
||||
os.unlink(link)
|
||||
os.symlink(outside, link)
|
||||
try:
|
||||
source, dest = self._seed("tempdir_xdev_src")
|
||||
# The receiver resolves a relative temp dir under the destination
|
||||
# root; a symlink there points the scratch at the second filesystem.
|
||||
link = os.path.join(dest, "xdev_scratch")
|
||||
os.symlink(scratch, link)
|
||||
result, _ = run_client(source, dest, flags=["--temp-dir", "xdev_scratch"],
|
||||
result, _ = run_client(source, dest, flags=["--temp-dir", "escape_scratch"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, f"cross-fs temp-dir failed: {result.stderr[:300]}"
|
||||
assert result.returncode != 0, "an escaping --temp-dir symlink must be refused"
|
||||
received = get_dest_received_dir(dest, source)
|
||||
mismatches, missing = verify_transfer(source, received)
|
||||
assert not missing, f"Missing: {missing}"
|
||||
assert not mismatches, f"Mismatch: {mismatches}"
|
||||
assert os.listdir(scratch) == [], "temp files left behind in the cross-fs scratch"
|
||||
assert not os.path.exists(os.path.join(received, "f.txt")), \
|
||||
"the receiver must not fall back to writing the file"
|
||||
assert os.listdir(outside) == [], "receiver wrote outside the authorized root"
|
||||
finally:
|
||||
shutil.rmtree(scratch, ignore_errors=True)
|
||||
shutil.rmtree(outside, ignore_errors=True)
|
||||
|
||||
|
||||
class TestRemoteOptionTransport:
|
||||
@@ -2655,7 +2786,12 @@ class TestItemizeChanges:
|
||||
flags=["--preserve", "-i", "--incremental"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, f"incremental itemize failed: {result.stderr[:200]}"
|
||||
itemized = [line for line in result.stdout.splitlines() if line and line[0] in ">.<c"]
|
||||
# -i also emits the transfer-root and directory lines; only FILE entries
|
||||
# matter here, so drop any line whose name has a trailing '/'.
|
||||
itemized = [
|
||||
line for line in result.stdout.splitlines()
|
||||
if line and line[0] in ">.<c" and not line.rsplit(" ", 1)[-1].endswith("/")
|
||||
]
|
||||
assert itemized == [], f"unchanged files were itemized: {itemized[:5]}"
|
||||
|
||||
def test_multithreaded_emits_same_itemize_lines(self, shared_server):
|
||||
@@ -2837,6 +2973,41 @@ class TestDelayUpdates:
|
||||
assert not os.path.isdir(os.path.join(delay_dest, self.STAGING)), \
|
||||
"staging directory left behind after a successful delayed transfer"
|
||||
|
||||
@pytest.mark.skipif(shutil.which("rsync") is None, reason="rsync not installed")
|
||||
def test_delay_updates_staging_name_collision_residual(self):
|
||||
"""Documented residual (RSYNC_COMPAT.md `--delay-updates` row): FastSync
|
||||
uses a fixed `.fastsync-stage` staging name and wipes a pre-existing tree
|
||||
of that name at the start of a delayed run (crash-leftover cleanup),
|
||||
even without `--delete`; rsync leaves a genuine destination entry of that
|
||||
name untouched. Pins the divergence that keeps the row Divergent."""
|
||||
source = self._make_source("delay_collide_src")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "delay_collide_rdst")
|
||||
fdst = os.path.join(TEST_DATA_DIR, "delay_collide_fdst")
|
||||
clean_dir(rdst)
|
||||
clean_dir(fdst)
|
||||
for root in (rdst, fdst):
|
||||
with open(os.path.join(root, "top.txt"), "wb") as fh:
|
||||
fh.write(b"old\n")
|
||||
stage = os.path.join(root, self.STAGING)
|
||||
os.makedirs(stage, exist_ok=True)
|
||||
with open(os.path.join(stage, "keepme.txt"), "wb") as fh:
|
||||
fh.write(b"genuine user data\n")
|
||||
|
||||
r = subprocess.run(["rsync", "-a", "--delay-updates", source + "/", rdst + "/"],
|
||||
capture_output=True, text=True,
|
||||
env=dict(os.environ, LC_ALL="C"))
|
||||
assert r.returncode == 0, r.stderr
|
||||
assert os.path.exists(os.path.join(rdst, self.STAGING, "keepme.txt")), \
|
||||
"rsync removed an unrelated destination entry named like the staging dir"
|
||||
|
||||
with ServerManager() as server:
|
||||
server.start()
|
||||
result, _ = run_client(source, fdst, flags=["--delay-updates"],
|
||||
port=server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
assert not os.path.exists(os.path.join(fdst, self.STAGING)), \
|
||||
"FastSync did not wipe the reserved staging name (residual changed)"
|
||||
|
||||
@pytest.mark.parametrize("mt", [False, True])
|
||||
def test_delay_updates_incremental_rerun_no_leftovers(self, shared_server, mt):
|
||||
source = self._make_source("delay_rerun_src")
|
||||
@@ -3523,23 +3694,25 @@ class TestMissingArgs:
|
||||
|
||||
|
||||
class TestNoImpliedDirs:
|
||||
"""--no-implied-dirs (only meaningful with -R + --files-from) refuses to
|
||||
place a listed file whose parent directory is not itself listed."""
|
||||
"""--no-implied-dirs (meaningful with -R) omits the source metadata of a
|
||||
listed path's implied parent directories but still creates those parents
|
||||
with default attributes, matching rsync 3.4.1."""
|
||||
|
||||
def _make(self):
|
||||
return _make_relative_source("noimplied_src")
|
||||
|
||||
@pytest.mark.parametrize("mt", [False, True])
|
||||
def test_implied_dir_only_fails_entry(self, shared_server, mt):
|
||||
def test_implied_dir_created_with_default_attrs(self, shared_server, mt):
|
||||
source = self._make()
|
||||
dest = os.path.join(TEST_DATA_DIR, "noimplied_dst")
|
||||
clean_dir(dest)
|
||||
lst = _write_rel_list(b"a/b.txt\n") # "a" itself is not listed
|
||||
flags = ["--files-from", lst, "-R", "--no-implied-dirs"] + (["--threads"] if mt else [])
|
||||
result, _ = run_client(source, dest, flags=flags, port=shared_server.port)
|
||||
assert result.returncode != 0, "implied parent directory was not rejected"
|
||||
assert "--no-implied-dirs" in (result.stderr or result.stdout)
|
||||
assert not os.path.exists(os.path.join(dest, "a", "b.txt"))
|
||||
assert result.returncode == 0, \
|
||||
f"implied parent directory was not created: {result.stderr[:200]}"
|
||||
assert os.path.isdir(os.path.join(dest, "a")), "implied parent 'a' was not created"
|
||||
assert _read_file(os.path.join(dest, "a", "b.txt")) == b"nested\n"
|
||||
|
||||
@pytest.mark.parametrize("mt", [False, True])
|
||||
def test_listed_dir_allows_file(self, shared_server, mt):
|
||||
@@ -3666,9 +3839,10 @@ class TestDirs:
|
||||
files.extend(os.path.relpath(os.path.join(root, n), mirror) for n in names)
|
||||
assert files == [], f"--dirs descended into contents: {files}"
|
||||
|
||||
def test_dirs_listed_dir_colliding_with_file_fails(self, shared_server):
|
||||
"""A listed directory that already exists as a regular file at the
|
||||
destination fails the transfer cleanly instead of clobbering the file."""
|
||||
def test_dirs_listed_dir_replaces_blocking_file(self, shared_server):
|
||||
"""rsync parity: a listed directory replaces a regular file already at
|
||||
its destination path (rsync removes the non-directory and creates the
|
||||
directory)."""
|
||||
source = self._make()
|
||||
dest = os.path.join(TEST_DATA_DIR, "dirs_coll_dst")
|
||||
clean_dir(dest)
|
||||
@@ -3678,8 +3852,10 @@ class TestDirs:
|
||||
lst = _write_rel_list(b"dir1\n")
|
||||
result, _ = run_client(source, dest, flags=["--files-from", lst, "--dirs", "-R"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode != 0, "dir entry over an existing file did not fail"
|
||||
assert os.path.isfile(blocker), "blocking regular file was clobbered"
|
||||
assert result.returncode == 0, \
|
||||
f"dir entry over an existing file failed: {(result.stderr or result.stdout)[:300]}"
|
||||
assert os.path.isdir(blocker) and not os.path.islink(blocker), \
|
||||
"blocking regular file was not replaced by the incoming directory"
|
||||
|
||||
|
||||
class TestMkpath:
|
||||
@@ -3836,15 +4012,16 @@ class TestDeleteTiming:
|
||||
assert _read_file(os.path.join(received, "sub", "deep.txt")) == b"deeply nested file\n", \
|
||||
f"{flag}: nested file was not written after the early deletion"
|
||||
|
||||
@pytest.mark.parametrize("flag", ["--delete", "--delete-after"])
|
||||
@pytest.mark.parametrize("flag", ["--delete-commit", "--delete-after"])
|
||||
@pytest.mark.parametrize("mt", [False, True])
|
||||
def test_late_flags_commit_only_after_success(self, flag, mt):
|
||||
"""Plain --delete/--delete-after defer deletion until the whole transfer
|
||||
"""--delete-commit/--delete-after defer deletion until the whole transfer
|
||||
succeeds: a mid-transfer write failure must leave every extra in place
|
||||
(commit-style safety). The -m receiver must also keep the extras: the
|
||||
deferred keep-set is committed by the server only after the disk-writer
|
||||
thread has finished, and a failing writer means the manifest is freed,
|
||||
never applied."""
|
||||
(commit-style safety). Plain --delete no longer defers (it defaults to
|
||||
delete-during), so only the explicitly late timings are exercised here.
|
||||
The --threads receiver must also keep the extras: the deferred keep-set is
|
||||
committed by the server only after the disk-writer thread has finished,
|
||||
and a failing writer means the manifest is freed, never applied."""
|
||||
source = self._seed("late")
|
||||
dest = os.path.join(TEST_DATA_DIR, "deltiming_late_dst")
|
||||
clean_dir(dest)
|
||||
@@ -4463,11 +4640,11 @@ class TestDeletePolicy:
|
||||
@pytest.mark.parametrize("mt", [False, True])
|
||||
@pytest.mark.setpriv
|
||||
def test_ignore_errors_keeps_deletion_active_on_scan_error(self, mt):
|
||||
"""A source I/O error (unreadable subdirectory) aborts the run so no
|
||||
deletion happens by default; --ignore-errors continues, still transfers
|
||||
the readable tree and still deletes, single-threaded and under -m. Run
|
||||
as an unprivileged user so the mode-000 directory is genuinely
|
||||
unreadable."""
|
||||
"""rsync's --ignore-errors semantics: a source I/O error (unreadable
|
||||
subdirectory) makes the run continue and transfer the readable tree, but
|
||||
the default suppresses deletion ("IO error encountered -- skipping file
|
||||
deletion"); --ignore-errors lets deletion proceed. Both exit 23. Run as
|
||||
an unprivileged user so the mode-000 directory is genuinely unreadable."""
|
||||
if os.geteuid() != 0 or shutil.which("setpriv") is None:
|
||||
pytest.skip("requires root + setpriv to drop privileges for the client")
|
||||
tag = f"ioerr_{os.getpid()}_{mt}"
|
||||
@@ -4487,11 +4664,15 @@ class TestDeletePolicy:
|
||||
try:
|
||||
os.chmod(os.path.join(source, "locked"), 0)
|
||||
|
||||
# Default: scan error aborts the run; nothing is deleted.
|
||||
# Default: the scan continues past the unreadable dir and the
|
||||
# readable tree transfers, but deletion is skipped (exit 23).
|
||||
self._write(os.path.join(received, "extra.txt"), b"extra\n")
|
||||
flags = ["--delete"] + (["--threads"] if mt else [])
|
||||
result = self._run_client_as_nobody(source, dest, server.port, flags)
|
||||
assert result.returncode != 0, "unreadable source dir did not fail the run"
|
||||
assert result.returncode == 23, \
|
||||
f"unreadable source dir should exit 23 (got {result.returncode})"
|
||||
assert os.path.exists(os.path.join(received, "top.txt")), \
|
||||
"readable tree did not transfer past the I/O error"
|
||||
assert os.path.exists(os.path.join(received, "extra.txt")), \
|
||||
"default run deleted although the scan hit an I/O error"
|
||||
|
||||
@@ -4499,6 +4680,8 @@ class TestDeletePolicy:
|
||||
self._write(os.path.join(received, "extra.txt"), b"extra\n")
|
||||
flags = ["--delete", "--ignore-errors"] + (["--threads"] if mt else [])
|
||||
result = self._run_client_as_nobody(source, dest, server.port, flags)
|
||||
assert result.returncode == 23, \
|
||||
f"--ignore-errors run should still exit 23 (got {result.returncode})"
|
||||
assert not os.path.exists(os.path.join(received, "extra.txt")), \
|
||||
f"--ignore-errors did not keep deletion active: {result.stderr[:300]}"
|
||||
assert not os.path.exists(os.path.join(received, "locked")), \
|
||||
@@ -4543,11 +4726,11 @@ class TestDeletePolicy:
|
||||
finally:
|
||||
os.chmod(source, 0o755)
|
||||
|
||||
def test_delete_excluded_protection_is_sender_derived(self):
|
||||
def test_delete_protection_reapplied_on_receiver(self):
|
||||
"""Plain --delete protects destination mirrors of files the SOURCE scan
|
||||
excluded, but a destination-only file that merely matches an exclude
|
||||
rule is still an extra and is removed (protection never re-applies rules
|
||||
to the destination)."""
|
||||
excluded, and (track 4a) also protects a destination-only file matching
|
||||
an exclude rule because the compiled rule set is re-applied on the
|
||||
receiver, matching rsync."""
|
||||
source = os.path.join(TEST_DATA_DIR, "senderderived_src")
|
||||
clean_dir(source)
|
||||
self._write(os.path.join(source, "keep.txt"), b"kept\n")
|
||||
@@ -4567,8 +4750,8 @@ class TestDeletePolicy:
|
||||
f"delete sync failed: {(result.stderr or result.stdout)[:300]}"
|
||||
assert os.path.exists(os.path.join(received, "secret.log")), \
|
||||
"source-excluded mirror was deleted under plain --delete"
|
||||
assert not os.path.exists(os.path.join(received, "stray.log")), \
|
||||
"destination-only file matching the exclude rule was left (should be deleted)"
|
||||
assert os.path.exists(os.path.join(received, "stray.log")), \
|
||||
"destination-only file matching the exclude rule must be protected like rsync"
|
||||
|
||||
|
||||
def _pin_mtime(path, ts):
|
||||
@@ -4588,11 +4771,12 @@ class TestBasisDestDirs:
|
||||
STAGING = ".fastsync-stage"
|
||||
TS = 1577836800 # 2020-01-01 00:00:00 UTC, used to pin matching mtimes
|
||||
|
||||
# fixture files: source and basis share the mtime pin, so a basis "match"
|
||||
# is decided purely by content (xxHash). unchanged.txt is byte-identical;
|
||||
# changed.txt is byte-DIFFERENT but has the SAME SIZE as the source (and
|
||||
# the same pinned mtime), which is what forces the content-hash gate;
|
||||
# added.txt does not exist in the basis at all.
|
||||
# fixture files: source and basis share the mtime pin, so the DEFAULT
|
||||
# (rsync-parity) quick-check is a size+mtime match and trusts the basis even
|
||||
# when the body differs. unchanged.txt is byte-identical; changed.txt is
|
||||
# byte-DIFFERENT but has the SAME SIZE as the source (and the same pinned
|
||||
# mtime), which is what the FastSync-only --verify-basis content gate
|
||||
# rejects; added.txt does not exist in the basis at all.
|
||||
UNCHANGED = "unchanged.txt"
|
||||
CHANGED = "changed.txt"
|
||||
ADDED = "added.txt"
|
||||
@@ -4632,18 +4816,19 @@ class TestBasisDestDirs:
|
||||
}
|
||||
|
||||
def _basis_tree(self, prefix):
|
||||
# unchanged.txt is identical to the source; changed.txt has the SAME
|
||||
# byte size and pinned mtime but a different body (equal size forces
|
||||
# the xxHash gate); added.txt is missing from the basis.
|
||||
# unchanged.txt is identical to the source; changed.txt has a DIFFERENT
|
||||
# size (and body) so the size leg of the quick-check fails and it is
|
||||
# transferred normally; added.txt is missing from the basis.
|
||||
return {
|
||||
self.UNCHANGED: b"stable content v1\n",
|
||||
self.CHANGED: b"CHANGED CONTENT NOW\n",
|
||||
self.CHANGED: b"CHANGED CONTENT NOW AND LONGER\n",
|
||||
}
|
||||
|
||||
def test_same_size_different_content_is_not_a_basis_match(self, shared_server):
|
||||
# Core safety property: equal size + pinned mtime but different content
|
||||
# must NEVER be hard-linked or copied from the basis -- the xxHash gate
|
||||
# rejects it and the sender's data is transferred instead.
|
||||
def test_same_size_different_content_default_trusts_quick_check(self, shared_server):
|
||||
# Default rsync-parity behavior: equal size + pinned mtime is a basis
|
||||
# match, so the basis body is materialized/linked without reading it.
|
||||
# This mirrors rsync 3.4.1's quick check (differential-tested in
|
||||
# test_differential_parity.py::test_verify_basis_restores_strict_content).
|
||||
for flag, basis_dir in (("--link-dest", "szlb"), ("--copy-dest", "szcp"),
|
||||
("--compare-dest", "szcmp")):
|
||||
source = self._make_source("basis_same_size_src",
|
||||
@@ -4655,7 +4840,36 @@ class TestBasisDestDirs:
|
||||
result, _ = run_client(source, dest, flags=[f"{flag}={basis_dir}"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, \
|
||||
f"{flag} same-size mismatch failed: {result.stderr[:300]}"
|
||||
f"{flag} same-size quick-check failed: {result.stderr[:300]}"
|
||||
received = get_dest_received_dir(dest, source)
|
||||
dest_file = os.path.join(received, self.UNCHANGED)
|
||||
if flag == "--compare-dest":
|
||||
assert not os.path.exists(dest_file), \
|
||||
f"{flag}: compare-dest must leave a matching file sparse"
|
||||
else:
|
||||
assert _read_file(dest_file) == b"SAME LENGTH BODY!", \
|
||||
f"{flag}: default quick-check did not trust the basis body"
|
||||
if flag == "--link-dest":
|
||||
assert os.stat(dest_file).st_ino == os.stat(basis_file).st_ino, \
|
||||
f"{flag}: basis was not hard-linked"
|
||||
|
||||
def test_verify_basis_rejects_same_size_different_content(self, shared_server):
|
||||
# FastSync-only --verify-basis: the whole-file digest gate rejects the
|
||||
# same-size/different-content basis, so the source data is transferred
|
||||
# instead of the wrong basis bytes.
|
||||
for flag, basis_dir in (("--link-dest", "vszlb"), ("--copy-dest", "vszcp"),
|
||||
("--compare-dest", "vszcmp")):
|
||||
source = self._make_source("basis_verify_src",
|
||||
{self.UNCHANGED: b"same length body\n"})
|
||||
dest = os.path.join(TEST_DATA_DIR, f"basis_verify_dst_{basis_dir}")
|
||||
clean_dir(dest)
|
||||
basis_file = self._seed_basis_file(dest, source, basis_dir, self.UNCHANGED,
|
||||
b"SAME LENGTH BODY!")
|
||||
result, _ = run_client(source, dest,
|
||||
flags=[f"{flag}={basis_dir}", "--verify-basis"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, \
|
||||
f"{flag} --verify-basis failed: {result.stderr[:300]}"
|
||||
received = get_dest_received_dir(dest, source)
|
||||
dest_file = os.path.join(received, self.UNCHANGED)
|
||||
assert _read_file(dest_file) == b"same length body\n", \
|
||||
@@ -4687,11 +4901,13 @@ class TestBasisDestDirs:
|
||||
self._source_tree("c")[self.ADDED], "added file not transferred"
|
||||
|
||||
@pytest.mark.ci
|
||||
def test_dry_run_compare_dest_does_not_read_basis(self, shared_server):
|
||||
# A dry-run --compare-dest must never read/hash the basis file: doing so
|
||||
# is a 1-bit content oracle against the client-supplied digest. Even a
|
||||
# byte-identical basis with a matching size+mtime is therefore reported
|
||||
# as would-transfer, and nothing is created.
|
||||
def test_dry_run_compare_dest_quick_check_does_not_read_basis(self, shared_server):
|
||||
# A dry-run --compare-dest must never read/hash the basis file. Under
|
||||
# the default metadata quick-check a matching basis is reported as a
|
||||
# skip (matching rsync) without reading it; nothing is created. Under
|
||||
# --verify-basis, which would require hashing, the dry-run cannot
|
||||
# confirm the hit (that would be a 1-bit content oracle) and reports
|
||||
# would-transfer instead.
|
||||
source = self._make_source("basis_dry_src", {self.UNCHANGED: b"stable content v1\n"})
|
||||
dest = os.path.join(TEST_DATA_DIR, "basis_dry_dst")
|
||||
clean_dir(dest)
|
||||
@@ -4702,11 +4918,30 @@ class TestBasisDestDirs:
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, \
|
||||
f"dry-run compare-dest failed: {result.stderr[:300]}"
|
||||
assert self.UNCHANGED in result.stdout, (
|
||||
"dry-run compare-dest silently skipped: receiver read the basis content"
|
||||
assert self.UNCHANGED not in result.stdout, (
|
||||
"dry-run compare-dest did not honor the metadata quick-check "
|
||||
"(reported would-transfer for a matching basis)"
|
||||
)
|
||||
assert _snapshot_tree(dest) == before, "dry-run compare-dest mutated the destination"
|
||||
|
||||
# --verify-basis: the hit needs the basis content, which a dry-run must
|
||||
# not read, so the file is reported as would-transfer.
|
||||
dest2 = os.path.join(TEST_DATA_DIR, "basis_dry_verify_dst")
|
||||
clean_dir(dest2)
|
||||
self._seed_basis(dest2, source, "drybasis", {self.UNCHANGED: b"stable content v1\n"})
|
||||
before2 = _snapshot_tree(dest2)
|
||||
result, _ = run_client(source, dest2,
|
||||
flags=["--compare-dest=drybasis", "--dry-run",
|
||||
"--verify-basis"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, \
|
||||
f"dry-run --verify-basis compare-dest failed: {result.stderr[:300]}"
|
||||
assert self.UNCHANGED in result.stdout, (
|
||||
"dry-run --verify-basis must not read the basis to confirm a hit"
|
||||
)
|
||||
assert _snapshot_tree(dest2) == before2, \
|
||||
"dry-run --verify-basis compare-dest mutated the destination"
|
||||
|
||||
def test_compare_dest_content_mismatch_forces_transfer(self, shared_server):
|
||||
# The basis holds a file with a DIFFERENT body: even though it shares
|
||||
# the mtime pin, the xxHash check fails and the data must be sent.
|
||||
@@ -4945,27 +5180,36 @@ class TestBasisDestDirs:
|
||||
assert os.stat(dest_file).st_ino != os.stat(basis_file).st_ino, \
|
||||
"--ignore-times must not hard-link to a basis file"
|
||||
|
||||
def test_basis_refuses_file_above_whole_file_limit(self, shared_server):
|
||||
# Every whole-file payload path in FastSync (basis dirs included) is
|
||||
# bounded by MAX_RECEIVE_WHOLE_FILE_SIZE. rsync supports basis dirs for
|
||||
# arbitrary sizes; FastSync refuses such a run up front with a clear
|
||||
# diagnostic instead of letting the receiver abort the whole transfer
|
||||
# mid-stream with no client-side explanation.
|
||||
def test_basis_handles_file_above_whole_file_limit(self, shared_server):
|
||||
# Track 5a: a basis hit streams the copy (and the --verify-basis digest
|
||||
# streams the basis), so a source larger than the whole-file payload
|
||||
# bound is supported for basis dirs exactly like rsync. A basis MISS
|
||||
# still falls back to the normal transfer, which keeps its own bound.
|
||||
source = self._make_source("basis_oversize_src", {"small.txt": b"ok\n"})
|
||||
big = os.path.join(source, "huge.bin")
|
||||
with open(big, "wb") as fh:
|
||||
os.ftruncate(fh.fileno(), 256 * 1024 * 1024 + 4096)
|
||||
dest = os.path.join(TEST_DATA_DIR, "basis_oversize_dst")
|
||||
clean_dir(dest)
|
||||
result, _ = run_client(source, dest, flags=["--link-dest=nope"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode != 0, \
|
||||
"basis run with an over-limit file unexpectedly succeeded"
|
||||
assert "larger than" in result.stderr, \
|
||||
f"no clear over-limit diagnostic: {result.stderr[:300]}"
|
||||
received = get_dest_received_dir(dest, source)
|
||||
assert not os.path.exists(received), \
|
||||
"over-limit basis run transferred files before failing"
|
||||
rel = os.path.relpath(received, dest)
|
||||
basis_big = os.path.join(dest, "ob", rel, "huge.bin")
|
||||
os.makedirs(os.path.dirname(basis_big), exist_ok=True)
|
||||
shutil.copyfile(big, basis_big)
|
||||
os.utime(basis_big, (self.TS, self.TS))
|
||||
os.utime(big, (self.TS, self.TS))
|
||||
|
||||
result, _ = run_client(source, dest,
|
||||
flags=["--link-dest=ob", "--incremental"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, \
|
||||
f"over-limit basis run failed: {result.stderr[:300]}"
|
||||
dest_big = os.path.join(received, "huge.bin")
|
||||
assert os.path.exists(dest_big), "over-limit basis hit was not materialized"
|
||||
assert os.path.getsize(dest_big) == 256 * 1024 * 1024 + 4096
|
||||
assert os.stat(dest_big).st_ino == os.stat(basis_big).st_ino, \
|
||||
"over-limit --link-dest did not hard-link to the basis"
|
||||
assert _read_file(os.path.join(received, "small.txt")) == b"ok\n"
|
||||
|
||||
|
||||
def _random_payloads(size=2 * 1024 * 1024, changed=64 * 1024, seed=1234):
|
||||
@@ -5568,6 +5812,35 @@ class TestStandaloneSuperDefault:
|
||||
"standalone server accepted --copy-as without --allow-super"
|
||||
)
|
||||
|
||||
@pytest.mark.skipif(
|
||||
os.geteuid() != 0,
|
||||
reason="root triggers the SUPER_MODE_OFF default and can create setuid sources",
|
||||
)
|
||||
def test_special_bits_masked_without_allow_super(self):
|
||||
"""A root standalone server without --allow-super forces SUPER_MODE_OFF,
|
||||
so client-supplied setuid/setgid/sticky bits must be stripped even under
|
||||
-p (they are super-user activities just like device-node creation)."""
|
||||
source = os.path.join(TEST_DATA_DIR, "super_default_mode_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "super_default_mode_dst")
|
||||
clean_dir(source)
|
||||
clean_dir(dest)
|
||||
src_file = os.path.join(source, "priv.sh")
|
||||
with open(src_file, "wb") as f:
|
||||
f.write(b"#!/bin/sh\necho hi\n")
|
||||
os.chmod(src_file, 0o4755)
|
||||
server = ServerManager()
|
||||
server.start() # deliberately no --allow-super -> SUPER_MODE_OFF as root
|
||||
try:
|
||||
result, _ = run_client(source, dest, flags=["-p"], port=server.port)
|
||||
finally:
|
||||
server.stop()
|
||||
assert result.returncode == 0, f"exit {result.returncode}: {(result.stderr or '')[:200]}"
|
||||
received = get_dest_received_dir(dest, source)
|
||||
mode = stat.S_IMODE(os.stat(os.path.join(received, "priv.sh")).st_mode)
|
||||
assert (mode & (stat.S_ISUID | stat.S_ISGID | stat.S_ISVTX)) == 0, \
|
||||
f"--no-super receiver kept a privileged bit: {oct(mode)}"
|
||||
assert (mode & 0o777) == 0o755, f"ordinary permission bits lost: {oct(mode)}"
|
||||
|
||||
@pytest.mark.skipif(os.geteuid() != 0, reason="root can create the source device node")
|
||||
def test_devices_skipped_without_allow_super(self):
|
||||
"""Root standalone server without --allow-super must skip device-node
|
||||
@@ -6482,6 +6755,31 @@ class TestExtendedAttributes:
|
||||
assert os.getxattr(received, "user.rootdir") == b"r"
|
||||
assert os.getxattr(os.path.join(received, "sub"), "user.subdir") == b"s"
|
||||
|
||||
@pytest.mark.ci
|
||||
@pytest.mark.parametrize("mt", [False, True])
|
||||
def test_dirs_directory_xattr_applied(self, shared_server, mt):
|
||||
"""#286.3: -d/-X must apply a transferred directory's user.* xattr at the
|
||||
destination through the --dirs STATUS_MKDIR path (both the
|
||||
single-threaded and -m/--threads receiver paths)."""
|
||||
source, dest = self._source_and_dest("dirsxattr")
|
||||
sub = os.path.join(source, "sub")
|
||||
os.makedirs(sub)
|
||||
if not _xattr_supported(sub):
|
||||
pytest.skip("filesystem does not support user xattrs")
|
||||
os.setxattr(sub, "user.dirsdir", b"dirs-value")
|
||||
lst = os.path.join(TEST_DATA_DIR, "dirs_xattr_list.txt")
|
||||
with open(lst, "wb") as fh:
|
||||
fh.write(b"sub\n")
|
||||
|
||||
flags = ["--files-from", lst, "--dirs", "-R", "-X"] + (["--threads"] if mt else [])
|
||||
result, _ = run_client(source, dest, flags=flags, port=shared_server.port)
|
||||
assert result.returncode == 0, \
|
||||
f"--dirs -X sync failed: {(result.stderr or result.stdout)[:300]}"
|
||||
received = os.path.join(dest, "sub")
|
||||
assert os.path.isdir(received), "--dirs directory entry was not created"
|
||||
assert os.getxattr(received, "user.dirsdir") == b"dirs-value", \
|
||||
"the --dirs directory's user.* xattr was not applied at the destination"
|
||||
|
||||
@pytest.mark.ci
|
||||
def test_directory_default_acl_preserved(self, shared_server):
|
||||
"""#286.3: -aA must preserve a directory's default POSIX ACL (the
|
||||
@@ -6690,11 +6988,10 @@ class TestDirectoryAndSymlinkTimes:
|
||||
|
||||
@pytest.mark.ci
|
||||
@pytest.mark.parametrize("mt", [False, True])
|
||||
def test_preserve_does_not_create_empty_source_dir(self, shared_server, mt):
|
||||
"""P7 Wave D #1: a captured-but-EMPTY source directory is never created
|
||||
at the destination. The scanner records its time (it is transmitted via
|
||||
STATUS_DIR_TIMES), but the receiver treats that entry as record-only, so
|
||||
`-a` keeps the documented "empty dirs are never transferred" behavior."""
|
||||
def test_preserve_creates_empty_source_dir(self, shared_server, mt):
|
||||
"""rsync parity: a recursive `-a` transfer recreates an empty source
|
||||
directory at the destination (the scanner emits it as an explicit
|
||||
directory entry)."""
|
||||
source = os.path.join(TEST_DATA_DIR, f"empty_dir_{'m' if mt else 's'}_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, f"empty_dir_{'m' if mt else 's'}_dst")
|
||||
clean_dir(source)
|
||||
@@ -6705,8 +7002,8 @@ class TestDirectoryAndSymlinkTimes:
|
||||
flags = ["-a"] + (["--threads"] if mt else [])
|
||||
received = self._run(source, dest, flags, shared_server)
|
||||
assert os.path.isfile(os.path.join(received, "keep.txt")), "regular file missing"
|
||||
assert not os.path.lexists(os.path.join(received, "empty_sub")), \
|
||||
f"-a created an empty source directory at {received}/empty_sub"
|
||||
assert os.path.isdir(os.path.join(received, "empty_sub")), \
|
||||
f"-a did not recreate the empty source directory at {received}/empty_sub"
|
||||
|
||||
@pytest.mark.ci
|
||||
@pytest.mark.parametrize("mt", [False, True])
|
||||
@@ -6729,10 +7026,10 @@ class TestDirectoryAndSymlinkTimes:
|
||||
|
||||
@pytest.mark.ci
|
||||
@pytest.mark.parametrize("mt", [False, True])
|
||||
def test_collision_at_dir_time_path_does_not_abort(self, shared_server, mt):
|
||||
"""P7 Wave D #1: a pre-existing regular file at a source-empty-dir's
|
||||
mirror path must not abort the transfer (the old mkdir failed and failed
|
||||
the run) and must not be clobbered."""
|
||||
def test_collision_at_empty_dir_path_replaces_blocker(self, shared_server, mt):
|
||||
"""rsync parity: a pre-existing regular file at a source empty-dir's
|
||||
mirror path is replaced by the incoming directory (rsync removes the
|
||||
non-directory and creates the directory); the run succeeds."""
|
||||
source = os.path.join(TEST_DATA_DIR, f"dirtime_collide_{'m' if mt else 's'}_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, f"dirtime_collide_{'m' if mt else 's'}_dst")
|
||||
clean_dir(source)
|
||||
@@ -6751,10 +7048,8 @@ class TestDirectoryAndSymlinkTimes:
|
||||
assert result.returncode == 0, \
|
||||
f"-a aborted on a pre-existing file at an empty-dir path: " \
|
||||
f"{(result.stderr or result.stdout)[:400]}"
|
||||
assert os.path.isfile(blocker) and not os.path.islink(blocker), \
|
||||
"the pre-existing blocker was replaced by a directory"
|
||||
with open(blocker, "rb") as fh:
|
||||
assert fh.read() == b"pre-existing blocker\n", "the blocker file was clobbered"
|
||||
assert os.path.isdir(blocker) and not os.path.islink(blocker), \
|
||||
"the pre-existing blocker was not replaced by the incoming directory"
|
||||
assert os.path.isfile(os.path.join(received, "keep.txt")), "regular file missing"
|
||||
|
||||
|
||||
|
||||
@@ -1,10 +1,12 @@
|
||||
"""--iconv=CONVERT_SPEC file-NAME charset conversion integration tests.
|
||||
|
||||
The client converts every source file name from LOCAL to REMOTE before it goes
|
||||
on the wire, and the receiver converts it back from REMOTE to LOCAL, so a
|
||||
source tree using one charset can be written into a destination tree using
|
||||
another (rsync compatibility; content bytes are never touched).
|
||||
rsync's spec is ``--iconv=LOCAL,REMOTE`` (the order is the same push or pull).
|
||||
The sender converts each source name from LOCAL to REMOTE for the wire, and on
|
||||
a PUSH the receiver's charset is the spec's REMOTE half, so it writes the wire
|
||||
bytes verbatim (only a server with its own ``--iconv`` declares a different
|
||||
destination charset and re-converts). Content bytes are never touched.
|
||||
"""
|
||||
import codecs
|
||||
import os
|
||||
import shutil
|
||||
|
||||
@@ -16,6 +18,11 @@ LATIN1_NAME = b"caf\xe9.txt"
|
||||
UTF8_NAME = "caf\u00e9.txt".encode("utf-8")
|
||||
|
||||
|
||||
def _to_utf8(name_bytes):
|
||||
"""The UTF-8 encoding of a name that is stored as ISO-8859-1 bytes."""
|
||||
return codecs.encode(codecs.decode(name_bytes, "iso-8859-1"), "utf-8")
|
||||
|
||||
|
||||
def _make(tag):
|
||||
source = os.path.join(TEST_DATA_DIR, f"iconv_{tag}_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, f"iconv_{tag}_dst")
|
||||
@@ -41,10 +48,10 @@ def _dest_file(source, dest, name):
|
||||
|
||||
|
||||
@pytest.mark.ci
|
||||
def test_iconv_latin1_roundtrip(shared_server):
|
||||
"""A source file whose name is ISO-8859-1 bytes is transferred with
|
||||
--iconv=iso-8859-1,utf-8 and lands on the destination with the ORIGINAL
|
||||
latin1 name (the wire carried it as UTF-8)."""
|
||||
def test_iconv_latin1_to_utf8_dest(shared_server):
|
||||
"""rsync push parity: --iconv=iso-8859-1,utf-8 converts a latin1 source name
|
||||
to the spec's REMOTE (UTF-8) on the wire and the default receiver writes it
|
||||
verbatim, so the destination name is UTF-8 (not the source's latin1)."""
|
||||
source, dest = _make("latin1")
|
||||
_place_bytes(source, LATIN1_NAME)
|
||||
|
||||
@@ -53,8 +60,10 @@ def test_iconv_latin1_roundtrip(shared_server):
|
||||
)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:400]
|
||||
|
||||
dst = _dest_file(source, dest, LATIN1_NAME)
|
||||
assert os.path.exists(dst), f"dest latin1-named file not found under {dest}"
|
||||
dst = _dest_file(source, dest, UTF8_NAME)
|
||||
assert os.path.exists(dst), f"dest UTF-8-named file not found under {dest}"
|
||||
assert not os.path.exists(_dest_file(source, dest, LATIN1_NAME)), \
|
||||
"destination kept the latin1 name instead of the wire (UTF-8) charset"
|
||||
|
||||
|
||||
@pytest.mark.ci
|
||||
@@ -153,7 +162,7 @@ def test_iconv_expanding_name_growth(shared_server):
|
||||
)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:400]
|
||||
|
||||
assert os.path.exists(_dest_file(source, dest, name_bytes))
|
||||
assert os.path.exists(_dest_file(source, dest, _to_utf8(name_bytes)))
|
||||
|
||||
|
||||
def test_iconv_symlink_path_and_target(shared_server):
|
||||
@@ -170,11 +179,13 @@ def test_iconv_symlink_path_and_target(shared_server):
|
||||
)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:400]
|
||||
|
||||
dst_target = _dest_file(source, dest, target)
|
||||
dst_link = _dest_file(source, dest, b"link\xe9")
|
||||
assert os.path.exists(dst_target), "dest latin1 target file missing"
|
||||
assert os.path.islink(dst_link), "dest latin1 symlink missing"
|
||||
assert os.readlink(dst_link) == target, "symlink target not preserved/decoded"
|
||||
utf8_target = _to_utf8(target)
|
||||
utf8_link = _to_utf8(b"link\xe9")
|
||||
dst_target = _dest_file(source, dest, utf8_target)
|
||||
dst_link = _dest_file(source, dest, utf8_link)
|
||||
assert os.path.exists(dst_target), "dest UTF-8 target file missing"
|
||||
assert os.path.islink(dst_link), "dest UTF-8 symlink missing"
|
||||
assert os.readlink(dst_link) == utf8_target, "symlink target not wire-converted"
|
||||
with open(dst_link, "rb") as fh:
|
||||
assert fh.read() == b"t\n"
|
||||
|
||||
@@ -198,8 +209,8 @@ def test_iconv_hardlink_path_and_target(shared_server):
|
||||
)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:400]
|
||||
|
||||
dst_a = _dest_file(source, dest, a)
|
||||
dst_b = _dest_file(source, dest, b)
|
||||
dst_a = _dest_file(source, dest, _to_utf8(a))
|
||||
dst_b = _dest_file(source, dest, _to_utf8(b))
|
||||
assert os.path.exists(dst_a) and os.path.exists(dst_b)
|
||||
assert os.stat(dst_a).st_ino == os.stat(dst_b).st_ino, \
|
||||
"hard-link relationship not preserved across the transfer"
|
||||
@@ -221,16 +232,16 @@ def test_iconv_delete_manifest_consistent(shared_server):
|
||||
flags = ["--iconv=iso-8859-1,utf-8"]
|
||||
result, _ = run_client(source, dest, flags=flags, port=server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:400]
|
||||
assert os.path.exists(_dest_file(source, dest, keep))
|
||||
assert os.path.exists(_dest_file(source, dest, gone))
|
||||
assert os.path.exists(_dest_file(source, dest, _to_utf8(keep)))
|
||||
assert os.path.exists(_dest_file(source, dest, _to_utf8(gone)))
|
||||
|
||||
os.remove(os.path.join(os.fsencode(source), gone))
|
||||
result, _ = run_client(
|
||||
source, dest, flags=flags + ["--delete"], port=server.port
|
||||
)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:400]
|
||||
assert os.path.exists(_dest_file(source, dest, keep)), "kept file deleted"
|
||||
assert not os.path.exists(_dest_file(source, dest, gone)), \
|
||||
assert os.path.exists(_dest_file(source, dest, _to_utf8(keep))), "kept file deleted"
|
||||
assert not os.path.exists(_dest_file(source, dest, _to_utf8(gone))), \
|
||||
"missing file was not deleted"
|
||||
|
||||
|
||||
@@ -247,4 +258,4 @@ def test_iconv_chunk_serialization_blob(shared_server):
|
||||
)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:400]
|
||||
|
||||
assert os.path.exists(_dest_file(source, dest, name))
|
||||
assert os.path.exists(_dest_file(source, dest, _to_utf8(name)))
|
||||
@@ -0,0 +1,538 @@
|
||||
"""Differential parity tests for the option wave (bwlimit, --info=*, -M,
|
||||
--ignore-errors, --filter protect).
|
||||
|
||||
Every differential here runs the SAME scenario with real ``rsync 3.4.1`` and
|
||||
with fastsync and compares the observable result, so the modules are skipped
|
||||
when rsync is unavailable. The privilege-dependent --ignore-errors differential
|
||||
drops the client to an unprivileged uid so a mode-000 source directory is
|
||||
genuinely unreadable; it is marked ``setpriv`` (run as root locally, excluded
|
||||
from the root PR gate exactly like the other privilege tests).
|
||||
"""
|
||||
import os
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
|
||||
import pytest
|
||||
|
||||
sys.path.insert(0, os.path.dirname(__file__))
|
||||
from common import ( # noqa: E402
|
||||
CLIENT_CMD,
|
||||
TEST_DATA_DIR,
|
||||
ServerManager,
|
||||
clean_dir,
|
||||
get_dest_received_dir,
|
||||
run_client,
|
||||
)
|
||||
|
||||
RSYNC = shutil.which("rsync")
|
||||
requires_rsync = pytest.mark.skipif(RSYNC is None, reason="rsync 3.4.1 not installed")
|
||||
|
||||
|
||||
def _rsync(args, timeout=120, as_nobody=False):
|
||||
env = dict(os.environ, LC_ALL="C")
|
||||
cmd = [RSYNC] + args
|
||||
if as_nobody:
|
||||
cmd = ["setpriv", "--reuid=65534", "--regid=65534", "--clear-groups"] + cmd
|
||||
return subprocess.run(cmd, capture_output=True, text=True, env=env, timeout=timeout)
|
||||
|
||||
|
||||
def _write(path, content):
|
||||
os.makedirs(os.path.dirname(path), exist_ok=True)
|
||||
with open(path, "wb") as fh:
|
||||
fh.write(content)
|
||||
|
||||
|
||||
class TestBwlimitParity:
|
||||
"""--bwlimit must accept rsync 3.4.1's spellings and pace like it."""
|
||||
|
||||
ACCEPTED = ["100", "0", "1.5", "100K", "100KB", "100KiB", "1M", "1MB", "1m", "1G", "512"]
|
||||
REJECTED = ["-1", "abc", "1x", "1 000"]
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_parse_acceptance_matches_rsync(self, shared_server):
|
||||
source = os.path.join(TEST_DATA_DIR, "bwp_src")
|
||||
clean_dir(source)
|
||||
_write(os.path.join(source, "f.txt"), b"payload\n")
|
||||
|
||||
for value in self.ACCEPTED + self.REJECTED:
|
||||
rdst = os.path.join(TEST_DATA_DIR, "bwp_rdst")
|
||||
clean_dir(rdst)
|
||||
rsync_result = _rsync(["-a", "--bwlimit=" + value, source + "/", rdst + "/"])
|
||||
dest = os.path.join(TEST_DATA_DIR, "bwp_dst")
|
||||
clean_dir(dest)
|
||||
result, _ = run_client(source, dest, flags=["-a", "--bwlimit=" + value],
|
||||
port=shared_server.port)
|
||||
assert (result.returncode == 0) == (rsync_result.returncode == 0), (
|
||||
f"--bwlimit={value}: fastsync rc={result.returncode} "
|
||||
f"({(result.stderr or result.stdout)[:120]!r}) "
|
||||
f"rsync rc={rsync_result.returncode} ({rsync_result.stderr[:120]!r})"
|
||||
)
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_throttle_rate_matches_rsync(self, shared_server):
|
||||
"""A 4 MiB transfer at --bwlimit=2048 (2 MiB/s) must take about the same
|
||||
wall-clock time for both tools (~2 s with rsync's leaky bucket)."""
|
||||
source = os.path.join(TEST_DATA_DIR, "bwt_src")
|
||||
clean_dir(source)
|
||||
_write(os.path.join(source, "big.bin"), os.urandom(4 * 1024 * 1024))
|
||||
dest = os.path.join(TEST_DATA_DIR, "bwt_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "bwt_rdst")
|
||||
|
||||
clean_dir(rdst)
|
||||
start = time.monotonic()
|
||||
rsync_result = _rsync(["-a", "--bwlimit=2048", source + "/", rdst + "/"])
|
||||
rsync_secs = time.monotonic() - start
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
|
||||
clean_dir(dest)
|
||||
result, fast_secs = run_client(source, dest, flags=["-a", "--bwlimit=2048"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
|
||||
# 4 MiB at 2 MiB/s rendezvous near 2 s. Use a coarse band on each side
|
||||
# (an unthrottled transfer finishes well under 1.5 s) plus a generous
|
||||
# cross-tolerance so a loaded CI runner cannot flake the parity assert.
|
||||
lo, hi = 1.5, 4.5
|
||||
assert lo <= fast_secs <= hi, f"fastsync throttle out of band: {fast_secs:.2f}s"
|
||||
assert lo <= rsync_secs <= hi, f"rsync throttle out of band: {rsync_secs:.2f}s"
|
||||
assert abs(fast_secs - rsync_secs) < 2.0, (
|
||||
f"fastsync {fast_secs:.2f}s vs rsync {rsync_secs:.2f}s"
|
||||
)
|
||||
|
||||
|
||||
def _output_tree(root):
|
||||
clean_dir(root)
|
||||
os.makedirs(os.path.join(root, "sub"))
|
||||
_write(os.path.join(root, "a.txt"), b"top\n")
|
||||
_write(os.path.join(root, "sub", "b.txt"), b"nested\n")
|
||||
os.symlink("a.txt", os.path.join(root, "link"))
|
||||
|
||||
|
||||
class TestInfoParity:
|
||||
"""The --info categories that map to a FastSync event must print rsync's
|
||||
line format."""
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_info_flist_matches_rsync(self, shared_server):
|
||||
source = os.path.join(TEST_DATA_DIR, "inf_fl_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "inf_fl_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "inf_fl_rdst")
|
||||
_output_tree(source)
|
||||
clean_dir(dest)
|
||||
clean_dir(rdst)
|
||||
rsync_result = _rsync(["-a", "--info=flist", source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
result, _ = run_client(source, dest, flags=["-a", "--info=flist"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
|
||||
assert "sending incremental file list" in result.stdout
|
||||
assert "sending incremental file list" in rsync_result.stdout
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_info_name_matches_rsync(self, shared_server):
|
||||
source = os.path.join(TEST_DATA_DIR, "inf_nm_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "inf_nm_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "inf_nm_rdst")
|
||||
_output_tree(source)
|
||||
clean_dir(dest)
|
||||
clean_dir(rdst)
|
||||
rsync_result = _rsync(["-a", "--info=name", source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
result, _ = run_client(source, dest, flags=["-a", "--info=name"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
|
||||
|
||||
def entries(text):
|
||||
# Compare the transferred entries only: rsync also prints the
|
||||
# transfer-root `./` and every directory (FastSync records dirs),
|
||||
# which are a separate documented divergence.
|
||||
out = []
|
||||
for line in text.splitlines():
|
||||
if not line or line.startswith("sending ") or line.startswith("created "):
|
||||
continue
|
||||
if line == "./" or line.endswith("/"):
|
||||
continue
|
||||
out.append(line)
|
||||
return sorted(out)
|
||||
|
||||
assert entries(result.stdout) == entries(rsync_result.stdout), (
|
||||
f"rsync={entries(rsync_result.stdout)} fastsync={entries(result.stdout)}"
|
||||
)
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_info_name_root_line_matches_rsync(self, shared_server):
|
||||
"""A fresh destination: rsync prints `created directory`, then the
|
||||
transfer-root `./` name line before the entries; FastSync must emit the
|
||||
same `./` line."""
|
||||
source = os.path.join(TEST_DATA_DIR, "inf_root_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "inf_root_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "inf_root_rdst")
|
||||
clean_dir(source)
|
||||
_write(os.path.join(source, "f.bin"), b"payload\n")
|
||||
clean_dir(dest)
|
||||
shutil.rmtree(rdst, ignore_errors=True)
|
||||
rsync_result = _rsync(["-a", "--info=name", source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
result, _ = run_client(source, dest, flags=["-a", "--info=name"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
|
||||
|
||||
def names(text):
|
||||
return [l for l in text.splitlines()
|
||||
if l and not l.startswith("created directory")
|
||||
and not (l.endswith("/") and l != "./")]
|
||||
|
||||
assert names(rsync_result.stdout) == ["./", "f.bin"], names(rsync_result.stdout)
|
||||
assert names(result.stdout) == ["./", "f.bin"], names(result.stdout)
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_info_name2_uptodate_matches_rsync(self, shared_server):
|
||||
"""--info=name2 prints `NAME is uptodate` for entries the receiver
|
||||
already has, matching rsync byte-for-byte."""
|
||||
source = os.path.join(TEST_DATA_DIR, "inf_up_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "inf_up_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "inf_up_rdst")
|
||||
clean_dir(source)
|
||||
os.makedirs(os.path.join(source, "sub"))
|
||||
_write(os.path.join(source, "a.txt"), b"a\n")
|
||||
_write(os.path.join(source, "sub", "b.txt"), b"b\n")
|
||||
clean_dir(rdst)
|
||||
assert _rsync(["-a", source + "/", rdst + "/"]).returncode == 0
|
||||
rsync_result = _rsync(["-a", "--info=name2", source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
|
||||
clean_dir(dest)
|
||||
seed, _ = run_client(source, dest, flags=["-a", "--incremental"],
|
||||
port=shared_server.port)
|
||||
assert seed.returncode == 0, (seed.stderr or seed.stdout)[:200]
|
||||
result, _ = run_client(source, dest, flags=["-a", "--incremental", "--info=name2"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
|
||||
rsync_lines = sorted(l for l in rsync_result.stdout.splitlines()
|
||||
if l.endswith("is uptodate"))
|
||||
fast_lines = sorted(l for l in result.stdout.splitlines()
|
||||
if l.endswith("is uptodate"))
|
||||
assert fast_lines == rsync_lines, (rsync_lines, fast_lines)
|
||||
assert fast_lines == ["a.txt is uptodate", "sub/b.txt is uptodate"], fast_lines
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_info_nonreg_matches_rsync(self, shared_server):
|
||||
source = os.path.join(TEST_DATA_DIR, "inf_nr_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "inf_nr_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "inf_nr_rdst")
|
||||
clean_dir(source)
|
||||
os.mkfifo(os.path.join(source, "fifo"))
|
||||
_write(os.path.join(source, "a.txt"), b"a\n")
|
||||
clean_dir(dest)
|
||||
clean_dir(rdst)
|
||||
rsync_result = _rsync(["-rlt", "--info=nonreg", source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
result, _ = run_client(source, dest, flags=["-rlt", "--info=nonreg"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
|
||||
rsync_lines = sorted(l for l in rsync_result.stdout.splitlines()
|
||||
if l.startswith("skipping non-regular"))
|
||||
fast_lines = sorted(l for l in result.stdout.splitlines()
|
||||
if l.startswith("skipping non-regular"))
|
||||
assert fast_lines == rsync_lines, (rsync_lines, fast_lines)
|
||||
assert fast_lines, "no non-regular skip line emitted"
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_info_del_real_matches_rsync(self, shared_server):
|
||||
source = os.path.join(TEST_DATA_DIR, "inf_dl_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "inf_dl_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "inf_dl_rdst")
|
||||
clean_dir(source)
|
||||
_write(os.path.join(source, "keep.txt"), b"keep\n")
|
||||
clean_dir(rdst)
|
||||
_write(os.path.join(rdst, "extra.txt"), b"x\n")
|
||||
_write(os.path.join(rdst, "extra2.txt"), b"y\n")
|
||||
rsync_result = _rsync(["-a", "--delete", "--info=del", source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
rsync_lines = sorted(l for l in rsync_result.stdout.splitlines()
|
||||
if l.startswith("deleting "))
|
||||
|
||||
clean_dir(dest)
|
||||
received = get_dest_received_dir(dest, source)
|
||||
_write(os.path.join(received, "extra.txt"), b"x\n")
|
||||
_write(os.path.join(received, "extra2.txt"), b"y\n")
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
result, _ = run_client(source, dest, flags=["-a", "--delete", "--info=del"],
|
||||
port=server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
|
||||
fast_lines = sorted(l for l in result.stdout.splitlines()
|
||||
if l.startswith("deleting "))
|
||||
assert fast_lines == rsync_lines, (rsync_lines, fast_lines)
|
||||
assert fast_lines, "no deletion lines emitted"
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_info_del_itemize_real_matches_rsync(self, shared_server):
|
||||
source = os.path.join(TEST_DATA_DIR, "inf_di_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "inf_di_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "inf_di_rdst")
|
||||
clean_dir(source)
|
||||
_write(os.path.join(source, "keep.txt"), b"keep\n")
|
||||
clean_dir(rdst)
|
||||
_write(os.path.join(rdst, "extra.txt"), b"x\n")
|
||||
rsync_result = _rsync(["-a", "-i", "--delete", source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
rsync_lines = sorted(l for l in rsync_result.stdout.splitlines()
|
||||
if l.startswith("*deleting"))
|
||||
|
||||
clean_dir(dest)
|
||||
received = get_dest_received_dir(dest, source)
|
||||
_write(os.path.join(received, "extra.txt"), b"x\n")
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
result, _ = run_client(source, dest, flags=["-a", "-i", "--delete"],
|
||||
port=server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
|
||||
fast_lines = sorted(l for l in result.stdout.splitlines()
|
||||
if l.startswith("*deleting"))
|
||||
assert fast_lines == rsync_lines, (rsync_lines, fast_lines)
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_info_del_dry_run_matches_rsync(self, shared_server):
|
||||
source = os.path.join(TEST_DATA_DIR, "inf_dd_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "inf_dd_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "inf_dd_rdst")
|
||||
clean_dir(source)
|
||||
_write(os.path.join(source, "keep.txt"), b"keep\n")
|
||||
clean_dir(rdst)
|
||||
_write(os.path.join(rdst, "extra.txt"), b"x\n")
|
||||
rsync_result = _rsync(["-a", "-n", "--delete", "--info=del", source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
rsync_lines = sorted(l for l in rsync_result.stdout.splitlines()
|
||||
if l.startswith("deleting "))
|
||||
|
||||
clean_dir(dest)
|
||||
received = get_dest_received_dir(dest, source)
|
||||
_write(os.path.join(received, "extra.txt"), b"x\n")
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
result, _ = run_client(source, dest, flags=["-a", "-n", "--delete", "--info=del"],
|
||||
port=server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
|
||||
fast_lines = sorted(l for l in result.stdout.splitlines()
|
||||
if l.startswith("deleting "))
|
||||
assert fast_lines == rsync_lines, (rsync_lines, fast_lines)
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_info_remove_matches_rsync(self, shared_server):
|
||||
tag = "inf_rm"
|
||||
rsync_src = os.path.join(TEST_DATA_DIR, f"{tag}_rsrc")
|
||||
rsync_dst = os.path.join(TEST_DATA_DIR, f"{tag}_rdst")
|
||||
fast_src = os.path.join(TEST_DATA_DIR, f"{tag}_fsrc")
|
||||
fast_dst = os.path.join(TEST_DATA_DIR, f"{tag}_fdst")
|
||||
for root in (rsync_src, rsync_dst, fast_src, fast_dst):
|
||||
clean_dir(root)
|
||||
_write(os.path.join(rsync_src, "a.txt"), b"a\n")
|
||||
_write(os.path.join(rsync_src, "sub", "b.txt"), b"b\n")
|
||||
_write(os.path.join(fast_src, "a.txt"), b"a\n")
|
||||
_write(os.path.join(fast_src, "sub", "b.txt"), b"b\n")
|
||||
|
||||
rsync_result = _rsync(["-a", "--remove-source-files", "--info=remove",
|
||||
rsync_src + "/", rsync_dst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
rsync_lines = sorted(l for l in rsync_result.stdout.splitlines()
|
||||
if l.startswith("sender removed "))
|
||||
|
||||
result, _ = run_client(fast_src, fast_dst,
|
||||
flags=["-a", "--remove-source-files", "--info=remove"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
|
||||
fast_lines = sorted(l for l in result.stdout.splitlines()
|
||||
if l.startswith("sender removed "))
|
||||
assert fast_lines == rsync_lines, (rsync_lines, fast_lines)
|
||||
assert fast_lines, "no source-removal lines emitted"
|
||||
|
||||
|
||||
class TestIgnoreErrorsParity:
|
||||
"""--ignore-errors: a source I/O error skips deletion by default; the flag
|
||||
lets deletion proceed. Both exit 23. Run the client as an unprivileged user
|
||||
so the mode-000 directory is genuinely unreadable."""
|
||||
|
||||
@pytest.mark.setpriv
|
||||
def test_delete_after_io_error_matches_rsync(self):
|
||||
if os.geteuid() != 0 or shutil.which("setpriv") is None:
|
||||
pytest.skip("requires root + setpriv to drop privileges for the client")
|
||||
tag = f"ie_{os.getpid()}"
|
||||
source = os.path.join(TEST_DATA_DIR, f"{tag}_src")
|
||||
rsync_dst = os.path.join(TEST_DATA_DIR, f"{tag}_rdst")
|
||||
dest = os.path.join(TEST_DATA_DIR, f"{tag}_dst")
|
||||
clean_dir(source)
|
||||
clean_dir(rsync_dst)
|
||||
clean_dir(dest)
|
||||
_write(os.path.join(source, "top.txt"), b"top\n")
|
||||
_write(os.path.join(source, "locked", "blocked.txt"), b"blocked\n")
|
||||
os.chmod(os.path.join(source, "locked"), 0)
|
||||
os.chmod(TEST_DATA_DIR, 0o777)
|
||||
os.chmod(source, 0o755)
|
||||
os.chmod(rsync_dst, 0o777)
|
||||
os.chmod(dest, 0o777)
|
||||
try:
|
||||
for ignore in (False, True):
|
||||
flags = ["-a", "--delete-after"] + (["--ignore-errors"] if ignore else [])
|
||||
# rsync side
|
||||
_write(os.path.join(rsync_dst, "extra.txt"), b"x\n")
|
||||
os.chmod(os.path.join(rsync_dst, "extra.txt"), 0o666)
|
||||
rres = _rsync(flags + [source + "/", rsync_dst + "/"], as_nobody=True)
|
||||
rsync_extra = os.path.exists(os.path.join(rsync_dst, "extra.txt"))
|
||||
|
||||
# fastsync side
|
||||
received = get_dest_received_dir(dest, source)
|
||||
_write(os.path.join(received, "extra.txt"), b"x\n")
|
||||
os.chmod(os.path.join(received, "extra.txt"), 0o666)
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
fflags = (["--delete", "--ignore-errors"] if ignore else ["--delete"])
|
||||
cmd = CLIENT_CMD + ["--source-dir", source, "--dest-dir", dest,
|
||||
"--save-to-disk", "--server-port", str(server.port)] + fflags
|
||||
fres = subprocess.run(
|
||||
["setpriv", "--reuid=65534", "--regid=65534", "--clear-groups"] + cmd,
|
||||
text=True, capture_output=True)
|
||||
fast_extra = os.path.exists(os.path.join(received, "extra.txt"))
|
||||
|
||||
assert rres.returncode == 23, (ignore, rres.returncode, rres.stderr[:200])
|
||||
assert fres.returncode == 23, (ignore, fres.returncode, fres.stderr[:200])
|
||||
assert rsync_extra == fast_extra, (
|
||||
f"ignore_errors={ignore}: rsync extra={rsync_extra} fastsync extra={fast_extra}"
|
||||
)
|
||||
assert fast_extra is (not ignore), (ignore, fast_extra)
|
||||
finally:
|
||||
os.chmod(os.path.join(source, "locked"), 0o755)
|
||||
|
||||
|
||||
class TestRemoteOptionDaemon:
|
||||
"""rsync forwards -M/--remote-option to its remote process over a daemon
|
||||
connection; FastSync's daemon has no per-connection argv channel and rejects
|
||||
it. This pins the documented divergence with evidence."""
|
||||
|
||||
@requires_rsync
|
||||
def test_rsync_forwards_M_over_daemon_and_fastsync_rejects(self, tmp_path):
|
||||
import socket
|
||||
|
||||
with socket.socket() as probe:
|
||||
probe.bind(("127.0.0.1", 0))
|
||||
port = probe.getsockname()[1]
|
||||
|
||||
module_root = tmp_path / "mod"
|
||||
module_root.mkdir()
|
||||
os.chmod(module_root, 0o777)
|
||||
source = tmp_path / "src"
|
||||
source.mkdir()
|
||||
(source / "a.txt").write_bytes(b"hello\n")
|
||||
conf = tmp_path / "rsyncd.conf"
|
||||
conf.write_text(
|
||||
f"port = {port}\nuse chroot = no\n[m]\npath = {module_root}\nread only = no\n"
|
||||
)
|
||||
daemon = subprocess.Popen(
|
||||
[RSYNC, "--daemon", "--no-detach", "--port", str(port), "--config", str(conf)],
|
||||
stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
|
||||
try:
|
||||
deadline = time.monotonic() + 5
|
||||
while time.monotonic() < deadline:
|
||||
try:
|
||||
with socket.create_connection(("127.0.0.1", port), timeout=0.3):
|
||||
break
|
||||
except OSError:
|
||||
time.sleep(0.05)
|
||||
else:
|
||||
pytest.skip("rsync daemon did not start")
|
||||
|
||||
# A well-formed -M option is forwarded and accepted by the daemon...
|
||||
ok = _rsync(["-a", "-M--safe-links", source.as_posix() + "/",
|
||||
f"rsync://127.0.0.1:{port}/m/"])
|
||||
# ...and a bogus one is rejected ON THE REMOTE with "unknown option",
|
||||
# which proves the option reached the daemon's parser.
|
||||
bogus = _rsync(["-a", "-M--totally-bogus", source.as_posix() + "/",
|
||||
f"rsync://127.0.0.1:{port}/m/"])
|
||||
assert bogus.returncode != 0
|
||||
assert "unknown option" in (bogus.stderr + bogus.stdout), bogus.stderr
|
||||
del ok
|
||||
finally:
|
||||
daemon.terminate()
|
||||
try:
|
||||
daemon.wait(timeout=5)
|
||||
except subprocess.TimeoutExpired:
|
||||
daemon.kill()
|
||||
|
||||
# FastSync rejects -M for a non-SSH transport up front.
|
||||
dest = os.path.join(TEST_DATA_DIR, "ro_dst")
|
||||
clean_dir(dest)
|
||||
result, _ = run_client(source.as_posix(), dest, flags=["-a", "-M--safe-links"])
|
||||
assert result.returncode != 0
|
||||
assert "remote-option" in (result.stderr + result.stdout)
|
||||
|
||||
|
||||
class TestFilterProtect:
|
||||
"""Receiver-derived delete protection: a `protect`/`P` rule is compiled by
|
||||
the sender and sent on the config frame, so the receiver shields a
|
||||
destination-only entry that never appeared on the sender, matching rsync."""
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_protect_dest_only_matches_rsync(self, shared_server):
|
||||
source = os.path.join(TEST_DATA_DIR, "fpd_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "fpd_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "fpd_rdst")
|
||||
clean_dir(source)
|
||||
_write(os.path.join(source, "keep.txt"), b"keep\n")
|
||||
clean_dir(rdst)
|
||||
_write(os.path.join(rdst, "extra.log"), b"extra\n")
|
||||
_write(os.path.join(rdst, "other.txt"), b"other\n")
|
||||
|
||||
rsync_result = _rsync(["-a", "--delete", "--filter=P *.log", source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
assert os.path.exists(os.path.join(rdst, "extra.log")), "rsync did not protect extra.log"
|
||||
assert not os.path.exists(os.path.join(rdst, "other.txt")), "rsync did not delete other.txt"
|
||||
|
||||
clean_dir(dest)
|
||||
received = get_dest_received_dir(dest, source)
|
||||
_write(os.path.join(received, "extra.log"), b"extra\n")
|
||||
_write(os.path.join(received, "other.txt"), b"other\n")
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
result, _ = run_client(source, dest,
|
||||
flags=["-a", "--delete", "--filter=P *.log"],
|
||||
port=server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:200]
|
||||
assert os.path.exists(os.path.join(received, "extra.log")), (
|
||||
"FastSync must protect a destination-only P match like rsync")
|
||||
assert not os.path.exists(os.path.join(received, "other.txt"))
|
||||
|
||||
@pytest.mark.ci
|
||||
def test_protect_dest_only_dry_run_enumeration(self, shared_server):
|
||||
source = os.path.join(TEST_DATA_DIR, "fpd_nd_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "fpd_nd_dst")
|
||||
clean_dir(source)
|
||||
_write(os.path.join(source, "keep.txt"), b"keep\n")
|
||||
received = get_dest_received_dir(dest, source)
|
||||
clean_dir(received)
|
||||
_write(os.path.join(received, "keep.txt"), b"keep\n")
|
||||
_write(os.path.join(received, "extra.log"), b"extra\n")
|
||||
_write(os.path.join(received, "other.txt"), b"other\n")
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
result, _ = run_client(source, dest,
|
||||
flags=["-a", "-n", "--delete", "--out-format=%n",
|
||||
"--filter=P *.log"],
|
||||
port=server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
assert "other.txt" in result.stdout, result.stdout
|
||||
assert "extra.log" not in result.stdout, result.stdout
|
||||
assert os.path.exists(os.path.join(received, "extra.log"))
|
||||
assert os.path.exists(os.path.join(received, "other.txt"))
|
||||
@@ -26,6 +26,17 @@ def _rsync(args):
|
||||
)
|
||||
|
||||
|
||||
def _file_entry_line(text):
|
||||
"""The file entry line for a single-file transfer.
|
||||
|
||||
-i/--out-format emit the transfer-root (and directory) lines too, so the
|
||||
file entry is not necessarily the first line; for the one-file corpora used
|
||||
by the wire-counter tests it is the last non-empty line.
|
||||
"""
|
||||
lines = [line for line in text.splitlines() if line.strip()]
|
||||
return lines[-1] if lines else ""
|
||||
|
||||
|
||||
def _make_selection_tree(root):
|
||||
clean_dir(root)
|
||||
os.makedirs(os.path.join(root, "sub"))
|
||||
@@ -184,6 +195,51 @@ class TestItemizeParity:
|
||||
)
|
||||
assert fast_lines == rsync_lines, f"rsync={rsync_lines} fastsync={fast_lines}"
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_itemize_directory_lines_match_rsync(self, shared_server):
|
||||
"""#292: -i/--out-format emit rsync's directory lines (including the
|
||||
transfer root) in rsync's depth-first order."""
|
||||
source = os.path.join(TEST_DATA_DIR, "out_itemdir_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "out_itemdir_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "out_itemdir_rdst")
|
||||
clean_dir(source)
|
||||
os.makedirs(os.path.join(source, "sub", "deep"))
|
||||
os.makedirs(os.path.join(source, "emptydir"))
|
||||
with open(os.path.join(source, "a.txt"), "wb") as fh:
|
||||
fh.write(b"hello\n")
|
||||
with open(os.path.join(source, "sub", "b.txt"), "wb") as fh:
|
||||
fh.write(b"world\n")
|
||||
with open(os.path.join(source, "sub", "deep", "d.txt"), "wb") as fh:
|
||||
fh.write(b"deep\n")
|
||||
clean_dir(dest)
|
||||
clean_dir(rdst)
|
||||
|
||||
def dir_lines(text):
|
||||
# Any line whose name ends with '/' is a directory entry.
|
||||
return sorted(
|
||||
line for line in text.splitlines()
|
||||
if line.rsplit(" ", 1)[-1].endswith("/")
|
||||
)
|
||||
|
||||
for fmt in (None, "%i %n%L"):
|
||||
rsync_flags = ["-a", "-i"] if fmt is None else ["-a", "--out-format=" + fmt]
|
||||
fast_flags = rsync_flags
|
||||
clean_dir(rdst)
|
||||
clean_dir(dest)
|
||||
rsync_result = _rsync(rsync_flags + [source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
result, _ = run_client(source, dest, flags=fast_flags,
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, result.stderr[:300]
|
||||
expected = [l for l in dir_lines(rsync_result.stdout)
|
||||
if not l.rsplit(" ", 1)[-1] == "./"]
|
||||
fast = dir_lines(result.stdout)
|
||||
assert [l for l in fast if not l.rsplit(" ", 1)[-1] == "./"] == expected, (
|
||||
f"fmt={fmt} rsync={rsync_result.stdout!r} fastsync={result.stdout!r}"
|
||||
)
|
||||
assert ".d..t...... ./" in fast, f"missing root line: {result.stdout!r}"
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_itemize_modified_file_matches_rsync(self, shared_server):
|
||||
@@ -289,6 +345,49 @@ def _make_one_file(root, name="f.bin", size=100):
|
||||
fh.write(bytes((i * 7 + 3) & 0xFF for i in range(size)))
|
||||
|
||||
|
||||
def _make_multidir_tree(root):
|
||||
"""Multi-directory corpus for the --progress file-list tests: nested files,
|
||||
a directory-only branch, an empty directory and a symlink."""
|
||||
clean_dir(root)
|
||||
for rel, data in (("a.txt", b"alpha\n"), ("b.txt", b"bravo\n"),
|
||||
("sub1/c.txt", b"charlie\n"), ("sub1/deep/d.txt", b"delta\n"),
|
||||
("sub2/e.txt", b"echo\n")):
|
||||
path = os.path.join(root, rel)
|
||||
os.makedirs(os.path.dirname(path), exist_ok=True)
|
||||
with open(path, "wb") as fh:
|
||||
fh.write(data)
|
||||
os.symlink("a.txt", os.path.join(root, "link1"))
|
||||
os.makedirs(os.path.join(root, "emptydir"), exist_ok=True)
|
||||
|
||||
|
||||
def _parse_progress(text):
|
||||
"""Name lines and the `to-chk` denominators from a --progress run."""
|
||||
names = []
|
||||
totals = set()
|
||||
for line in text.splitlines():
|
||||
line = line.rstrip()
|
||||
if not line or line == "sending incremental file list":
|
||||
continue
|
||||
if "%" in line:
|
||||
match = re.search(r"to-chk=\d+/(\d+)", line)
|
||||
if match:
|
||||
totals.add(int(match.group(1)))
|
||||
continue
|
||||
if line == "./": # root-line trigger is a separate documented residual
|
||||
continue
|
||||
names.append(line)
|
||||
return sorted(names), totals
|
||||
|
||||
|
||||
def _pick_stats(text, keys):
|
||||
out = {}
|
||||
for line in text.splitlines():
|
||||
for key in keys:
|
||||
if line.startswith(key + ":"):
|
||||
out[key] = line
|
||||
return out
|
||||
|
||||
|
||||
class TestWireStatsParity:
|
||||
"""Wire-counter output parity: --out-format %b/%c/%C, --progress and
|
||||
--stats versus real rsync 3.4.1."""
|
||||
@@ -344,8 +443,8 @@ class TestWireStatsParity:
|
||||
result, _ = run_client(source, dest, flags=["-a", "--out-format=" + fmt],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, result.stderr[:300]
|
||||
rb, rl = (int(x) for x in rsync_result.stdout.split()[:2])
|
||||
fb, fl = (int(x) for x in result.stdout.split()[:2])
|
||||
rb, rl = (int(x) for x in _file_entry_line(rsync_result.stdout).split()[:2])
|
||||
fb, fl = (int(x) for x in _file_entry_line(result.stdout).split()[:2])
|
||||
assert rl == fl == 5000, (rsync_result.stdout, result.stdout)
|
||||
assert rb > rl, f"rsync %b must include framing: {rsync_result.stdout!r}"
|
||||
assert fb > fl, f"fastsync %b must include framing: {result.stdout!r}"
|
||||
@@ -378,7 +477,9 @@ class TestWireStatsParity:
|
||||
assert file_lines(result.stdout) == file_lines(rsync_result.stdout), (
|
||||
f"rsync={rsync_result.stdout!r} fastsync={result.stdout!r}"
|
||||
)
|
||||
assert result.stdout.split()[0] == rsync_result.stdout.split()[0] == "16", (
|
||||
fs_c = _file_entry_line(result.stdout).split()[0]
|
||||
rs_c = _file_entry_line(rsync_result.stdout).split()[0]
|
||||
assert fs_c == rs_c == "16", (
|
||||
f"%c must be rsync's 16-byte sum header: {result.stdout!r}"
|
||||
)
|
||||
|
||||
@@ -407,8 +508,8 @@ class TestWireStatsParity:
|
||||
"--out-format=" + fmt],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, result.stderr[:300]
|
||||
rs_c = int(rsync_result.stdout.split()[0])
|
||||
fs_c = int(result.stdout.split()[0])
|
||||
rs_c = int(_file_entry_line(rsync_result.stdout).split()[0])
|
||||
fs_c = int(_file_entry_line(result.stdout).split()[0])
|
||||
# No basis exists, so rsync still reports only its sum header.
|
||||
assert rs_c == 16, rsync_result.stdout
|
||||
# FastSync reports its own handshake bytes and is not aligned.
|
||||
@@ -447,6 +548,99 @@ class TestWireStatsParity:
|
||||
assert fast_frames[0] == rsync_frames[0], (rsync_frames[0], fast_frames[0])
|
||||
assert "(xfr#1," in fast_frames[-1], fast_frames[-1]
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_progress_leading_root_line_and_to_chk_match_rsync(self, shared_server):
|
||||
"""A single-file transfer: rsync emits the transfer-root `./` name line
|
||||
and a `to-chk=0/2` denominator that counts that root entry. Both must
|
||||
match FastSync byte-for-byte for the deterministic frames."""
|
||||
source = os.path.join(TEST_DATA_DIR, "wire_pgroot_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "wire_pgroot_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "wire_pgroot_rdst")
|
||||
_make_one_file(source, "f.bin", 100)
|
||||
clean_dir(dest)
|
||||
# rsync prints the `./` root line only when the transfer root itself is
|
||||
# created, so make the rsync destination absent. The "created directory"
|
||||
# line it then emits has no FastSync counterpart (different mirror
|
||||
# layout), so only the name/frame lines are compared.
|
||||
shutil.rmtree(rdst, ignore_errors=True)
|
||||
rsync_result = _rsync(["-a", "--progress", source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
result, _ = run_client(source, dest, flags=["-a", "--progress"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, result.stderr[:300]
|
||||
|
||||
# subprocess text mode normalizes \r to \n (universal newlines).
|
||||
def lines_of(text):
|
||||
return [ln for ln in text.splitlines() if ln and not ln.startswith("created directory")]
|
||||
|
||||
rsync_lines = lines_of(rsync_result.stdout)
|
||||
fast_lines = lines_of(result.stdout)
|
||||
rsync_names = [ln for ln in rsync_lines if "%" not in ln]
|
||||
fast_names = [ln for ln in fast_lines if "%" not in ln]
|
||||
|
||||
assert rsync_names == ["sending incremental file list", "./", "f.bin"], rsync_names
|
||||
assert fast_names == rsync_names, (rsync_names, fast_names)
|
||||
# The final frame's to-chk denominator must include the source-root entry.
|
||||
assert "to-chk=0/2" in fast_lines[-1], fast_lines[-1]
|
||||
assert fast_lines[-1] == rsync_lines[-1], (rsync_lines[-1], fast_lines[-1])
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
@pytest.mark.parametrize("mt", [False, True])
|
||||
def test_progress_multidir_file_list_matches_rsync(self, shared_server, mt):
|
||||
"""A multi-directory tree: the paths-only pre-count must reproduce
|
||||
rsync's file-list set and `to-chk` denominator. Per-directory name
|
||||
lines are emitted for directories, symlinks and the empty directory; the
|
||||
name set and the denominator (every entry plus the transfer root) match
|
||||
rsync, while the emitted *order* remains a documented residual (rsync
|
||||
sorts depth-first, FastSync streams in readdir/BFS order)."""
|
||||
source = os.path.join(TEST_DATA_DIR, "wire_pgmd_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "wire_pgmd_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "wire_pgmd_rdst")
|
||||
_make_multidir_tree(source)
|
||||
clean_dir(dest)
|
||||
clean_dir(rdst)
|
||||
|
||||
rsync_result = _rsync(["-a", "--progress", source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
flags = ["-a", "--progress"] + (["--threads"] if mt else [])
|
||||
result, _ = run_client(source, dest, flags=flags, port=shared_server.port)
|
||||
assert result.returncode == 0, result.stderr[:300]
|
||||
|
||||
rsync_names, rsync_totals = _parse_progress(rsync_result.stdout)
|
||||
fast_names, fast_totals = _parse_progress(result.stdout)
|
||||
assert sorted(rsync_names) == [
|
||||
"a.txt", "b.txt", "emptydir/", "link1 -> a.txt", "sub1/",
|
||||
"sub1/c.txt", "sub1/deep/", "sub1/deep/d.txt", "sub2/", "sub2/e.txt",
|
||||
], rsync_names
|
||||
assert fast_names == rsync_names, (rsync_names, fast_names)
|
||||
# 10 entries + the transfer-root "." counted by rsync's file list.
|
||||
assert rsync_totals == {11}, rsync_totals
|
||||
assert fast_totals == rsync_totals, (rsync_totals, fast_totals)
|
||||
|
||||
@pytest.mark.ci
|
||||
def test_progress_delete_during_reuses_pre_scan(self):
|
||||
"""--delete-during + --progress reuses the keep-set pre-scan instead of
|
||||
walking the tree a second time: the file-list total and directory name
|
||||
lines are identical to a plain --progress run."""
|
||||
source = os.path.join(TEST_DATA_DIR, "wire_pgdel_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "wire_pgdel_dst")
|
||||
_make_multidir_tree(source)
|
||||
clean_dir(dest)
|
||||
server = ServerManager()
|
||||
server.start(extra_args=["--allow-super", "--allow-delete"])
|
||||
try:
|
||||
result, _ = run_client(source, dest, flags=["-a", "--progress", "--delete-during"],
|
||||
port=server.port)
|
||||
finally:
|
||||
server.stop()
|
||||
assert result.returncode == 0, result.stderr[:300]
|
||||
names, totals = _parse_progress(result.stdout)
|
||||
assert totals == {11}, totals
|
||||
assert "sub1/" in names and "sub1/deep/" in names and "emptydir/" in names, names
|
||||
assert "link1 -> a.txt" in names, names
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
@pytest.mark.parametrize("mt", [False, True])
|
||||
@@ -459,6 +653,9 @@ class TestWireStatsParity:
|
||||
_make_one_file(source, "f.bin", 6000)
|
||||
clean_dir(dest)
|
||||
clean_dir(rdst)
|
||||
# Start both tools from the same state: rsync's destination root exists,
|
||||
# so pre-create FastSync's mirrored logical root as well.
|
||||
os.makedirs(get_dest_received_dir(dest, source), exist_ok=True)
|
||||
rsync_result = _rsync(["-a", "--stats", source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
flags = ["-a", "--stats"] + (["--threads"] if mt else [])
|
||||
@@ -488,24 +685,19 @@ class TestWireStatsParity:
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_stats_file_count_breakdown_residual(self, shared_server):
|
||||
"""Residual (row #3): rsync prints the `Number of files` and
|
||||
`Number of created files` lines with a per-type breakdown
|
||||
(`(reg: X, dir: Y, link: Z)`).
|
||||
|
||||
FastSync cannot reproduce it from what the sender currently knows: the
|
||||
scanner does not put directory entries in the transfer list (directories
|
||||
are created implicitly), and without a per-entry destination-probe the
|
||||
sender cannot tell which entries the receiver newly created. So FastSync
|
||||
prints the bare transferred-entry count. This test pins the divergence
|
||||
explicitly -- the row must not be marked ✅.
|
||||
"""
|
||||
def test_stats_file_count_breakdown_matches_rsync(self, shared_server):
|
||||
"""`Number of files` and `Number of created files` both carry rsync's
|
||||
per-type breakdown (protocol 2.28.0 reports the receiver-created
|
||||
reg/dir/link/special split over STATUS_STATS)."""
|
||||
source = os.path.join(TEST_DATA_DIR, "wire_stc_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "wire_stc_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "wire_stc_rdst")
|
||||
_make_one_file(source, "f.bin", 6000)
|
||||
clean_dir(dest)
|
||||
clean_dir(rdst)
|
||||
# Start both tools from the same state: rsync's destination root exists,
|
||||
# so pre-create FastSync's mirrored logical root as well.
|
||||
os.makedirs(get_dest_received_dir(dest, source), exist_ok=True)
|
||||
rsync_result = _rsync(["-a", "--stats", source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
result, _ = run_client(source, dest, flags=["-a", "--stats"],
|
||||
@@ -523,14 +715,144 @@ class TestWireStatsParity:
|
||||
f_files = stats_line(result.stdout, "Number of files")
|
||||
f_created = stats_line(result.stdout, "Number of created files")
|
||||
|
||||
# rsync always carries the type breakdown (the source root counts as a
|
||||
# directory; the single regular file as reg).
|
||||
assert re.match(r"Number of files: 2 \(reg: 1, dir: 1\)$", r_files), r_files
|
||||
assert r_files == f_files, (r_files, f_files)
|
||||
assert re.match(r"Number of created files: 1 \(reg: 1\)$", r_created), r_created
|
||||
# FastSync prints only the bare count: no directory accounting and no
|
||||
# per-entry "created" knowledge.
|
||||
assert re.fullmatch(r"Number of files: 1", f_files), f_files
|
||||
assert re.fullmatch(r"Number of created files: 1", f_created), f_created
|
||||
assert f_created == r_created, (r_created, f_created)
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
@pytest.mark.parametrize("mt", [False, True])
|
||||
def test_stats_r_directory_breakdown_matches_rsync(self, shared_server, mt):
|
||||
"""A recursive `-r` scan (no -t/-p) exposes no directory metadata, but
|
||||
rsync still counts every directory in `Number of files`; the sender's
|
||||
lightweight directory counter must reproduce the `dir: N` category."""
|
||||
source = os.path.join(TEST_DATA_DIR, "wire_stdir_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "wire_stdir_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "wire_stdir_rdst")
|
||||
clean_dir(source)
|
||||
clean_dir(dest)
|
||||
clean_dir(rdst)
|
||||
os.makedirs(os.path.join(source, "sub", "deep"))
|
||||
os.makedirs(os.path.join(source, "empty"))
|
||||
for rel in ("a.txt", os.path.join("sub", "b.txt"), os.path.join("sub", "deep", "c.txt")):
|
||||
with open(os.path.join(source, rel), "wb") as fh:
|
||||
fh.write(b"x\n")
|
||||
os.makedirs(get_dest_received_dir(dest, source), exist_ok=True)
|
||||
|
||||
rsync_result = _rsync(["-r", "--stats", source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
flags = ["-r", "--stats"] + (["--threads"] if mt else [])
|
||||
result, _ = run_client(source, dest, flags=flags, port=shared_server.port)
|
||||
assert result.returncode == 0, result.stderr[:300]
|
||||
|
||||
def stats_line(text, key):
|
||||
for line in text.splitlines():
|
||||
if line.startswith(key + ":"):
|
||||
return line
|
||||
return None
|
||||
|
||||
r_files = stats_line(rsync_result.stdout, "Number of files")
|
||||
f_files = stats_line(result.stdout, "Number of files")
|
||||
# 3 regular files, 4 directories (root, sub, sub/deep, empty).
|
||||
assert re.match(r"Number of files: 7 \(reg: 3, dir: 4\)$", r_files), r_files
|
||||
assert f_files == r_files, (r_files, f_files)
|
||||
assert (stats_line(result.stdout, "Number of regular files transferred") ==
|
||||
stats_line(rsync_result.stdout, "Number of regular files transferred"))
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
@pytest.mark.parametrize("mt", [False, True])
|
||||
def test_stats_created_and_literal_fresh_update_delta(self, shared_server, mt):
|
||||
"""The receiver-observed counters must match rsync for the three
|
||||
transfer shapes: a fresh create (created breakdown + whole-file literal),
|
||||
an update (created == 0, whole-file literal), and a delta update (only
|
||||
the literal delta fragments are counted, not the whole file)."""
|
||||
source = os.path.join(TEST_DATA_DIR, "wire_stcd_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "wire_stcd_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "wire_stcd_rdst")
|
||||
clean_dir(source)
|
||||
clean_dir(dest)
|
||||
clean_dir(rdst)
|
||||
os.makedirs(source, exist_ok=True)
|
||||
os.makedirs(get_dest_received_dir(dest, source), exist_ok=True)
|
||||
with open(os.path.join(source, "big.bin"), "wb") as fh:
|
||||
fh.write(bytes(range(256)) * 4096) # 1 MiB
|
||||
mt_flag = ["--threads"] if mt else []
|
||||
|
||||
def compare(tag):
|
||||
# Pin the delta block size on both ends: rsync's adaptive block size
|
||||
# would otherwise make the literal/matched split non-comparable.
|
||||
rsync_result = _rsync(["-a", "--stats", "--no-whole-file", "-B8192",
|
||||
source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
result, _ = run_client(
|
||||
source, dest,
|
||||
flags=["-a", "--stats", "--incremental", "--delta", "-B8192"] + mt_flag,
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, result.stderr[:300]
|
||||
keys = ("Number of created files", "Literal data", "Matched data",
|
||||
"Total transferred file size")
|
||||
r = _pick_stats(rsync_result.stdout, keys)
|
||||
f = _pick_stats(result.stdout, keys)
|
||||
assert r == f, f"{tag}: rsync={r} fastsync={f}"
|
||||
return r
|
||||
|
||||
fresh = compare("fresh")
|
||||
assert re.match(r"Number of created files: 1 \(reg: 1\)$",
|
||||
fresh["Number of created files"]), fresh
|
||||
|
||||
# Update the source and re-run: the destination already exists.
|
||||
sleep_mtime = os.path.getmtime(os.path.join(source, "big.bin")) + 2
|
||||
with open(os.path.join(source, "big.bin"), "r+b") as fh:
|
||||
fh.seek(100)
|
||||
fh.write(b"XXXXXXXXXX")
|
||||
os.utime(os.path.join(source, "big.bin"), (sleep_mtime, sleep_mtime))
|
||||
update = compare("update")
|
||||
assert update["Number of created files"] == "Number of created files: 0", update
|
||||
|
||||
# Second delta update: change bytes far apart, so rsync ships only the
|
||||
# literal fragments and FastSync must report the same Literal data.
|
||||
sleep_mtime = os.path.getmtime(os.path.join(source, "big.bin")) + 2
|
||||
with open(os.path.join(source, "big.bin"), "r+b") as fh:
|
||||
fh.seek(500000)
|
||||
fh.write(b"YYYYYYYYYY")
|
||||
os.utime(os.path.join(source, "big.bin"), (sleep_mtime, sleep_mtime))
|
||||
delta = compare("delta")
|
||||
assert delta["Number of created files"] == "Number of created files: 0", delta
|
||||
lit = int(delta["Literal data"].split(":", 1)[1].strip().split()[0].replace(",", ""))
|
||||
assert 0 < lit < 1024 * 1024, delta
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
@pytest.mark.parametrize("choice", ["xxh128", "xxh64", "xxh3", "md5", "md4", "sha1", "none"])
|
||||
def test_out_format_C_selected_algorithm_matches_rsync(self, shared_server, choice):
|
||||
"""`%C` must use the algorithm selected by --checksum-choice, not always
|
||||
xxh128, and render it exactly like rsync (big-endian for the 64-bit
|
||||
hashes, high-then-low for xxh128, standard hex for md5/md4/sha1)."""
|
||||
source = os.path.join(TEST_DATA_DIR, f"wire_cc_{choice}_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, f"wire_cc_{choice}_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, f"wire_cc_{choice}_rdst")
|
||||
_make_one_file(source, "f.bin", 200000)
|
||||
clean_dir(dest)
|
||||
clean_dir(rdst)
|
||||
fmt = "%C %l %n"
|
||||
rsync_result = _rsync(["-a", "--checksum-choice=" + choice,
|
||||
"--out-format=" + fmt, source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
result, _ = run_client(source, dest,
|
||||
flags=["-a", "--checksum-choice=" + choice,
|
||||
"--out-format=" + fmt],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, result.stderr[:300]
|
||||
|
||||
def file_lines(text):
|
||||
return [line for line in text.splitlines()
|
||||
if line and not line.rsplit(" ", 1)[-1].endswith("/")]
|
||||
|
||||
assert file_lines(result.stdout) == file_lines(rsync_result.stdout), (
|
||||
f"choice={choice}: rsync={rsync_result.stdout!r} fastsync={result.stdout!r}"
|
||||
)
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
|
||||
@@ -0,0 +1,198 @@
|
||||
"""Differential rsync-parity coverage for two residuals closed on this branch.
|
||||
|
||||
* A4 -- ``--compare-dest``/``--copy-dest``/``--link-dest`` relative-DIR
|
||||
resolution: rsync resolves a relative DIR against the destination directory
|
||||
and appends the file's TRANSFER-RELATIVE name. FastSync's default transfer
|
||||
mirrors the absolute source path below its receive root, so a naive relative
|
||||
DIR used to probe a different tree. These tests seed the basis at rsync's
|
||||
spelling and assert FastSync finds it (byte-exact / hard-linked / sparse),
|
||||
matching real rsync 3.4.1.
|
||||
|
||||
* A5 -- ``-y``/``--fuzzy`` candidate eligibility: rsync's ``find_fuzzy`` has no
|
||||
delta-size gate, so it reuses an oversized (>10x) or sub-16-KiB sibling;
|
||||
FastSync used to decline both. These tests assert FastSync now uses the same
|
||||
sibling as rsync (observable as ``Matched data``) with a byte-exact result.
|
||||
|
||||
Every test skips cleanly when rsync is absent.
|
||||
"""
|
||||
import os
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
import pytest
|
||||
|
||||
sys.path.insert(0, os.path.dirname(__file__))
|
||||
from common import ( # noqa: E402
|
||||
TEST_DATA_DIR,
|
||||
clean_dir,
|
||||
get_dest_received_dir,
|
||||
run_client,
|
||||
)
|
||||
|
||||
RSYNC = shutil.which("rsync")
|
||||
requires_rsync = pytest.mark.skipif(RSYNC is None, reason="rsync 3.4.1 not installed")
|
||||
|
||||
OLD_MTIME = 1_500_000_000
|
||||
|
||||
|
||||
def _write(path, content, mtime=None):
|
||||
os.makedirs(os.path.dirname(path), exist_ok=True)
|
||||
with open(path, "wb") as fh:
|
||||
fh.write(content)
|
||||
if mtime is not None:
|
||||
os.utime(path, (mtime, mtime))
|
||||
|
||||
|
||||
def _read(path):
|
||||
with open(path, "rb") as fh:
|
||||
return fh.read()
|
||||
|
||||
|
||||
def _rsync(args):
|
||||
env = dict(os.environ, LC_ALL="C")
|
||||
return subprocess.run([RSYNC] + args, capture_output=True, text=True, env=env, timeout=120)
|
||||
|
||||
|
||||
def _stat_bytes(text, label):
|
||||
for line in text.splitlines():
|
||||
if line.startswith(label + ":"):
|
||||
return int(line.split(":", 1)[1].strip().split()[0].replace(",", ""))
|
||||
return None
|
||||
|
||||
|
||||
class TestRelativeBasisDirResolution:
|
||||
"""A4: a relative basis DIR must resolve to the same tree as rsync's."""
|
||||
|
||||
_FILES = {
|
||||
"root.txt": b"root-basis-content\n",
|
||||
"sub/nested.txt": b"nested-basis-content\n",
|
||||
}
|
||||
|
||||
def _seed_source(self, source):
|
||||
clean_dir(source)
|
||||
for rel, data in self._FILES.items():
|
||||
_write(os.path.join(source, rel), data, OLD_MTIME)
|
||||
return self._FILES
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.parametrize("flag", ["--compare-dest", "--link-dest"])
|
||||
def test_relative_dir_resolves_like_rsync(self, shared_server, flag):
|
||||
tag = flag.lstrip("-")
|
||||
source = os.path.join(TEST_DATA_DIR, f"relbasis_{tag}_src")
|
||||
rdst = os.path.join(TEST_DATA_DIR, f"relbasis_{tag}_rdst")
|
||||
fdst = os.path.join(TEST_DATA_DIR, f"relbasis_{tag}_fdst")
|
||||
self._seed_source(source)
|
||||
|
||||
# rsync: relative DIR -> dest/basis/<transfer-relative name>.
|
||||
clean_dir(rdst)
|
||||
for rel, data in self._FILES.items():
|
||||
_write(os.path.join(rdst, "basis", rel), data, OLD_MTIME)
|
||||
rs = _rsync(["-a", f"{flag}=basis", source + "/", rdst + "/"])
|
||||
assert rs.returncode == 0, rs.stderr
|
||||
|
||||
# FastSync: the SAME relative spelling seeded at the SAME
|
||||
# transfer-relative location under its destination root.
|
||||
clean_dir(fdst)
|
||||
for rel, data in self._FILES.items():
|
||||
_write(os.path.join(fdst, "basis", rel), data, OLD_MTIME)
|
||||
result, _ = run_client(source, fdst,
|
||||
flags=["-a", f"{flag}=basis", "--incremental"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
received = get_dest_received_dir(fdst, source)
|
||||
|
||||
for rel, data in self._FILES.items():
|
||||
rfile = os.path.join(rdst, rel)
|
||||
ffile = os.path.join(received, rel)
|
||||
basis = os.path.join(fdst, "basis", rel)
|
||||
if flag == "--compare-dest":
|
||||
# compare-dest never copies: both destinations stay sparse.
|
||||
assert not os.path.exists(rfile), f"rsync copied {rel}"
|
||||
assert not os.path.exists(ffile), (
|
||||
f"FastSync did not resolve the relative basis DIR at {basis!r} "
|
||||
f"(expected {rel!r} to stay sparse like rsync)")
|
||||
else:
|
||||
# link-dest hard-links; a basis miss would transfer a new file.
|
||||
assert os.path.exists(ffile), f"FastSync lost {rel}"
|
||||
assert _read(ffile) == data
|
||||
assert os.stat(ffile).st_ino == os.stat(basis).st_ino, (
|
||||
f"FastSync did not hard-link {rel!r} to the relative basis at "
|
||||
f"{basis!r} (basis not resolved like rsync)")
|
||||
|
||||
@requires_rsync
|
||||
def test_relative_dir_copy_dest_content(self, shared_server):
|
||||
source = os.path.join(TEST_DATA_DIR, "relbasis_copy_src")
|
||||
fdst = os.path.join(TEST_DATA_DIR, "relbasis_copy_fdst")
|
||||
self._seed_source(source)
|
||||
clean_dir(fdst)
|
||||
for rel, data in self._FILES.items():
|
||||
_write(os.path.join(fdst, "basis", rel), data, OLD_MTIME)
|
||||
result, _ = run_client(source, fdst,
|
||||
flags=["-a", "--copy-dest=basis", "--incremental"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
received = get_dest_received_dir(fdst, source)
|
||||
for rel, data in self._FILES.items():
|
||||
ffile = os.path.join(received, rel)
|
||||
assert os.path.exists(ffile), f"copy-dest did not materialize {rel}"
|
||||
assert _read(ffile) == data
|
||||
assert os.stat(ffile).st_ino != os.stat(os.path.join(fdst, "basis", rel)).st_ino
|
||||
|
||||
|
||||
class TestFuzzyEligibilityWindow:
|
||||
"""A5: --fuzzy candidate eligibility must match rsync's uncapped window."""
|
||||
|
||||
BASE = b"the quick brown fox jumps over the lazy dog\n" * 4000
|
||||
|
||||
def _run_pair(self, shared_server, tag, payload, sibling):
|
||||
source = os.path.join(TEST_DATA_DIR, f"fzw_{tag}_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, f"fzw_{tag}_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, f"fzw_{tag}_rdst")
|
||||
clean_dir(source)
|
||||
clean_dir(dest)
|
||||
clean_dir(rdst)
|
||||
_write(os.path.join(source, "report_v2.txt"), payload)
|
||||
for root in (rdst, get_dest_received_dir(dest, source)):
|
||||
_write(os.path.join(root, "report_v1.txt"), sibling)
|
||||
|
||||
rs = _rsync(["-a", "--no-whole-file", "--fuzzy", "--stats",
|
||||
source + "/", rdst + "/"])
|
||||
assert rs.returncode == 0, rs.stderr
|
||||
result, _ = run_client(
|
||||
source, dest,
|
||||
flags=["-a", "--incremental", "--delta", "--fuzzy", "--stats"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
|
||||
# The reconstructed file is byte-exact in every case.
|
||||
assert _read(os.path.join(get_dest_received_dir(dest, source),
|
||||
"report_v2.txt")) == payload
|
||||
return rs, result
|
||||
|
||||
@requires_rsync
|
||||
def test_oversized_sibling_eligible_like_rsync(self, shared_server):
|
||||
"""A sibling 20x the source is used by rsync; FastSync must too (its old
|
||||
10x delta-size gate declined it)."""
|
||||
n = 65536
|
||||
payload = (self.BASE * ((n // len(self.BASE)) + 1))[:n]
|
||||
sibling = (self.BASE * 200)[: n * 20]
|
||||
rs, result = self._run_pair(shared_server, "big", payload, sibling)
|
||||
assert _stat_bytes(rs.stdout, "Matched data") > 0, \
|
||||
"rsync should use a >10x fuzzy basis"
|
||||
assert _stat_bytes(result.stdout, "Matched data") > 0, (
|
||||
"FastSync's fuzzy eligibility must accept a >10x sibling like rsync "
|
||||
f"(Matched data={_stat_bytes(result.stdout, 'Matched data')})")
|
||||
|
||||
@requires_rsync
|
||||
def test_small_source_sibling_eligible_like_rsync(self, shared_server):
|
||||
"""A sub-16-KiB source with an identical sibling is used by rsync;
|
||||
FastSync's old 16 KiB delta minimum declined it."""
|
||||
n = 8192
|
||||
payload = (self.BASE * ((n // len(self.BASE)) + 1))[:n]
|
||||
rs, result = self._run_pair(shared_server, "small", payload, payload)
|
||||
assert _stat_bytes(rs.stdout, "Matched data") > 0, \
|
||||
"rsync applies --fuzzy below 16 KiB"
|
||||
assert _stat_bytes(result.stdout, "Matched data") > 0, (
|
||||
"FastSync's fuzzy eligibility must accept a sub-16-KiB source like "
|
||||
f"rsync (Matched data={_stat_bytes(result.stdout, 'Matched data')})")
|
||||
@@ -0,0 +1,105 @@
|
||||
"""`--debug=FLAGS` natural-event categories (no-wire).
|
||||
|
||||
FastSync maps the rsync `--debug` categories that correspond to a real event it
|
||||
already performs (``flist``, ``del``, ``hash``/``deltasum``, ``recv``,
|
||||
``filter`` and ``send``) onto debug output. A normal run prints none of it.
|
||||
"""
|
||||
import os
|
||||
import sys
|
||||
|
||||
import pytest
|
||||
|
||||
sys.path.insert(0, os.path.dirname(__file__))
|
||||
from common import TEST_DATA_DIR, run_client, clean_dir, get_dest_received_dir, ServerManager
|
||||
|
||||
|
||||
def _make_tree(root):
|
||||
clean_dir(root)
|
||||
os.makedirs(os.path.join(root, "sub"))
|
||||
with open(os.path.join(root, "a.txt"), "wb") as fh:
|
||||
fh.write(b"alpha\n")
|
||||
with open(os.path.join(root, "keep.log"), "wb") as fh:
|
||||
fh.write(b"log\n")
|
||||
with open(os.path.join(root, "sub", "b.txt"), "wb") as fh:
|
||||
fh.write(b"beta\n")
|
||||
|
||||
|
||||
@pytest.mark.ci
|
||||
def test_debug_flist_and_send_emit_output(shared_server):
|
||||
"""`--debug=flist,send` produces category-tagged debug output."""
|
||||
source = os.path.join(TEST_DATA_DIR, "dbg_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "dbg_dst")
|
||||
_make_tree(source)
|
||||
clean_dir(dest)
|
||||
result, _ = run_client(source, dest, flags=["-a", "--debug=flist,send"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
assert "flist: scanning" in result.stdout, result.stdout
|
||||
assert "send: " in result.stdout, result.stdout
|
||||
|
||||
|
||||
@pytest.mark.ci
|
||||
def test_debug_filter_emits_excluded_entry(shared_server):
|
||||
source = os.path.join(TEST_DATA_DIR, "dbg_filter_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "dbg_filter_dst")
|
||||
_make_tree(source)
|
||||
clean_dir(dest)
|
||||
result, _ = run_client(source, dest,
|
||||
flags=["-a", "--debug=filter", "--exclude=*.log"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
assert "filter: excluded keep.log" in result.stdout, result.stdout
|
||||
|
||||
|
||||
@pytest.mark.ci
|
||||
def test_debug_hash_and_recv_emit_on_incremental(shared_server):
|
||||
source = os.path.join(TEST_DATA_DIR, "dbg_hash_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "dbg_hash_dst")
|
||||
_make_tree(source)
|
||||
clean_dir(dest)
|
||||
result, _ = run_client(source, dest,
|
||||
flags=["-a", "--incremental", "--checksum",
|
||||
"--debug=hash,recv"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
assert "hash: " in result.stdout, result.stdout
|
||||
assert "recv: " in result.stdout, result.stdout
|
||||
|
||||
|
||||
@pytest.mark.ci
|
||||
def test_debug_del_emits_deleted_path():
|
||||
"""`--debug=del` reports the paths the receiver actually removed.
|
||||
|
||||
A deletion-capable server is required (the shared fixture refuses
|
||||
client-requested deletion)."""
|
||||
source = os.path.join(TEST_DATA_DIR, "dbg_del_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "dbg_del_dst")
|
||||
_make_tree(source)
|
||||
clean_dir(dest)
|
||||
seeded = get_dest_received_dir(dest, source)
|
||||
os.makedirs(seeded)
|
||||
with open(os.path.join(seeded, "extra.tmp"), "wb") as fh:
|
||||
fh.write(b"stale\n")
|
||||
server = ServerManager()
|
||||
server.start(extra_args=["--allow-super", "--allow-delete"])
|
||||
try:
|
||||
result, _ = run_client(source, dest, flags=["-a", "--delete", "--debug=del"],
|
||||
port=server.port)
|
||||
finally:
|
||||
server.stop()
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
assert "del: " in result.stdout and "extra.tmp" in result.stdout, result.stdout
|
||||
assert not os.path.exists(os.path.join(seeded, "extra.tmp"))
|
||||
|
||||
|
||||
@pytest.mark.ci
|
||||
def test_normal_run_has_no_debug_output(shared_server):
|
||||
source = os.path.join(TEST_DATA_DIR, "dbg_quiet_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "dbg_quiet_dst")
|
||||
_make_tree(source)
|
||||
clean_dir(dest)
|
||||
result, _ = run_client(source, dest, flags=["-a"], port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
assert "[DEBUG]" not in result.stdout
|
||||
assert "flist: scanning" not in result.stdout
|
||||
assert "send: " not in result.stdout
|
||||
@@ -0,0 +1,174 @@
|
||||
"""Differential parity for `--info=mount` and `--info=stats` (no-wire).
|
||||
|
||||
Both behaviours are compared against real rsync 3.4.1:
|
||||
|
||||
* `--info=mount` prints rsync's ``[sender] skipping mount-point dir NAME`` line
|
||||
when ``-xx`` drops a mount-point directory. Plain ``-x`` keeps the empty
|
||||
directory and stays silent, exactly like rsync.
|
||||
* `--info=stats` requests the same transfer-statistics block as `--stats`
|
||||
(rsync spells the full block ``--info=stats2``/``--stats``).
|
||||
|
||||
The tests are skipped when rsync is unavailable.
|
||||
"""
|
||||
import os
|
||||
import re
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
import pytest
|
||||
|
||||
sys.path.insert(0, os.path.dirname(__file__))
|
||||
from common import TEST_DATA_DIR, run_client, clean_dir, get_dest_received_dir
|
||||
|
||||
RSYNC = shutil.which("rsync")
|
||||
requires_rsync = pytest.mark.skipif(RSYNC is None, reason="rsync 3.4.1 not installed")
|
||||
|
||||
|
||||
def _rsync(args):
|
||||
env = dict(os.environ, LC_ALL="C")
|
||||
return subprocess.run([RSYNC] + args, capture_output=True, text=True, env=env, timeout=120)
|
||||
|
||||
|
||||
def _cross_device_mount_tree(source):
|
||||
"""Build a source whose ``nested_link`` is a symlink onto a tmpfs directory.
|
||||
|
||||
``--copy-links`` dereferences it so ``-x`` sees a mount-point directory on a
|
||||
different device. Returns the probe path to remove, or skips the test when
|
||||
no cross-device filesystem is available.
|
||||
"""
|
||||
local = os.stat(".")
|
||||
shm = "/dev/shm"
|
||||
try:
|
||||
shm_stat = os.stat(shm)
|
||||
except OSError:
|
||||
pytest.skip("/dev/shm not available")
|
||||
if shm_stat.st_dev == local.st_dev:
|
||||
pytest.skip("no cross-device filesystem available")
|
||||
|
||||
clean_dir(source)
|
||||
with open(os.path.join(source, "keep.txt"), "wb") as fh:
|
||||
fh.write(b"keep\n")
|
||||
probe = os.path.join(shm, f"fastsync_info_mount_{os.getpid()}")
|
||||
shutil.rmtree(probe, ignore_errors=True)
|
||||
os.makedirs(probe)
|
||||
with open(os.path.join(probe, "inside.txt"), "wb") as fh:
|
||||
fh.write(b"cross\n")
|
||||
try:
|
||||
os.symlink(probe, os.path.join(source, "nested_link"))
|
||||
except OSError:
|
||||
shutil.rmtree(probe, ignore_errors=True)
|
||||
pytest.skip("cannot create symlink")
|
||||
return probe
|
||||
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_info_mount_xx_matches_rsync(shared_server):
|
||||
"""`-xx --info=mount` drops the mount-point dir and prints rsync's line."""
|
||||
source = os.path.join(TEST_DATA_DIR, "info_mount_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "info_mount_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "info_mount_rdst")
|
||||
probe = _cross_device_mount_tree(source)
|
||||
clean_dir(dest)
|
||||
clean_dir(rdst)
|
||||
flags = ["-a", "--copy-links", "-xx", "--info=mount"]
|
||||
try:
|
||||
rsync_result = _rsync(flags + [source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
result, _ = run_client(source, dest, flags=flags, port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
|
||||
expected = "[sender] skipping mount-point dir nested_link"
|
||||
assert expected in rsync_result.stdout, rsync_result.stdout
|
||||
assert expected in result.stdout, (result.stdout, result.stderr)
|
||||
|
||||
received = get_dest_received_dir(dest, source)
|
||||
assert os.path.exists(os.path.join(received, "keep.txt"))
|
||||
# -xx omits the mount-point directory entirely.
|
||||
assert not os.path.exists(os.path.join(received, "nested_link"))
|
||||
assert not os.path.exists(os.path.join(rdst, "nested_link"))
|
||||
finally:
|
||||
shutil.rmtree(probe, ignore_errors=True)
|
||||
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_info_mount_single_x_is_silent(shared_server):
|
||||
"""Plain `-x` keeps the empty mount-point directory and prints no line."""
|
||||
source = os.path.join(TEST_DATA_DIR, "info_mount1_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "info_mount1_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "info_mount1_rdst")
|
||||
probe = _cross_device_mount_tree(source)
|
||||
clean_dir(dest)
|
||||
clean_dir(rdst)
|
||||
flags = ["-a", "--copy-links", "-x", "--info=mount"]
|
||||
try:
|
||||
rsync_result = _rsync(flags + [source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
result, _ = run_client(source, dest, flags=flags, port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
|
||||
assert "skipping mount-point dir" not in rsync_result.stdout
|
||||
assert "skipping mount-point dir" not in result.stdout
|
||||
|
||||
received = get_dest_received_dir(dest, source)
|
||||
assert os.path.isdir(os.path.join(received, "nested_link"))
|
||||
assert not os.path.exists(os.path.join(received, "nested_link", "inside.txt"))
|
||||
assert os.path.isdir(os.path.join(rdst, "nested_link"))
|
||||
assert not os.path.exists(os.path.join(rdst, "nested_link", "inside.txt"))
|
||||
finally:
|
||||
shutil.rmtree(probe, ignore_errors=True)
|
||||
|
||||
|
||||
def _make_stats_tree(root):
|
||||
clean_dir(root)
|
||||
os.makedirs(os.path.join(root, "sub"))
|
||||
with open(os.path.join(root, "a.txt"), "wb") as fh:
|
||||
fh.write(b"alpha\n")
|
||||
with open(os.path.join(root, "sub", "b.txt"), "wb") as fh:
|
||||
fh.write(b"beta\n")
|
||||
|
||||
|
||||
def _pick_stats(text):
|
||||
keys = ("Number of files", "Number of regular files transferred", "Total file size",
|
||||
"Total transferred file size", "Literal data", "Matched data")
|
||||
out = {}
|
||||
for line in text.splitlines():
|
||||
for key in keys:
|
||||
if line.startswith(key + ":"):
|
||||
out[key] = line
|
||||
return out
|
||||
|
||||
|
||||
@requires_rsync
|
||||
@pytest.mark.ci
|
||||
def test_info_stats_emits_full_stats_block(shared_server):
|
||||
"""`--info=stats` is the same full block as `--stats` and matches rsync."""
|
||||
source = os.path.join(TEST_DATA_DIR, "info_stats_src")
|
||||
dest = os.path.join(TEST_DATA_DIR, "info_stats_dst")
|
||||
rdst = os.path.join(TEST_DATA_DIR, "info_stats_rdst")
|
||||
dest2 = os.path.join(TEST_DATA_DIR, "info_stats_dst2")
|
||||
_make_stats_tree(source)
|
||||
for path in (dest, rdst, dest2):
|
||||
clean_dir(path)
|
||||
os.makedirs(get_dest_received_dir(path, source), exist_ok=True)
|
||||
|
||||
rsync_result = _rsync(["-a", "--stats", source + "/", rdst + "/"])
|
||||
assert rsync_result.returncode == 0, rsync_result.stderr
|
||||
|
||||
info_result, _ = run_client(source, dest, flags=["-a", "--info=stats"],
|
||||
port=shared_server.port)
|
||||
assert info_result.returncode == 0, (info_result.stderr or info_result.stdout)[:300]
|
||||
stats_result, _ = run_client(source, dest2, flags=["-a", "--stats"],
|
||||
port=shared_server.port)
|
||||
assert stats_result.returncode == 0, (stats_result.stderr or stats_result.stdout)[:300]
|
||||
|
||||
# --info=stats must print the same block as --stats...
|
||||
assert _pick_stats(info_result.stdout) == _pick_stats(stats_result.stdout), (
|
||||
f"info={info_result.stdout} stats={stats_result.stdout}")
|
||||
# ...and the protocol-independent counters must match real rsync.
|
||||
assert _pick_stats(info_result.stdout) == _pick_stats(rsync_result.stdout), (
|
||||
f"rsync={_pick_stats(rsync_result.stdout)} fastsync={_pick_stats(info_result.stdout)}")
|
||||
assert re.search(r"^Number of files: \d+ \(reg: 2, dir: 2\)$", info_result.stdout,
|
||||
re.MULTILINE), info_result.stdout
|
||||
@@ -0,0 +1,185 @@
|
||||
"""Differential rsync-parity coverage for FastSync's transfer/delete ORDER.
|
||||
|
||||
rsync walks a source tree in its sorted flist order: within each directory the
|
||||
non-directories come first (ascending name), then the subdirectories (ascending
|
||||
name), each subdirectory immediately followed by its own subtree (depth-first).
|
||||
The sequential scanner now reproduces that order, which makes both the
|
||||
``--info=name`` stream and the ``--delete-during`` deletion sequence match real
|
||||
``rsync 3.4.1`` exactly. ``--threads`` has no rsync analogue and is unordered.
|
||||
|
||||
Every test skips cleanly when rsync is absent.
|
||||
"""
|
||||
import os
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
import pytest
|
||||
|
||||
sys.path.insert(0, os.path.dirname(__file__))
|
||||
from common import ( # noqa: E402
|
||||
TEST_DATA_DIR,
|
||||
ServerManager,
|
||||
clean_dir,
|
||||
get_dest_received_dir,
|
||||
run_client,
|
||||
)
|
||||
|
||||
RSYNC = shutil.which("rsync")
|
||||
requires_rsync = pytest.mark.skipif(RSYNC is None, reason="rsync 3.4.1 not installed")
|
||||
|
||||
MTIME = 1_500_000_000
|
||||
|
||||
_TREE = {
|
||||
"a.txt": b"a\n",
|
||||
"b.txt": b"b\n",
|
||||
"z.txt": b"z\n",
|
||||
"a_dir/f.txt": b"f\n",
|
||||
"a_dir/deep/g.txt": b"g\n",
|
||||
"m_dir/h.txt": b"h\n",
|
||||
"Z_dir/i.txt": b"i\n",
|
||||
}
|
||||
|
||||
|
||||
def _write(path, data):
|
||||
os.makedirs(os.path.dirname(path), exist_ok=True)
|
||||
with open(path, "wb") as fh:
|
||||
fh.write(data)
|
||||
os.utime(path, (MTIME, MTIME))
|
||||
|
||||
|
||||
def _rsync(args):
|
||||
env = dict(os.environ, LC_ALL="C")
|
||||
return subprocess.run([RSYNC] + args, capture_output=True, text=True, env=env, timeout=120)
|
||||
|
||||
|
||||
def _deleting(text):
|
||||
out = []
|
||||
for line in text.splitlines():
|
||||
stripped = line.strip()
|
||||
if stripped.startswith("*deleting") or stripped.startswith("deleting"):
|
||||
out.append(stripped.split()[-1])
|
||||
return out
|
||||
|
||||
|
||||
class TestTransferOrderParity:
|
||||
@requires_rsync
|
||||
def test_info_name_file_order_matches_rsync(self, shared_server):
|
||||
source = os.path.join(TEST_DATA_DIR, "order_name_src")
|
||||
clean_dir(source)
|
||||
for rel, data in _TREE.items():
|
||||
_write(os.path.join(source, rel), data)
|
||||
|
||||
rdst = os.path.join(TEST_DATA_DIR, "order_name_rdst")
|
||||
clean_dir(rdst)
|
||||
rs = _rsync(["-a", "--info=name", source + "/", rdst + "/"])
|
||||
assert rs.returncode == 0, rs.stderr
|
||||
# rsync also names the directories (trailing '/'); FastSync names the
|
||||
# transferred entries. Compare the file/symlink sequence, which is what
|
||||
# the traversal order determines.
|
||||
rsync_files = [l for l in rs.stdout.splitlines() if l.strip() and not l.endswith("/")]
|
||||
|
||||
fdst = os.path.join(TEST_DATA_DIR, "order_name_fdst")
|
||||
clean_dir(fdst)
|
||||
result, _ = run_client(source, fdst, flags=["-a", "--info=name"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
fsync_files = [
|
||||
l for l in result.stdout.splitlines()
|
||||
if l.strip() and l.strip() != "./" and not l.startswith("sending")
|
||||
]
|
||||
assert fsync_files == rsync_files, (
|
||||
f"transfer order differs\nrsync={rsync_files}\nfastsync={fsync_files}")
|
||||
|
||||
|
||||
class TestDeleteOrderParity:
|
||||
_EXTRA = {
|
||||
"a_extra.txt": b"a\n",
|
||||
"z_extra.txt": b"z\n",
|
||||
"a_extra_dir/f": b"f\n",
|
||||
"z_extra_dir/f": b"f\n",
|
||||
"a_extra_dir/sub/g": b"g\n",
|
||||
}
|
||||
|
||||
def _seed_source(self):
|
||||
source = os.path.join(TEST_DATA_DIR, "order_del_src")
|
||||
clean_dir(source)
|
||||
_write(os.path.join(source, "keep.txt"), b"k\n")
|
||||
_write(os.path.join(source, "keepdir", "x.txt"), b"x\n")
|
||||
_write(os.path.join(source, "keep2", "y.txt"), b"y\n")
|
||||
return source
|
||||
|
||||
def _assert_order(self, timing, dry_run=False):
|
||||
source = self._seed_source()
|
||||
rdst = os.path.join(TEST_DATA_DIR, f"order_{timing}_rdst")
|
||||
clean_dir(rdst)
|
||||
for rel, data in self._EXTRA.items():
|
||||
_write(os.path.join(rdst, rel), data)
|
||||
rs_flags = ["-a", "-n"] if dry_run else ["-a"]
|
||||
rs = _rsync(rs_flags + [timing, "--info=del", source + "/", rdst + "/"])
|
||||
assert rs.returncode == 0, rs.stderr
|
||||
|
||||
fdst = os.path.join(TEST_DATA_DIR, f"order_{timing}_fdst")
|
||||
clean_dir(fdst)
|
||||
received = get_dest_received_dir(fdst, source)
|
||||
for rel, data in self._EXTRA.items():
|
||||
_write(os.path.join(received, rel), data)
|
||||
fs_flags = ["-a", "-n"] if dry_run else ["-a"]
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
result, _ = run_client(source, fdst, flags=fs_flags + [timing, "--info=del"],
|
||||
port=server.port)
|
||||
assert result.returncode == 0, (result.stderr or result.stdout)[:300]
|
||||
|
||||
rsync_order = _deleting(rs.stdout)
|
||||
fsync_order = _deleting(result.stdout)
|
||||
assert sorted(fsync_order) == sorted(rsync_order), (
|
||||
f"{timing} deleted set differs\nrsync={rsync_order}\nfastsync={fsync_order}")
|
||||
assert fsync_order == rsync_order, (
|
||||
f"{timing} deletion order differs\nrsync={rsync_order}\nfastsync={fsync_order}")
|
||||
|
||||
@requires_rsync
|
||||
def test_delete_during_deletion_order_matches_rsync(self):
|
||||
self._assert_order("--delete-during")
|
||||
|
||||
@requires_rsync
|
||||
def test_delete_delay_deletion_order_matches_rsync(self):
|
||||
self._assert_order("--delete-delay")
|
||||
|
||||
@requires_rsync
|
||||
def test_dry_run_delete_order_matches_rsync(self):
|
||||
self._assert_order("--delete", dry_run=True)
|
||||
|
||||
@requires_rsync
|
||||
def test_partial_max_delete_survivor_order_matches_rsync(self):
|
||||
"""With the exact removal order matching rsync, a --max-delete cap stops
|
||||
after the same entries, so the survivor set is identical too."""
|
||||
source = os.path.join(TEST_DATA_DIR, "order_maxdel_src")
|
||||
clean_dir(source)
|
||||
_write(os.path.join(source, "keep.txt"), b"k\n")
|
||||
extra = {f"e{i}.txt": b"x\n" for i in range(6)}
|
||||
extra["ed/f"] = b"f\n"
|
||||
extra["ed/g"] = b"g\n"
|
||||
|
||||
rdst = os.path.join(TEST_DATA_DIR, "order_maxdel_rdst")
|
||||
clean_dir(rdst)
|
||||
for rel, data in extra.items():
|
||||
_write(os.path.join(rdst, rel), data)
|
||||
rs = _rsync(["-a", "--delete-during", "--max-delete=3", "--info=del",
|
||||
source + "/", rdst + "/"])
|
||||
assert rs.returncode in (0, 25), (rs.returncode, rs.stderr)
|
||||
|
||||
fdst = os.path.join(TEST_DATA_DIR, "order_maxdel_fdst")
|
||||
clean_dir(fdst)
|
||||
received = get_dest_received_dir(fdst, source)
|
||||
for rel, data in extra.items():
|
||||
_write(os.path.join(received, rel), data)
|
||||
with ServerManager() as server:
|
||||
server.start(extra_args=["--allow-delete"])
|
||||
result, _ = run_client(source, fdst,
|
||||
flags=["-a", "--delete-during", "--max-delete=3", "--info=del"],
|
||||
port=server.port)
|
||||
assert result.returncode in (0, 25), (result.returncode, result.stderr[:300])
|
||||
assert _deleting(result.stdout) == _deleting(rs.stdout), (
|
||||
f"partial --max-delete survivor order differs\n"
|
||||
f"rsync={_deleting(rs.stdout)}\nfastsync={_deleting(result.stdout)}")
|
||||
@@ -719,13 +719,18 @@ class TestVerifyAndFlip:
|
||||
source = self._src("cmpd")
|
||||
dest = self._dst("cmpd")
|
||||
rdst = self._dst("cmpd_r")
|
||||
# Pin the mtime so rsync's size+mtime quick-check (and FastSync's
|
||||
# default) matches deterministically across a second boundary.
|
||||
OLD = 1_500_000_000
|
||||
with open(os.path.join(source, "f.txt"), "wb") as fh:
|
||||
fh.write(b"basis-content\n")
|
||||
os.utime(os.path.join(source, "f.txt"), (OLD, OLD))
|
||||
# rsync resolves --compare-dest relative to the destination dir; its
|
||||
# basis file sits at the transfer-relative path.
|
||||
os.makedirs(os.path.join(rdst, "basis"), exist_ok=True)
|
||||
with open(os.path.join(rdst, "basis", "f.txt"), "wb") as fh:
|
||||
fh.write(b"basis-content\n")
|
||||
os.utime(os.path.join(rdst, "basis", "f.txt"), (OLD, OLD))
|
||||
rs = _rsync(["-a", "--compare-dest=basis", source + "/", rdst + "/"])
|
||||
assert rs.returncode == 0, rs.stderr
|
||||
assert not os.path.exists(os.path.join(rdst, "f.txt")), \
|
||||
@@ -738,6 +743,7 @@ class TestVerifyAndFlip:
|
||||
os.makedirs(basis, exist_ok=True)
|
||||
with open(os.path.join(basis, "f.txt"), "wb") as fh:
|
||||
fh.write(b"basis-content\n")
|
||||
os.utime(os.path.join(basis, "f.txt"), (OLD, OLD))
|
||||
received = get_dest_received_dir(dest, source)
|
||||
result, _ = run_client(source, dest,
|
||||
flags=["--compare-dest=basis", "--incremental"],
|
||||
@@ -751,14 +757,17 @@ class TestVerifyAndFlip:
|
||||
def test_link_dest_hardlinks_matches_rsync(self, shared_server):
|
||||
source = self._src("linkd")
|
||||
dest = self._dst("linkd")
|
||||
OLD = 1_500_000_000
|
||||
with open(os.path.join(source, "f.txt"), "wb") as fh:
|
||||
fh.write(b"link-basis-content\n")
|
||||
os.utime(os.path.join(source, "f.txt"), (OLD, OLD))
|
||||
rel = os.path.abspath(source).lstrip(os.sep)
|
||||
basis = os.path.join(dest, "basis", rel)
|
||||
os.makedirs(basis, exist_ok=True)
|
||||
basis_file = os.path.join(basis, "f.txt")
|
||||
with open(basis_file, "wb") as fh:
|
||||
fh.write(b"link-basis-content\n")
|
||||
os.utime(basis_file, (OLD, OLD))
|
||||
received = get_dest_received_dir(dest, source)
|
||||
result, _ = run_client(source, dest,
|
||||
flags=["--link-dest=basis", "--incremental"],
|
||||
@@ -769,6 +778,162 @@ class TestVerifyAndFlip:
|
||||
assert os.stat(dest_file).st_ino == os.stat(basis_file).st_ino, \
|
||||
"--link-dest must hard-link to the basis file"
|
||||
|
||||
@requires_rsync
|
||||
def test_basis_dir_size_only_content_residual(self, shared_server):
|
||||
"""rsync parity (default): a basis hit is decided by the metadata
|
||||
quick-check alone. With `--size-only`, a same-size, different-content
|
||||
basis is trusted, so rsync links the basis content and FastSync must now
|
||||
do the same instead of xxHash-verifying it. `--verify-basis` restores
|
||||
the stricter content equality (covered by the differential test)."""
|
||||
source = self._src("basissz")
|
||||
rdest = self._dst("basissz_r")
|
||||
fdest = self._dst("basissz_f")
|
||||
with open(os.path.join(source, "f.txt"), "wb") as fh:
|
||||
fh.write(b"AAAA\n")
|
||||
OLD = 1_400_000_000
|
||||
# rsync basis at the transfer-relative path (relative to the dest dir).
|
||||
os.makedirs(os.path.join(rdest, "basis"), exist_ok=True)
|
||||
with open(os.path.join(rdest, "basis", "f.txt"), "wb") as fh:
|
||||
fh.write(b"BBBB\n")
|
||||
os.utime(os.path.join(rdest, "basis", "f.txt"), (OLD, OLD))
|
||||
rs = _rsync(["-a", "--size-only", "--link-dest=basis", source + "/", rdest + "/"])
|
||||
assert rs.returncode == 0, rs.stderr
|
||||
with open(os.path.join(rdest, "f.txt"), "rb") as fh:
|
||||
assert fh.read() == b"BBBB\n", "rsync --size-only did not trust the basis size"
|
||||
|
||||
# FastSync basis is relative to the receive root; the file mirrors the
|
||||
# source path.
|
||||
rel = os.path.abspath(source).lstrip(os.sep)
|
||||
basis = os.path.join(fdest, "basis", rel)
|
||||
os.makedirs(basis, exist_ok=True)
|
||||
with open(os.path.join(basis, "f.txt"), "wb") as fh:
|
||||
fh.write(b"BBBB\n")
|
||||
os.utime(os.path.join(basis, "f.txt"), (OLD, OLD))
|
||||
received = get_dest_received_dir(fdest, source)
|
||||
result, _ = run_client(source, fdest,
|
||||
flags=["-a", "--size-only", "--link-dest=basis", "--incremental"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, result.stderr[:300]
|
||||
with open(os.path.join(received, "f.txt"), "rb") as fh:
|
||||
assert fh.read() == b"BBBB\n", \
|
||||
"FastSync must trust the metadata quick-check exactly like rsync"
|
||||
|
||||
@requires_rsync
|
||||
def test_verify_basis_restores_content_check(self, shared_server):
|
||||
"""FastSync-only `--verify-basis`: a same-size, same-mtime basis with
|
||||
different content is rejected by the whole-file digest, so the source is
|
||||
transferred instead of installing the wrong basis bytes. The default
|
||||
(no flag) installs the basis content, matching rsync."""
|
||||
source = self._src("vbasis")
|
||||
fdest = self._dst("vbasis_f")
|
||||
with open(os.path.join(source, "f.txt"), "wb") as fh:
|
||||
fh.write(b"AAAA\n")
|
||||
OLD = 1_400_000_000
|
||||
os.utime(os.path.join(source, "f.txt"), (OLD, OLD))
|
||||
rel = os.path.abspath(source).lstrip(os.sep)
|
||||
basis = os.path.join(fdest, "basis", rel)
|
||||
os.makedirs(basis, exist_ok=True)
|
||||
with open(os.path.join(basis, "f.txt"), "wb") as fh:
|
||||
fh.write(b"BBBB\n")
|
||||
os.utime(os.path.join(basis, "f.txt"), (OLD, OLD))
|
||||
received = get_dest_received_dir(fdest, source)
|
||||
result, _ = run_client(source, fdest,
|
||||
flags=["-a", "--link-dest=basis", "--incremental",
|
||||
"--verify-basis"],
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, result.stderr[:300]
|
||||
with open(os.path.join(received, "f.txt"), "rb") as fh:
|
||||
assert fh.read() == b"AAAA\n", \
|
||||
"--verify-basis must reject the same-size/different-content basis"
|
||||
|
||||
|
||||
def _stat_bytes(output, key):
|
||||
"""Parse a --stats byte counter (e.g. ``Matched data: 65,536 bytes``)."""
|
||||
for line in output.splitlines():
|
||||
if line.startswith(key + ":"):
|
||||
raw = line.split(":", 1)[1].strip().split()[0]
|
||||
return int(raw.replace(",", ""))
|
||||
return None
|
||||
|
||||
|
||||
class TestFuzzy:
|
||||
"""Track 5b: `-y`/`--fuzzy` is an internal bandwidth optimization with a
|
||||
byte-exact result. FastSync ports rsync 3.4.1's weighted-Levenshtein name
|
||||
heuristic, so where both delta engines admit the candidate the tools pick
|
||||
the same basis (the ``fuzzy_basis`` differential asserts the tree and the
|
||||
Matched/Literal counters match with the block size pinned). Candidate
|
||||
ELIGIBILITY is now rsync's too: the fuzzy search no longer inherits the
|
||||
ordinary delta engine's 16 KiB minimum or 10x size-ratio bound, so an
|
||||
oversized or sub-16-KiB sibling is reused exactly as rsync reuses it.
|
||||
These tests pin that window on both sides."""
|
||||
|
||||
_BASE = b"the quick brown fox jumps over the lazy dog\n" * 4000
|
||||
|
||||
def _src(self, tag):
|
||||
source = os.path.join(TEST_DATA_DIR, f"fz_{tag}_src")
|
||||
clean_dir(source)
|
||||
return source
|
||||
|
||||
def _dst(self, tag):
|
||||
d = os.path.join(TEST_DATA_DIR, f"fz_{tag}_dst")
|
||||
clean_dir(d)
|
||||
return d
|
||||
|
||||
def _run_both(self, shared_server, source, dest, rdst, payload, sibling,
|
||||
rs_extra=(), fs_extra=()):
|
||||
with open(os.path.join(source, "report_v2.txt"), "wb") as fh:
|
||||
fh.write(payload)
|
||||
for root in (rdst, get_dest_received_dir(dest, source)):
|
||||
os.makedirs(root, exist_ok=True)
|
||||
with open(os.path.join(root, "report_v1.txt"), "wb") as fh:
|
||||
fh.write(sibling)
|
||||
rs = _rsync(["-a", "--no-whole-file", "--fuzzy", "--stats"] +
|
||||
list(rs_extra) + [source + "/", rdst + "/"])
|
||||
assert rs.returncode == 0, rs.stderr[:300]
|
||||
result, _ = run_client(
|
||||
source, dest,
|
||||
flags=["-a", "--incremental", "--delta", "--fuzzy", "--stats"] +
|
||||
list(fs_extra),
|
||||
port=shared_server.port)
|
||||
assert result.returncode == 0, result.stderr[:300]
|
||||
_assert_same_tree(rdst, get_dest_received_dir(dest, source), "(--fuzzy)")
|
||||
return rs, result
|
||||
|
||||
@requires_rsync
|
||||
def test_fuzzy_above_size_window_matches_rsync(self, shared_server):
|
||||
"""A sibling >10x the source is used by rsync and by FastSync: fuzzy
|
||||
eligibility is rsync's, not the ordinary delta size-ratio gate; both
|
||||
destinations stay byte-identical and both reuse the basis."""
|
||||
n = 65536
|
||||
payload = (self._BASE * ((n // len(self._BASE)) + 1))[:n]
|
||||
sibling = (self._BASE * 200)[: n * 20]
|
||||
source, dest, rdst = (self._src("big"), self._dst("big"),
|
||||
self._dst("big_r"))
|
||||
rs, result = self._run_both(shared_server, source, dest, rdst,
|
||||
payload, sibling)
|
||||
assert _stat_bytes(rs.stdout, "Matched data") > 0, \
|
||||
"rsync should use a >10x fuzzy basis"
|
||||
assert _stat_bytes(result.stdout, "Matched data") > 0, \
|
||||
"FastSync must accept a >10x fuzzy basis like rsync"
|
||||
assert _stat_bytes(result.stdout, "Literal data") < n
|
||||
|
||||
@requires_rsync
|
||||
def test_fuzzy_below_delta_minimum_matches_rsync(self, shared_server):
|
||||
"""A sibling below the 16 KiB delta minimum is used by rsync and by
|
||||
FastSync: fuzzy eligibility no longer inherits the delta engine's
|
||||
minimum; both trees stay byte-identical and both reuse the basis."""
|
||||
n = 8192
|
||||
payload = (self._BASE * ((n // len(self._BASE)) + 1))[:n]
|
||||
source, dest, rdst = (self._src("small"), self._dst("small"),
|
||||
self._dst("small_r"))
|
||||
rs, result = self._run_both(shared_server, source, dest, rdst,
|
||||
payload, payload)
|
||||
assert _stat_bytes(rs.stdout, "Matched data") > 0, \
|
||||
"rsync applies --fuzzy below 16 KiB"
|
||||
assert _stat_bytes(result.stdout, "Matched data") > 0, \
|
||||
"FastSync must apply --fuzzy below 16 KiB like rsync"
|
||||
assert _stat_bytes(result.stdout, "Literal data") < n
|
||||
|
||||
|
||||
class TestIgnoreExistingShortCircuit:
|
||||
"""#9: --ignore-existing is decided by the receiver during the per-file
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user