334 Commits
Author SHA1 Message Date
TapTap ef76c9034d Merge pull request 'Release v2.26.0' (#284) from dev into main
CI / lint (push) Successful in 1m40s
CI / sanitizers (undefined) (push) Successful in 44s
CI / sanitizers (address) (push) Successful in 50s
CI / build-and-test (push) Successful in 58s
CI / coverage (push) Successful in 40s
CI / fuzz-build (push) Successful in 46s
CI / valgrind (push) Successful in 2m12s
Reviewed-on: #284
2026-09-18 19:05:51 +02:00
TapTap cee281ff55 Merge pull request 'fix(test): stop the valgrind hang, init per_dir_filter_count' (#299) from fix/valgrind-hang into dev
CI / lint (push) Successful in 1m41s
CI / lint (pull_request) Successful in 1m40s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / sanitizers (undefined) (push) Successful in 44s
CI / sanitizers (address) (push) Successful in 50s
CI / build-and-test (push) Successful in 54s
CI / fuzz-build (push) Successful in 45s
CI / coverage (push) Successful in 41s
CI / build-and-test (pull_request) Successful in 44s
CI / valgrind (push) Successful in 2m12s
2026-09-17 23:47:47 +02:00
TapTap 5c509831b8 fix(test): robust valgrind detection + per-suite io-fd reset; init per_dir_filter_count
CI / lint (pull_request) Successful in 1m41s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / build-and-test (pull_request) Successful in 46s
The post-merge valgrind job on dev hangs.  Root cause: the CI valgrind step
exports FASTSYNC_UNDER_VALGRIND=1, but nothing read it, and the
/proc/self/maps "vgpreload" probe is unreliable on valgrind 3.22 (the guest's
maps no longer list the tool's own libraries).  So the fork-based unit tests
ran under valgrind anyway; tests that call io_set_fds() left the thread-local
read/write descriptors pointing at a closed test pipe, and a later
send_n_data()/receive_n_data() call was silently redirected to those stale fds
(legacy_session() prefers the globals, which the stdin/stdout SSH server
requires).  Later tests only worked by fd-reuse luck; under valgrind the fd
numbers no longer coincide, so the read blocked forever on an empty pipe.

- test_utils.h: honor FASTSYNC_UNDER_VALGRIND (already set by ci.yaml) and keep
  the maps scan as a best-effort fallback.  Reset io_set_fds(-1, -1) at the
  start of every RUN_TEST so one suite cannot leak descriptor redirection into
  the next.
- test_iconv.c: skip the forking wire-string roundtrip under valgrind like the
  other fork-based tests.
- config.c: initialize per_dir_filter_count in config_set_defaults.  The field
  was never initialized, so -F/-FF counting read uninitialized heap (valgrind:
  conditional jump on uninitialised value at client_cli.c:1464) and could count
  from garbage.

Verified with the CI-equivalent command (FASTSYNC_UNDER_VALGRIND=1 valgrind
--leak-check=full --show-leak-kinds=definite --error-exitcode=1): completes
with 0 errors (previously hung >80 min).  Unit 43/43; full integration 729
passed; cppcheck and clang-format clean.
2026-09-17 23:44:41 +02:00
TapTap ee57aeea4a Merge pull request 'rsync 3.4.1 drop-in parity (#285-#297) + parity completion (protocol 2.26.0)' (#298) from feat/rsync-parity into dev
CI / lint (pull_request) Successful in 1m46s
CI / lint (push) Successful in 1m47s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / build-and-test (pull_request) Successful in 47s
CI / sanitizers (address) (push) Successful in 50s
CI / build-and-test (push) Successful in 59s
CI / sanitizers (undefined) (push) Successful in 47s
CI / fuzz-build (push) Successful in 45s
CI / coverage (push) Successful in 44s
CI / valgrind (push) Failing after 3h0m1s
2026-09-17 20:33:16 +02:00
TapTap 803c1d3385 fix: review pass — uninit stats, append crash, filter rollback, ASan leak
CI / lint (pull_request) Successful in 1m45s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / build-and-test (pull_request) Successful in 45s
Address findings from the four-agent review of PR #298:

- file_create: zero the new File.matched_bytes.  It was uninitialized
  malloc memory, so the receiver could sum a garbage value into
  STATUS_STATS "Matched data" (nondeterministic --stats divergence and
  an uninitialized-heap disclosure on the wire).
- send_append: load the source into memory before hashing/copying the
  prefix and tail.  Files >64 MiB without compression (and --sendfile
  runs) are streamed without loading, so --append/--append-verify
  dereferenced a NULL data->data and crashed.
- filter_file_append: clamp the rollback to the surviving rule count.
  A "clear" rule in a merge file frees every rule including the
  caller's; the old rollback rewound count to rules_before and
  resurrected freed pointers for a double free / UAF.  Also roll back
  when set_rule_owner fails instead of leaving owner-less rules.
- filter_rule_parse: reject the xattr-name filter modifier (x), which
  was parsed and silently reinterpreted as a filename rule (affecting
  what --delete protects).  The p modifier stays accepted (existing
  grammar test).
- Remove two dead functions: compression_default_algo and
  change_render_itemize_code.
- tests: free ctx->would_delete in the two test_multiprocessing manual
  teardowns (ASan leak, 1648 bytes/run).
- docs: correct the RSYNC_COMPAT/HANDOFF tally (156 rows: 106/27/23),
  downgrade --info/--debug to caveat with their silent categories, add
  %C-vs-xxh64 and --delete-delay count caveats, refresh stale xattr
  mode comments, and document -p special-bit (setuid/setgid/sticky)
  parity plus its mitigations.
2026-09-17 20:30:15 +02:00
TapTap b8ec62beef fix(file): create implicit directories with 0777 & ~umask (rsync parity)
CI / lint (pull_request) Successful in 1m58s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / build-and-test (pull_request) Successful in 45s
Implicit parent directories were created with a hardcoded 0755, diverging from rsync's 0777 & ~umask whenever the process umask is not 022 (the CI runner uses 0). Matches rsync under any umask; verified with umask 0.
2026-09-17 19:39:21 +02:00
TapTap 59bfd32b9e docs(usage): correct --temp-dir help to the confined receive-root behavior
CI / lint (pull_request) Successful in 1m40s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / build-and-test (pull_request) Failing after 57s
2026-09-17 19:33:16 +02:00
TapTap 3432a33d9a docs(parity): finalize rsync-parity docs for protocol 2.26.0
Reclassify the RSYNC_COMPAT matrix after the parity-completion wave (protocol
2.23.0 -> 2.26.0): 9 already-parity rows to parity, 17 inherently non-rsync
rows to divergent, and the genuine fixes to parity, recounting to
109 parity / 25 caveat / 23 divergent of 157 rows. Add rows for --bwlimit,
--partial, --partial-dir, --no-whole-file, --inc-recursive/--no-inc-recursive,
--protect-args and --msgs2stderr; fix the documented -f/filter, -F/.rsync-filter,
empty --files-from, and --preallocate/--sparse precedence bugs; add the Parity
Completion Wave section.

Refresh README/HANDOFF/release skill/cmake-expert to protocol 2.26.0 and the
zlib/lz4 + md4/sha1/none codecs, add the CHANGELOG 2.26.0 entry, and correct
the stale --max-delete --help wording.
2026-09-17 19:25:06 +02:00
TapTap a1b081d328 Merge branch 'fix/parity-tests' into feat/parity-completion 2026-09-17 01:38:19 +02:00
TapTap bbecff9c04 fix(delete): keep the empty-scan safety guard file-only
Adding every traversed directory to the per-directory plan keep set must not
make an I/O-errored partial scan look non-empty.  Count only transmitted file
entries for delete_plan_sender_empty(), so a scan that hit an unreadable
directory and found no files still refuses to delete.
2026-09-17 01:34:26 +02:00
TapTap c1553bd5d6 style(delete): simplify redundant root[0] check (cppcheck) 2026-09-17 01:31:19 +02:00
TapTap 1a550bda24 test: pin delta-mode %c divergence against rsync
Add test_out_format_c_delta_mode_divergence documenting that FastSync's
delta %c (its own handshake bytes) cannot match rsync's (16-byte sum header
plus per-block checksums); only the whole-file case is aligned.  The test
pins both values so a future parity improvement is noticed.
2026-09-17 01:30:09 +02:00
TapTap 902f86192d test: stats file-count residual, --threads coverage, deterministic delete timing
- Output parity: add `File list size` to the strict --stats differential and
  document the row-#3 residual with test_stats_file_count_breakdown_residual
  (rsync's `Number of files`/`Number of created files` type breakdown is not
  reproducible from what the sender knows: no directory accounting and no
  per-entry destination-created state).  The row stays a caveat.
- Add --threads variants for the --stats and --progress/-P differentials.
  The `-n --delete --threads` variant is a documented xfail: the threaded
  dry-run path does not consume the receiver's STATUS_STATS delete list yet.
- Replace the 0.2s sleep flake in the delete-timing proxy with a socket
  barrier: the hook now fires only after the server sends a reply (proving it
  processed the preceding per-directory delete plan), using --incremental +
  --ignore-times to guarantee a mid-transfer handshake reply.
2026-09-17 01:29:29 +02:00
TapTap 6a40ac86e5 test(parity): cover --threads for -R delete scope and empty-dir retention 2026-09-17 01:25:37 +02:00
TapTap 410ba6e992 style: clang-format receiver.c and server.c 2026-09-17 01:23:47 +02:00
TapTap 6c6f02e5dd fix(parity): empty-dir delete, per-dir filter errors, -R protect, stats parser
Blockers addressed together (shared scanner/delete-plan plumbing):

* #10: an empty in-scope source directory produced no plan keep entry, so the
  receiver deleted the destination directory itself.  The scanner now records
  every traversed directory into a delete-plan sink, the plan sender keeps them,
  and any directory whose plan the data stream never triggered is emitted after
  the data so its extras are still removed.  Differential tests cover
  --delete-during and --delete-delay.
* #8: an invalid per-directory filter file was silently ignored when an earlier
  merge file in the same directory existed; key the failure off the error text
  (both sequential and parallel scanners) and fail the scan.
* #9: -R + --files-from receiver-protect rules recorded the source-relative
  path; record the bare relative wire path in both scanners so the protected
  destination mirror survives --delete.
* #5: the STATUS_STATS would-delete parser now validates each retained path and
  enforces the shared MAX_MANIFEST_BYTES budget, and the --out-format dry-run
  delete line is escaped like the itemize line.
* #11: drop the unused DELETE_PLAN_MAX_NAMES macro, log the delete-limit
  warning once per session, roll back dir-merge names from a per-directory file
  that fails to parse, and guard every filter error snprintf against err==NULL.

#10 leaves the empty directory itself kept and its extras removed, matching
rsync's final state on both per-directory timings.
2026-09-17 01:21:06 +02:00
TapTap d9006d1fda client: report rsync's 16-byte %c sum header for whole-file transfers
rsync's %c counts the block-checksum bytes received: even a whole-file
transfer with no basis receives rsync's 16-byte sum header (append and
inplace included), while a dry run receives nothing.  FastSync's
whole-file path has no equivalent header, so report 16 for parity when
delta is inactive, keep 0 for dry runs, and keep the real received bytes
when delta is active (FastSync's signature framing differs, so delta %c
stays divergent).  Add a strict %c/%l/%n differential against rsync and
turn the %b check into a real rsync differential (semantics: both count
wire bytes and exceed %l; the exact values are protocol-specific).
2026-09-17 01:19:07 +02:00
TapTap 36375010d3 client: accept rsync 3.4.1 --info/--debug category vocabulary
Accept the full rsync 3.4.1 --info (backup, del, flist, mount, nonreg,
progress, remove, symsafe) and --debug (acl, backup, bind, chdir, connect,
cmd, del, deltasum, dup, exit, filter, flist, fuzzy, genr, hash, hlink,
iconv, nstr, own, recv, send, time) vocabularies, plus the historical
syms/hl/owner aliases, with level suffixes.  Categories FastSync already
emits (copy/name/misc/skip/stats and io/proto/pack/util) still set their
log flags; the rest are accepted but silent.  Unknown names remain
rejected by name, matching rsync.  Update the CLI unit tests and the
rsync-parity integration tests (previously they required del/filter to be
rejected).
2026-09-17 01:16:39 +02:00
TapTap 5efa0dba7c test: differential st_blocks for --sparse/--preallocate and --ignore-existing wire volume
- Assert FastSync's sparse/preallocate block accounting matches rsync 3.4.1
  for a hole file (preallocate wins over sparse, exactly like rsync).
- Assert --ignore-existing is answered by the receiver before the sender
  transmits the payload (CountingProxy wire volume near-zero).
- CountingProxy.run gains a bounded join_timeout so tests that only need the
  client->server count do not wait for the server's idle socket.
- Fix the stale protocol 2.24.0 comment in test_config.c (golden is 2.26.0).
2026-09-17 01:13:05 +02:00
TapTap e32733fbf6 feat(stats): populate receiver wire counters on both receive paths
The single-threaded and -m receivers never populated ReceiverStats.matched_data
or .deleted_files, so --stats always printed 0 for both even when rsync
reported nonzero.  Track the bytes reconstructed from the basis file while
applying a delta, and tally the delete-commit counts (manifest and
per-directory sessions) into the receiver stats.  The -m pipeline now carries
its own stats/would-delete fields and emits the STATUS_STATS frame before the
terminal success, so --threads finally reports the counters and renders
-n --delete  lines.

Also normalize the -n --delete would-delete enumeration's absolute basis
prefixes exactly like the real commit path (fixing an over-report) and fix the
basis_delete_relative off-by-one when the receive root is '/'.  Unit tests
cover the root mapping and the basis protection; integration tests cover
matched/deleted stats for both receivers and the --threads dry-run delete
lines.
2026-09-17 01:08:32 +02:00
TapTap b02799327d fix(delete): guard the per-directory delete commit against dry-run
delete_plan_session_commit() lacked the central no-mutation guard that
manifest_delete_all() has, so a server-contacting -n run (or a hostile plan
frame) could still remove --delete-missing-args mirrors on the per-directory
timing path.  Return DELETE_COMMIT_OK immediately when the session is a
dry-run, and gate the receiver/server commit call sites too.  Add a unit test
that streams a plan naming an existing destination file and asserts it
survives.
2026-09-17 01:02:22 +02:00
TapTap 946aa934cc fix(delete): scope -R per-directory delete walk to the transferred prefix
The -R prefix marker installed in synced_dirs was discarded when finalizing
the per-directory delete sender (--delete-during/--delete-delay), so the
up-front root plan was the receive root '.', whose keep list only held the
first prefix component.  The receiver then deleted destination content
outside the transferred prefix (e.g. unrelated/keep.txt), a data-loss bug;
rsync keeps it.

Confine the walk to the -R prefix: send that prefix's plan as the root plan,
only transmit plans at or below it, and never emit the receive root plan for
a scoped run.  Add a differential test covering both --delete-during and
--delete-delay.
2026-09-17 01:00:55 +02:00
TapTap 125921c11b Merge branch 'feat/parity-codecs' into feat/parity-completion
# Conflicts:
#	src/shared/checksum.h
#	src/shared/config.h
#	tests/integration/test_fault_injection.py
#	tests/integration/test_preflight.py
#	tests/test_client_cli.c
#	tests/test_config.c
#	tests/test_fuzz_smoke.c
2026-09-16 23:49:41 +02:00
TapTap 9dd5288381 Merge branch 'feat/parity-wirestats' into feat/parity-completion
# Conflicts:
#	src/server/receiver_pipeline.c
#	src/shared/config.h
#	src/shared/protocol.h
#	tests/integration/test_fault_injection.py
#	tests/integration/test_preflight.py
#	tests/test_client_cli.c
#	tests/test_config.c
2026-09-16 23:46:28 +02:00
TapTap 3f2c74dd9e Merge branch 'feat/parity-deltiming2' into feat/parity-completion 2026-09-16 23:44:13 +02:00
TapTap 51e41dee2a Merge branch 'feat/parity-leftovers' into feat/parity-completion
# Conflicts:
#	src/client/scanner.c
#	src/client/scanner.h
2026-09-16 23:44:13 +02:00
TapTap dbf1b39d47 fix(delete): share the per-frame manifest byte budget across delete-plan sections 2026-09-16 23:36:46 +02:00
TapTap 845f20a28d docs(delete): describe per-directory timing in usage and config comments 2026-09-16 23:30:26 +02:00
TapTap a5083776da test(delete): cover per-dir timings for files-from scope, max-delete, excluded protection, refuse-delete, type conflicts 2026-09-16 23:29:25 +02:00
TapTap 76a81f1684 test(codec): differential coverage vs rsync and wire-golden updates
Add tests/integration/test_codecs.py (accept/reject matrix and byte
differential against rsync 3.4.1 for every algorithm), extend the
checksum/compression unit tests with MD4/SHA1/none vectors and per-codec
round-trips, pin the new config golden (2.26.0), and update the version
strings and codec acceptance expectations.
2026-09-16 23:27:44 +02:00
TapTap 24b81c7e5a feat(codec): negotiate checksum/compression algorithms (protocol 2.26.0)
Accept the full rsync 3.4.1 --compress-choice set (zstd/lz4/zlib/zlibx/
none/auto) and the two-name --checksum-choice TRANSFER,PRE-TRANSFER form,
including rsync's 'none' rules (rejected with --checksum at exit 4, and
forcing --whole-file as the transfer half) and unknown names at exit 4.
The checksum default becomes the auto-negotiated xxh128.

Negotiation is deterministic and symmetric: both peers run the same
preference resolver (rsync's --version order).  The resolved
compression_algo crosses the wire as a new trailing config-frame int so
the receiver validates and installs the exact codec; an unsupported
choice is refused before STATUS_OK like rsync's failed negotiation.
Bump PROTOCOL_VERSION to 2.26.0.
2026-09-16 23:27:40 +02:00
TapTap 0e33f84f38 feat(codec): implement md4/sha1/none digests and lz4/zlib/zlibx codecs
Add real implementations for the rsync 3.4.1 checksum and compression
breadth: a self-contained MD4 (RFC 1320), OpenSSL-backed SHA1, a
no-digest mode, and LZ4/zlib codecs alongside zstd.  Compressed buffers
are now self-describing (a leading codec id), so every existing
decompression call site keeps working through a process-global codec
selection.  zlibx shares the zlib codec because FastSync compresses only
delta/token bytes (never matched file data), matching the 'x' intent.
2026-09-16 23:27:34 +02:00
TapTap c7ac039523 docs(parity): correct implied-dir walk comment 2026-09-16 23:26:22 +02:00
TapTap e771cc9da6 fix(delete): guard per-dir missing-args by server policy; fall back for --dirs
- Only honor the --delete-missing-args exact paths when the server's
  --allow-delete policy left delete_missing_args set.
- -d/--dirs does not recurse, so a per-directory plan would carry no child
  information and could delete the contents of an untraversed directory; fall
  back to the whole-tree end-of-transfer commit for that mode.
2026-09-16 23:26:14 +02:00
opencode 7f9f82a068 style: clang-format and cppcheck fixes; document filter grammar in usage 2026-09-16 23:24:36 +02:00
TapTap 695b5c8c25 fix(parity): protect -R prefix-relative excluded and size-skipped mirrors
A -R source prune (--exclude/--max-size) must record the destination wire
path below the reconstructed prefix so --delete protects it; the parallel
root scan and the sequential skip path used the source path instead.
2026-09-16 23:21:58 +02:00
TapTap 5b2188d909 fix(parity): scope -R --delete to the transferred prefix subtree
A general -R transfer places its files below the reconstructed prefix, so
marking the whole receive root as the delete scope deleted unrelated
sibling directories (data loss; rsync keeps them).  Use the prefix itself
as the root marker when it is non-empty, in both the single-threaded and
multithreaded pipelines.
2026-09-16 23:20:46 +02:00
TapTap e5da916d54 test(parity): cover -d one-level listing, empty --files-from and negation rejection 2026-09-16 23:19:18 +02:00
TapTap de640bba1b feat(parity): report --stats during server-contacting dry-run
rsync prints the --stats block (with the (DRY RUN) suffix) for -n; route
the dry-run path through report_transfer_stats using the wire counters and
the STATUS_STATS receiver report.
2026-09-16 23:15:53 +02:00
TapTap f0f5719be0 docs(protocol): correct STATUS_STATS field description 2026-09-16 23:14:45 +02:00
TapTap 6200b298ac test(parity): ignore the untransferred source-root line in %C diff 2026-09-16 23:13:22 +02:00
opencode 394a9aae22 feat(parity): rsync filter grammar (merge/dir-merge/hide/show/protect/risk/clear + modifiers) and -F click semantics 2026-09-16 23:10:17 +02:00
TapTap 12d4af1b89 docs(usage): list %c/%C in --out-format help 2026-09-16 23:07:40 +02:00
TapTap 9d7c55d3c0 fix(parity): read STATUS_STATS before --remove-source-files acks
The receiver emits the wire-stats frame before the per-file acks and the
terminal status; the client must consume it in that order or a combined
--stats --remove-source-files run desynchronizes.
2026-09-16 23:06:54 +02:00
TapTap 5a104bfd88 fix(receiver): initialize new sink fields in -m pipeline 2026-09-16 23:05:46 +02:00
TapTap 2ada8f9ad5 style: clang-format wire-stats changes 2026-09-16 23:02:53 +02:00
TapTap 1493f1806d feat(parity): rsync-style per-file --progress and differential tests
Replace the aggregate stderr progress with rsync 3.4.1's per-file progress
block (name, 32 KiB first frame, final frame with (xfr#N, to-chk=X/Y)).
Add differential tests against real rsync for --out-format %C/%b, the
--progress frames, selected --stats lines and -n --delete lines.
2026-09-16 23:00:39 +02:00
TapTap ea28e25535 feat(parity): receiver STATUS_STATS report and -n --delete lines
Add the STATUS_STATS end-of-transfer receiver report (matched/deleted
counters plus a would-delete path list) behind the report_stats wire
bool, and a read-only delete_extras_list walker.  --stats now renders
true wire byte totals and the receiver-reported deleted count; a
server-contacting -n --delete prints transfer-relative '*deleting' lines
matching rsync's itemize layout.
2026-09-16 22:55:26 +02:00
TapTap d6295d62ce test(delete): differential + timing regression tests for per-directory delete plans
- Compare --delete-during/--delete-delay final state against rsync 3.4.1.
- Force a mid-transfer failure through a byte-slicing proxy: --delete-during has
  removed the processed directory's extra, --delete-delay has not.
- Create a destination entry while the transfer is in flight: it survives
  --delete-delay's snapshot but is removed by --delete-after's fresh end scan.
- Cover the --delete-delay type-conflict case now matching rsync.
2026-09-16 22:50:51 +02:00
opencode a9f416ce44 feat(parity): rsync fuzzy distance/suffix heuristic + exact size+mtime pass 2026-09-16 22:45:57 +02:00
TapTap 448edc0432 feat(delete): per-directory delete plans for --delete-during/--delete-delay (protocol 2.24.0)
Stream one delete plan per source directory from sender to receiver instead of
a single whole-tree keep-set manifest:

- --delete-during applies each directory's extras as its plan arrives, before
  that directory's data (rsync's generator-order deletion).
- --delete-delay snapshots each directory's extras while the plan arrives and
  commits the removals only after a fully-successful transfer, so files created
  after the scan survive (matching rsync's delete-delay, not delete-after).
- Type conflicts (a destination file blocking a source directory, or vice
  versa) are cleared immediately in both modes, so the nested write succeeds.

The plan carries the destination-relative directory, its kept child directory
names and its kept child file names; the first frame also carries the global
protected prefixes, size-skipped prefixes and --delete-missing-args paths.
--delete-before keeps the existing whole-tree early manifest; plain --delete and
--delete-after keep the end-of-transfer manifest commit.

Preserves the existing safety surface: protected/size-skipped prefixes and the
--delay-updates/basis skips are honored at any depth, deletion is scoped to the
synchronized directories (--files-from), MAX_SERVER_DELETE_COUNT and
--max-delete (partial + exit 25) are shared across plans, symlinks are never
followed, and paths are confined to the receive root.
2026-09-16 22:45:26 +02:00
TapTap 4a7703b06a feat(parity): -d/--dirs one-level listing for dir/, dir/. and .
A trailing slash (or trailing '/.', or a bare '.') now lists the source's
immediate contents -- files transferred, subdirectories created empty --
without recursing, while a bare directory still sends only its own entry.
The -R prefix applies to the generated entries and to the root entry.
2026-09-16 22:44:22 +02:00
TapTap 6fc297544e test(parity): differential coverage for -R, --no-implied-dirs and client aliases 2026-09-16 22:42:30 +02:00
TapTap 583d3c8edb feat(parity): general -R/--relative path semantics and --no-implied-dirs
Reconstruct the destination-relative prefix from the source spec outside
--files-from: cut at rsync's first '/./' (or a leading './'), normalize
later '.' components and trailing slashes.  Apply it as each File's
send_path in the sequential and parallel scanners (root and worker paths,
files, one-file-system mount entries and directory-time capture).

Transmit the metadata of implied parent directories (prefix components
above the source root), suppressed by --no-implied-dirs, so parent attrs
match rsync in both the single-threaded and -m pipelines.
2026-09-16 22:41:28 +02:00
TapTap 9883757190 fix(parity): accept rsync client aliases, --iconv=. / - and lone -h
- --ignore-non-existing (alias of --existing)
- --protect-args (pre-3.2.6 --secluded-args no-op)
- --msgs2stderr / --no-msgs2stderr (deprecated --stderr=all/client)
- --iconv=. (locale codeset via nl_langinfo), --iconv=- and --no-iconv (disable)
- lone -h prints help and exits 0; -h elsewhere stays human-readable
2026-09-16 22:41:23 +02:00
opencode a690109975 feat(parity): resolve --chown TO names on receiver via map rules 2026-09-16 22:41:12 +02:00
TapTap 36d4d0e43e feat(parity): wire-stats protocol 2.25.0 + out-format %b/%c/%C
Bump PROTOCOL_VERSION to 2.25.0 and append a report_stats bool to the
config frame, add a STATUS_STATS status, and add process-wide wire byte
counters (protocol_bytes_written/read) for the client.

Render the rsync 3.4.1 --out-format %b (wire bytes sent) and %c (wire
bytes read back) tokens from per-file counter deltas, and %C (whole-file
xxh128 checksum, seed 0) via a new streaming checksum_digest_file().
2026-09-16 22:35:43 +02:00
opencode a0b9d9794b feat(parity): absolute basis dirs + link-dest relink of up-to-date dest 2026-09-16 22:34:43 +02:00
opencode 3e9f70d9ba fix(parity): receiver-side --ignore-existing short-circuit before payload 2026-09-16 22:29:46 +02:00
opencode 2b5aaef409 fix(parity): --preallocate wins over --sparse, prefer fallocate(2) 2026-09-16 22:26:58 +02:00
TapTap 478f80be9f test(parity): add --stop-at rsync date-form differential coverage 2026-09-16 22:22:56 +02:00
TapTap e674b25213 fix(parity): client quick wins for rsync 3.4.1 (copy-links exit 23, info/debug flags, empty files-from, -F ordering, delete edges) 2026-09-16 22:21:45 +02:00
TapTap 1116da9f64 docs: correct rsync-parity claims and stale facts (#297)
CI / lint (pull_request) Successful in 1m47s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / build-and-test (pull_request) Successful in 40s
Reclassify every rsync-compatibility row as parity / caveat / divergent
(replacing the misleading 143-OK / 0-divergence summary), and document the
protocol 2.23.0 behavior:

- Split the conflated `-M, --preserve` row: `-M` is `--remote-option`,
  `--preserve` is the FastSync `-p`+`-t` alias.
- Fix `MAX_CONNECTION_MEMORY` (256 MiB, not 1 GB), `--rsync-path`
  (client-only, never crosses the wire), and the `-p` mode behavior
  (strict rsync parity; no masking).
- `--specials` now recreates sockets, so `-D` is real parity; fake-super
  records the resolved owner and replays mode/time (never real-chowns).
- Document short options/clustering, checksum/compression choices, seed
  randomization, timeout/max-alloc defaults, temp-dir confinement + EXDEV,
  identity/map parity, verbatim symlinks, delete scoping, `--max-delete`
  partial + exit 25, `--chmod`, output caveats, and server `--port`.
- Bump version refs to 2.23.0 and add the 2.23.0 CHANGELOG entry.

Docs-only; no source changes.
2026-09-16 01:48:13 +02:00
TapTap 684153350a Merge branch 'fix/parity-chmod' into feat/rsync-parity 2026-09-16 01:24:05 +02:00
TapTap ec206b02d0 chmod: match rsync 3.4.1 --chmod and remove mode masking (#293)
- --chmod no longer implies --preserve-perms; repeated --chmod options
  accumulate, and D/F/X selectors plus s/t special bits are supported with
  rsync's exact parse_chmod/tweak_mode semantics.
- Stop masking group/other write and setuid/setgid/sticky: -p copies the
  source mode exactly, no-p new entries use source&~umask, directories keep
  setgid/sticky, and special nodes follow the same rules.
- Apply ownership before mode on the fd path so a chown cannot clear the
  setuid/setgid bits -p just restored (rsync order).
- Update unit and integration tests, including differential checks against
  rsync 3.4.1.
2026-09-16 01:23:45 +02:00
TapTap 88bdfeeb58 fix(parity): receiver temp-dir confinement, server I/O floor, delete budget
Address review findings on feat/rsync-parity:
- confine --temp-dir below the receive root (reject absolute/.. like
  backup-dir/partial-dir); keep EXDEV non-atomic fallback
- floor server session I/O deadlines at SERVER_IO_TIMEOUT_SEC (60s) and
  install it on the socket layer at startup (slow-loris)
- charge each --delete-missing-args directory removal once and clamp the
  extras-walk remaining budget so it can never underflow past --max-delete
- normalize --compress-choice=auto to zstd client-side and accept it on
  receive so auto transfers no longer fail
- map received --max-alloc=0 to MAX_SERVER_ALLOC (receive path only)
- zero File.dest_state; include log-file-format in report_dest_info;
  add STATUS_DELETE_LIMIT name; recognize --skip-compress as a
  separate-value option; OOM-guard send_list_only root entry; drop the
  dead -M= branch; record the bare relative protected prefix for -R
  size-prunes in both scanners; refresh delete-manifest comment
- pin the rsync tarball sha256 and bump integrator image to v11

Tests: temp-dir rejection/relative/cross-device, server timeout floor,
delete-missing dir budget regression, compress-choice=auto e2e,
max-alloc=0 receive mapping, dest_state, report_dest_info modes,
skip-compress dash value, -M short forms, -R root size-prune mirror
protection (rsync 3.4.1 confirmed).
2026-09-16 01:11:59 +02:00
TapTap 3f5b0250f4 fix: ASan out-of-bounds argv in cli test, cppcheck uninit rate buffer 2026-09-15 23:53:49 +02:00
TapTap c41bfb2cdb test: align trust-sender tests with rsync-parity symlink storage 2026-09-15 23:44:14 +02:00
TapTap 1b2632f968 test: fix merged parity branches (4-section manifest fixtures, timeout-teardown) 2026-09-15 23:38:15 +02:00
TapTap 7dbca70a4b Merge branch 'feat/parity-output' into feat/rsync-parity
# Conflicts:
#	src/shared/config.h
#	src/shared/protocol.h
#	tests/test_config.c
2026-09-15 23:25:34 +02:00
TapTap 1a053f06e5 Merge branch 'feat/parity-ownership' into feat/rsync-parity
# Conflicts:
#	src/shared/config.h
#	tests/test_config.c
2026-09-15 23:24:58 +02:00
TapTap 58b3a33e82 Merge branch 'feat/parity-delete' into feat/rsync-parity 2026-09-15 23:24:24 +02:00
TapTap 17b0632098 Merge branch 'feat/parity-network' into feat/rsync-parity 2026-09-15 23:24:21 +02:00
TapTap 376e6500ab fix(delete): count recursive missing-arg removals per entry (#290)
A non-empty --delete-missing-args directory removed under --force/--delete
now has its contents deleted entry-by-entry through the budgeted walker, so
every deleted file/dir counts toward --max-delete exactly like rsync (a
capped run leaves the remaining entries and exits 25).
2026-09-15 23:23:52 +02:00
TapTap 82a1d5e240 fix(delete): match rsync deletion semantics (#290)
- Scope the --delete extras walk to directories synchronized by the
  transfer: add a synchronized-directory section to the delete manifest
  (protocol 2.23.0) so --files-from subsets no longer delete untransmitted
  paths outside listed directory subtrees (data-loss fix).
- Separate --max-size/--min-size prune protection from --delete-excluded so
  size-pruned source mirrors survive (rsync parity).
- Unlink extraneous destination symlinks instead of skipping them.
- Make --max-delete partial (delete up to N, skip the rest) and exit 25;
  accept negative values as unlimited.
- Draw --delete-missing-args deletions from the shared --max-delete budget.
- Honor --force during --delay-updates publication.

Add unit and integration regression tests; update the pinned config wire
golden and version strings for the 2.23.0 manifest/status additions.
2026-09-15 23:12:03 +02:00
TapTap 3eec5a4cc3 feat(parity): rsync 3.4.1 checksum/timeout/temp-dir/connectivity parity (#289 #295 #296)
#289 checksum/compression:
- -c/--checksum now implies the incremental content quick-check (without
  implying -t), so an unchanged file is skipped like rsync.
- --checksum-choice/--cc accepts xxh64/xxhash, xxh3, xxh128, md5 and auto;
  md4/sha1/none and the two-name form are rejected by name.
- --compress-choice/--zc rejects lz4/zlib/zlibx by name (zstd/none/auto kept).
- --checksum-seed=0 is randomized per transfer and sent on the wire.
- --skip-compress uses rsync 3.4.1's default suffix list; slash separators and
  dot-less suffixes are accepted.
- add --no-whole-file.

#295 timeouts/alloc/temp-dir:
- --timeout default 0 (disabled), --contimeout default 60; 0 disables both,
  plus --no-timeout/--no-contimeout.
- --max-alloc=0 means no allocation limit (was rejected).
- --temp-dir accepts any dir, requires it to exist, and falls back to a
  non-atomic copy on EXDEV instead of aborting.

#296 connectivity/daemon:
- -M/--remote-option is rejected for daemon/TCP destinations (SSH-only).
- --trust-sender clarified as receiver-local; server-path tests added.
- --stop-at accepts rsync's full date form (y-m-dTh:m etc.).

Adds unit and integration coverage; no wire-field change, PROTOCOL_VERSION stays
2.22.0.
2026-09-15 22:20:02 +02:00
TapTap 6144c7fc7f fix(identity): rsync ownership parity for numeric-ids, dirs, maps, fake-super (#286, #294)
- #286: --numeric-ids is a mapping modifier only; it no longer activates
  chown by itself (identity_active_enabled/owner/group predicates), and
  --fake-super stores the resolved mapping instead of real-chowning.
- #286: apply owner/group to directories via the deferred directory
  metadata path; capture+transmit+apply directory xattrs/ACLs (-aX/-aA),
  including default ACLs, in STATUS_MKDIR/STATUS_DIR_TIMES.
- #294: --usermap/--groupmap support inclusive ranges, '*', empty FROM
  (unnamed ids), and receiver-side TO name resolution; --chown mixing with
  a same-side map is rejected like rsync.
- Protocol 2.22.0 -> 2.23.0 (map wire entry gains from_hi + to_name;
  dir frames gain a bounded xattr block).
2026-09-15 22:10:02 +02:00
TapTap ea4ab661b4 fix(parity): rsync 3.4.1 symlink and special-node semantics (#287, #288)
#287:
- --safe-links: keep safe in-tree links AS symlinks and skip unsafe
  (absolute or ".."-escaping) ones, mirroring rsync's unsafe_symlink().
  Skipped links are recorded as delete-protected so --delete does not
  remove their destination mirror (no silent data loss).
- --copy-unsafe-links: preserve safe links as symlinks and dereference
  only unsafe ones.
- --munge-links: receiver-side rewrite storing /rsyncd-munged/-prefixed
  targets (rsync parity), replacing the no-op #SYMLINK sender prefix.
- -l: store the target verbatim, including absolute and ".." targets
  (rsync -l parity); the old receiver containment silently dropped them.

#288:
- --specials: recreate unix-domain sockets via mknod(S_IFSOCK), which
  Linux permits unprivileged; keep EEXIST/EPERM skip behavior.
- --copy-devices: copy a device's content into a regular file when
  requested; skip unrequested non-regular entries like rsync's default.
2026-09-15 21:57:48 +02:00
TapTap 84827ca617 feat(output): rsync 3.4.1 selection and output parity (#291, #292)
#291:
- Compile --exclude/--include/--exclude-from/--include-from into the SAME
  ordered rule list as --filter/-f (first match wins), so the common
  `--include='*.txt' --exclude='*'` idiom and include-alone semantics match
  rsync. The legacy per-kind scanner arrays are no longer applied.
- -x/--one-file-system emits the cross-device mount-point directory entry
  (empty) instead of dropping it, in both the sequential and parallel scanners.
- Stop passing the legacy arrays to the scanner; document -f is --filter.

#292:
- New src/shared/format.c/.h: rsync "big_num" (comma-grouped integers) and
  decimal -h human sizes, %M/%t timestamp, and the STATUS_DEST_INFO codec.
- Receiver answers each STATUS_CHECK with a pre-transfer destination snapshot
  (new report_dest_info wire field + STATUS_DEST_INFO, PROTOCOL_VERSION
  2.23.0) so the sender can render true itemize columns.
- Itemize now emits rsync-correct update/type chars and c/s/t/p/o/g columns
  for files, dirs, symlinks and hard links, comparing size/time/perms/owner/
  group against the reported destination.
- --out-format gains %i %n %f %l %b %M %t %o %p %B %U %G %L; %f is the
  relative display path, %M the YYYY/MM/DD-HH:MM:SS form, %b the literal
  bytes sent.
- --list-only prints transfer-relative names, directory entries and ls-style
  grouped sizes.
- --stats prints rsync's multi-line block on stdout; -h uses decimal units.

Tests: unit tests for the filter ordering, format primitives, itemize
columns; integration + differential tests against real rsync 3.4.1 for
itemize/out-format/list-only/selection and -x. Golden wire len/hash and
protocol version strings updated for 2.23.0.
2026-09-15 21:57:39 +02:00
TapTap 23552e823d feat(cli): rsync short-option clustering and inline/attached values (#285)
Implement rsync 3.4.1 client-CLI parity:
- cluster boolean shorts (-av, -aAX, -rlpt) and accept attached values
  (-B1048576, -essh, -Mfoo); add the -r, -b, -L and -B short aliases
  (-r is a faithful no-op since FastSync is always recursive)
- stop OPT_NOOP (-s/--secluded-args, -r/--recursive) from swallowing the
  next argv
- add inline --opt=value for every value-taking long option, including
  --exclude/--include/--exclude-from/--include-from/--log-file (#291)
- accept --port on the server CLI in addition to -p (#296)
- reject unknown flags naming the flag and stating it is unsupported

Unit tests cover clustering, attached/inline values, the OPT_NOOP
argument-consumption fix and rejected shorts.
2026-09-15 21:12:53 +02:00
TapTap d1a567f7e3 ci: pin fastsync-ci:v11 with rsync 3.4.1 + acl/attr for parity tests
Add POSIX ACL/xattr tooling (acl, attr), zstd/lz4/xxhash dev libs and build rsync 3.4.1 from source so drop-in parity tests can run inside CI. Bump all workflow/agent image references v10 -> v11.
2026-09-15 20:37:50 +02:00
TapTap 93c1fc3c1f test: create fault-injection destination root in seeding fixture
CI / lint (push) Successful in 1m29s
CI / lint (pull_request) Successful in 1m29s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / sanitizers (undefined) (push) Successful in 1m4s
CI / sanitizers (address) (push) Successful in 1m10s
CI / fuzz-build (push) Successful in 39s
CI / coverage (push) Successful in 58s
CI / build-and-test (pull_request) Successful in 1m55s
CI / valgrind (push) Successful in 3m23s
CI / build-and-test (push) Successful in 5m11s
The captured_config fixture assumed fault_dst already existed, relying on earlier tests in the same xdist worker creating it via _recover. Under --dist=load a worker can receive the capture test first, so the receiver rejected a missing destination root and the capture run failed. Create DEST_DIR in the autouse seeding fixture so test order/distribution cannot matter.
2026-09-15 20:02:00 +02:00
TapTap 4815b1b281 test: account for group/other-write sanitization in new-dest mode expectation
CI / lint (push) Successful in 1m28s
CI / lint (pull_request) Successful in 1m28s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / sanitizers (undefined) (push) Successful in 1m1s
CI / sanitizers (address) (push) Successful in 1m6s
CI / fuzz-build (push) Successful in 38s
CI / coverage (push) Successful in 1m3s
CI / build-and-test (pull_request) Successful in 1m56s
CI / build-and-test (push) Failing after 4m50s
CI / valgrind (push) Successful in 3m23s
2026-09-15 19:41:12 +02:00
TapTap ad7bc3348b Merge branch 'feat/preserve-attr-split' into dev
CI / lint (push) Successful in 1m29s
CI / lint (pull_request) Successful in 1m28s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / sanitizers (undefined) (push) Successful in 1m4s
CI / sanitizers (address) (push) Successful in 1m10s
CI / fuzz-build (push) Successful in 38s
CI / coverage (push) Successful in 57s
CI / build-and-test (pull_request) Failing after 1m55s
CI / build-and-test (push) Failing after 4m55s
CI / valgrind (push) Successful in 3m23s
2026-09-15 19:32:12 +02:00
TapTap 34970b961c feat: per-attribute preservation flags -p/-t/-o/-g with --no-* negations (protocol 2.22.0)
Split FastSync's single use_metadata bundle into four independent rsync-parity attributes: preserve_perms, preserve_times, preserve_owner, preserve_group. use_metadata is now a derived transport bit (config_derived_use_metadata).

CLI: real -p/--perms, -t/--times, -o/--owner, -g/--group plus --no-perms/--no-times/--no-owner/--no-group (short and long) and --no-preserve; -a is now rsync -rlptgoD; --preserve = -pt; -A implies -p; -X does not; --chmod implies -p; --usermap/--groupmap/--chown imply owner/group per side; --incremental/--delta still auto-preserve unless negated.

Receiver: per-attribute FileAttrPolicy gating for files, dirs (modes applied at end of transfer), symlinks and specials; rsync -E read-bit rule; new files get source_mode & ~umask sanitized (no group/other write); per-side identity resolution; deferred directory metadata; batch dir-metadata replay; daemon modules without 'client owner = yes' no longer refuse plain -a but force super off (no ownership) with a warning.

Wire: PROTOCOL_VERSION 2.21.0 -> 2.22.0 (four appended config bools, golden 653 / 95530566005420798). FileMetadata/chunk/batch framing unchanged. Docs/CHANGELOG/CMake updated to 2.22.0.
2026-09-15 19:32:02 +02:00
TapTap b3f7cad4db Merge docs/handoff: session handoff document
CI / lint (push) Successful in 1m29s
CI / lint (pull_request) Successful in 1m29s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / sanitizers (undefined) (push) Successful in 59s
CI / sanitizers (address) (push) Successful in 1m6s
CI / fuzz-build (push) Successful in 38s
CI / coverage (push) Successful in 56s
CI / build-and-test (pull_request) Successful in 1m55s
CI / valgrind (push) Successful in 3m24s
CI / build-and-test (push) Successful in 5m35s
2026-09-14 19:36:28 +02:00
TapTap cd8a84c0a2 docs: add session handoff (status, next steps, deferred security items) 2026-09-14 19:36:28 +02:00
TapTap 09c384d7d0 Merge branch 'docs/readme-refresh' into dev
CI / lint (push) Successful in 1m25s
CI / lint (pull_request) Successful in 1m24s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / sanitizers (undefined) (push) Successful in 1m1s
CI / sanitizers (address) (push) Successful in 1m7s
CI / fuzz-build (push) Successful in 36s
CI / coverage (push) Successful in 56s
CI / build-and-test (pull_request) Successful in 1m54s
CI / valgrind (push) Successful in 3m18s
CI / build-and-test (push) Successful in 5m34s
# Conflicts:
#	README.md
2026-09-14 18:52:54 +02:00
TapTap 81ad313ee5 docs: refresh README against implementation and guard against drift
Bring README.md and RSYNC_COMPAT.md in line with the actual code/CLI and add
an automated guard so they cannot silently drift again.

Waves A-E:
- Correct stale compatibility claims: archive is `-rlptD` (owner/group are
  opt-in via identity flags, not implied), and symlinks, hard links, xattrs,
  ACLs and `--dirs` are implemented.
- Remove documented-but-nonexistent features: the six unread FASTSYNC_* env
  vars, and `--client-cn` (server-only) from the client table.
- Repair the corrupted "Implementation Details" section (broken list numbering
  and emphasis) and correct it against the source.
- Sync the client and server option tables with usage.c / server_cli.c, and
  document server-contacting `--dry-run` (protocol 2.21.0).
- Hygiene: `# FastSync` heading, real build commands, consistent binary names,
  runnable TLS examples, daemon module keys.

Also align the client `--help` / archive log wording and the RSYNC_COMPAT
archive rows with the opt-in ownership model, and add
tests/integration/test_readme_consistency.py (marked `ci`) asserting every
documented FASTSYNC_* var is read in src/ and every documented client/server
flag appears in the corresponding `--help`.
2026-09-14 18:48:42 +02:00
TapTap 919a729206 Release v2.21.0
CI / lint (push) Successful in 1m25s
CI / lint (pull_request) Successful in 1m25s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / sanitizers (address) (push) Successful in 1m6s
CI / sanitizers (undefined) (push) Successful in 1m0s
CI / fuzz-build (push) Successful in 34s
CI / coverage (push) Successful in 55s
CI / build-and-test (pull_request) Successful in 1m49s
CI / valgrind (push) Successful in 3m19s
CI / build-and-test (push) Successful in 5m23s
- Protocol 2.21.0: STATUS_ERROR_DETAIL rejection reasons and server-contacting --dry-run
- Daemon per-module/per-host caps and cross-process auth lockout
- Config X-macro serialization, authorized_root single-owner, Data charge ownership, receiver pipeline move
- Security audit hardening (SSH injection, FIFO/inplace, zstd DoS, TLS, dry-run oracle, bounds)
- Pre-auth basis_count NULL-deref fix; benchmark and nix-shell improvements
- Tested: unit, integration, ASan/UBSan, valgrind, fuzz, coverage (CI green)
2026-09-14 18:16:42 +02:00
TapTap 8cd2b550d9 Merge dev environment fix and push-only documentation
CI / lint (push) Successful in 1m25s
CI / sanitizers (undefined) (push) Successful in 1m2s
CI / sanitizers (address) (push) Successful in 1m9s
CI / fuzz-build (push) Successful in 36s
CI / coverage (push) Successful in 56s
CI / valgrind (push) Successful in 3m18s
CI / build-and-test (push) Successful in 5m25s
2026-09-14 18:08:51 +02:00
TapTap 99c0fd8016 Merge benchmark improvements: accurate data mix, transfer verification, warm mode 2026-09-14 18:08:51 +02:00
TapTap a2200f039a chore(dev): fix nix-shell environment; document push-only direction 2026-09-14 18:08:46 +02:00
TapTap 1437c6dc6b bench: fix data mix, verify transfers, robust netem, warm mode
- generate_bench_data now writes exactly (1-random_ratio)*target bytes of
  genuinely compressible repeated content instead of only the small fixed
  STRUCTURED_FILES set; measured composition is reported and --dry-run prints
  it for scaling checks
- verify each transfer against the source (paths/sizes/byte compare) before
  recording timing; add --no-verify; failed runs are counted as invalid
- correct p50/p95 with linear-interpolation percentile (was int(len*0.95))
- tc/netem: run tc directly as root, else sudo; clear error when tc/iproute2
  is missing or qdisc setup fails; netem_reset is always safe
- build into dedicated build-bench/ via --build-dir (Release), never reconfigure
  the user's build/
- parse --configs with shlex.split
- add MB/s throughput column and throughput_mbps JSON field
- add --warm incremental mode: untimed full seed then measure add/change deltas
2026-09-14 18:07:37 +02:00
TapTap 38356ecc1e test: fix valgrind definite leak in forked compression truncation test
CI / lint (push) Successful in 1m24s
CI / sanitizers (undefined) (push) Successful in 1m1s
CI / sanitizers (address) (push) Successful in 1m6s
CI / fuzz-build (push) Successful in 35s
CI / coverage (push) Successful in 55s
CI / valgrind (push) Successful in 3m18s
CI / build-and-test (push) Successful in 5m25s
2026-09-14 17:54:20 +02:00
TapTap 7badac7f97 test(compression): free inherited Data in forked truncation test (valgrind) 2026-09-14 17:54:20 +02:00
TapTap 6ccf16b650 Merge security hardening wave: parser/compression, receiver confinement, server/transport/TLS
CI / lint (push) Successful in 1m26s
CI / sanitizers (undefined) (push) Successful in 1m2s
CI / sanitizers (address) (push) Successful in 1m7s
CI / fuzz-build (push) Successful in 35s
CI / coverage (push) Successful in 56s
CI / valgrind (push) Failing after 3m19s
CI / build-and-test (push) Successful in 5m26s
2026-09-14 17:40:29 +02:00
TapTap 1174993d6b Merge branch 'fix/sec-server' into fix/sec-integration 2026-09-14 17:38:23 +02:00
TapTap d5fcfa2c5c style(ssh): drop redundant condition flagged by cppcheck 2026-09-14 17:38:23 +02:00
TapTap e79d2b47b0 Merge branch 'fix/sec-server' into fix/sec-integration 2026-09-14 17:27:38 +02:00
TapTap 7c24a365cf Merge branch 'fix/sec-receiver' into fix/sec-integration 2026-09-14 17:27:38 +02:00
TapTap 5893de4a34 fix(receiver): close re-review findings — dry-run basis oracle, ACL capture, fsync reopen
Follow-up to a237043 addressing three security/correctness re-review findings.

(1) MEDIUM: a server-contacting --dry-run with --compare-dest/--copy-dest/
    --link-dest still read and hashed the basis file and compared it with the
    client-supplied digest, a 1-bit content oracle. basis_match_find() gains a
    hash_content parameter; the dry-run shortcut passes false and returns no
    match without touching basis bytes, so an otherwise-matching entry is
    reported as would-transfer. The real (non-dry-run) path is unchanged.

(2) LOW: xattr_capture_path() hardcoded preserve_acls=true, so the receiver's
    hard-link copy fallback re-applied system.posix_acl_* even when -A was not
    negotiated. The function now takes preserve_acls and members.* is
    unaffected; scanner and receiver callers thread the negotiated flag.

(3) INFO: the --fsync --link-dest temp reopen now uses O_NONBLOCK and treats
    a raced-in FIFO's ENXIO as a benign fsync-skip instead of blocking.

Tests: dry-run + basis unit test (asserts would-transfer, no content read) and
integration test; xattr capture ACL-filter test. Verified strict build, ASan,
clang-format, cppcheck, and the CI integration subset.
2026-09-14 17:19:27 +02:00
TapTap 825ba69753 fix(server): reject --allow-super with --stdio, fix module host-list append
Re-review findings on the C3/C4 hardening branch:

- --stdio is the SSH transport whose remote argv is composed by the client
  (including via --remote-option), so accepting --allow-super there let a
  client defeat the C3 secure default for a root receiver.  Reject it at CLI
  parse time (standalone TCP only) and force the process-global flag off for
  --stdio as defense in depth.  Correct the help text and README/RSYNC_COMPAT:
  the --stdio argv is client-composed, super stays off, and a forced command is
  needed if the default must hold.
- daemon_conf: the per-module 'hosts allow'/'hosts deny' call sites passed
  module_name and replace in the wrong order, so multiple lines replaced
  instead of appended and the empty-value error omitted the module name.  Pass
  (module->name, false) like the global keys; add a unit test for two
  per-module allow/deny lines appending.
- tls: read the client CN via ASN1_STRING_to_UTF8 so an exactly-required-length
  name is accepted and only actual over-length CNs are rejected.
2026-09-14 17:16:01 +02:00
TapTap 34abaadb9a fix: address low/informational sec-parser follow-ups
- client_cli: capture errno before output_escape() in
  read_patterns_from_file() so an over-long line is still reported as
  EFBIG instead of the (possibly malloc-clobbered) errno.
- file_list: guard string_list_add() capacity doubling against
  overflow (capacity > INT_MAX / 2), matching filter_rule_list_add();
  callers already surface the false as a memory-allocation error.
- compression: ZSTD_isError() is true for ZSTD_CONTENTSIZE_UNKNOWN,
  which made the 3x unknown-size fallback dead code.  Test the
  CONTENTSIZE_ERROR/UNKNOWN sentinels explicitly so unknown-size frames
  reach the estimate path (still bounded by the existing hard limit)
  while invalid frames are rejected.  Known-size frames and the 100 MB
  ceiling/overflow checks are unchanged.
- tests: add an unknown-content-size-frame decompression test.

Tests: ./build/tests and ./build-asan/tests all pass (42/42);
clang-format + cppcheck clean.
2026-09-14 17:11:02 +02:00
TapTap 10c4ffebdf fix(client): harden CLI args, log escaping, and local artifact opens
- parse_ull_arg() rejects a leading '-'/'+' (strtoull would silently wrap
  -1 to ULLONG_MAX) and --chunk-size/--delta-max enforce their upper bounds.
- Escape local untrusted paths before logging (client_send, scanner,
  --filter rule, pattern-file reads) with output_escape(..., 8-bit mode).
- Read --exclude-from/--include-from through the bounded line reader.
- Open --log-file with O_NOFOLLOW|O_CLOEXEC, mode 0600, via open+fdopen;
  create --write-batch with O_NOFOLLOW|O_CLOEXEC, mode 0600.
- Reject --dry-run together with --write-batch (dry-run must not write the
  batch file), alongside the existing --read-batch/--only-write-batch rules.

Tests: signed/oversized numeric rejection, over-long pattern file, dry-run +
write-batch unit and integration coverage.
2026-09-14 16:39:19 +02:00
TapTap dfa2a42028 fix(protocol): retry EINTR on receive and clamp SSL_write length
protocol_receive_n_data_until() aborted on a signal-interrupted plaintext
read (and on SSL_ERROR_SYSCALL with errno==EINTR); retry both, matching the
send path and protocol_read_status_until().  Also clamp each SSL_write() to
INT_MAX so a >INT_MAX size_t request can never truncate into a partial write.
2026-09-14 16:39:19 +02:00
TapTap 5b0ec5880d fix(filter): bound .rsync-filter lines and guard capacity growth
Read per-directory filter files through the bounded reader, guard the rule
list's capacity doubling against INT_MAX/2 overflow, and escape the local
directory path before logging a read failure.
2026-09-14 16:39:14 +02:00
TapTap 1a26bde2d4 fix(file_list): bound entry length and reject embedded NUL bytes
Read list files through utils_getdelim_bounded() so a single multi-gigabyte
line can no longer force unbounded allocation; over-long entries fail with a
clear error.  Also add the documented memchr() NUL-byte check (excluding the
NUL delimiter in NUL-separated mode).

Tests: an over-long entry is rejected with an 'exceeds' diagnostic.
2026-09-14 16:39:14 +02:00
TapTap 58a28334b8 fix(utils): bound glob matching and line reads
Replace the recursive glob matcher with an iterative O(pattern*string)
dynamic program.  The old recursion explored exponentially many paths for
overlapping '*'/'**' wildcards (e.g. '*a*a*...*b' against a long run of
'a'), a CPU DoS reachable from --exclude/--include patterns and
.rsync-filter.  A differential fuzz against the original matcher confirms
identical results.  Doc: has_path_traversal() is a lexical '..' check only.

Add utils_getdelim_bounded(): a getdelim-style reader that never allocates
beyond UTILS_MAX_LINE_LEN, used to cap untrusted list/filter line reads.

Tests: pathological glob completes quickly; bounded reader returns EFBIG on
an over-long record.
2026-09-14 16:39:10 +02:00
TapTap e48f19ee2b fix(compression): fail truncated zstd frames instead of spinning
data_decompress_limited() looped while ZSTD_decompressStream() returned a
positive hint.  A truncated frame keeps returning that hint with all input
consumed, so a malformed/truncated payload spun forever (CPU DoS).  Detect
input exhaustion with an incomplete frame and fail via the existing cleanup,
skipping the check when the output buffer merely needs to grow first.

Add a fork+alarm regression test that truncates a valid frame and asserts
decompression returns NULL promptly.
2026-09-14 16:39:05 +02:00
TapTap a2370433b2 fix(receiver): non-blocking receiver opens, inplace type gate, dry-run/B4/B5/B6
Address confirmed receiver security findings B1-B6:

B1 (HIGH): add O_NONBLOCK to the three receiver read-opens that opened an
existing destination/basis entry before the S_ISREG gate
(incremental_check_open_destination, basis_open_regular, hardlink_read_source)
so a client-planted FIFO can no longer block the receive thread forever while
the post-open type gate still rejects it.

B2 (HIGH/MED): --inplace now fstatat(AT_SYMLINK_NOFOLLOW)-probes the target and
refuses any existing non-regular entry, opens with O_NONBLOCK, and re-checks
S_ISREG on the opened fd.  This stops a FIFO from hanging the open and stops a
char/block device from being written directly (bypassing --write-devices).

B3 (MED): under --dry-run the incremental quick-skip no longer reads/hashes the
destination file for --checksum/--delta; it decides from metadata only and
reports would-transfer when the comparison is inconclusive, closing the
read-only-module content-hash oracle.

B4 (LOW): xattr_name_appliable() now gates the two system.posix_acl_* names on
preserve_acls (--acls), not the derived use_xattrs (--xattrs OR --acls).  The
receiver drops (never applies) ACL entries when -A was not negotiated while
keeping user.* working for -X.

B5 (INFO): receive_manifest_section() charges a per-entry overhead against
MAX_MANIFEST_BYTES and the aggregate entry count across all three sections is
capped at MAX_MANIFEST_ENTRIES.

B6 (MED): data_charge_session() reserves decompressed/chunk-copy bytes against
the owning ProtocolSession (MAX_CONNECTION_MEMORY) and records them on the Data
so data_destroy() releases them via the Data.owner path.  Applied to the
whole-file/append/delta decompression sites and chunk_deserialize() per-file
copies; a missing session owner degrades to the previous uncharged behavior.

Tests: FIFO destination/basis non-hang (with alarm), --inplace FIFO/device
refusal, dry-run no-read oracle test plus updated metadata-only dry-run tests,
ACL-without--acls drop, manifest total-entry cap, and chunk session charging.
2026-09-14 16:19:26 +02:00
TapTap 9da5a0a9ed fix(server): gate --force by --allow-delete and secure root super default
C2: --force is deletion authority (an incoming regular file may remove a
non-empty destination directory tree, and --delete-missing-args may
remove a non-empty directory mirror), but it was not masked by the
operator --allow-delete policy.  The handler now clears
config->force_delete unless --allow-delete was given, exactly like
--delete and --delete-missing-args.

C3: a standalone TCP / --stdio server running as root defaulted to
SUPER_MODE_AUTO, so an untrusted client --devices/--write-devices/
--super could make it create device nodes, write raw devices, or apply
client-chosen ownership.  A privileged standalone receiver now forces
SUPER_MODE_OFF unless the operator opts in with the new server-only
--allow-super flag.  Non-root receivers are unchanged, and the daemon
path keeps its per-module `client owner = yes` gate.  --allow-super is
rejected with --no-super or --daemon.

C6: tls_client_identity_allowed now rejects a CN whose reported length
reached the buffer bound, so a truncated over-long CN cannot be matched
by a required --client-cn prefix.

Tests: an integration regression proving --force cannot replace a
destination directory without --allow-delete; standalone-default tests
for --copy-as refusal and (root-only) skipped device creation; a CLI
unit test for the new flag.  The integration shared_server fixture opts
in with --allow-super so the existing root-only ownership/device/copy-as
tests continue to exercise the opted-in configuration.  README and
RSYNC_COMPAT document the flag and the force/delete gating.
2026-09-14 16:09:44 +02:00
TapTap 80c1ff321c fix(credentials): length-check before legacy-hex scan (C9)
secret_is_legacy_hex indexed s[0..63] without first checking the string
length, reading out of bounds for a shorter secret.  Require
strlen(s) == 64 before scanning, and add a unit test that short and
63-hex-digit secrets are rejected as ordinary malformed verifiers (never
misreported as legacy).
2026-09-14 16:09:25 +02:00
TapTap 551c187005 fix(tls): AEAD-only 1.2 suites, server preference, TOCTOU key load, IP SAN
C5: restrict the TLS 1.2 and below cipher list to ECDHE AEAD suites
(ECDHE+AESGCM:ECDHE+CHACHA20, minus NULL/eNULL/MD5/RC4/3DES) instead of
HIGH (which includes CBC), and set SSL_OP_CIPHER_SERVER_PREFERENCE so the
server's order decides the negotiated cipher.  Client and server share
create_ssl_ctx, so both are updated.

C7: load the private key through an O_RDONLY|O_NOFOLLOW|O_CLOEXEC fd,
fstat that fd and validate owner/mode (now also rejecting group/other
execute bits), then load from the fd via BIO_new_fd.  This removes the
stat-to-load TOCTOU race while keeping the exact-owner/0600 policy.

C8: verify an IP-literal client hostname against the certificate IP SAN
with X509_VERIFY_PARAM_set1_ip_asc instead of SSL_set1_host (a DNS
check), falling back to SSL_set1_host for real names.

Unit tests assert the server-preference option, the absence of CBC/RC4/
3DES suites, and that context creation still succeeds.
2026-09-14 16:09:21 +02:00
TapTap 0d6c1f784f fix(daemon-conf): reject empty hosts/auth allow-lists (C4)
A present hosts allow/hosts deny/auth users key with an empty or
separator-only value produced a zero-length list, silently meaning no
ACL / no auth and contradicting the strict-parse contract.

store_host_list and the auth users parser now track how many entries a
present key actually added and fail the load with a clear error when it
is zero, so a restrictive directive can never silently become open.
Unit tests cover empty, whitespace-only and comma-only values.
2026-09-14 16:09:16 +02:00
TapTap f75a69f96a fix(ssh): reject option-injection destinations (C1)
A remote destination's user@host token is passed to ssh in option
position, so a host beginning with '-' (e.g. -oProxyCommand=...) was
parsed by ssh as an option, allowing arbitrary command execution.

- config_parse_ssh_dest now validates the user@host prefix and returns
  -1 (with a clear logged error) for an empty host or a user/host that
  starts with '-'; config_parse_transport_dest propagates the failure.
- transport_ssh.c's parse_remote_dest applies the same validation as
  defense-in-depth, and ssh_build_client_argv inserts a '--'
  end-of-options marker before the destination token.
- Unit tests cover -oProxyCommand=... / -prefixed hosts / empty host
  rejection and the argv shape.
2026-09-14 16:08:56 +02:00
TapTap df887c73b1 Merge Wave 9: error-detail frame and server-contacting dry-run (protocol 2.21.0)
CI / lint (push) Successful in 1m21s
CI / sanitizers (undefined) (push) Successful in 1m1s
CI / sanitizers (address) (push) Successful in 1m7s
CI / fuzz-build (push) Successful in 36s
CI / coverage (push) Successful in 56s
CI / valgrind (push) Successful in 3m17s
CI / build-and-test (push) Successful in 5m16s
2026-09-13 13:39:21 +02:00
TapTap 5d39619a8a Merge branch 'feat/w9-dryrun' into fix/w9-integration 2026-09-13 13:27:32 +02:00
TapTap 6269ae54e5 test(dry-run): strengthen no-mutation coverage and refresh docs
Extend _snapshot_tree to record mode, inode, xattrs, directories and
special nodes, and add coverage proving a server-contacting --dry-run
leaves the destination structurally identical for --delay-updates,
--backup, symlinks, hardlinks, FIFOs, and daemon modules (including a
read-only module).  Add a regression test for the --read-batch --dry-run
refusal and for a missing/non-directory receive root failing a dry-run
exactly like a real run.

Fix stale version comments (2.20.0/633 -> 2.21.0/637) and RSYNC_COMPAT's
current --protocol value, and add a unit assertion that
--server-port/--port (and --server-host) set the dry-run routing bit.
2026-09-13 12:57:37 +02:00
TapTap 07f7555c1d fix(server): skip dry-run per-file outcome bookkeeping
receiver_save_file appended to context->outcomes for --remove-source-files
without the !dry_run guard the multithreaded pipeline has, so a hostile
dry-run client could grow outcomes unbounded (raw, uncharged realloc) and
force a per-frame ack.  Guard the append on !dry_run.
2026-09-13 12:57:33 +02:00
TapTap 5b8aca5799 fix(receive): enforce dry-run no-mutation centrally
--dry-run --read-batch=FILE still wrote to the destination because
batch_read_apply -> file_save_to_disk_full bypassed the per-caller
!dry_run guards.  Guard file_save_to_disk_full and manifest_delete_all
directly (return SKIPPED/no-op) so every save/delete path is mutation-free
in dry-run, and keep the per-caller guards.  Reject --dry-run combined with
--read-batch/--only-write-batch at CLI validation with a clear error (a
dry-run of a local batch apply is not meaningful).
2026-09-13 12:57:30 +02:00
TapTap a1eaa93357 fix(client): route explicit remote dry-run targets, fail closed on stray status
--dry-run --server-host=H (or TLS / source-bind --address) silently ran the
client-side manifest even though a real run contacts the server.  Add a
client-only, never-serialized server_host_set bit (alongside the existing
server_port_set) and extend dry_run_targets_server so every explicit remote
target contacts the receiver.

Also make incremental_check return the dry-run code (4) only when the
session actually requested dry-run; a stray STATUS_DRY_RUN_TRANSFER from a
hostile/buggy peer is now a logged protocol error (STATUS_ERROR) instead of
falling through to send file data and desync.  Both normal send_single_file
callers handle rc == 4 explicitly as an abort.
2026-09-13 12:57:26 +02:00
TapTap 99df0a8a6d fix(server): fail-closed dry-run destination-root precondition
A wire dry_run bit must not relax the destination-root precondition:
previously handlers skipped ensure_receive_root entirely in dry-run, so a
client could dry-run against a nonexistent/regular-file root a real session
rejects.  Split the existence check (receive_root_exists, never creates)
from the create path and apply the precondition unconditionally: dry-run
runs the existence/directory check only, reports the failure, and creates
nothing (no --mkpath).

Also allow a `read only = yes` daemon module for a dry-run session (a
server-contacting dry-run IS a read-only wire operation) while still
refusing it for real writes, and update the stale read-only comments.
2026-09-13 12:57:23 +02:00
TapTap 334fc5b3e8 fix(protocol): harden STATUS_ERROR_DETAIL receive path
Address review/security findings in the 2.21.0 error-detail feature:

- Keepalive drain no longer erases the terminal detail: capture/clear is
  skipped for STATUS_KEEPALIVE so the reason the peer just sent survives the
  owed keepalive replies.
- Replace the capture path with a dedicated protocol_receive_error_detail:
  the declared length is validated against MAX_ERROR_DETAIL_BYTES before any
  allocation, over-cap bodies are drained through a fixed scratch buffer (so
  the stream never desyncs), in-cap bodies read straight into the thread-local
  detail buffer, and session->max_alloc is never raised.  Lengths beyond
  MAX_STRING_SIZE are treated as a fatal framing error.
- The detail body now honors the caller's deadline (timed/keepalive paths) and
  polls the abort callback between drain chunks.
- Escape peer-controlled detail text with output_escape before logging it in
  client_send.c and config.c.
- Clear io_error_detail in io_set_fds so a new connection on the same thread
  cannot inherit a stale reason.
- Add unit tests for the keepalive-survival, over-cap drain, absurd-length
  fatal framing, and deadline-clamped body read cases.
2026-09-13 12:40:03 +02:00
TapTap 88aee6ce94 feat(protocol): add optional STATUS_ERROR_DETAIL rejection reason (2.21.0)
Today a server rejection sends a bare STATUS_ERROR and the reason only
reaches the server log, so the client cannot say why a transfer was
refused.  Add an optional, bounded server->client error-detail frame:

  - Status gains STATUS_ERROR_DETAIL appended LAST so existing wire
    values are unchanged.
  - send_error_detail(fd, msg) sends STATUS_ERROR_DETAIL followed by the
    existing length-prefixed string primitive, slicing over-long messages
    to MAX_ERROR_DETAIL_BYTES (4096).
  - receive_status() (and the timed/keepalive status readers) always
    consume the detail body and map the status back to STATUS_ERROR,
    capturing the text into a thread-local buffer exposed by
    protocol_last_error(); a bare STATUS_ERROR leaves it cleared.  Every
    existing call site keeps working and the stream cannot desync.
  - Upgrade the daemon module gate / config validation (config.c), the
    final transfer failure (server.c) and receiver-side path/node
    validation (file_receive.c) to send a concrete reason; surface it on
    the client in client_send.c/config.c.
  - Bump PROTOCOL_VERSION to 2.21.0 (CMake VERSION, CHANGELOG, docs) and
    update the pinned config wire golden hash / CLI-version tests.
  - Add tests/test_protocol_error.c covering mapping+capture, the
    over-long bound, bare-error clearing, and thread-locality.
2026-09-13 12:19:46 +02:00
TapTap f6e8b6ddc4 refactor(file_receive): split receive_incremental_check into helpers
The per-file STATUS_CHECK fast path was a single 534-line function that
was hard to review.  Extract it into small static helpers called in order
by a short linear orchestrator:

  - incremental_check_receive_request  (receive/validate request frame)
  - incremental_check_open_destination (secure open + stat)
  - incremental_check_quick_skip       (metadata/content skip decision)
  - incremental_check_try_basis        (compare/copy/link-dest)
  - incremental_check_try_append_resume(--append tail resume)
  - incremental_check_try_delta        (block delta)
  - incremental_check_try_fuzzy        (--fuzzy basis)
  - incremental_check_receive_full     (STATUS_NEXT + whole file)

Pure refactor: the ordered sequence of wire operations
(send_status/send_n_data/receive_n_data/receive_status/receive_wire_str)
is byte-for-byte identical to the original, and every resource cleanup
is preserved (a unified idempotent cleanup replaces the duplicated
per-path close/free blocks).  No functional changes.
2026-09-13 12:04:31 +02:00
TapTap 6f974eff19 feat(dry-run): server-contacting --dry-run (protocol 2.21.0)
--dry-run now handshakes with a remote/daemon receiver and reports what
WOULD transfer/skip based on receiver state, mutating nothing on either
side.

- Serialize Config.dry_run into the wire config frame and append
  STATUS_DRY_RUN_TRANSFER to the status enum (no renumbering); bump
  PROTOCOL_VERSION/CMake VERSION/CHANGELOG/golden wire to 2.21.0.
- Receiver: receive_incremental_check_ex runs the normal read-only
  decision and answers STATUS_OK (skip) or STATUS_DRY_RUN_TRANSFER
  (would transfer) with no basis materialization/append/delta/full
  transfer.  All mutation sites are guarded by !dry_run: file store,
  manifest deletes, --mkpath root creation, --delay-updates staging,
  publication, directory-time application, and outcome acks.
- Client: send_dry_run_remote connects, sends the config, checks each
  regular file and prints the would-transfer set + trailer; no file data
  or delete manifest is sent.  Plain local destinations keep the
  client-side manifest.
2026-09-13 11:56:05 +02:00
TapTap 72cccaa256 Merge Wave 8: config X-macro, authorized_root dedup, daemon limits, protocol_charge ownership
CI / lint (push) Successful in 1m31s
CI / sanitizers (undefined) (push) Successful in 1m0s
CI / sanitizers (address) (push) Successful in 1m6s
CI / fuzz-build (push) Successful in 34s
CI / coverage (push) Successful in 53s
CI / valgrind (push) Successful in 3m14s
CI / build-and-test (push) Successful in 5m33s
2026-09-13 11:19:55 +02:00
TapTap eb71b29d1c Merge branch 'fix/w8-authroot' into fix/w8-integration 2026-09-13 11:13:40 +02:00
TapTap 4d5befedfe Merge branch 'fix/w8-charge' into fix/w8-integration 2026-09-13 11:13:40 +02:00
TapTap f00844cf9a Merge branch 'fix/w8-daemonlim' into fix/w8-integration 2026-09-13 11:13:40 +02:00
TapTap 2ec17e821c fix(daemon): harden bounded per-source registry races
Stamp host_last_use before publishing a bucket key and treat an unstamped
(last_use == 0) bucket as live, so a just-claimed bucket can no longer be
stolen by a concurrent reclaimer.

After a successful eviction CAS, re-scan for the interned key and, when an
earlier bucket already holds it, zero the duplicate's active count and
return the canonical bucket, preventing orphaned per-host counts and cap
overshoot under full-table concurrency.

Add a message-carrying EXPECT_FAIL primitive and use it for the daemon-conf
buffer-overflow guard, and add a fork-based test that records auth failures
from forked children and asserts the parent observes the shared lockout.
2026-09-13 11:11:23 +02:00
TapTap 6d47d93fd7 fix(config): validate received counts before publishing them
The config_receive_{basis,skip,idmap}_count helpers wrote the
peer-controlled int through the Config member before range-checking it.
An over-cap basis_count therefore left config->basis_count huge while
config->basis_dirs was still NULL; config_receive()'s error path then
called config_delete(), whose basis loop dereferenced NULL and crashed
the daemon before authentication.

Read each count into a local, validate, and only then assign, leaving the
member untouched on failure.  config_delete() also guards the basis loop
with the array pointer as defense in depth.

Add a regression test that feeds over-cap basis/idmap/skip counts and
asserts rejection without crashing, plus a direct config_delete() check
on the partial (count set, array NULL) state.
2026-09-13 11:05:57 +02:00
TapTap 6ea966781f test(config): add receive-side golden oracle and sharpen fixture
Address low-severity review findings on the X-macro config refactor:

1. The golden test only hashed config_send_wire_block(), so a
   receive-side KIND that reads a different width/order could still
   round-trip symmetrically.  Add test_config_wire_golden_receive():
   capture the same hash-pinned 633-byte frame and feed it through
   config_receive(), asserting every field (config_wire_equal) plus the
   derived use_delta/use_xattrs bits and representative bounded kinds.
   Add test_config_wire_receive_bounds() for bounds the symmetric
   round-trip cannot reach: an out-of-range BOOL (hand-built frame),
   RAW_MAXALLOC zero, a malformed STR_MODULE, an over-cap
   INT_IDMAPCOUNT, and an out-of-range INT_IDENTITY chown_uid.

2. golden_config_populate() set long runs of booleans to all-1, so an
   adjacent swap within a run produced identical bytes.  Alternate the
   boolean values and make the fixture receiver-valid (chmod grammar
   "u=rwx,go=rx" is the same 11 bytes; delta_max_file_size inside the
   bound).  Re-pin the golden: len stays 633, hash is now
   9160991280011164139 (computed, not guessed).

3. Document in config.h and client_cli.c that the CLI option tables
   remain hand-maintained and are deliberately not generated from the
   wire-field X-macro (client-only fields, flag/alias/negation
   semantics).  No CLI-table rewrite.

PROTOCOL_VERSION stays "2.20.0"; src/shared/config.c is untouched and
the wire bytes are unchanged apart from the fixture's own new values.
2026-09-13 10:56:34 +02:00
TapTap 0a7f5faea6 test: replace strcat with a bounds-checked append in daemon-conf test 2026-09-13 10:51:01 +02:00
TapTap 25909110ac fix(daemon): exempt trusted loopback peers from per-host limits
Every client on loopback shares the 127.0.0.1 identity, so counting them
against 'max connections per host' or the default-on auth lockout lets one
local client deny service to all the others (and makes a shared-NAT/proxy
address a natural DoS vector for remote clients).  Use
utils_fd_peer_is_local (fail-closed) in the daemon gate to exempt a
provably local peer from the per-source cap and the auth lockout while
keeping the per-module and global caps.  Remote peers are unchanged.

Document the shared-NAT/proxy identity limitation and the loopback
exemption in README/RSYNC_COMPAT/CHANGELOG, update the integration test to
assert the exemption, and fix the README 'auth failure delay' cap (5000,
not 60000).
2026-09-13 10:50:58 +02:00
TapTap bd43448af2 fix(daemon): bound per-source table lifetime and recompute occupancy
The per-source host table only grew: once its fixed open-addressed table
filled, host_intern returned -1 and the per-host cap plus the shared auth
lockout silently failed open forever.  Add a bounded-lifetime eviction
policy: track a per-bucket last-use time and, when no empty bucket exists,
atomically repurpose the first bucket that has no active connection and
either has an expired lockout or has been idle, resetting its counters.
Warn (rate-limited) on the genuine fail-open path.

A child SIGKILLed mid-registration could also leak a module/host count
because the parent only decremented on a REGISTERED slot.  Make the slot
table the source of truth: after the SIGCHLD reap the parent recomputes
module_active[]/host_active[] from the surviving REGISTERED slots (atomics
only, async-signal-safe) so any leaked increment is erased.

Also clamp module_count to DAEMON_LIMITS_MAX_MODULES and use one helper
for the sizing/register host-tracking condition (a lockout threshold with
duration 0 is a no-op and must not intern hosts).
2026-09-13 10:50:53 +02:00
TapTap 264964411c Merge PR #283: chore(opencode): fix drifted agent/skill docs and repo hygiene
CI / lint (push) Successful in 1m29s
CI / sanitizers (undefined) (push) Successful in 55s
CI / sanitizers (address) (push) Successful in 1m1s
CI / fuzz-build (push) Successful in 34s
CI / coverage (push) Successful in 52s
CI / valgrind (push) Successful in 3m14s
CI / build-and-test (push) Successful in 5m32s
2026-09-13 10:44:57 +02:00
TapTap 18d1b84246 refactor(protocol): guard session release, clarify Data.owner contract
Add a NULL guard to protocol_release_memory_for_session so it no-ops like
the sibling session setters.  Correct the Data.owner doc comment, which
implied a non-zero protocol_charge always has an owner; document that
owner may be NULL for uncharged/ownerless Data, that any such charge
falls back to the bound session, and that a charged Data must not outlive
its owning session.  Note the lifetime contract on the release API too.

Extend tests/test_protocol.c to cover destroying a charged Data with no
session bound (the other half of the original bug) and to assert that
data_create/data_create_reserve start with owner == NULL and
protocol_charge == 0.
2026-09-13 10:42:45 +02:00
TapTap 0f95f48899 fix(opencode): correct remaining agent/skill doc drift
CI / lint (pull_request) Successful in 1m31s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / build-and-test (pull_request) Successful in 1m44s
- pr-review: replace invalid 'tea pr comment' with 'tea comment' (the
  former is not a tea subcommand)
- integrator: drop stray '-M' from client examples (-M is now
  --remote-option and requires an argument), use the canonical pytest
  integration command, and bump the CI image tag to v10
- test-writer: build fuzz targets via -DENABLE_FUZZ=ON instead of
  hand-rolled -fsanitize flags; fix the fuzz binary path
- cmake-expert: document -DSANITIZER=undefined, which is now live in
  CMakeLists.txt
- README: add --allow-unauthenticated to the plain-TCP server example,
  use --preserve for metadata (not -M), and use the canonical
  integration command
- AGENTS.md: use the canonical integration command
2026-09-13 10:41:23 +02:00
TapTap c78a21de57 docs(shared): clarify authorized_root accessor contracts
Document on utils_get_authorized_root_path() that the returned pointer is
borrowed and invalidated by the next authorized-root setter, that the fd
and path are not read atomically (non-reentrant), and that the fd remains
caller-owned.  Add a matching single-threaded/set-before-threads note at
the accessor definitions in utils.c.

In server.c, drop the redundant utils_set_authorized_root(-1, NULL) after
a failed utils_set_authorized_root(): the setter already fail-closes the
state on allocation failure.  The following close(root_fd) is unchanged.
2026-09-13 10:38:21 +02:00
TapTap 5c8970c64f chore(opencode): fix drifted agent/skill docs and repo hygiene
CI / lint (pull_request) Successful in 1m29s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / build-and-test (pull_request) Successful in 1m45s
The agent and skill definitions had drifted badly from the current
codebase and tooling, repeating the same class of bug as the benchmark
tool (references to nonexistent scripts and invented flags):

- Replace the removed `python3 test.py` with the real integration
  command (`python3 -m pytest tests/integration/ -n 4 --dist=load
  -m "not setpriv"`) across agents and skills.
- Fix `feature-scout`'s fabricated CLI flag list (--host, --server-mode,
  --use-* etc.) using the authoritative src/client/usage.c flags.
- Fix `perf-analyst` benchmark flags (-m -c -> -j -z) and point at
  benchmark/bench.py instead of stale numbers.
- Correct `code-explainer` (no getopt_long; --sendfile not -f) and
  version drift in the release skill (1.1.0 -> 2.20.0).
- Replace GitHub/`gh` workflows with Gitea/`tea` (PRs target dev; issues
  via tea; branch strategy updated in all agents).
- Use the built-in `-DSANITIZER=address|thread` CMake option instead of
  hand-rolled -fsanitize flags.
- Add `-p 8080 --allow-unauthenticated` to plain-TCP server examples.
- Merge the redundant security-screener into security-auditor; drop the
  duplicate (16 agents remain).

Repo hygiene: gitignore `root/` and `test_partial_install_tmp/`, remove
the empty leftover trees, delete the tracked scratch scripts tmux.sh and
to_one_file.py, and note the compile_commands.json symlink in README.
2026-09-13 10:34:59 +02:00
TapTap 4e918a1b69 test(config): pin wire bytes and round-trip every field
test_config_wire_golden() serializes a fully-populated Config through
config_send_wire_block() and pins the exact frame to len=633 and FNV-1a
hash 6163263374908258816, captured from the pre-X-macro implementation.
Any field reorder, resize or codec change fails the test.

test_config_wire_roundtrip_all_fields() serializes/deserializes a defaults
Config and a fully-populated Config over a socketpair and compares every
serialized field.  The comparison is itself generated from
CONFIG_WIRE_FIELDS (one CONFIG_CMP_<KIND> per table entry), so a new table
entry automatically extends coverage; it cannot fall out of sync.  It
normalizes the receiver's NULL/"" canonicalization, the max_alloc server
clamp and the derived use_delta/use_xattrs bits.
2026-09-13 10:28:31 +02:00
TapTap 87f6cb0243 refactor(config): single X-macro table for serialized fields
Every Config field that crosses the wire was declared in up to six places
(struct member, config_set_defaults, send_*, receive_*, and the two CLI
option tables) and could drift silently.  Add CONFIG_WIRE_FIELDS in
config.h: one ordered per-segment table where each serialized field is
declared once with its C type, default and wire codec (KIND).

config.h now expands the table to declare the struct members;
config_set_defaults() expands it to assign the defaults; and
config_send_wire_block()/config_receive() expand the per-segment lists to
emit/consume the frame.  The per-segment function names, call order and
segment boundaries are preserved exactly.

Fields with genuinely custom logic keep dedicated helpers but are still
declared once in the table: the protocol-version handshake (HEADER), daemon
SCRAM auth (STR_REDACTED_AUTH), the daemon module name (STR_MODULE), the
repeated count+array blocks (BLOCK_SKIP_SUFFIXES/BLOCK_BASIS/BLOCK_IDMAP),
--copy-as presence/ids (COPY_AS_*), and the derived --delta / use_xattrs
bits (DERIVED_DELTA, BOOL_XATTR_DERIVE).  The version field remains a
special header (validated before any other field is parsed) and is sent by
config_send_wire_block() explicitly.

No public field is renamed and PROTOCOL_VERSION stays "2.20.0".  Because
the struct declaration order is no longer the wire order, the wire order is
now enforced solely by the table and by a byte-exact golden test
(follow-up commit).  Add config_send_wire_block() so that test can hash the
frame body without the STATUS_OK handshake.
2026-09-13 10:28:26 +02:00
TapTap e1f8f75e7c docs: document daemon per-module/per-host caps and shared auth lockout 2026-09-13 10:24:09 +02:00
TapTap 4c17122b00 feat(daemon): enforce per-module/per-host caps and shared auth lockout
Wire the shared registry into the accept loop (parent claims a slot before
fork, blocks SIGCHLD across fork+pid publication, and reclaims the dead
child's slot from the SIGCHLD handler so per-module/per-source counts are
released even on SIGKILL). The connection child records the selected module
and normalized peer IP once the config frame names them: an over-cap module
or source is refused at the config gate with an audit log, and a source
that exceeded the auth-failure threshold is refused before a SCRAM
challenge (the counter is shared across children and cleared on success).
The existing global cap and host ACLs are untouched.
2026-09-13 10:24:05 +02:00
TapTap 0abaa62193 feat(daemon): parse per-host cap and auth lockout config keys
Add global keys `max connections per host` (default 0 = unlimited),
`auth lockout threshold` (default 10, 0 disables) and
`auth lockout duration` (default 300 s, 0 disables). Module
`max connections` now accepts 0 as unlimited. Bound the number of
[module] sections (DAEMON_CONF_MAX_MODULES) so the shared registry's
per-module counter array stays fixed-size; absent keys keep their
defaults so old configs still load.
2026-09-13 10:24:01 +02:00
TapTap 5334397b81 feat(daemon): add shared cross-process connection registry
The daemon forks one child per accepted connection, so per-module and
per-source accounting must live in state shared across the children. Add a
fixed-size registry carved from an anonymous shared mapping
(mmap(MAP_SHARED|MAP_ANONYMOUS)) created before the accept loop: a slot
lifecycle (FREE/CLAIMED/REGISTERED) with parent claim/reclaim and a
lock-free, open-addressed per-source table for the per-host occupancy and
the shared auth-failure counter. C11 atomics only; no pthread locks across
fork.

Unit tests cover slot exhaustion, the module/host caps, pid reclaim and
fork-shared visibility.
2026-09-13 10:23:58 +02:00
TapTap 6968ff6734 Merge PR #282: fix(benchmark): use real FastSync flags and Release builds
CI / lint (push) Successful in 1m31s
CI / sanitizers (undefined) (push) Successful in 54s
CI / sanitizers (address) (push) Successful in 1m1s
CI / fuzz-build (push) Successful in 33s
CI / coverage (push) Successful in 52s
CI / valgrind (push) Successful in 3m14s
CI / build-and-test (push) Successful in 5m31s
2026-09-13 10:18:21 +02:00
TapTap 5aca91ab22 fix(benchmark): use real FastSync flags and Release builds
CI / lint (pull_request) Successful in 1m29s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / build-and-test (pull_request) Successful in 1m44s
The benchmark tool used stale rsync-style spellings that map to

different FastSync options, so it never enabled the features it

claimed to measure:

  -c -> --checksum (not compression)

  -m -> --prune-empty-dirs (not multithreading)

  -s -> --secluded-args, a no-op (not chunk serialization)

  -f -> --filter, needs an argument (not sendfile)

Replace them with the real flags (-z, -j, --chunk-serialization,

--sendfile), force CMAKE_BUILD_TYPE=Release, route informational

output to stderr so --output json emits valid JSON, surface

client/rsync failures instead of silently dropping them, and widen

the results table for the longer config names. Update the benchmark

skill to match (correct flags, server invocation, and replace the

nonexistent test.py --full with benchmark/bench.py).
2026-09-13 10:14:37 +02:00
TapTap 3260a39ab4 refactor(shared): single owner for authorized_root state 2026-09-13 10:06:04 +02:00
TapTap 5d3c43305e fix(protocol): release Data charge to its owning session
Data charged against a ProtocolSession kept only the charge amount, so
data_destroy released it from whatever session was thread-locally bound
at destroy time. Destroying a received Data on another thread, after the
session was unbound, or while a different session was bound leaked the
originating session's budget and underflowed the other's.

Add Data.owner, set it whenever protocol_receive_data_limited charges a
session, and have data_destroy release against that owner directly via
the newly-exported protocol_release_memory_for_session. Uncharged Data
(owner NULL) keeps the previous bound-session fallback.

Add a unit test proving a Data acquired on session A is released to A
even when unrelated session B is bound at destroy time.
2026-09-13 10:05:38 +02:00
TapTap 9242e86772 Merge Wave 7: fix shared/server layering and explicit CMake targets
CI / lint (push) Successful in 1m30s
CI / sanitizers (undefined) (push) Successful in 56s
CI / sanitizers (address) (push) Successful in 1m3s
CI / fuzz-build (push) Successful in 34s
CI / coverage (push) Successful in 52s
CI / valgrind (push) Successful in 3m14s
CI / build-and-test (push) Successful in 5m33s
2026-09-13 07:30:33 +02:00
TapTap 0155902d95 docs: update stale PipelineContextReceiver reference 2026-09-13 07:30:28 +02:00
TapTap 3499baf80b build: explicit CMake targets; move receiver pipeline out of shared 2026-09-13 07:20:28 +02:00
TapTap 83eacf3151 Merge Wave 6: client features (--port, --threads=N, abort, keepalive) and test coverage
CI / lint (push) Successful in 1m30s
CI / sanitizers (undefined) (push) Successful in 1m4s
CI / sanitizers (address) (push) Successful in 1m9s
CI / fuzz-build (push) Successful in 36s
CI / coverage (push) Successful in 51s
CI / valgrind (push) Successful in 3m14s
CI / build-and-test (push) Successful in 5m37s
2026-09-13 06:59:08 +02:00
TapTap c2df0347ef fix(client,protocol): EINTR-safe sends, armed abort, keepalive drain grace, TLS WANT_WRITE 2026-09-13 06:59:02 +02:00
TapTap dd44537b44 Merge branch 'fix/w6-tests' into fix/w6-integration 2026-09-13 06:25:05 +02:00
TapTap 2854a9d149 test: fuzz manifest/protocol/xattr, hardlink unit, fault injection 2026-09-13 06:24:46 +02:00
TapTap 1fa2fbd266 feat(client): --port alias, --threads=N, graceful abort, keepalive 2026-09-13 06:22:28 +02:00
TapTap 10b18ab2d2 Merge Wave 5a: dead-code removal, scanner options embed, shared config invariants, bounded metadata API
CI / lint (push) Successful in 1m31s
CI / sanitizers (undefined) (push) Successful in 1m0s
CI / sanitizers (address) (push) Successful in 1m7s
CI / fuzz-build (push) Successful in 30s
CI / coverage (push) Successful in 51s
CI / valgrind (push) Successful in 3m12s
CI / build-and-test (push) Successful in 5m32s
2026-09-13 05:48:32 +02:00
TapTap dcc78c14c5 docs,fuzz: fix ownership/alloc comments; fuzz chunk metadata path 2026-09-13 05:48:27 +02:00
TapTap 5a829adb85 Merge branch 'fix/w5-scanner' into fix/w5-integration 2026-09-13 05:28:50 +02:00
TapTap 42c72030fb Merge branch 'fix/w5-config' into fix/w5-integration 2026-09-13 05:28:50 +02:00
TapTap 99f8045105 refactor(scanner,send): embed scanner options; unify stats and config ownership 2026-09-13 05:28:32 +02:00
TapTap eea66a7848 refactor(config,metadata): shared invariants; length-bounded metadata parser 2026-09-13 05:24:23 +02:00
TapTap c944787e03 refactor: remove dead file_store subsystem and unused wrappers 2026-09-13 05:19:18 +02:00
TapTap 57ce6d04f0 Merge Wave 4: performance (packed metadata 2.20.0, indexed lookups, zstd reuse, byte-bounded queues)
CI / lint (push) Successful in 1m30s
CI / sanitizers (undefined) (push) Successful in 1m1s
CI / sanitizers (address) (push) Successful in 1m7s
CI / fuzz-build (push) Successful in 29s
CI / coverage (push) Successful in 50s
CI / valgrind (push) Successful in 3m12s
CI / build-and-test (push) Successful in 5m22s
2026-09-13 05:01:21 +02:00
TapTap 3cf2e2c91f docs(version): align 2.20.0 artifacts; fix protocol-bump rationale 2026-09-13 05:01:15 +02:00
TapTap 301cb0dbaf fix(utils,file-list): bound keep/files-from indexes to O(M) memory 2026-09-13 04:52:10 +02:00
TapTap a4f4110397 Merge branch 'fix/w4-queue' into fix/w4-integration 2026-09-13 04:19:12 +02:00
TapTap 24fe8c5583 Merge branch 'fix/w4-compress' into fix/w4-integration 2026-09-13 04:19:12 +02:00
TapTap d274bdff4e Merge branch 'fix/w4-hash' into fix/w4-integration 2026-09-13 04:19:12 +02:00
TapTap 6c636a19e6 perf(compression,tcp): reuse zstd contexts; enable TCP_NODELAY 2026-09-13 04:18:55 +02:00
TapTap 1a83e284c4 perf(utils,file-list): index delete keep-set and --files-from lookups 2026-09-13 04:16:28 +02:00
TapTap 317d5d081a perf(protocol): pack metadata into one frame (PROTOCOL 2.20.0) 2026-09-13 04:08:36 +02:00
TapTap 69fe7f3c9f perf(send,scanner): byte-bound sender queues; drop redundant stat 2026-09-13 04:06:30 +02:00
TapTap ddc71a7df5 Merge Wave 3b: configurable protocol timeout and idle/session bounds
CI / lint (push) Successful in 1m31s
CI / sanitizers (undefined) (push) Successful in 1m1s
CI / sanitizers (address) (push) Successful in 1m8s
CI / fuzz-build (push) Successful in 30s
CI / coverage (push) Successful in 50s
CI / build-and-test (push) Successful in 4m37s
CI / valgrind (push) Successful in 3m12s
2026-09-13 03:46:25 +02:00
TapTap ffa1d24625 fix(receiver): harden idle-progress definition, single error frame, sendfile timeout 2026-09-13 03:46:20 +02:00
TapTap b16349b81e fix(protocol): honor --timeout for protocol I/O; bound idle/session time 2026-09-13 03:29:01 +02:00
TapTap 4bc84fe954 Merge Wave 3a: daemon host ACL, configurable max connections, peer audit, auth-failure delay
CI / lint (push) Successful in 1m31s
CI / sanitizers (undefined) (push) Successful in 57s
CI / sanitizers (address) (push) Successful in 1m5s
CI / fuzz-build (push) Successful in 29s
CI / coverage (push) Successful in 49s
CI / build-and-test (push) Successful in 4m35s
CI / valgrind (push) Successful in 3m10s
2026-09-13 02:51:11 +02:00
TapTap fc560246c1 fix(daemon): close ACL fail-opens (v4-mapped peers, invalid patterns) and cap auth delay 2026-09-13 02:51:07 +02:00
TapTap dff6609976 feat(daemon): host ACL, configurable max connections, peer audit, auth-failure delay 2026-09-13 02:36:15 +02:00
TapTap 1acb66628d Merge Wave 2: thread-safety fixes (signals, fd ownership, handler epilogue, logging, scanner leak)
CI / lint (push) Successful in 1m30s
CI / sanitizers (undefined) (push) Successful in 59s
CI / sanitizers (address) (push) Successful in 1m6s
CI / fuzz-build (push) Successful in 28s
CI / coverage (push) Successful in 49s
CI / build-and-test (push) Successful in 4m31s
CI / valgrind (push) Successful in 3m10s
2026-09-13 02:13:27 +02:00
TapTap ba1c7a369f fix(server,log): non-socket shutdown fallback, drop redundant delay cleanup, unlock logging I/O 2026-09-13 02:13:22 +02:00
TapTap d28489d83c Merge branch 'fix/w2-scan' into fix/w2-integration 2026-09-13 01:49:44 +02:00
TapTap 312ed05170 Merge branch 'fix/w2-log' into fix/w2-integration 2026-09-13 01:49:44 +02:00
TapTap fecbe2c90c fix(server): child-safe signals, single fd owner, handler cleanup epilogue 2026-09-13 01:49:27 +02:00
TapTap c8f5d80fcb fix(log): serialize message emission; clear log_fp before close; use logger 2026-09-13 01:44:59 +02:00
TapTap 8147ff7b50 fix(scanner): free chunk_data on chunk-create failure 2026-09-13 01:36:51 +02:00
TapTap b7fbb56289 test(file): silence cppcheck constVariablePointer in empty-path test
CI / lint (push) Successful in 1m32s
CI / sanitizers (undefined) (push) Successful in 57s
CI / sanitizers (address) (push) Successful in 1m4s
CI / fuzz-build (push) Successful in 30s
CI / coverage (push) Successful in 49s
CI / build-and-test (push) Successful in 4m28s
CI / valgrind (push) Successful in 3m10s
2026-09-13 01:25:29 +02:00
TapTap 08063b6d73 Merge Wave 1: critical/High fixes (UAF, DoS caps, leaks, hardening)
CI / lint (push) Failing after 1m32s
CI / build-and-test (push) Skipped
CI / sanitizers (address) (push) Skipped
CI / sanitizers (undefined) (push) Skipped
CI / fuzz-build (push) Skipped
CI / coverage (push) Skipped
CI / valgrind (push) Skipped
2026-09-13 01:18:58 +02:00
TapTap ea2f76cd7a fix(receiver): charge per-entry DirTimeList cost; cap client --skip-compress 2026-09-13 01:18:54 +02:00
TapTap 76eeba1773 Merge branch 'fix/w1d-hardening' into fix/w1-integration 2026-09-13 00:59:35 +02:00
TapTap b72ab298ab Merge branch 'fix/w1c-wire' into fix/w1-integration 2026-09-13 00:59:35 +02:00
TapTap 446a714ef8 Merge branch 'fix/w1b-receiver' into fix/w1-integration 2026-09-13 00:59:35 +02:00
TapTap a90e234eb3 harden: overflow guards, auth-user validation, TLS1.3 policy, build hardening 2026-09-13 00:59:21 +02:00
TapTap f8252cf3e7 fix(protocol): bound pre-auth config string memory 2026-09-13 00:57:32 +02:00
TapTap 4557924972 fix(receiver): cap DirTimeList growth and fix placeholder Data leaks 2026-09-13 00:57:32 +02:00
TapTap 59ce174d22 fix(client-send): UAF in basis preflight and missing_args leak 2026-09-13 00:57:09 +02:00
TapTap 2a8941ee5c Merge feat/ref-integration: SuperMode enum, dir-time gate dedup, parse_args/server_module_gate splits
CI / lint (push) Successful in 1m30s
CI / sanitizers (undefined) (push) Successful in 59s
CI / sanitizers (address) (push) Successful in 1m6s
CI / fuzz-build (push) Successful in 29s
CI / coverage (push) Successful in 49s
CI / valgrind (push) Successful in 2m9s
CI / build-and-test (push) Successful in 4m31s
2026-09-12 21:11:03 +02:00
TapTap 37037a6ee7 refactor(client-cli): drop unused CliParseCtx positional fields (cppcheck) 2026-09-12 21:07:14 +02:00
TapTap 061e9ad43f Merge feat/ref-modulegate: split server_module_gate into helpers 2026-09-12 20:56:43 +02:00
TapTap f928879755 Merge feat/ref-parseargs: split parse_args into focused helpers 2026-09-12 20:56:43 +02:00
TapTap 84b7cb0de3 refactor(server): split server_module_gate into ordered helper stages 2026-09-12 20:56:33 +02:00
TapTap 921472b8b3 refactor(client-cli): split parse_args into focused option handlers
Break the ~700-line parse_args god function into cohesive static helpers
grouped by concern: output controls, pre-negation, range/time options, the
OPTION_TABLE dispatcher, flag/meta handlers, IO/network options, filter and
logging options, checksum/socket options, remote/basis/identity options,
positional handling, and a final lowering step.

A file-local CliParseCtx carries the config, cursor, positional buffers and
the mutable parse flags, so each handler stays focused. The dispatcher calls
the handlers in the original recognition order and preserves the exact
return contract (0/1/negative), error messages, log levels and control flow.

Behavior preserved; no functional changes.
2026-09-12 20:55:04 +02:00
TapTap 082ac2645d Merge feat/ref-dirtime: dir-time capture gate dedup 2026-09-12 20:43:42 +02:00
TapTap eefbd1e849 Merge feat/ref-supermode: SuperMode enum 2026-09-12 20:43:42 +02:00
TapTap 1fd462cca8 refactor(config): replace SUPER_MODE_* macros with SuperMode enum
Type Config.super_mode as SuperMode (a proper C enum) instead of a bare
int.  The wire boundary still carries the mode as an int: send casts the
enum explicitly and receive reads a temporary int, validates the
AUTO..OFF range, then casts.  Emitted bytes and accepted values are
unchanged.  ModuleGateContext.super_mode_override keeps its -1 sentinel
as int with an explicit cast at the apply site.

Behavior preserved.
2026-09-12 20:43:24 +02:00
TapTap 4ac37c4d8a refactor(dir-times): extract dir_times_should_capture predicate
Deduplicate the repeated directory-time capture gate
(`config->use_metadata && !config->omit_dir_times`) used by the
sender-side (multiprocessing.c) and receiver-side (receiver.c) sinks
into a single predicate declared next to the DirTimeList machinery in
file_receive.h and defined in file_receive.c.

Behavior preserved: identical short-circuit condition and semantics,
no signature or protocol changes.
2026-09-12 20:42:47 +02:00
TapTap 08af945bd6 Merge main back into dev after v2.19.0 release
CI / lint (push) Successful in 1m32s
CI / sanitizers (address) (push) Successful in 55s
CI / fuzz-build (push) Successful in 30s
CI / sanitizers (undefined) (push) Successful in 53s
CI / coverage (push) Successful in 47s
CI / build-and-test (push) Successful in 4m28s
CI / valgrind (push) Successful in 2m11s
2026-09-12 20:22:54 +02:00
TapTap 378d881ca7 Merge release/v2.19.0 into main: FastSync v2.19.0
CI / lint (push) Successful in 1m33s
CI / sanitizers (undefined) (push) Successful in 59s
CI / sanitizers (address) (push) Successful in 1m4s
CI / fuzz-build (push) Successful in 28s
CI / coverage (push) Successful in 49s
CI / valgrind (push) Successful in 2m10s
CI / build-and-test (push) Successful in 4m29s
2026-09-12 20:22:50 +02:00
TapTap cc27ee1b83 Release v2.19.0
- Protocol version 2.19.0 (SCRAM-SHA-256 daemon auth replacing the replayable digest)
- Salted PBKDF2 verifier store + --hash-credentials; legacy store hard-rejected
- Persistent anti-enumeration dummy key (<store>.dummykey)
- Verified TLS / opted-in loopback transport required for auth modules
- Secret wiping; carried-over hardening from the security phases
- Add CHANGELOG.md and set the CMake project version
2026-09-12 20:22:46 +02:00
TapTap d653c3e151 Merge feat/tls-dummy-integration: verified/Local-only transport for daemon auth + persistent dummy key
CI / lint (push) Successful in 1m38s
CI / sanitizers (undefined) (push) Successful in 1m1s
CI / sanitizers (address) (push) Successful in 1m7s
CI / fuzz-build (push) Successful in 30s
CI / coverage (push) Successful in 49s
CI / valgrind (push) Successful in 2m14s
CI / build-and-test (push) Successful in 4m32s
2026-09-12 19:55:28 +02:00
TapTap 1b90ee2449 fix(review): close loopback TLS auth bypass; align docs and wrong-CN test
- server gate: the --allow-unauthenticated loopback allowance now requires
  an actual plaintext connection (!gate_ctx->ssl), so a loopback TLS client
  whose cert fails the --client-cn check is refused before any SCRAM
  challenge instead of falling through the plaintext opt-in.  Keep the
  invalid-fd guard as belt-and-braces (unreachable after the policy check).
- test: rewrote test_wrong_client_cn_refused_before_auth_challenge to run
  deterministically over 127.0.0.1 with --tls + --allow-unauthenticated and
  a CA-valid wrong-CN client cert, asserting the gate refusal log and an
  unchanged module tree (no skip).
- docs: --client-cn is mandatory with --tls; dummykey sidecar is secret
  material; document all transient-fallback reasons; qualify
  --allow-unauthenticated in README and --help so it cannot read as
  permitting remote plaintext auth.
- credentials.h: drop stale restrictive-umask claim (fchmod forces exact
  0600; only create/write/fsync/link/fchmod failure degrades to ephemeral).
2026-09-12 19:50:52 +02:00
TapTap 89b967f29a Merge feat/dummy-key: persist anti-enumeration dummy key across restarts
# Conflicts:
#	RSYNC_COMPAT.md
2026-09-12 19:30:35 +02:00
TapTap 8f06ae5262 Merge feat/tls-auth: require verified/local transport for daemon auth modules 2026-09-12 19:29:35 +02:00
TapTap ac3c4c7c72 fix(a7-auth): harden dummy-key temp creation
- Make the atomic-publish temp name unpredictable by appending 16 random
  hex chars to the pid, so a leftover/planted temp cannot be targeted.
- On EEXIST, unlink the stale temp and retry the O_EXCL create once
  (bounded), so a crash leftover or reused pid cannot silently defeat
  sidecar persistence.
- fchmod the temp fd to 0600 after creation (umask can clear owner bits)
  and treat failure as a create failure, so the published sidecar is
  always exactly 0600.
- Clarify comments: the sidecar requires exact 0600 while the store and
  password files only reject group/other bits.
- Add a unit test that a restrictive umask still yields an exact 0600
  sidecar; clean random-suffixed temps in tests.
2026-09-12 19:29:23 +02:00
TapTap d53614d06b fix(a7-3/s1): fail closed on non-loopback peers; require plaintext opt-in before challenge
utils_fd_peer_is_local now returns true only when getpeername SUCCEEDS and the
peer address classifies as loopback. A non-socket descriptor (pipe/socketpair)
or any getpeername error is NOT local, so the daemon auth gate fails closed
instead of treating an untestable --stdio pipe as trusted (daemon auth modules
are --daemon-only and the stdio path never loads a daemon config).

server_module_gate now requires --allow-unauthenticated for the loopback
plaintext auth path: a plaintext loopback connection without the operator
opt-in is refused at the config gate BEFORE server_auth_handshake, so no SCRAM
challenge is sent. Remote peers still require verified TLS regardless of the
flag; the handler keeps its defense-in-depth checks.

Docs state the exact policy (verified TLS with matching --client-cn, or
operator-opted-in loopback plaintext), drop the SSH/stdio auth-transport claim
(they are daemon-only), and add the loopback trust-boundary relay caveat and
the CN-only (no SAN) residual. Adds a unit-test negative for pipe/socketpair
and an integration test where a relay observes no challenge when the flag is
absent.
2026-09-12 19:18:29 +02:00
TapTap f0381a6b8e fix(a7-auth): publish dummy-key sidecar atomically and harden reads
Address review findings on the persistent dummy-key sidecar:

- Publish atomically: write a private same-directory temp file
  (<store>.dummykey.tmp.<pid>, 0600), fsync, then link(2) into place;
  fsync the containing directory and drop the temp name. A concurrent
  starter can no longer observe a zero/partial sidecar and fail closed.
  On EEXIST adopt the winner's sidecar; otherwise warn and use a
  transient ephemeral key.
- Harden the read path (initial and EEXIST-adopt) with
  O_RDONLY|O_NOFOLLOW|O_NONBLOCK|O_CLOEXEC: reject planted symlinks
  (ELOOP fails closed) and never block on a planted FIFO.
- Require the exact owner-only mode (st_mode & 07777) == 0600 and make
  the rejection message truthful.
- Report a clear "short write" instead of a stale strerror(errno) when
  write() returns 0.
- Document the artifact and its creation-failure caveat (FIFO store
  path, read-only filesystem, missing directory) in README.md and
  RSYNC_COMPAT.md.
- Tests: known-key sidecar adoption (dummy salt KAT + reload), symlink
  rejection, and the exact-0600 rule (0400 now rejected).
2026-09-12 19:16:11 +02:00
TapTap a7a1930e88 fix(a7-3/s1): require TLS or local transport for daemon auth
Daemon modules that declare 'auth users' no longer accept credentials over a
remote plaintext connection: server_module_gate refuses at the config gate,
before any SCRAM challenge is sent, unless the connection is verified TLS with
a client certificate matching --client-cn, or a local/SSH transport (loopback
TCP peer or the --stdio pipe). --allow-unauthenticated does not relax this.

The TLS client-CN comparison now uses credentials_secure_equal (S2). Clients
sending --password-file to a non-loopback daemon must use --tls; validate_config
rejects the plaintext case before any network I/O.

Adds utils_sockaddr_is_loopback / utils_fd_peer_is_local / utils_host_is_loopback
helpers with unit tests, a client validation unit test, and integration tests
for the client-side plaintext rejection and the wrong-CN gate refusal.
2026-09-12 19:02:05 +02:00
TapTap 42f01c0968 fix(a7-auth): persist dummy key in owner-only sidecar
The store-wide dummy key was regenerated on every credentials_load, so an
unknown user's dummy salt changed across daemon restarts while a real user's
stored salt stayed stable -- a restart-gated username-enumeration oracle.

Persist the 32-byte key in a 0600 <store>.dummykey sidecar next to the
credential store.  An absent sidecar is created with O_EXCL and fsynced; a
present sidecar is read only when it is an owner-only regular file of exactly
32 bytes (otherwise the load fails closed).  If the sidecar cannot be created
(read-only mount, missing directory) fall back to a transient per-run key with
a warning.  A NULL store path keeps the key ephemeral.
2026-09-12 18:40:21 +02:00
TapTap 2489d422e5 Merge feat/a7-integration: SCRAM-SHA-256 daemon auth + lazy protocol debug escaping
CI / lint (push) Successful in 1m33s
CI / sanitizers (undefined) (push) Successful in 59s
CI / sanitizers (address) (push) Successful in 1m4s
CI / fuzz-build (push) Successful in 29s
CI / coverage (push) Successful in 50s
CI / valgrind (push) Successful in 1m54s
CI / build-and-test (push) Successful in 4m45s
2026-09-12 18:21:13 +02:00
TapTap e2ddc0c07f Merge feat/hardening-misc: lazy protocol debug escaping, log_debug_enabled 2026-09-12 18:08:56 +02:00
TapTap 1480716304 Merge feat/a7-auth: SCRAM-SHA-256 daemon auth replacing replayable static digest 2026-09-12 18:08:56 +02:00
TapTap 1de1376e54 fix(a7-auth): final hardening pass on SCRAM auth
- burn the store-wide dummy_key in credentials_free()
- burn the local mac on hmac_sha256 failure in credentials_get_verifier()
- always run the O(store) constant-time scan, even for off-list users, to
  close the pre-existing off-list timing channel; select the real verifier
  only when on_list && match
- clarify the server_auth_handshake STATUS_AUTH_FAILED comment (failure
  before success vs. a dropped broken connection while writing the signature)
- document accepted anti-enumeration residuals (restart-gated dummy salt;
  pre-auth-observable iteration count)
2026-09-12 18:08:46 +02:00
TapTap eaf67f6257 fix(a7-auth): address SCRAM auth review findings A-G
- tests: pass CREDENTIAL_KEY_LEN to unhex for the 32-byte KAT proof/sig
  (sizeof(expect) is 348, over-reading the 65-byte hex literal under ASan)
- credentials: close the username-enumeration oracle with a store-wide
  dummy_key and a deterministic per-username dummy salt; make the store's
  iteration count uniform (reject intra-file and layered disagreements) and
  answer a miss with the store-wide count; run the constant-time key compare
  even when found=false and fold the decision with bitwise AND
- credentials_compute_keys: enforce [CREDENTIAL_MIN_ITERS, CREDENTIAL_MAX_ITERS]
- tests: recompute the whole KAT independently at CREDENTIAL_DEFAULT_ITERS
  (600000) and pin the golden store line; add non-uniform-store rejection,
  bound and deterministic-dummy-salt assertions
- server: send exactly one generic STATUS_AUTH_FAILED on every failure path
  (including credentials_get_verifier failure); route all handshake exits
  through one burn path
- credentials/server: burn the base64 decoders' scratch on error, the
  hash_store_line base64/line buffers on failure, and all handshake key/proof
  material
- fuzz: guard the auth-offset scan against size_t underflow and use a found flag
- docs: drop stale digest wording, use CREDENTIAL_MIN_ITERS as the --iterations
  bound, document 0600 output for --hash-credentials (plus a stderr warning on
  a group/other-accessible stdout file), and describe the deterministic dummy
  salt in the no-oracle claims
2026-09-12 17:56:52 +02:00
TapTap 8c94ec9886 feat(a7): SCRAM-SHA-256 daemon auth to replace replayable digest
Replace the challenge-less static-SHA-256 daemon bearer credential with a
SCRAM-SHA-256-style challenge/response and a salted PBKDF2 verifier store.
PROTOCOL_VERSION 2.18.0 -> 2.19.0; legacy user:SHA256HEX stores hard-reject.

- credentials: b64/rand/PBKDF2/HMAC primitives, verifier store parser,
  constant-time proof verify + ServerSignature, --hash-credentials helper
- config: auth block is now [present][username]; client runs the challenge
  exchange; config_burn_auth wipes plaintext/derived secrets (A7-4)
- server: gate drives the challenge, dummy verifier for unknown/off-list users
- tests: independent Python KAT, replay + legacy integration tests, fuzz paths
- docs: new store format, --hash-credentials, 2.19.0 bump

TLS verification behavior (A7-3/S1) is intentionally unchanged.
2026-09-12 17:19:33 +02:00
TapTap 87585e9881 perf(protocol): skip string debug escaping when proto debug is off
output_escape() was called on every send_str/receive_str even when
LOG_DEBUG_PROTO logging was disabled, allocating and scanning the whole
payload for a line that log_debug_message() then discarded.  Add a
log_debug_enabled(flag) gate mirroring log_debug_message()'s own filter and
check it before escaping.  Redacted (secret) strings still log the same
<redacted> marker; no observable log output changes.
2026-09-12 16:54:43 +02:00
TapTap 1ba6372017 Merge feat/p8-integration: P8 hardening batch (fs/transport, identity/server, CLI quality, fuzz)
CI / lint (push) Successful in 1m28s
CI / sanitizers (undefined) (push) Successful in 57s
CI / sanitizers (address) (push) Successful in 1m0s
CI / fuzz-build (push) Successful in 27s
CI / coverage (push) Successful in 47s
CI / valgrind (push) Successful in 40s
CI / build-and-test (push) Successful in 5m24s
2026-09-12 15:55:49 +02:00
TapTap 9f74b21c64 test(p8h): assert a --no-super daemon still refuses client --super 2026-09-12 15:55:44 +02:00
TapTap 108fee1e41 fix(p8h): restore daemon --super refusal, race-free secret-file check, remaining log escapes
- server_module_gate: refuse client-chosen ownership against the ORIGINAL config
  so an explicit --super is still refused under an operator --no-super veto
  (the veto must not turn a refusal into an accept).
- credentials: open-then-fstat the exact secret inode, require current-user
  ownership and no group/other bits, but continue to allow process-substitution
  FIFOs; removes the stat->fopen TOCTOU.
- file.c preallocate + protocol.c send-string debug logs escape attacker paths.
- usage/RSYNC_COMPAT updated for --old-args no-op and secret-file rules.
2026-09-12 15:50:14 +02:00
TapTap e1bb2e9233 chore(p8h): reconcile ssh old-args docs, secret-file perms docs; fix test cppcheck 2026-09-12 15:33:54 +02:00
TapTap 81fed86748 Merge branch 'feat/p8h-fuzz' into feat/p8-integration 2026-09-12 15:22:00 +02:00
TapTap 5844648fb2 Merge branch 'feat/p8h-cli' into feat/p8-integration 2026-09-12 15:22:00 +02:00
TapTap 4331bc4a8d Merge branch 'feat/p8h-core' into feat/p8-integration 2026-09-12 15:22:00 +02:00
TapTap d97e3982b4 test(fuzz): add config-frame receive and identity parser fuzz targets
Add two libFuzzer harnesses (GLOBbed from tests/fuzz/*.c) and deterministic
P8 config-frame receive tests:

- fuzz_config_receive.c drives config_receive() from arbitrary bytes. It
  captures one canonical valid frame with the production sender and feeds the
  receiver four shapes: raw bytes, valid-version-prefix + fuzz bytes, valid
  frame minus the P8 tail (super_mode + copy-as) + fuzz bytes, and valid frame
  minus the usermap count + fuzz bytes. This reaches the --super/--copy-as and
  huge/negative map-count paths that random bytes cannot get through the
  preceding wire-bool gate.
- fuzz_identity_parse.c fuzzes identity_parse_copy_as/map/chown plus the
  identity_wire_valid/identity_ownership_requested predicates on a fresh
  config per input.
- test_fuzz_smoke.c gains deterministic malformed-frame cases: out-of-range
  super_mode, negative/extreme copy-as ids, non-bool copy-as presence, tail
  truncation, huge/negative usermap counts, version mismatch and a
  wrong-order field after the version gate.

Unit build (STRICT_WARNINGS) and the fuzz build are clean; both targets run
3000+ iterations with no crash. No production code changed.
2026-09-12 15:21:33 +02:00
TapTap 6d32bc795b fix(p8h-core): escape log paths, fail closed on identity activation, tidy server gate
- A6: escape attacker-controlled file paths and the receive root in log
  lines (file_receive, server, protocol DEBUG) with output_escape()
- A8: identity_set_active() returns bool and fails closed when a requested
  usermap/groupmap cannot be deep-copied; handler refuses the connection
- remove the const cast and duplicate super_mode clamp from
  server_module_gate via an explicit override the handler applies once
- release the identity snapshot on the queue_create failure path
- refactor identity_parse_copy_as to a single cleanup tail and drop the
  duplicated group error format specifier
2026-09-12 15:03:54 +02:00
TapTap abad1664ba security(shared): fix -K TOCTOU, ssh old-args quoting, TLS opts, secret-file perms, sparse dedup
- file: open -K dirlink referents via a race-safe relative O_NOFOLLOW walk
  from the authorized-root fd instead of re-opening an absolute realpath()
  result (removes the intermediate-symlink swap TOCTOU).
- transport_ssh: always single-quote the server path, including --old-args,
  so no mode can inject shell metacharacters.
- transport_tls: set SSL_OP_NO_COMPRESSION and (guarded) SSL_OP_NO_RENEGOTIATION.
- credentials: reject --password-file/--early-input with any group/other
  permission bit; chmod 0600 the affected test fixtures.
- file_store: export file_store_write_sparse() and remove the verbatim
  file.c duplicate.
2026-09-12 15:01:48 +02:00
TapTap 7e45891257 refactor(cli): dedupe server/client option parsing and tighten CLI tests
- server_cli: handle --password-file/--early-input/--iconv via arg_has_value
  in one place, removing the unreachable duplicate separate-form arms while
  keeping both --opt VALUE and --opt=VALUE working
- client_cli: factor the triplicated --delta-block/--block-size range check
  into set_delta_block_size(); drop the redundant use_metadata assignment
  after identity_parse_copy_as (the parser already forces it)
- tests: cover both spellings of --iconv/--delta-block, make the archive
  short-form test actually call parse_args, add delta-block invalid cases
2026-09-12 14:54:47 +02:00
TapTap b8db810ee5 Merge feat/p8-security: P8 security hardening (daemon ownership opt-in, fake-super gating, no implicit numeric-ids, copy-as fail-fast)
CI / lint (push) Successful in 1m33s
CI / sanitizers (undefined) (push) Successful in 56s
CI / sanitizers (address) (push) Successful in 56s
CI / fuzz-build (push) Successful in 23s
CI / coverage (push) Successful in 47s
CI / valgrind (push) Successful in 40s
CI / build-and-test (push) Successful in 4m59s
2026-09-12 14:33:47 +02:00
TapTap ea0a0e2eaf fix(p8-security): make --copy-as directory ownership airtight; harden tests/logs
- file_ensure_directory_secure() now chowns a final directory it creates under
  --copy-as and fails on error; the symlink parent-creation call site propagates
  it.  The is_dir branch fails when the confined parent cannot be opened under
  --copy-as.  Closes the residual wrong-owner gap for synthesized/symlink
  parent directories.
- file_restore_symlink_metadata() early NULL return is copy-as-aware.
- Preserve errno across the implicit-parent failure cleanup.
- Neutral skip messages (the clamp, not --no-super, may be responsible).
- Daemon copy-as test tolerates the non-root privilege refusal; usage text lists
  --copy-as.
2026-09-12 14:33:42 +02:00
TapTap b216ed31fb fix(p8-security): close review gaps in the ownership gate and copy-as failure propagation
- H3: a daemon module without 'client owner = yes' now also has super-user
  device activity forced off (char/block mknod, --write-devices), so a root
  daemon can no longer be made to create/write raw devices under AUTO.  The
  entries are skipped, preserving ordinary -a pushes.
- H1/H2: propagate a failed required --copy-as chown from symlink metadata
  restore and implicitly-created parent directories, so the entry (and run)
  reports failure instead of a wrong-owner success.
- Docs/help/headers updated for A2/A3 and the device clamp; startup warning
  spells out the client-owner risk.
- Tests: daemon device clamp (skipped without opt-in, created with opt-in),
  updated --super/--fake-super expectations.
2026-09-12 14:18:03 +02:00
TapTap e3840c8326 fix(p8-security): enforce daemon ownership policy, gate fake-super replay, drop implicit numeric-ids, make copy-as failures per-entry
A1: daemon refuses every client-chosen ownership/super-user request
(--numeric-ids/--chown/--usermap/--groupmap/--fake-super/--copy-as/--super)
unless the selected module opts in with 'client owner = yes'.
A2: fake-super owner replay requires an explicit ownership identity policy.
A3: --super no longer implies --numeric-ids (ownership stays opt-in).
A5: a failed --copy-as chown marks the entry failed instead of reporting
success with the wrong owner.
2026-09-12 14:00:55 +02:00
TapTap 58409cee10 test(p7-privilege): skip copy-as refusal test when workspace is not traversable by the unprivileged uid
CI / lint (push) Successful in 1m30s
CI / sanitizers (undefined) (push) Successful in 58s
CI / sanitizers (address) (push) Successful in 58s
CI / fuzz-build (push) Successful in 23s
CI / coverage (push) Successful in 47s
CI / valgrind (push) Successful in 39s
CI / build-and-test (push) Successful in 5m3s
2026-09-12 13:08:22 +02:00
TapTap 549a23993f Merge feat/p7-privilege: Phase 7 Wave E (privilege: --super/--no-super + --copy-as safe subset; PROTOCOL 2.18.0; all rsync rows implemented) 2026-09-12 12:55:55 +02:00
TapTap a5f6899590 docs(p7-privilege): recount Summary to 143/0/4/0/0/0; Wave E shipping note 2026-09-12 12:55:47 +02:00
TapTap 938886829e fix(p7-privilege): close re-review gaps (implicit dir ownership, daemon --super, write-devices gate) 2026-09-12 12:54:45 +02:00
TapTap fdc0f238c9 fix(p7-privilege): harden copy-as/super gates, own dirs/specials
- fake-super owner replay honors --no-super and an active --copy-as
- copy-as/identity ownership now applied to directories and special nodes
- reject copy_as_set && !use_metadata (receiver + client --no-preserve)
- daemon refuses --copy-as; add server-side --no-super operator veto
- implement identity_copy_as_refused/identity_copy_as_active
- reject copy-as ids that overflow int32; escape spec in log errors
- copy-as chown EPERM/EACCES logged at ERROR (still non-fatal)
- identity_wire_valid copy-as bounds; CLI help and RSYNC_COMPAT docs
- add unit tests and root-gated integration coverage
2026-09-12 12:34:43 +02:00
TapTap e6a65d0980 Merge branch 'feat/p7-copy-as' into feat/p7-privilege
# Conflicts:
#	RSYNC_COMPAT.md
#	src/shared/config.c
#	src/shared/config.h
#	src/shared/identity.c
#	tests/integration/test_preflight.py
#	tests/test_client_cli.c
#	tests/test_config.c
2026-09-12 12:14:44 +02:00
TapTap a785ec13c4 feat(p7-super): implement --super/--no-super safe-subset privilege gate (protocol 2.18.0)
Add the receiver-side --super / --no-super tri-state (Config->super_mode)
under the safe-subset + clear-refusal privilege model: FastSync never
elevates privileges, it only permits super-user attempts that are already
confined fd-relative below the authorized receive root.

- identity: privilege_super_permitted() gate (OFF=false, ON=true, AUTO follows
  geteuid()==0); identity_apply_ownership/_link become no-ops when not
  permitted; --super with no explicit identity policy implies raw numeric-id
  preservation (explicit usermap/groupmap/chown/numeric-ids still win); warn
  exactly once when --super is requested by a non-root receiver.
- file_receive: gate char/block device-node creation on the gate; FIFO/socket
  handling is unchanged.
- wire: trailing super_mode int after the --iconv spec, validated 0..2 in
  receive_privilege_options and validate_received_config; PROTOCOL_VERSION
  2.17.0 -> 2.18.0; version-sensitive tests and docs updated.
- CLI: --super/--no-super parsed explicitly before the generic --no-* branch
  (malformed --super=x rejected); usage text added.
- tests: config wire round-trip + invalid-value rejection, privilege-gate mode
  unit test, CLI parse test, integration transfer + root-gated ownership
  suppression/appliance tests.
- docs: RSYNC_COMPAT --super row + Wave E note, protocol mentions, README.
2026-09-12 12:01:19 +02:00
TapTap 80dd64aae6 feat(identity): implement --copy-as USER[:GROUP] safe subset (P7 Wave E)
Force the receiver to apply the requested owner/group to every written
entry through the confined fd-relative identity path instead of switching
the process credentials (unsafe for the multithreaded receiver).  An
unprivileged receiver refuses the transfer up front in server_module_gate,
before STATUS_OK, so no data is written with the wrong ownership.

- new Config fields copy_as_set/copy_as_uid/copy_as_gid + defaults
- identity_parse_copy_as (name/@N/* resolution, primary-gid default,
  gid==uid fallback for numeric ids with no passwd entry); implies -M
- identity snapshot + highest-priority forcing in identity_resolve_targets
- identity_copy_as_refused() helper
- trailing config-frame block (presence int + two int32 ids, >=0 checked)
- PROTOCOL_VERSION 2.17.0 -> 2.18.0; version-sensitive tests updated
- unit tests for parse + wire round-trip/negative-id rejection
- integration TestCopyAs: unprivileged refusal + root chown assertion
- RSYNC_COMPAT.md --copy-as row updated (safe subset + divergence); README
  protocol version refreshed
2026-09-12 11:58:37 +02:00
TapTap f64d252faf docs: recount RSYNC_COMPAT to 141/0/4/0/0/2 after Phase 7 Waves B-D; refresh README short-option note
CI / lint (push) Successful in 1m42s
CI / sanitizers (undefined) (push) Successful in 1m12s
CI / sanitizers (address) (push) Successful in 1m14s
CI / fuzz-build (push) Successful in 23s
CI / coverage (push) Successful in 47s
CI / valgrind (push) Successful in 39s
CI / build-and-test (push) Successful in 4m44s
2026-09-12 11:30:06 +02:00
TapTap 003a5e8f2f Merge feat/p7-times: Phase 7 Wave D (real dir/symlink time preservation making -O/-J meaningful; secluded-args -> Impossible/Divergence; PROTOCOL 2.17.0) 2026-09-12 11:22:23 +02:00
TapTap 9d97e1d3c0 Merge feat/p7-devices: Phase 7 Wave C (device/special statuses; --copy-devices --sendfile non-regular fallback) 2026-09-12 11:21:09 +02:00
TapTap 402e829fef docs(p7-times): re-review nits (chunked STATUS_DIR_TIMES comment; -M/--metadata -> --preserve; -m sink -> -j/--threads) 2026-09-12 11:20:40 +02:00
TapTap 9749c7878c fix(p7-times): dir-time entries only record (never create dirs); bound/chunk dir-time frames; harden list add; docs+tests
Review fixes for Phase 7 Wave D.

#1 (HIGH): STATUS_DIR_TIMES entries no longer create directories. A new
receiver-only File.dir_time_only flag marks dir-time entries; file_save_to_disk_full
short-circuits them as FILE_SAVE_SKIPPED before any device/dir branch, so the sink
still accumulates metadata into the deferred DirTimeList but creates nothing. Empty
source dirs stay untransferred (-a), -m/--prune-empty-dirs semantics are preserved,
and a pre-existing regular file/symlink at an empty-dir mirror path no longer aborts
the transfer. dir_time_list_apply fstatat()s the leaf (AT_SYMLINK_NOFOLLOW) and skips
absent/non-directory paths QUIETLY; only a real existing directory is stamped.
Also initialize File.dir_time_only in file_create() (uninitialised garbage otherwise).

#2 (MED): send_dir_times() chunks entries into repeated STATUS_DIR_TIMES frames of at
most MAX_MANIFEST_ENTRIES, matching the receiver's per-frame bound; the tautological
> INT_MAX check is gone.

#3 (LOW): dir_time_list_add() assigns each grown array right after its realloc (no
dangling) and advances capacity only after both succeed.

#4 (LOW): RSYNC_COMPAT.md -- STATUS_MKDIR carries metadata, dir times are transmitted
via STATUS_DIR_TIMES and applied at the end, empty dirs are still never created; -m
rationale, -O row and Wave D notes updated. Summary counts untouched.

#5 (LOW): integration tests for the three #1 scenarios (empty-dir non-creation under
-a and -a -m, collision non-abort), scanner test now covers empty-dir capture, and
test_file_restore_symlink_metadata asserts the positive apply path when supported.

PROTOCOL_VERSION stays 2.17.0; config-frame layout unchanged.
2026-09-12 11:06:05 +02:00
TapTap 8ca2b74e79 fix(p7-devices): sendfile falls back to buffered read for non-regular sources (--copy-devices no longer hangs); test/doc hardening 2026-09-12 11:00:26 +02:00
TapTap 04514c5b70 Merge feat/p7-output-fs: Phase 7 Wave B (sparse hole preservation, partial retention, fake-super replay, block-size; crtimes/stderr -> Impossible/Divergence) 2026-09-12 10:32:03 +02:00
TapTap a539d8b2ba docs(p7-output-fs): align fake-super comments with silent EPERM/EACCES skip (re-review nit) 2026-09-12 10:31:56 +02:00
TapTap f1a447bb4a feat(p7-times): real directory/symlink time preservation; -O/-J meaningful
Wave D of Phase 7. Make -O/--omit-dir-times and -J/--omit-link-times real by
preserving directory and symlink times, and mark --secluded-args as an explicit
Impossible/Divergence no-op.

Wire: PROTOCOL_VERSION 2.16.0 -> 2.17.0. Adds a terminal STATUS_DIR_TIMES frame
(int count + (wire path, metadata) pairs) sent after all file data and the
optional delete manifest. STATUS_MKDIR also carries metadata for --dirs entries.
Config-frame layout is unchanged.

Sender: the recursive scanner captures every traversed source directory (both
DirectoryScanner and the parallel scanner root + workers, appends mutex-guarded)
into a shared list; the single-threaded and -m paths transmit it last.

Receiver: a DirTimeList accumulates received directory metadata and applies it
with fd-relative no-follow utimensat only at the very end -- after all children,
after the commit-style --delete, and after --delay-updates publication -- in the
single-threaded success frame and in server.c after the -m threads join. -O skips
the application. Symlink metadata is applied at link creation with
utimensat/fchownat/fchmodat AT_SYMLINK_NOFOLLOW; -J suppresses only link times.
identity_apply_ownership_link shares the identity resolver with the fd path.

Docs: -O/-J rows -> Implemented; --secluded-args -> Impossible/Divergence;
--protocol accepted/rejected values and Phase-6/7 notes updated.

Tests: unit (scanner dir capture, DirTimeList apply, symlink metadata, protocol
version values) and integration (dir mtime round-trip + -O, symlink mtime
round-trip + -J, independent suppression), parameterized over single/multithread.
2026-09-12 10:31:20 +02:00
TapTap 227d001092 feat(p7-devices): finalize device/special-file statuses (devices/copy/write -> implemented, specials -> Impossible/Divergence for sockets) + coverage 2026-09-12 10:10:21 +02:00
TapTap f2ba8211ce fix(p7-output-fs): address c-review (fake-super mode sanitization HIGH; sparse/preallocate precedence; partial no_replace guard; stronger tests)
- fake_super_restore_fd now sanitizes mode like metadata_mode (never grants
  S_IWGRP|S_IWOTH; 0666 -> 0644), fixing a privilege regression
- --sparse takes precedence over --preallocate (skip posix_fallocate when
  sparse) so holes are not re-allocated; docs corrected
- --partial retention disabled under --no_replace (ignore/existing) and only
  marks write_attempted after the write begins (no empty-temp retention)
- accept --block-size=SIZE / --delta-block=SIZE inline forms; neutral messages
- fake-super EPERM/EACCES skipped silently (docs aligned); EINVAL still logged
- sparse unit test now memcmp's the full buffer; TestBlockSize integration keeps
  the destination basis so delta is genuinely exercised
- unit 37/37, cppcheck 0, clang-format 0
2026-09-12 09:55:40 +02:00
TapTap 47de05d215 feat(p7-output-fs): sparse hole preservation (-S), partial retention (-P), fake-super replay; block-size verified; crtimes/stderr -> Impossible/Divergence (Wave B)
- write_all_sparse: skips all-zero runs >= 4096 bytes via lseek(SEEK_CUR) and
  ftruncates the final size, wired into the atomic temp+rename and --inplace
  paths with no wire change (full image already in memory).
- --partial retention: on a save failure after the temp held data, rename the
  already-written temp to the destination path (best-effort; falls through to
  unlink; never retains when --partial is off) so --append/--append-verify can
  resume; tested by forcing futimens EINVAL with an out-of-range nsec.
- --block-size aliases --delta-block; verified config->delta_block_size is
  honored by the delta engine end-to-end (unit + integration tests).
- fake_super_restore_fd: parses and re-applies user.fastsync.stat fd-relative
  (fchown best-effort/non-root skipped, fchmod, futimens); a save under
  --fake-super now both records and re-applies.
- -N/--crtimes and --stderr=client promoted to a new 'Impossible/Divergence'
  status bucket (Summary: 136/2/4/3/2 = 147).
- cppcheck/clang-format clean; unit 37/37; integration 407 passed.
2026-09-11 15:58:20 +02:00
TapTap f34eb34f87 Merge feat/p7-cli-namespace: Phase 7 Wave A CLI-namespace parity (renames colliding short flags to rsync parity)
CI / lint (push) Successful in 1m17s
CI / sanitizers (address) (push) Successful in 54s
CI / sanitizers (undefined) (push) Successful in 54s
CI / fuzz-build (push) Successful in 22s
CI / coverage (push) Successful in 46s
CI / valgrind (push) Successful in 38s
CI / build-and-test (push) Successful in 4m30s
2026-09-11 14:44:58 +02:00
TapTap 560ed601f9 docs(p7-cli-namespace): sweep remaining -m-as-multithreading refs to -j/--threads 2026-09-11 14:37:48 +02:00
TapTap d39ddab42c fix(p7-cli-namespace): address c-review (setfacl -m mangled to --threads; two -s tests lost intent; label/doc sweeps)
- revert over-eager replacement of setfacl -m in test_features.py
- use --chunk-serialization (+ set dirs) in the two append/append-verify
  chunk-serialization rejection unit tests so they exercise the real check
- rename archive-negation integration test (--no-preserve is inert under
  archive because devices/specials force metadata)
- update stale (-c)/(-m)/(-s)/(-f) display labels and RSYNC_COMPAT -c/-m refs
- document the --no-perms negation limitation in the Wave A note
2026-09-11 14:34:03 +02:00
TapTap 66e82f9384 feat(p7-cli-namespace): rename colliding short flags to rsync parity (Wave A)
-c -> --checksum, -m -> --prune-empty-dirs, -M -> --remote-option,
-f -> --filter, -s -> --secluded-args, -p -> --perms, -T -> --temp-dir;
-a/--archive is now real rsync -rlptgoD (links+metadata+devices+specials).

FastSync's own flags moved to long-form-only or new shorts:
-j/--threads (multithreading), --preserve (metadata), --sendfile,
--chunk-serialization, --timeout, --ssh-port. Client-side only; the
wire config fields are unchanged (no PROTOCOL_VERSION bump). The server
keeps -p as its port. Docs (README, RSYNC_COMPAT summary 129->132) and
unit/integration tests updated. 37/37 unit, 400-pass integration.
2026-09-11 14:20:15 +02:00
TapTap 15363c0018 docs: add Phase 7 plan (CLI namespace parity, output/filesystem completion, privilege) toward full rsync flag parity 2026-09-11 13:47:44 +02:00
TapTap c1eac6321c docs: recount RSYNC_COMPAT to 129/2 after --batch (Wave D); Phase 6 complete
CI / lint (push) Successful in 1m16s
CI / sanitizers (address) (push) Successful in 57s
CI / sanitizers (undefined) (push) Successful in 55s
CI / fuzz-build (push) Successful in 22s
CI / coverage (push) Successful in 46s
CI / valgrind (push) Successful in 38s
CI / build-and-test (push) Successful in 4m33s
2026-09-10 22:07:10 +02:00
TapTap f43f236f66 Merge feat/p6-batch: residual batch --write-batch/--only-write-batch/--read-batch 2026-09-10 22:06:23 +02:00
TapTap fc2d144492 fix(p6-batch): reject clean EOF on record read; add traversal/EOF security tests 2026-09-10 22:05:18 +02:00
TapTap c026176bb3 feat(p6-batch): residual-batch write/read driver + codec
Implements the client-only residual-batch feature end-to-end:
- src/shared/batch.{c,h}: self-contained single-file batch codec using the
  existing chunk_serialize/chunk_deserialize codec (byte-identical by
  construction).  Magic+format-version header (metadata mode is persisted into
  the header so a batch is self-describing across machines), length-prefixed
  chunk records, bounded reads that reject malformed/truncated/oversized
  records cleanly.
- src/client/client_send.c: write_batch_from_source (deterministic separate
  scan pass, loads every chunk's file images, emits header+records) and
  apply_batch_to_dest (local apply to a destination root via
  file_save_to_disk_full).  No wire change, no server involved.
- src/client/client_validation.c: --write-batch XOR --only-write-batch;
  --read-batch exclusive with both; --read-batch needs only a DEST,
  --only-write-batch only a SOURCE.
- src/client/client_cli.c: main() drives the three batch modes without
  connecting/transferring for read/only-write; --write-batch runs the live
  transfer (single-threaded so the config survives) then emits the batch.
- tests/test_batch.{c,h} (unit: byte-identical roundtrip with and without
  metadata; bad-magic/truncated/oversized rejection) + tests/integration/
  test_batch.py (only-write no-server, read-batch no-source roundtrip,
  --write-batch with a live transfer, conflict rejections).
- clang-format: realign PART-1 config.h comment block.

No PROTOCOL_VERSION bump, no config-frame field, no server flag.
2026-09-10 21:40:07 +02:00
TapTap 4930127312 feat(p6-batch): CLI/config plumbing for --write-batch/--only-write-batch/--read-batch 2026-09-10 21:14:30 +02:00
TapTap 52f45692b9 docs: recount RSYNC_COMPAT to 126/5 after --protocol (Wave C)
CI / lint (push) Successful in 1m15s
CI / sanitizers (undefined) (push) Successful in 54s
CI / sanitizers (address) (push) Successful in 56s
CI / fuzz-build (push) Successful in 21s
CI / coverage (push) Successful in 45s
CI / valgrind (push) Successful in 39s
CI / build-and-test (push) Successful in 4m41s
2026-09-10 20:45:32 +02:00
TapTap 40b0870514 Merge feat/p6-protocol: --protocol version-force flag (client-only) 2026-09-10 20:44:44 +02:00
TapTap 5262cc2597 test(p6-protocol): harden validation null-check; cover missing-arg --protocol 2026-09-10 20:44:33 +02:00
TapTap 9fe6d6c748 test(p6-protocol): unit + integration coverage for --protocol 2026-09-10 20:36:05 +02:00
TapTap 2b7bb2d523 feat(p6-protocol): --protocol version-force flag (client-only) 2026-09-10 20:36:05 +02:00
TapTap bb27b2af50 docs: correct RSYNC_COMPAT selection/update rows to implemented (no code change)
CI / lint (push) Successful in 1m15s
CI / sanitizers (undefined) (push) Successful in 55s
CI / sanitizers (address) (push) Successful in 56s
CI / fuzz-build (push) Successful in 21s
CI / coverage (push) Successful in 44s
CI / valgrind (push) Successful in 38s
CI / build-and-test (push) Successful in 4m43s
2026-09-10 20:16:03 +02:00
TapTap a2ac599298 docs: recount RSYNC_COMPAT to 121/10 after Phase 6 waves A-B (--stop-after/--stop-at, --iconv)
CI / lint (push) Successful in 1m15s
CI / sanitizers (address) (push) Successful in 53s
CI / sanitizers (undefined) (push) Successful in 52s
CI / fuzz-build (push) Successful in 22s
CI / coverage (push) Successful in 45s
CI / valgrind (push) Successful in 38s
CI / build-and-test (push) Successful in 4m40s
2026-09-10 18:18:17 +02:00
TapTap 282aebb7f5 Merge feat/p6-stop: --stop-after/--stop-at deadline stop 2026-09-10 18:16:55 +02:00
TapTap b2afcf2c65 Merge feat/p6-iconv: --iconv charset conversion + PROTOCOL 2.16.0 2026-09-10 18:16:51 +02:00
TapTap 09a07179c9 fix(p6-stop): init -m scan_stopped_early; gate delete warning; stabilize partial-stop test 2026-09-10 18:16:35 +02:00
TapTap cac805b661 fix(p6-iconv): prevent convert buffer overflow; validate both directions; reset on error; more tests 2026-09-10 17:49:29 +02:00
TapTap ac7e9e3bc1 fix(p6-stop): block delete-manifest on early stop; sync -m manifest access; overflow guard 2026-09-10 17:39:34 +02:00
TapTap 7d6665633d test(p6-iconv): iconv unit, config wire-roundtrip, integration tests 2026-09-10 17:13:06 +02:00
TapTap 6e02a24232 feat(p6-iconv): --iconv charset conversion + PROTOCOL 2.16.0 2026-09-10 17:13:01 +02:00
TapTap 37cff96537 test(p6-stop): unit and integration tests for stop deadlines 2026-09-10 16:27:49 +02:00
TapTap 24d1448246 feat(p6-stop): --stop-after/--stop-at deadline transfer stop 2026-09-10 16:27:49 +02:00
TapTap 6bb63c6f21 docs: recount RSYNC_COMPAT to 118/13 after daemon Wave C (--no-motd + MOTD)
CI / lint (push) Successful in 1m14s
CI / sanitizers (address) (push) Successful in 55s
CI / sanitizers (undefined) (push) Successful in 53s
CI / fuzz-build (push) Successful in 21s
CI / coverage (push) Successful in 45s
CI / valgrind (push) Successful in 38s
CI / build-and-test (push) Successful in 5m42s
2026-09-10 13:59:15 +02:00
TapTap 61eb08023a Merge feat/d5-daemon-motd: daemon MOTD display + --no-motd 2026-09-10 13:59:07 +02:00
TapTap 0c35cf2b82 test(d5-daemon-motd): assert no banner when no config key 2026-09-10 13:58:21 +02:00
TapTap cf5f730940 feat(d5-daemon-motd): daemon MOTD display + --no-motd 2026-09-10 13:52:24 +02:00
TapTap 8330f275a6 docs: recount RSYNC_COMPAT to 117/14 after daemon Wave B auth (--password-file, --early-input)
CI / lint (push) Successful in 1m13s
CI / sanitizers (undefined) (push) Successful in 54s
CI / sanitizers (address) (push) Successful in 56s
CI / fuzz-build (push) Successful in 20s
CI / coverage (push) Successful in 44s
CI / valgrind (push) Successful in 38s
CI / build-and-test (push) Successful in 5m10s
2026-09-10 13:27:39 +02:00
TapTap d3c9f1d218 Merge feat/d5-daemon-auth: daemon password auth (--password-file, --early-input) 2026-09-10 13:27:27 +02:00
TapTap accd34ad60 fix(d5-daemon-auth): address auth review findings (Wave B)
- test_credentials.c: NUL-terminate the overlong-line stack buffer before
  make_tmp_file's strlen() (was a stack-buffer-overflow READ under ASan);
  still exercises the overlong-rejection path.
- Add redacted protocol string variants (protocol_send_str_redacted /
  receive + fd send_str_redacted/receive_str_redacted) and use them for the
  daemon auth username/digest so --verbose / LOG_DEBUG_ALL never logs a
  replayable credential while other protocol strings keep their debug trace.
- credentials_verify/gate: replace byte-wise-short-circuiting strcmp with a
  fixed-length constant-time username compare (closes user-enumeration oracle);
  update doc comment to match.
- read_secret_file: preserve password exact bytes (only strip trailing CR/LF)
  and burn the stack line buffer; document the whitespace behavior.
- test_server_cli.c: note the parser zero-inits opts on failure.
- Add debug-level daemon test asserting the digest never appears under --verbose.

PROTOCOL_VERSION stays 2.15.0.
2026-09-10 13:20:42 +02:00
TapTap dd5ae60459 feat(d5-daemon-auth): password auth, --password-file, --early-input 2026-09-10 12:50:56 +02:00
TapTap 958eddf414 docs: recount RSYNC_COMPAT to 115/16 after daemon Wave A (--daemon/--config/--dparam/--no-detach)
CI / lint (push) Successful in 1m12s
CI / sanitizers (undefined) (push) Successful in 54s
CI / sanitizers (address) (push) Successful in 55s
CI / fuzz-build (push) Successful in 22s
CI / coverage (push) Successful in 46s
CI / valgrind (push) Successful in 38s
CI / build-and-test (push) Successful in 4m41s
2026-09-09 18:40:06 +02:00
TapTap 8440dbfdb8 Merge feat/d5-daemon-core: daemon lifecycle, module config, ::dest (PROTOCOL 2.15.0) 2026-09-09 18:39:37 +02:00
TapTap 7c55409a6b fix(d5-daemon-core): cppcheck const-correctness, wire module length cap, daemonize chdir/umask, daemon confinement tests
- daemon_conf.c/server.c/test_daemon_conf.c: const-qualify parse/loop pointers;
  scope user_path static inside its block (clears the 9-wave-A cppcheck findings)
- config.c receive_daemon_module: reject invalid/over-long wire module names
  (> DAEMON_MAX_MODULE_NAME) with a clean STATUS_ERROR; client side already
  enforced via daemon_module_name_valid in config_parse_daemon_dest
- server.c daemonize: chdir(/) and umask(0) so module paths resolve from /
  and config-requested file modes are honored; PROTOCOL_VERSION stays 2.15.0
- test_daemon.py: confinement (read-only/unknown no-write anywhere), module-less
  and dot-dot destination refusal, real daemon_detach double-fork path
2026-09-09 18:30:52 +02:00
TapTap b3d7d64347 feat(d5-daemon-core): daemon lifecycle, module config, ::dest, PROTOCOL 2.15.0 2026-09-09 17:48:07 +02:00
TapTap 9d17e951f1 docs: recount RSYNC_COMPAT to 111/20 after Phase 5 waves A-C; clang-format 18 reflow
CI / lint (push) Successful in 1m9s
CI / sanitizers (undefined) (push) Successful in 53s
CI / sanitizers (address) (push) Successful in 55s
CI / fuzz-build (push) Successful in 19s
CI / coverage (push) Successful in 43s
CI / valgrind (push) Successful in 37s
CI / build-and-test (push) Successful in 4m28s
2026-09-09 14:37:52 +02:00
TapTap 1e02ebcd32 Merge feat/p5-remote-option: --remote-option, --trust-sender
# Conflicts:
#	RSYNC_COMPAT.md
#	src/client/client_cli.c
#	src/client/client_send.c
#	src/shared/transport_ssh.c
#	src/shared/transport_ssh.h
#	tests/integration/test_ssh.py
#	tests/test_client_cli.c
#	tests/test_transport_ssh.c
2026-09-09 14:35:26 +02:00
TapTap b1c63ff947 Merge feat/p5-socket: --address, -4/-6, --sockopts, server bind options
# Conflicts:
#	RSYNC_COMPAT.md
#	tests/test_client_cli.c
2026-09-09 14:30:36 +02:00
TapTap 35f4297538 Merge feat/p5-rsh: --rsh/-e, --rsync-path, --blocking-io, --outbuf 2026-09-09 14:29:58 +02:00
TapTap f86ba7a556 fix(p5-socket): clang-format 18 reflow + cppcheck const-correctness 2026-09-09 14:29:46 +02:00
TapTap cdcaf21acd fix(p5-rsh): NULL-check argv tail str_dups in ssh_build_client_argv 2026-09-09 14:29:46 +02:00
TapTap ad228db915 fix(p5-remote-option): wire server --trust-sender, align save-layer gates, add hostile-sender test 2026-09-09 14:23:59 +02:00
TapTap 07d1dd84f7 feat(p5-socket): --address, -4/-6, --sockopts, server bind options 2026-09-09 13:41:28 +02:00
TapTap 858af3d63d feat(p5-rsh): --rsh/-e, --rsync-path, --blocking-io, --outbuf 2026-09-09 13:41:28 +02:00
TapTap 5e79d7d76b feat(p5-remote-option): --remote-option (probe 2.14.0), --trust-sender 2026-09-09 13:41:28 +02:00
TapTap 90ccf297d7 style: cppcheck — const-correctness in file_send_special and test_chunk
CI / lint (push) Successful in 1m7s
CI / sanitizers (address) (push) Successful in 53s
CI / sanitizers (undefined) (push) Successful in 53s
CI / fuzz-build (push) Successful in 21s
CI / coverage (push) Successful in 44s
CI / valgrind (push) Successful in 37s
CI / build-and-test (push) Successful in 4m41s
file_send_special now takes const File*; test_chunk declares the deserialized
Chunk* const.  cppcheck (--error-exitcode=1) now exits clean; clang-format
clean; unit 29/29.
2026-09-08 22:53:47 +02:00
TapTap 5a00a4fe2b style: cppcheck — drop redundant if(fd>=0) in test_xattr copy-fallback test
CI / lint (push) Failing after 1m6s
CI / build-and-test (push) Skipped
CI / sanitizers (address) (push) Skipped
CI / sanitizers (undefined) (push) Skipped
CI / fuzz-build (push) Skipped
CI / coverage (push) Skipped
CI / valgrind (push) Skipped
EXPECT_TRUE(fd >= 0) already guards; the enclosing if is flagged always-true
by cppcheck. read(-1) is safe.
2026-09-08 22:46:43 +02:00
TapTap 743b00ffdd docs: recount RSYNC_COMPAT summary after Phase-4 wave C (symlink-trust + devices + acl/xattr)
CI / lint (push) Failing after 1m8s
CI / build-and-test (push) Skipped
CI / sanitizers (address) (push) Skipped
CI / sanitizers (undefined) (push) Skipped
CI / fuzz-build (push) Skipped
CI / coverage (push) Skipped
CI / valgrind (push) Skipped
Wave C moved 12 rows: -l/--links now real (was Partial). New Implemented (7):
--munge-links, -k/--copy-dirlinks, -K/--keep-dirlinks, -D, -A/--acls, -X/--xattrs
(links was Partial->Implemented). New Partial (5): --devices, --specials,
--copy-devices, --write-devices, --fake-super. Summary: Implemented 94->101,
Partial 6->10, Not Implemented 41->30 (Total 147).

Final Phase-4 summary: 101/3/10/3/30 = 147; PROTOCOL_VERSION 2.13.0.
2026-09-08 22:42:02 +02:00
TapTap 94b6662018 Merge feat/p4-acl-xattr: -X/-A/--fake-super
# Conflicts:
#	src/shared/config.c
#	src/shared/file.c
#	src/shared/file_types.h
#	tests/integration/test_features.py
#	tests/test_client_cli.c
#	tests/test_config.c
2026-09-08 22:38:15 +02:00
TapTap 5bbdc5b450 Merge feat/p4-devices
# Conflicts:
#	src/client/client_send.c
#	src/server/receiver.c
#	src/shared/chunk.c
#	src/shared/file.c
#	src/shared/file_receive.c
#	src/shared/file_receive.h
#	src/shared/file_types.h
#	src/shared/protocol.h
#	tests/test_chunk.c
2026-09-08 22:33:52 +02:00
TapTap e600b56f10 Merge feat/p4-symlink-trust 2026-09-08 22:27:56 +02:00
TapTap 747946c318 acls/xattrs: -X/--xattrs, -A/--acls, --fake-super
CI / lint (pull_request) Failing after 55s
CI / build-and-test (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
New src/shared/xattr.{c,h}: capture user.* + POSIX ACL xattrs, transmit a bounded
per-file block, re-apply fd-relative. security.*/trusted.*/other system.* never
transmitted/applied (receiver re-validates). Bounds: name<=255 value<=1MiB count
<=256 total<=4MiB. --fake-super records uid:gid:mode:mtime in reserved
user.fastsync.stat (receiver-only). PROTOCOL_VERSION 2.12.0->2.13.0. Review
fixes: reserved key not forwardable, link/hardlink copy-fallback preserves
xattrs, no const-param mutation, per-file warning dedup.
2026-09-08 22:27:53 +02:00
TapTap 007e8f90f2 devices: --devices/--specials/-D/--copy-devices/--write-devices
CI / lint (pull_request) Failing after 56s
CI / build-and-test (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
Recreate char/block nodes via mknodat (privilege-gated, EPERM->warn+skip) and
FIFOs via mkfifoat; new STATUS_SPECIAL frame + validated rdev; sockets skipped;
-copy-devices copies st_size; -write-devices O_NOFOLLOW+O_NONBLOCK warn+skip.
preserve_specials/copy_devices/write_devices cross the wire. PROTOCOL_VERSION
2.12.0->2.13.0. Review fixes: -m source-removal keeps recreated specials, FIFO
ENXIO skip, rdev bounds at chunk_deserialize, STATUS_ERROR on receive branch,
scanner_prepare_special dedup.
2026-09-08 22:27:53 +02:00
TapTap 820188c2cc symlink-trust: -k/--copy-dirlinks, -K/--keep-dirlinks, --munge-links (+real -l/--links)
CI / lint (pull_request) Successful in 1m1s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / build-and-test (pull_request) Successful in 1m19s
Adds symlink-target transmission (File is_symlink+symlink_target, STATUS_SYMLINK
frame, chunk type 2), munge-links sender containment + receiver-side symmetric
target containment, keep-dirlinks confined dir-symlink following (O_NOFOLLOW
realpath-rechecked), and fixes -l to copy symlinks as symlinks. munge_links +
keep_dirlinks cross the wire; copy_dirlinks client-only. PROTOCOL_VERSION
2.12.0->2.13.0. Review fixes: receiver rejects absolute/.. targets, gated unmunge,
-K O_NOFOLLOW+re-fstat, keep_dirlinks set once at config-accept, rel_buf overflow
fails the walk.
2026-09-08 22:27:53 +02:00
TapTap 0de859b302 docs: recount RSYNC_COMPAT summary after Phase-4 wave B (metadata times + hard links)
CI / lint (push) Successful in 53s
CI / sanitizers (address) (push) Successful in 49s
CI / sanitizers (undefined) (push) Successful in 49s
CI / fuzz-build (push) Successful in 20s
CI / coverage (push) Successful in 43s
CI / valgrind (push) Successful in 37s
CI / build-and-test (push) Successful in 4m28s
Wave B moved 6 rows: -U/--atimes, --open-noatime, -H/--hard-links -> ✅;
-N/--crtimes -> ⚠️; -O/--omit-dir-times, -J/--omit-link-times -> 🔄 (no-ops).
Summary: ✅91->94, ⚠️5->6, 🔄1->3, ❌47->41 (Total 147).
2026-09-08 21:02:34 +02:00
TapTap 4e7e84f947 Merge feat/p4-metadata-capture: atimes/crtimes/open-noatime/omit-dir-times/omit-link-times
# Conflicts:
#	src/shared/config.c
#	tests/integration/test_features.py
2026-09-08 21:00:07 +02:00
TapTap cf7cc5b8ca Merge feat/p4-hard-links: -H/--hard-links 2026-09-08 20:59:04 +02:00
TapTap e80888ce7b metadata times: -U/--atimes, -N/--crtimes, --open-noatime, -O/-J
CI / lint (pull_request) Successful in 52s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / build-and-test (pull_request) Successful in 1m19s
Capture+transmit source atime (pre-read stat; O_NOATIME sender guard) and birth
time (statx STATX_BTIME); receiver restores atime with mtime (crtime not settable
portably -> transmitted, explicitly not applied). --open-noatime is client-only.
-O/-J documented as accepted no-ops (FastSync never preserves dir/symlink times).
Wire: metadata frame gains atime/crtime val+sec+nsec; PROTOCOL_VERSION
2.11.0->2.12.0. Review fixes: gate atime capture to Linux (no epoch clobber on
non-Linux), close fd on fdopen failure, honest -O/-J status (Compat no-op).
2026-09-08 20:58:56 +02:00
TapTap f891cd0a6a hard-links: -H/--hard-links preserves inode relationships
CI / lint (pull_request) Successful in 50s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / build-and-test (pull_request) Successful in 1m19s
Source files sharing (st_dev,st_ino) are recreated as hard links on the
destination; only the first member's data crosses the wire (siblings ride a
payload-less STATUS_HARDLINK frame). Ordering requires the single-FIFO-writer
receiver + forced sequential scan (documented). link()-failure falls back to a
byte-identical local copy. Rejects -s/--append. PROTOCOL_VERSION 2.11.0->2.12.0.
Review fixes: delete the dead HardLinkRegistry (ordering holds by FIFO writer),
and --existing no longer aborts when the first member is absent but the sibling
exists (leaves the sibling in place).
2026-09-08 20:58:56 +02:00
TapTap cbb09e41ab identity: fix use-after-free in --usermap/--groupmap parse error path
CI / lint (push) Successful in 48s
CI / sanitizers (undefined) (push) Successful in 53s
CI / sanitizers (address) (push) Successful in 53s
CI / fuzz-build (push) Successful in 18s
CI / coverage (push) Successful in 42s
CI / valgrind (push) Successful in 37s
CI / build-and-test (push) Successful in 4m37s
identity_parse_map freed the str_dup'd list before logging the offending rule
( points into that buffer), causing an invalid read caught by CI valgrind
(MSAN/MSAN-style; the ONLY definite valgrind error in the suite).  Log before
freeing.  valgrind now reports 0 errors / 0 definite leaks in both the parent
and the forked wire-roundtrip child.
2026-09-08 18:49:23 +02:00
TapTap 30048dddef style: clang-format 18 (CI lint) on Wave-A sources
CI / lint (push) Successful in 48s
CI / sanitizers (address) (push) Successful in 47s
CI / sanitizers (undefined) (push) Successful in 48s
CI / fuzz-build (push) Successful in 19s
CI / coverage (push) Successful in 44s
CI / valgrind (push) Failing after 37s
CI / build-and-test (push) Successful in 4m36s
Reflow usage.c, config.c, file.c, test_client_cli.c, test_file.c so the
clang-format check passes.  Whitespace-only; no behavior change.
2026-09-08 18:38:50 +02:00
TapTap ba82a0b6da docs: recount RSYNC_COMPAT summary after Phase-4 wave A (identity + preallocate)
CI / lint (push) Failing after 5s
CI / build-and-test (push) Skipped
CI / sanitizers (address) (push) Skipped
CI / sanitizers (undefined) (push) Skipped
CI / fuzz-build (push) Skipped
CI / coverage (push) Skipped
CI / valgrind (push) Skipped
5 rows moved ❌ -> ✅: --numeric-ids/--usermap/--groupmap/--chown and
--preallocate.  Summary: ✅86 -> ✅91, ❌52 -> ❌47 (Total 147 unchanged).
2026-09-08 18:37:45 +02:00
TapTap b283c8084c identity: keep --numeric-ids in the activate set (-M --numeric-ids)
CI / lint (push) Failing after 4s
CI / build-and-test (push) Skipped
CI / sanitizers (address) (push) Skipped
CI / sanitizers (undefined) (push) Skipped
CI / fuzz-build (push) Skipped
CI / coverage (push) Skipped
CI / valgrind (push) Skipped
identity_active_enabled() only gates identity_apply_ownership, which runs only
when metadata is present, so --numeric-ids must stay in the set: combined with
-M it activates raw-id application, while a standalone --numeric-ids (no
ownership-affecting flag) carries no metadata and correctly stays inert.  My
earlier review fix removed it and broke 'owner not applied' for -M --numeric-ids
(uid 0 instead of the source ids).  Revert that removal.
2026-09-08 18:37:07 +02:00
TapTap 279fc8468a Merge feat/p4-preallocate: --preallocate
# Conflicts:
#	tests/test_client_cli.c
#	tests/test_config.c
2026-09-08 18:25:49 +02:00
TapTap 8defaf5e8d Merge feat/p4-identity-mapping: --numeric-ids / --usermap / --groupmap / --chown 2026-09-08 18:23:52 +02:00
TapTap 362a6a5488 preallocate: --preallocate allocates dest space up front
CI / lint (pull_request) Failing after 3s
CI / build-and-test (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
Receiver allocates the destination file's full size before streaming data
(posix_fallocate preferred, ftruncate fallback on EOPNOTSUPP/ENOSYS) so an
out-of-space transfer fails fast instead of partway. Additive config bool
crossing the wire; PROTOCOL_VERSION 2.10.0 -> 2.11.0. Threaded through all
store paths (atomic, inplace, partial/delay-updates staging, link-dest copy
fallback). Review hardening: explicit lseek(0) before the data write so
correctness does not depend on posix_fallocate leaving the fd offset unchanged.
2026-09-08 18:23:48 +02:00
TapTap 53ce00b830 identity mapping: --numeric-ids / --usermap / --groupmap / --chown
CI / lint (pull_request) Failing after 3s
CI / build-and-test (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
Receiver-side ownership application, opt-in and privilege-gated:
- OFF for every existing transfer (plain -M/--preserve still never applies
  ownership); only triggers on an explicit identity flag + receiver permission.
- EPERM/EACCES warn-and-continue (never aborts); other fchown errors escalate.
- fd-relative fchown after the file is written (symlink-safe, confined).
- New src/shared/identity.{c,h}; config fields numeric_ids / chown uid/gid /
  usermap + groupmap id-pair tables cross the wire; PROTOCOL_VERSION 2.10.0
  -> 2.11.0. CLI in client_cli.c; per-connection snapshot in server.c.
- Review fixes: EPERM/EACCES-only warn-and-continue, prominent root-receiver
  notice, identity_clear_active on early server error paths, --numeric-ids
  kept inert standalone (removed from activation trigger set).
2026-09-08 18:23:43 +02:00
201 changed files with 54374 additions and 5707 deletions

No files matched your search

+12 -12
View File
@@ -9,10 +9,10 @@ on:
jobs:
lint:
runs-on: ubuntu-latest
container: gitea.tap-tap.win/taptap/fastsync-ci:v10
container: gitea.tap-tap.win/taptap/fastsync-ci:v11
steps:
- name: Checkout
uses: actions/checkout@v4
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: clang-format check
run: find src/ tests/ -name '*.c' -o -name '*.h' | xargs clang-format --dry-run --Werror
@@ -26,11 +26,11 @@ jobs:
# suite) run on merge to dev/main, so PR CI stays well under ~3 minutes.
build-and-test:
runs-on: ubuntu-latest
container: gitea.tap-tap.win/taptap/fastsync-ci:v10
container: gitea.tap-tap.win/taptap/fastsync-ci:v11
needs: lint
steps:
- name: Checkout
uses: actions/checkout@v4
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Configure
run: cmake -B build -S . -DSTRICT_WARNINGS=ON
@@ -51,7 +51,7 @@ jobs:
sanitizers:
runs-on: ubuntu-latest
container: gitea.tap-tap.win/taptap/fastsync-ci:v10
container: gitea.tap-tap.win/taptap/fastsync-ci:v11
needs: lint
if: github.event_name == 'push'
strategy:
@@ -59,7 +59,7 @@ jobs:
sanitizer: [address, undefined]
steps:
- name: Checkout
uses: actions/checkout@v4
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Configure
run: cmake -B build-${{ matrix.sanitizer }} -S . -DSANITIZER=${{ matrix.sanitizer }}
@@ -72,12 +72,12 @@ jobs:
fuzz-build:
runs-on: ubuntu-latest
container: gitea.tap-tap.win/taptap/fastsync-ci:v10
container: gitea.tap-tap.win/taptap/fastsync-ci:v11
needs: lint
if: github.event_name == 'push'
steps:
- name: Checkout
uses: actions/checkout@v4
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Configure (clang + fuzz)
run: CC=clang CXX=clang++ cmake -B build-fuzz -S . -DENABLE_FUZZ=ON
@@ -94,12 +94,12 @@ jobs:
coverage:
runs-on: ubuntu-latest
container: gitea.tap-tap.win/taptap/fastsync-ci:v10
container: gitea.tap-tap.win/taptap/fastsync-ci:v11
needs: lint
if: github.event_name == 'push'
steps:
- name: Checkout
uses: actions/checkout@v4
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Configure
run: cmake -B build -S . -DENABLE_COVERAGE=ON
@@ -118,12 +118,12 @@ jobs:
valgrind:
runs-on: ubuntu-latest
container: gitea.tap-tap.win/taptap/fastsync-ci:v10
container: gitea.tap-tap.win/taptap/fastsync-ci:v11
needs: lint
if: github.event_name == 'push'
steps:
- name: Checkout
uses: actions/checkout@v4
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Configure
run: cmake -B build -S . -DSTRICT_WARNINGS=ON
+4
View File
@@ -8,3 +8,7 @@ build-*/
build2/
build3/
build_docker2/
# Test/run artifacts
root/
test_partial_install_tmp/
+1 -1
View File
@@ -128,7 +128,7 @@ Do not wait for the user to tell you CI failed — check proactively. The user s
## Branch Strategy
Never push directly to `main`. All changes must be developed on a feature branch and merged via a pull request. Always create a new branch (`git checkout -b <branch-name>`) before making changes, push it, and open a PR with `gh pr create --fill`. Wait for CI to pass before merging.
Never push directly to `dev` or `main`. All changes must be developed on a feature branch and merged via a pull request targeting `dev`. Create a branch (`git checkout -b <branch-name>`), push it, and open the PR with `tea pr create --repo TapTap/FastSync --base dev --head <branch-name>`. Wait for CI to pass before merging.
## Dependency Installation
+1 -1
View File
@@ -92,7 +92,7 @@ When using `tea` (the task execution agent) to run CI or tests, always set a suf
## Branch Strategy
Never push directly to `main`. All changes must be developed on a feature branch and merged via a pull request. Always create a new branch (`git checkout -b <branch-name>`) before making changes, push it, and open a PR with `gh pr create --fill`. Wait for CI to pass before merging.
Never push directly to `dev` or `main`. All changes must be developed on a feature branch and merged via a pull request targeting `dev`. Create a branch (`git checkout -b <branch-name>`), push it, and open the PR with `tea pr create --repo TapTap/FastSync --base dev --head <branch-name>`. Wait for CI to pass before merging.
## Dependency Installation
+34 -17
View File
@@ -27,16 +27,19 @@ FetchContent_Declare(xxhash GIT_REPOSITORY https://github.com/Cyan4973/xxHash GI
FetchContent_MakeAvailable(xxhash)
# Sanitizer option
set(SANITIZER "none" CACHE STRING "Sanitizer to enable (address, thread, none)")
set_property(CACHE SANITIZER PROPERTY STRINGS address thread none)
set(SANITIZER "none" CACHE STRING "Sanitizer to enable (address, thread, undefined, none)")
set_property(CACHE SANITIZER PROPERTY STRINGS address thread undefined none)
if(SANITIZER STREQUAL "address")
add_compile_options(-fsanitize=address -fno-omit-frame-pointer -g)
add_link_options(-fsanitize=address)
elseif(SANITIZER STREQUAL "thread")
add_compile_options(-fsanitize=thread -fno-omit-frame-pointer -g)
add_link_options(-fsanitize=thread)
elseif(SANITIZER STREQUAL "undefined")
add_compile_options(-fsanitize=undefined -fno-omit-frame-pointer -g)
add_link_options(-fsanitize=undefined)
elseif(NOT SANITIZER STREQUAL "none")
message(FATAL_ERROR "Unknown sanitizer: ${SANITIZER}. Supported values: address, thread, none")
message(FATAL_ERROR "Unknown sanitizer: ${SANITIZER}. Supported values: address, thread, undefined, none")
endif()
option(STRICT_WARNINGS "Enable strict warnings" OFF)
@@ -52,6 +55,16 @@ if(NOT ZSTD_LIBRARY)
message(FATAL_ERROR "zstd library not found. Ensure it is in your nix-shell!")
endif()
find_library(ZLIB_LIBRARY z)
if(NOT ZLIB_LIBRARY)
message(FATAL_ERROR "zlib library not found. Ensure zlib1g-dev / nix zlib is available!")
endif()
find_library(LZ4_LIBRARY lz4)
if(NOT LZ4_LIBRARY)
message(FATAL_ERROR "lz4 library not found. Ensure liblz4-dev / nix lz4 is available!")
endif()
find_package(OpenSSL REQUIRED)
file(GLOB SHARED_SRCS "src/shared/*.c")
@@ -61,15 +74,15 @@ file(GLOB TEST_SRCS "tests/*.c")
add_executable(server ${SERVER_SRCS} ${SHARED_SRCS})
target_include_directories(server PRIVATE src/shared src/server src/client)
target_link_libraries(server PRIVATE Threads::Threads ${ZSTD_LIBRARY} OpenSSL::SSL OpenSSL::Crypto xxhash)
target_link_libraries(server PRIVATE Threads::Threads ${ZSTD_LIBRARY} ${ZLIB_LIBRARY} ${LZ4_LIBRARY} OpenSSL::SSL OpenSSL::Crypto xxhash)
add_executable(client ${CLIENT_SRCS} ${SHARED_SRCS})
target_include_directories(client PRIVATE src/shared src/server src/client)
target_link_libraries(client PRIVATE Threads::Threads ${ZSTD_LIBRARY} OpenSSL::SSL OpenSSL::Crypto xxhash)
target_link_libraries(client PRIVATE Threads::Threads ${ZSTD_LIBRARY} ${ZLIB_LIBRARY} ${LZ4_LIBRARY} OpenSSL::SSL OpenSSL::Crypto xxhash)
add_executable(tests ${TEST_SRCS} ${SHARED_SRCS} src/client/scanner.c)
target_include_directories(tests PRIVATE tests src/shared src/server src/client)
target_link_libraries(tests PRIVATE Threads::Threads ${ZSTD_LIBRARY} OpenSSL::SSL OpenSSL::Crypto xxhash)
target_link_libraries(tests PRIVATE Threads::Threads ${ZSTD_LIBRARY} ${ZLIB_LIBRARY} ${LZ4_LIBRARY} OpenSSL::SSL OpenSSL::Crypto xxhash)
```
### Source Layout
@@ -82,19 +95,25 @@ tests/integration/ — Python pytest integration tests
```
### Dependencies
- **zstd** — found via `find_library(ZSTD_LIBRARY zstd)`
- **zstd** — found via `find_library(ZSTD_LIBRARY zstd)` (default compression codec)
- **zlib** — found via `find_library(ZLIB_LIBRARY z)` (the `zlib`/`zlibx` codecs)
- **lz4** — found via `find_library(LZ4_LIBRARY lz4)` (the `lz4` codec)
- **OpenSSL** — found via `find_package(OpenSSL REQUIRED)` (TLS 1.2+ transport)
- **xxHash** — fetched via `FetchContent` from GitHub (delta transfer hashing, v0.8.3)
- **xxHash** — fetched via `FetchContent` from the upstream repository (delta transfer hashing, v0.8.3)
- **pthreads** — found via `find_package(Threads REQUIRED)`
- **C11 standard** — required
- **CMake 3.22+** — minimum version
The codec matrix (protocol 2.26.0) uses zstd/zlib/lz4 for compression and
xxHash/OpenSSL for the `xxh128`/`xxh3`/`xxh64`/`md5`/`md4`/`sha1` checksums
(`none` needs no library); both codec families are negotiated per transfer.
## Conventions
- Use `file(GLOB ...)` for source collection (existing pattern).
- All targets link `Threads::Threads`, `${ZSTD_LIBRARY}`, `OpenSSL::SSL`, `OpenSSL::Crypto`, and `xxhash`.
- All targets link `Threads::Threads`, `${ZSTD_LIBRARY}`, `${ZLIB_LIBRARY}`, `${LZ4_LIBRARY}`, `OpenSSL::SSL`, `OpenSSL::Crypto`, and `xxhash`.
- Include directories: `src/shared`, `src/server`, `src/client`, `tests` (for test target).
- Sanitizer support: pass `-DSANITIZER=address` or `-DSANITIZER=thread` to cmake (live option in CMakeLists.txt).
- Sanitizer support: pass `-DSANITIZER=address`, `-DSANITIZER=thread`, or `-DSANITIZER=undefined` to cmake (live option in CMakeLists.txt).
- Build with `cmake -B build -S . && cmake --build build -j$(nproc)`.
- For CI, dependencies are provided by the project's custom Docker image (repo-root `Dockerfile`, same image CI uses). For local development, use `nix-shell`. Never add `apt-get install` / `pip install` to CI workflows. See `AGENTS.md`.
@@ -105,7 +124,7 @@ tests/integration/ — Python pytest integration tests
3. Add new dependencies with `find_package` or `find_library`.
4. When adding a new executable target, follow the pattern of existing targets.
5. When adding a new library (static/shared), use `add_library` and follow the project's naming.
6. For sanitizer builds, pass `-DSANITIZER=address` or `-DSANITIZER=thread` to cmake (matching CI's matrix strategy).
6. For sanitizer builds, pass `-DSANITIZER=address`, `-DSANITIZER=thread`, or `-DSANITIZER=undefined` to cmake (matching CI's matrix strategy).
7. Always verify the build compiles after changes.
## Sanitizer Configurations
@@ -119,11 +138,9 @@ cmake -B build -S . -DSANITIZER=thread # ThreadSanitizer (race conditions)
cmake --build build -j$(nproc)
```
For UndefinedBehaviorSanitizer (no `-DSANITIZER=undefined` option in CMakeLists.txt yet), use the manual flag approach:
UndefinedBehaviorSanitizer uses the same built-in option:
```bash
cmake -B build -S . \
-DCMAKE_C_FLAGS="-fsanitize=undefined -fno-omit-frame-pointer -g" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=undefined"
cmake -B build -S . -DSANITIZER=undefined
cmake --build build -j$(nproc)
```
@@ -159,7 +176,7 @@ cmake -B build -S . -DCMAKE_BUILD_TYPE=RelWithDebInfo
```bash
cmake -B build -S .
cmake --build build -j$(nproc)
./build/server
./build/server -p 8080 --allow-unauthenticated
./build/client
./build/tests
```
@@ -187,7 +204,7 @@ When using `tea` (the task execution agent) to run CI or tests, always set a suf
## Branch Strategy
Never push directly to `main`. All changes must be developed on a feature branch and merged via a pull request. Always create a new branch (`git checkout -b <branch-name>`) before making changes, push it, and open a PR with `gh pr create --fill`. Wait for CI to pass before merging.
Never push directly to `dev` or `main`. All changes must be developed on a feature branch and merged via a pull request targeting `dev`. Create a branch (`git checkout -b <branch-name>`), push it, and open the PR with `tea pr create --repo TapTap/FastSync --base dev --head <branch-name>`. Wait for CI to pass before merging.
## Dependency Installation
+5 -5
View File
@@ -27,7 +27,7 @@ FastSync is a file synchronization tool (like rsync, but faster). It transfers f
cmake -B build -S . && cmake --build build -j$(nproc)
# Server (TCP mode)
./build/server
./build/server -p 8080 --allow-unauthenticated
# Client (TCP mode)
./build/client --source-dir /path/to/send --dest-dir /path/to/receive --save-to-disk
@@ -37,13 +37,13 @@ cmake -B build -S . && cmake --build build -j$(nproc)
# Run tests
./build/tests # unit tests
python3 test.py # integration tests
python3 -m pytest tests/integration/ -n 4 --dist=load -m "not setpriv" # integration tests
```
## Code Walkthrough
### Client Entry Point (`src/client/client_cli.c`)
- Parses CLI arguments using `getopt_long`
- Parses CLI arguments using a custom option-table parser (`OPTION_TABLE` in `src/client/client_cli.c`); there is no `getopt*` usage
- Creates `Config` struct with all options
- Detects SSH destinations (contains `:`)
- Calls into `client_send.c` for the actual transfer
@@ -109,7 +109,7 @@ Collection of files for batch transfer. Serialized with file count, then per-fil
zstd streaming compression via `ZSTD_compressStream2`/`ZSTD_decompressStream`. Compression happens per-chunk in the sender stage. Level 1-22 (default 5). Streaming means memory usage stays bounded regardless of file size.
### "How does sendfile() work?"
On Linux, `sendfile()` copies data directly from kernel file buffer to socket, bypassing userspace. ~2x faster for large files. Enabled with `-f` flag. Only works with TCP (not SSH, not compression).
On Linux, `sendfile()` copies data directly from kernel file buffer to socket, bypassing userspace. ~2x faster for large files. Enabled with `--sendfile` (long form only). Only works with TCP (not SSH, not compression).
### "How does incremental sync work?"
Client sends file metadata (path, size, mtime) to server. Server checks if destination file has same size+mtime. If match, server responds `STATUS_OK` (skip). If mismatch, server responds `STATUS_NEXT` (send).
@@ -138,7 +138,7 @@ When using `tea` (the task execution agent) to run CI or tests, always set a suf
## Branch Strategy
Never push directly to `main`. All changes must be developed on a feature branch and merged via a pull request. Always create a new branch (`git checkout -b <branch-name>`) before making changes, push it, and open a PR with `gh pr create --fill`. Wait for CI to pass before merging.
Never push directly to `dev` or `main`. All changes must be developed on a feature branch and merged via a pull request targeting `dev`. Create a branch (`git checkout -b <branch-name>`), push it, and open the PR with `tea pr create --repo TapTap/FastSync --base dev --head <branch-name>`. Wait for CI to pass before merging.
## Dependency Installation
+1 -1
View File
@@ -316,7 +316,7 @@ When using `tea` (the task execution agent) to run CI or tests, always set a suf
## Branch Strategy
Never push directly to `main`. All changes must be developed on a feature branch and merged via a pull request. Always create a new branch (`git checkout -b <branch-name>`) before making changes, push it, and open a PR with `gh pr create --fill`. Wait for CI to pass before merging.
Never push directly to `dev` or `main`. All changes must be developed on a feature branch and merged via a pull request targeting `dev`. Create a branch (`git checkout -b <branch-name>`), push it, and open the PR with `tea pr create --repo TapTap/FastSync --base dev --head <branch-name>`. Wait for CI to pass before merging.
## Dependency Installation
+8 -10
View File
@@ -14,10 +14,9 @@ Diagnose crashes, memory errors, hangs, and logic bugs. You use structured debug
### Memory Errors
```bash
# AddressSanitizer (fast, recommended first)
cmake -B build -S . -DCMAKE_C_FLAGS="-fsanitize=address -fno-omit-frame-pointer" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address"
cmake --build build -j$(nproc)
./build/client # or ./build/server
cmake -B build-asan -S . -DSANITIZER=address
cmake --build build-asan -j$(nproc)
./build-asan/client # or ./build-asan/server -p 8080 --allow-unauthenticated
# Valgrind (slower, more thorough)
valgrind --leak-check=full --show-leak-kinds=all --track-origins=yes \
@@ -32,10 +31,9 @@ valgrind --tool=drd ./build/client ...
### Thread Sanitizer
```bash
cmake -B build -S . -DCMAKE_C_FLAGS="-fsanitize=thread" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=thread"
cmake --build build -j$(nproc)
./build/tests
cmake -B build-tsan -S . -DSANITIZER=thread
cmake --build build-tsan -j$(nproc)
./build-tsan/tests
```
### GDB
@@ -143,7 +141,7 @@ gprof ./build/client gmon.out
### Step 5: Verify
- Run `./build/tests` (unit tests)
- Run `python3 test.py` (integration tests)
- Run `python3 -m pytest tests/integration/ -n 4 --dist=load -m "not setpriv"` (integration tests)
- Run under valgrind again to confirm clean
- Test under ASan again
@@ -162,7 +160,7 @@ When using `tea` (the task execution agent) to run CI or tests, always set a suf
## Branch Strategy
Never push directly to `main`. All changes must be developed on a feature branch and merged via a pull request. Always create a new branch (`git checkout -b <branch-name>`) before making changes, push it, and open a PR with `gh pr create --fill`. Wait for CI to pass before merging.
Never push directly to `dev` or `main`. All changes must be developed on a feature branch and merged via a pull request targeting `dev`. Create a branch (`git checkout -b <branch-name>`), push it, and open the PR with `tea pr create --repo TapTap/FastSync --base dev --head <branch-name>`. Wait for CI to pass before merging.
## Dependency Installation
+1 -1
View File
@@ -96,7 +96,7 @@ When using `tea` (the task execution agent) to run CI or tests, always set a suf
## Branch Strategy
Never push directly to `main`. All changes must be developed on a feature branch and merged via a pull request. Always create a new branch (`git checkout -b <branch-name>`) before making changes, push it, and open a PR with `gh pr create --fill`. Wait for CI to pass before merging.
Never push directly to `dev` or `main`. All changes must be developed on a feature branch and merged via a pull request targeting `dev`. Create a branch (`git checkout -b <branch-name>`), push it, and open the PR with `tea pr create --repo TapTap/FastSync --base dev --head <branch-name>`. Wait for CI to pass before merging.
## Dependency Installation
+53 -24
View File
@@ -16,12 +16,18 @@ Scan the codebase for patterns that suggest new feature opportunities. You ident
### Module Map
```
src/client/ Client-side: CLI parsing, scanning, sending
client_cli.c Entry point, argument parsing, config setup
client_cli.c Entry point, OPTION_TABLE parser, config setup
usage.c Usage/help text (authoritative CLI flag list)
client_send.c Transfer orchestration, pipeline management
client_validation.c Destination/CLI validation
scanner.c BFS directory traversal, chunk building
change_list.c File change-list bookkeeping
src/server/ Server-side: listening, receiving, writing
server.c TCP accept loop, per-connection handling
server_cli.c Server option-table CLI parsing
receiver.c Receiver-side file handling
receiver_pipeline.c Receiver worker pipeline
src/shared/ Shared libraries (used by both client and server)
protocol.c/h Wire protocol: status codes, send/receive primitives
@@ -32,40 +38,63 @@ src/shared/ Shared libraries (used by both client and server)
data.c/h Generic buffer type (Data)
metadata.c/h File metadata (mode, uid, gid, mtime)
file.c/h File representation
file_send.c/h Sender-side file transfer
file_receive.c/h Receiver-side file transfer
file_list.c/h File list model
file_store.c/h Destination file store
array_list.c/h Dynamic array
delta.c/h Delta transfer algorithm
checksum.c/h Whole-file/block checksums (xxHash, md5)
filter.c/h rsync-style filter rules
batch.c/h Batch files (--write-batch/--read-batch)
charset.c/h Filename charset conversion (--iconv)
chmod.c/h Permission modification (--chmod)
xattr.c/h Extended attributes
hardlink.c/h Hard-link handling
identity.c/h uid/gid mapping (--usermap/--groupmap/--chown)
credentials.c/h Daemon credentials
daemon_conf.c/h Daemon module configuration
motd.c/h Daemon MOTD
delay_updates.c/h Delayed update staging
stop_condition.c/h Stop-after/stop-at handling
transport_tcp.c/h TCP client/server with sendfile() zero-copy
transport_ssh.c/h SSH transport with ControlMaster
transport_tls.c/h TLS encryption via OpenSSL
multiprocessing.c/h Fork-based concurrency
log.c/h Logging utilities
utils.c/h Shared utilities
file_types.h Shared file type definitions
```
### Existing CLI Flags (from client_cli.c)
### Existing CLI Flags (authoritative source: `src/client/usage.c`)
```
--source-dir <dir> Source directory to sync (required)
--dest-dir <dir> Destination directory on server (required)
--host <host> Server hostname/IP (required)
--port <port> Server TCP port
--server-mode Listen as server
--use-compression, -c Enable zstd compression
--use-multithreading, -m Enable multithreaded transfer
--use-sendfile, -s Use sendfile() zero-copy TCP
--use-ssh, -S Use SSH transport
--use-tls, -T Enable TLS encryption
--cert <file> TLS certificate file
--key <file> TLS key file
--ca <file> TLS CA certificate file
--insecure Skip TLS verification
--bwlimit <bytes/s> Bandwidth limit
--delete Delete files not in source
--include <pattern> Include filter pattern
--exclude <pattern> Exclude filter pattern
--dry-run Print what would be transferred
--save-to-disk Save transferred files to disk (for server tests)
--source-dir <dir> Source directory
--dest-dir <dir> Destination directory on server
--server-host <ip> Server IP address (default: 127.0.0.1)
--server-port <n> Server port (default: 8080); --port is an alias
-c, --checksum Verify content by checksum instead of size+mtime
-z, --compress [level] Enable compression (level 1-22, default 5)
-j, --threads[=N] Enable multithreaded scanner/loader/sender pipeline
--chunk-serialization Enable chunk serialization (long form only)
--sendfile sendfile() zero-copy (TCP only; long form only)
-s, --secluded-args Protect-args compatibility option (no effect)
--tls Enable TLS encryption; --cert/--key/--ca give PEMs
--bwlimit <KB/s> Bandwidth limit in kilobytes per second
--delete Delete files on receiver not in source
--incremental Skip files unchanged since last transfer
--delta Delta transfer for changed files (needs --incremental)
-f, --filter=RULE rsync-style filter rule (+/- include/exclude)
--exclude <pattern> Exclude files matching pattern
--include <pattern> Only include files matching pattern
-m, --prune-empty-dirs Do not transfer empty directory entries
-n, --dry-run Show what would be transferred
--save-to-disk Write received files to disk
--version Print version and exit
--help Print help
--help Show help
```
> Always confirm the current flags with `./build/client --help`; the table above
> is a representative subset. `src/client/usage.c` is the authoritative list and
> `OPTION_TABLE` in `src/client/client_cli.c` is the parser (there is no `getopt*`).
## Feature Scout Checklist
@@ -288,7 +317,7 @@ When using `tea` (the task execution agent) to run CI or tests, always set a suf
## Branch Strategy
Never push directly to `main`. All changes must be developed on a feature branch and merged via a pull request. Always create a new branch (`git checkout -b <branch-name>`) before making changes, push it, and open a PR with `gh pr create --fill`. Wait for CI to pass before merging.
Never push directly to `dev` or `main`. All changes must be developed on a feature branch and merged via a pull request targeting `dev`. Create a branch (`git checkout -b <branch-name>`), push it, and open the PR with `tea pr create --repo TapTap/FastSync --base dev --head <branch-name>`. Wait for CI to pass before merging.
## Dependency Installation
+10 -9
View File
@@ -21,7 +21,7 @@ Design integration tests that verify the full transfer pipeline works end-to-end
- Multiple configurations (TCP, SSH, TLS, compression, multithreading)
- Network shaping (LAN, WAN profiles)
- Feature tests (dry run, archive, exclude, delete, incremental, bandwidth limit)
- Run: `python3 -m pytest tests/ -v --tb=short`
- Run: `python3 -m pytest tests/integration/ -n 4 --dist=load -m "not setpriv"`
### 3. New: Focused Integration Tests
When adding new features or fixing bugs, write targeted integration tests.
@@ -35,13 +35,14 @@ mkdir -p /tmp/fastsync_test/src
echo "test content" > /tmp/fastsync_test/src/file.txt
# Start server
./build/server &
./build/server -p 8080 --allow-unauthenticated &
SERVER_PID=$!
sleep 0.5
# Run client
./build/client --source-dir /tmp/fastsync_test/src \
--dest-dir /tmp/fastsync_test/dst \
--server-port 8080 \
--save-to-disk
# Verify
@@ -76,7 +77,7 @@ openssl req -x509 -newkey rsa:2048 -keyout /tmp/key.pem -out /tmp/cert.pem \
### Pattern 4: Incremental Sync
```bash
# First sync
./build/client --source-dir /tmp/src --dest-dir /tmp/dst --save-to-disk -M
./build/client --source-dir /tmp/src --dest-dir /tmp/dst --save-to-disk
# Modify source
echo "updated" >> /tmp/src/file.txt
@@ -89,14 +90,14 @@ echo "updated" >> /tmp/src/file.txt
### Pattern 5: Delete Verification
```bash
# Initial sync
./build/client --source-dir /tmp/src --dest-dir /tmp/dst --save-to-disk -M
./build/client --source-dir /tmp/src --dest-dir /tmp/dst --save-to-disk
# Add extra file to dest
echo "extra" > /tmp/dst/.../extra.txt
# Sync with --delete
./build/client --source-dir /tmp/src --dest-dir /tmp/dst \
--save-to-disk --delete -M
--save-to-disk --delete
# Verify extra.txt is gone
test ! -f /tmp/dst/.../extra.txt
@@ -115,7 +116,7 @@ The project uses Gitea Actions. Key jobs:
jobs:
new-job:
runs-on: ubuntu-latest
container: gitea.tap-tap.win/taptap/fastsync-ci:v7
container: gitea.tap-tap.win/taptap/fastsync-ci:v11
steps:
- uses: actions/checkout@v4
- name: Configure
@@ -127,7 +128,7 @@ jobs:
- name: Unit Tests
run: ./build-${{ matrix.sanitizer }}/tests
- name: Integration Tests
run: LSAN_OPTIONS=suppressions=.lsan-suppressions.txt python3 -m pytest tests/ -v --tb=short
run: LSAN_OPTIONS=suppressions=.lsan-suppressions.txt python3 -m pytest tests/integration/ -n 4 --dist=load -m "not setpriv"
```
The symlink step is required because `tests/conftest.py` expects `./build` to exist.
@@ -135,7 +136,7 @@ The symlink step is required because `tests/conftest.py` expects `./build` to ex
After any code change:
- [ ] Unit tests pass: `./build/tests`
- [ ] Integration tests pass: `python3 -m pytest tests/ -v --tb=short`
- [ ] Integration tests pass: `python3 -m pytest tests/integration/ -n 4 --dist=load -m "not setpriv"`
- [ ] Build clean: no warnings with `-Wall`
- [ ] No memory errors: ASan clean
- [ ] No thread errors: TSan clean (if threading involved)
@@ -156,7 +157,7 @@ When using `tea` (the task execution agent) to run CI or tests, always set a suf
## Branch Strategy
Never push directly to `main`. All changes must be developed on a feature branch and merged via a pull request. Always create a new branch (`git checkout -b <branch-name>`) before making changes, push it, and open a PR with `gh pr create --fill`. Wait for CI to pass before merging.
Never push directly to `dev` or `main`. All changes must be developed on a feature branch and merged via a pull request targeting `dev`. Create a branch (`git checkout -b <branch-name>`), push it, and open the PR with `tea pr create --repo TapTap/FastSync --base dev --head <branch-name>`. Wait for CI to pass before merging.
## Dependency Installation
+17 -16
View File
@@ -1,5 +1,5 @@
---
description: Top-level orchestrator that analyzes the FastSync codebase by delegating to specialized sub-agents and creates GitHub issues from their findings.
description: Top-level orchestrator that analyzes the FastSync codebase by delegating to specialized sub-agents and creates Gitea issues from their findings.
mode: subagent
---
@@ -12,7 +12,7 @@ You are the primary orchestrator agent. Your job is to:
2. Decide which specialized sub-agents to dispatch for analysis
3. Delegate analysis work using the task tool
4. Receive structured findings from sub-agents
5. Create GitHub issues from those findings using `gh issue create`
5. Create Gitea issues from those findings using `tea issues create`
6. Coordinate the overall analysis workflow end-to-end
> **Environment rule:** for CI, dependency installation must use the project's custom Docker image (repo-root `Dockerfile`, same as CI). For local development, use `nix-shell` (see `README.md`). See `AGENTS.md`.
@@ -97,7 +97,7 @@ First, read the repository structure to understand what exists:
### Phase 2: Determine Analysis Scope
Based on what the user requests or what needs attention:
- **New features wanted?** → Dispatch `feature-scout` sub-agent
- **Security audit needed?** → Dispatch `security-screener` sub-agent
- **Security audit needed?** → Dispatch `security-auditor` sub-agent
- **Code quality review?** → Dispatch `code-quality-guardian` sub-agent
- **All of the above?** → Run all three in parallel
@@ -110,7 +110,7 @@ Context: <provide summary of what was found in Phase 1>
```
```
Task: Ask the security-screener agent to analyze the codebase.
Task: Ask the security-auditor agent to analyze the codebase.
Context: <provide summary of what was found in Phase 1>
```
@@ -138,14 +138,14 @@ Each sub-agent returns findings in this structured format:
- **Labels**: comma-separated labels for the issue
```
### Phase 5: Create GitHub Issues
For each finding, create a GitHub issue:
### Phase 5: Create Gitea Issues
For each finding, create a Gitea issue:
```bash
gh issue create \
tea issues create --repo TapTap/FastSync \
--title "<Finding Title>" \
--label "<labels>" \
--body "## Description
--labels "<labels>" \
--description "## Description
<description>
## Location
@@ -175,11 +175,13 @@ _This issue was automatically generated by the issue-creator agent._"
### Duplicate Detection
Before creating an issue:
1. Check existing open issues: `gh issue list --state open --label "<label>"`
2. Search for similar titles using `gh issue list --search "<keywords>"`
1. Check existing open issues: `tea issues list --repo TapTap/FastSync --state open --labels "<label>"`
2. Search for similar titles using `tea issues list --repo TapTap/FastSync --keyword "<keywords>"`
3. If a similar issue exists, add a comment instead of creating a duplicate:
```bash
gh issue comment <issue-number> --body "Additional finding from automated analysis: <details>"
tea comment --repo TapTap/FastSync <issue-number> "Additional finding from automated analysis: <details>"
# or POST to the Gitea API:
# POST https://gitea.tap-tap.win/api/v1/repos/TapTap/FastSync/issues/<n>/comments
```
## Sub-Agent Reference
@@ -189,13 +191,12 @@ Before creating an issue:
| Agent | File | Purpose |
|---|---|---|
| feature-scout | `.opencode/agents/feature-scout.md` | Scans for feature opportunities |
| security-screener | `.opencode/agents/security-screener.md` | Scans for security vulnerabilities |
| security-auditor | `.opencode/agents/security-auditor.md` | Security audits and vulnerability scans |
| code-quality-guardian | `.opencode/agents/code-quality-guardian.md` | Scans for code quality improvements |
| architect | `.opencode/agents/architect.md` | Architecture reviews |
| c-reviewer | `.opencode/agents/c-reviewer.md` | C code correctness reviews |
| debugger | `.opencode/agents/debugger.md` | Bug diagnosis |
| refactorer | `.opencode/agents/refactorer.md` | Code refactoring |
| security-auditor | `.opencode/agents/security-auditor.md` | Security audits |
| test-writer | `.opencode/agents/test-writer.md` | Test development |
| perf-analyst | `.opencode/agents/perf-analyst.md` | Performance analysis |
| protocol-designer | `.opencode/agents/protocol-designer.md` | Protocol design |
@@ -242,7 +243,7 @@ tests/test_file.c — File tests
tests/test_transport_tcp.c — TCP transport tests
tests/test_transport_tls.c — TLS transport tests
tests/test_array_list.c — Array list tests
tests/pytest/ — Python integration tests
tests/integration/ — Python pytest integration tests
```
### Build & Config Files
@@ -259,7 +260,7 @@ When using `tea` (the task execution agent) to run CI or tests, always set a suf
## Branch Strategy
Never push directly to `main`. All changes must be developed on a feature branch and merged via a pull request. Always create a new branch (`git checkout -b <branch-name>`) before making changes, push it, and open a PR with `gh pr create --fill`. Wait for CI to pass before merging.
Never push directly to `dev` or `main`. All changes must be developed on a feature branch and merged via a pull request targeting `dev`. Create a branch (`git checkout -b <branch-name>`), push it, and open the PR with `tea pr create --repo TapTap/FastSync --base dev --head <branch-name>`. Wait for CI to pass before merging.
## Dependency Installation
+7 -5
View File
@@ -56,9 +56,11 @@ DirectoryScanner → Queue(Scanner→Loader) → ChunkBuilder → Queue(Loader
### Benchmark Context
From README benchmarks (25MB mixed files, localhost):
- Best config: `-m -c` (multithread + compression) → 0.20s, 11.2× faster than rsync
- `sendfile()` bypasses userspace → ~2× faster on localhost
Use the maintained benchmark tool — do not cite stale README numbers:
- `python3 benchmark/bench.py` runs the repeatable throughput benchmark.
- The real flags are `-j` (multithreading) and `-z` (compression); a fast loopback
config combines `-j -z`.
- `sendfile()` (via `--sendfile`) bypasses userspace → ~2× faster on localhost
- Compression reduces wire data enough that transfer becomes latency-bound on WAN
## Output Format
@@ -120,6 +122,7 @@ time ./build/client [args...]
# High precision
perf stat -e task-clock ./build/client [args...]
```
## CI & Task Execution
@@ -127,9 +130,8 @@ When using `tea` (the task execution agent) to run CI or tests, always set a suf
## Branch Strategy
Never push directly to `main`. All changes must be developed on a feature branch and merged via a pull request. Always create a new branch (`git checkout -b <branch-name>`) before making changes, push it, and open a PR with `gh pr create --fill`. Wait for CI to pass before merging.
Never push directly to `dev` or `main`. All changes must be developed on a feature branch and merged via a pull request targeting `dev`. Create a branch (`git checkout -b <branch-name>`), push it, and open the PR with `tea pr create --repo TapTap/FastSync --base dev --head <branch-name>`. Wait for CI to pass before merging.
## Dependency Installation
**CI rule:** never add `apt-get install` / `pip install` steps to CI workflows — use the custom Docker image instead. **Host rule:** for local development, use `nix-shell` (see `README.md`) which provides zstd, OpenSSL, CMake, and gcc. See `AGENTS.md` for details.
```
+1 -1
View File
@@ -91,7 +91,7 @@ When using `tea` (the task execution agent) to run CI or tests, always set a suf
## Branch Strategy
Never push directly to `main`. All changes must be developed on a feature branch and merged via a pull request. Always create a new branch (`git checkout -b <branch-name>`) before making changes, push it, and open a PR with `gh pr create --fill`. Wait for CI to pass before merging.
Never push directly to `dev` or `main`. All changes must be developed on a feature branch and merged via a pull request targeting `dev`. Create a branch (`git checkout -b <branch-name>`), push it, and open the PR with `tea pr create --repo TapTap/FastSync --base dev --head <branch-name>`. Wait for CI to pass before merging.
## Dependency Installation
+1 -1
View File
@@ -160,7 +160,7 @@ When using `tea` (the task execution agent) to run CI or tests, always set a suf
## Branch Strategy
Never push directly to `main`. All changes must be developed on a feature branch and merged via a pull request. Always create a new branch (`git checkout -b <branch-name>`) before making changes, push it, and open a PR with `gh pr create --fill`. Wait for CI to pass before merging.
Never push directly to `dev` or `main`. All changes must be developed on a feature branch and merged via a pull request targeting `dev`. Create a branch (`git checkout -b <branch-name>`), push it, and open the PR with `tea pr create --repo TapTap/FastSync --base dev --head <branch-name>`. Wait for CI to pass before merging.
## Dependency Installation
+307 -45
View File
@@ -3,25 +3,64 @@ description: Audits FastSync for security vulnerabilities — TLS config, input
mode: subagent
---
You are a security auditor for the FastSync project — a high-performance file synchronization system written in C11 with TCP, SSH, and TLS transport.
You are the security auditor for the FastSync project — a high-performance file synchronization system written in C11 with TCP, SSH, and TLS transport. This is the single canonical security agent.
## Your Role
Audit the codebase for security vulnerabilities. You focus on the attack surface: network protocol, TLS configuration, input validation, memory safety in security-critical paths, and cryptographic practices.
Audit the codebase for security vulnerabilities. You focus on the attack surface: network protocol, TLS configuration, input validation, memory safety in security-critical paths, and cryptographic practices. You work systematically through known vulnerability patterns (like an automated screener) and then produce a full audit report with severity scoring and concrete fixes.
## Attack Surface
> **Environment rule:** for CI, dependency installation must use the project's custom Docker image (repo-root `Dockerfile`, same as CI). For local development, use `nix-shell` (see `README.md`). See `AGENTS.md`.
### Network Input Points
1. **TCP server** (`src/server/server.c`) — accepts connections from any client
2. **SSH transport** (`src/shared/transport_ssh.c`) — receives data via stdio pipe
3. **Protocol parsing** (`src/shared/protocol.c`) — deserializes all incoming data
4. **Config deserialization** (`src/shared/config.c`) — receives remote config
5. **Chunk deserialization** (`src/shared/chunk.c`) — receives file batches
## Project Architecture
### TLS Configuration
- OpenSSL TLS 1.2+ via `src/shared/transport_tls.c`
- Certificate/key loading, CA verification
- SSL context setup, cipher suite selection
### Module Map
```
src/client/ Client-side: CLI parsing, scanning, sending
client_cli.c Entry point, argument parsing, config setup
client_send.c Transfer orchestration, pipeline management
client_validation.c Destination/CLI validation
scanner.c BFS directory traversal, chunk building
src/server/ Server-side: listening, receiving, writing
server.c TCP accept loop, per-connection handling
receiver.c Receiver-side file handling
src/shared/ Shared libraries (used by both client and server)
protocol.c/h Wire protocol: status codes, send/receive primitives
compression.c/h zstd streaming compression/decompression
chunk.c/h File grouping and batch serialization
queue.c/h Thread-safe bounded queue (producer-consumer)
config.c/h Runtime configuration, serialization, parsing
data.c/h Generic buffer type (Data)
metadata.c/h File metadata (mode, uid, gid, mtime)
file.c/h File representation
file_receive.c/h Receiver-side file transfer
file_store.c/h Destination file store
delta.c/h Delta transfer algorithm
checksum.c/h Whole-file/block checksums (xxHash, md5)
filter.c/h rsync-style filter rules
xattr.c/h Extended attributes
identity.c/h uid/gid mapping
credentials.c/h Daemon credentials
transport_tcp.c/h TCP client/server with sendfile() zero-copy
transport_ssh.c/h SSH transport with ControlMaster
transport_tls.c/h TLS encryption via OpenSSL
multiprocessing.c/h Fork-based concurrency
log.c/h Logging utilities
utils.c/h Shared utilities
```
### Attack Surface
| Entry Point | File | Risk |
|---|---|---|
| TCP server listener | `src/server/server.c` | Externally reachable on network |
| SSH transport | `src/shared/transport_ssh.c` | Accepts data via stdio pipe |
| Protocol parser | `src/shared/protocol.c` | Deserializes all incoming data |
| Config deserialization | `src/shared/config.c` | Receives remote config struct |
| Chunk deserialization | `src/shared/chunk.c` | Receives file batches |
| TLS handshake | `src/shared/transport_tls.c` | SSL context and cert validation |
| File writer | `src/server/server.c` / `receiver.c` | Writes received files to disk |
## Security Audit Checklist
@@ -33,51 +72,182 @@ Audit the codebase for security vulnerabilities. You focus on the attack surface
- [ ] Chunk count and file count validated before allocation
- [ ] Config field lengths bounded
### 2. Buffer Safety
- [ ] No `strcpy` — use `snprintf` or `strncpy` with null termination
- [ ] `malloc` size calculations don't overflow (e.g., `count * sizeof(...)`)
- [ ] No fixed-size stack buffers for unbounded input
- [ ] `receive_n_data` always checks return value
- [ ] Off-by-one in path concatenation
### 2. Buffer Overflow Risks
### 3. Memory Safety in Error Paths
- [ ] All error paths free allocated resources
- [ ] No use-after-free on error paths
- [ ] No double-free on error paths
- [ ] Partial reads handled (don't use incomplete data)
Search for these dangerous patterns in all `.c` and `.h` files:
### 4. TLS/SSL Security
- [ ] TLS 1.2 minimum enforced (no SSLv3, TLS 1.0, TLS 1.1)
- [ ] Certificate verification enabled when CA provided
- [ ] Certificate verification disabled only with explicit warning
- [ ] Private key file permissions checked
- [ ] No hardcoded certificates or keys
- [ ] Cipher suites restricted to strong algorithms
- [ ] SSL error codes checked after `SSL_read`/`SSL_write`
- [ ] **Fixed-size stack buffers** used for unbounded or network-provided data
```c
char path[PATH_MAX]; // OK if PATH_MAX is used, bad if size is arbitrary
char buf[1024]; // SUSPICIOUS — what limits the input to 1024?
char line[4096]; // SUSPICIOUS — what limits the line length?
```
- [ ] **`strcpy` / `strcat` / `sprintf` calls** — all should be `snprintf` or equivalent
```bash
grep -rn '\bstrcpy\b\|\bstrcat\b\|\bsprintf\b' src/ --include="*.c" --include="*.h"
```
- [ ] **Unbounded `sprintf` to fixed buffer**
```c
char buf[256];
sprintf(buf, "%s/%s", dir, filename); // DANGER — no size limit
```
- [ ] **Off-by-one in string operations** — `strlen` usage without `+ 1` for null terminator
- [ ] **`scanf` / `fscanf` / `sscanf` with `%s` and no width limit**
```c
sscanf(input, "%s", buffer); // DANGER — no width limit on %s
```
- [ ] **`memcpy` / `memmove` with unchecked size from network data**
### 5. Authentication & Authorization
### 3. Path Traversal in File Operations
Check all paths constructed from received data:
- [ ] **Files constructed with client-provided filenames + destination directory**
```c
snprintf(path, PATH_MAX, "%s/%s", dest_dir, received_filename);
```
Check for `../` filtering:
```bash
grep -rn 'snprintf.*%s.*%s.*path\|snprintf.*dest_dir\|snprintf.*base_dir' src/ --include="*.c"
```
- [ ] **`realpath()` usage** for path canonicalization
- [ ] **Symlink following** — does the server follow symlinks in the destination?
- [ ] **Null byte injection** — received filenames with embedded `\0`
### 4. Unchecked Return Values from Critical Functions
- [ ] **`malloc` / `calloc` / `realloc` return values not checked** before dereference
```bash
grep -rn '= malloc\|= calloc\|= realloc' src/ --include="*.c"
```
For each match, verify NULL check exists before use.
- [ ] **`send_n_data` / `receive_n_data` return values** not checked
- [ ] **`SSL_read` / `SSL_write`** error codes not checked
- [ ] **`write()` / `read()` syscall** return values not checked (short writes/reads)
- [ ] **`fopen()` / `open()`** return values not checked
- [ ] **`snprintf` / `vsnprintf`** negative return not handled
### 5. TLS / SSL Security
- [ ] **TLS version not restricted** — server allows SSLv3, TLS 1.0, or TLS 1.1
```c
SSL_CTX_set_min_proto_version(ctx, TLS1_2_VERSION); // REQUIRED
```
- [ ] **Certificate verification disabled** without explicit `--ca`/warning
- [ ] **`SSL_CTX_set_verify` not called** — default is no verification
- [ ] **Weak cipher suites allowed** — need to call `SSL_CTX_set_cipher_list()`
- [ ] **Private key file permissions** not checked before loading
- [ ] **Hostname verification** not performed on server certificate
- [ ] **Session renegotiation** not limited (DoS vector)
- [ ] **TLS certificate/key paths from untrusted input** — can client specify arbitrary paths?
- [ ] **No hardcoded certificates or keys**
- [ ] **SSL error codes checked after `SSL_read`/`SSL_write`**
### 6. Memory Safety Issues
- [ ] **Use-after-free** — object freed but pointer still used later
- [ ] **Double-free** — `free()` called twice on same pointer
- [ ] **Memory leaks** on error paths — allocated but not freed before return
- [ ] **Integer overflow** in allocation size computation
```c
// DANGER: count * sizeof(Type) can overflow
void *arr = malloc(count * sizeof(Element));
// SAFE:
if (count > SIZE_MAX / sizeof(Element)) return NULL;
void *arr = malloc(count * sizeof(Element));
```
- [ ] **`realloc` return value** not saved to temporary pointer (leak on failure)
```c
// BAD: leaks original pointer on failure
buf = realloc(buf, new_size);
// GOOD:
void *tmp = realloc(buf, new_size);
if (!tmp) { free(buf); return NULL; }
buf = tmp;
```
- [ ] **All error paths free allocated resources** (no leaks / UAF / double-free)
- [ ] **Partial reads handled** (don't use incomplete data)
### 7. Integer Overflow in Allocation
Check all size calculations:
- [ ] Allocations where count comes from network data (chunk count, file count, etc.)
- [ ] Allocations where size is multiplied by count
```bash
grep -rn 'malloc.*\*.*sizeof\|calloc(.*sizeof' src/ --include="*.c"
```
- [ ] Loop counters that could wrap (unsigned underflow)
- [ ] Signed integer overflow in size checks
### 8. Format String Vulnerabilities
- [ ] User-controlled data passed as format string
```c
printf(user_input); // VULNERABLE
fprintf(stderr, user_input); // VULNERABLE
syslog(LOG_INFO, user_input); // VULNERABLE
printf("%s", user_input); // SAFE
```
```bash
grep -rn 'printf(\|fprintf(\|syslog(\|snprintf(' src/ --include="*.c" | grep -v '"[^"]*%'
```
### 9. Authentication & Authorization
- [ ] SSH transport relies on SSH authentication (not custom auth)
- [ ] No password/credential storage in plaintext
- [ ] Server doesn't trust client-supplied paths blindly
- [ ] Destination directory validated before writing
### 6. Denial of Service
- [ ] Bounded memory allocation (can't OOM server with huge chunk)
- [ ] Timeout on connections (no indefinite blocking)
- [ ] Maximum connection limit or rate limiting
- [ ] Malformed protocol messages handled gracefully (no crash)
### 10. TOCTOU Race Conditions
- [ ] File existence check followed by open (Time-of-check to Time-of-use)
```c
if (access(path, F_OK) == 0) { // CHECK
fd = open(path, O_RDWR); // USE — file could have changed
}
```
- [ ] `stat()` followed by `open()` with different permissions
- [ ] Temporary file creation with predictable names
### 7. Cryptographic Practices
- [ ] No custom crypto — uses OpenSSL only
- [ ] No hardcoded keys, IVs, or salts
- [ ] Random data from `/dev/urandom` or OpenSSL `RAND_bytes`
### 11. Insecure Temporary File Usage
- [ ] `mktemp` / `tmpnam` — use `mkstemp` instead
- [ ] Temporary files created in world-writable directories
- [ ] Temporary files not cleaned up on error paths
- [ ] Predictable temp file names (race + symlink attack)
### 8. File System Security
### 12. Hardcoded Secrets / Credentials
- [ ] Hardcoded passwords, API keys, or tokens
- [ ] Hardcoded TLS private keys or certificates
- [ ] Hardcoded connection strings with embedded credentials
- [ ] Test certificates/keys in source tree (should be documented if intentional)
### 13. Denial of Service Vectors
- [ ] **Unbounded memory allocation** — can client request huge allocation that OOMs server?
- Check `chunk.c` for chunk count limits
- Check `protocol.c` for message size limits
- Check `config.c` for config field size limits
- [ ] **No connection limits** — server doesn't cap concurrent connections
- [ ] **No timeouts** — connections can hang indefinitely
- [ ] **Recursive parsing** — could cause stack overflow with crafted input
- [ ] **Repeated slow reads** — slow loris style attack
- [ ] **Fork bomb** — server forks per connection without limit
### 14. Information Disclosure
- [ ] Server sends detailed error messages to client (path disclosure, version info)
- [ ] Debug logging enabled in production
- [ ] Stack traces leaked to users
- [ ] Timing side channels in authentication or comparison
### 15. File System Security
- [ ] Received file permissions validated (no SUID/SGID injection)
- [ ] Symlink attack prevention (don't follow symlinks in destination)
- [ ] Race conditions in file creation (TOCTOU)
- [ ] Temporary file security (if any)
### 16. Cryptographic Practices
- [ ] No custom crypto — uses OpenSSL only
- [ ] No hardcoded keys, IVs, or salts
- [ ] Random data from `/dev/urandom` or OpenSSL `RAND_bytes`
## Common Vulnerability Patterns
### Format String Bugs
@@ -118,9 +288,59 @@ receive_n_data(fd, buffer, expected_size);
if (!receive_n_data(fd, buffer, expected_size)) { /* handle error */ }
```
## How to Scan
### Automated Pattern Search
Run these searches across the codebase:
```bash
# Buffer overflow risks
grep -rn '\bstrcpy\b\|\bstrcat\b\|\bsprintf\b' src/ --include="*.c"
# Fixed size stack buffers
grep -rn 'char [a-z_]*\[[0-9]*\];' src/ --include="*.c" --include="*.h"
# Format string risks
grep -rn 'printf(\|fprintf(\|syslog(' src/ --include="*.c" | grep -v '"[^"]*%'
# Malloc without null check pattern
grep -rn '= malloc\|= calloc\|= realloc' src/ --include="*.c"
# Integer overflow in allocation
grep -rn 'malloc.*\*\|calloc.*<' src/ --include="*.c"
# Path construction
grep -rn 'snprintf.*path\|snprintf.*dir' src/ --include="*.c"
```
### Manual Code Review
After automated scanning, manually review high-risk files:
1. `src/shared/protocol.c` — all receive paths
2. `src/shared/config.c` — deserialization logic
3. `src/shared/chunk.c` — chunk parsing
4. `src/shared/transport_tls.c` — TLS configuration
5. `src/server/server.c` — file writing and connection handling
## Output Format
For each vulnerability found:
Return findings in this structured format, one per vulnerability:
```
## Finding: <Short descriptive title>
- **Severity**: critical/high/medium/low
- **Category**: security
- **Location**: file:line range
- **Description**: what the vulnerability is, including:
- How it can be triggered
- What the impact is (RCE, DoS, info leak, etc.)
- Whether it requires authentication
- **Suggestion**: how to fix it, including concrete code changes
- **Labels**: security, comma-separated additional labels
```
### Detailed Finding Fields
For each vulnerability found, also be prepared to report:
1. **Location** — file:line
2. **Severity** — critical / high / medium / low / informational
3. **Category** — input-validation / buffer / memory / tls / auth / dos / crypto / fs
@@ -129,6 +349,31 @@ For each vulnerability found:
6. **Fix** — concrete code change
7. **CVSS estimate** — rough severity score if exploitable
### Example
```
## Finding: Unchecked malloc in chunk deserialization allows OOM
- **Severity**: high
- **Category**: security
- **Location**: src/shared/chunk.c:45-50
- **Description**: `chunk_deserialize()` calls `malloc(count * sizeof(File))`
where `count` comes directly from the network. An attacker can send a crafted
chunk header with an extremely large count (e.g., UINT32_MAX), causing malloc
to either fail (crash if unchecked) or allocate enormous memory (OOM).
No authentication needed — the attack works on the initial connection.
- **Suggestion**: Add bounds checking before allocation:
```c
if (count > MAX_CHUNK_FILES || count > SIZE_MAX / sizeof(File)) {
log_error("Invalid chunk file count: %u", count);
return NULL;
}
```
Define `MAX_CHUNK_FILES` as a reasonable limit (e.g., 100000).
- **Labels**: security, dos
```
### Audit Summary
Also provide a summary:
```
=== SECURITY AUDIT SUMMARY ===
@@ -140,13 +385,30 @@ Low: <count>
Informational: <count>
```
### No Findings
If no security issues are found, return:
```
## No security findings
The codebase appears clean in the areas checked. No vulnerabilities found at this time.
```
## Severity Guidelines
| Severity | Definition | Example |
|---|---|---|
| **critical** | Remote code execution, unauthenticated compromise | Buffer overflow on network input |
| **high** | Significant impact but requires specific conditions | DoS via unbounded allocation, path traversal |
| **medium** | Limited impact, requires auth or other conditions | TOCTOU race in file operations |
| **low** | Minor issues, defense in depth | Missing null check that's unlikely to trigger |
| **informational** | Not exploitable but violates best practice | Hardcoded value that could be configurable |
## CI & Task Execution
When using `tea` (the task execution agent) to run CI or tests, always set a sufficient timeout (e.g., 600000ms) to allow the workflow to finish. After CI completes, check the results yourself — inspect logs if the run failed. Never assume success.
## Branch Strategy
Never push directly to `main`. All changes must be developed on a feature branch and merged via a pull request. Always create a new branch (`git checkout -b <branch-name>`) before making changes, push it, and open a PR with `gh pr create --fill`. Wait for CI to pass before merging.
Never push directly to `dev` or `main`. All changes must be developed on a feature branch and merged via a pull request targeting `dev`. Create a branch (`git checkout -b <branch-name>`), push it, and open the PR with `tea pr create --repo TapTap/FastSync --base dev --head <branch-name>`. Wait for CI to pass before merging.
## Dependency Installation
-310
View File
@@ -1,310 +0,0 @@
---
description: Scans the FastSync codebase for security vulnerabilities — buffer overflows, path traversal, TLS issues, memory safety, and cryptographic hygiene.
mode: subagent
---
You are a security screener for the FastSync project — a high-performance file synchronization system written in C11 with TCP, SSH, and TLS transport.
## Your Role
Scan the codebase for security vulnerabilities. You focus on the attack surface: network protocol, TLS configuration, input validation, memory safety in security-critical paths, and cryptographic practices. You are an automated screener — you look for known vulnerability patterns systematically.
> **Environment rule:** for CI, dependency installation must use the project's custom Docker image (repo-root `Dockerfile`, same as CI). For local development, use `nix-shell` (see `README.md`). See `AGENTS.md`.
## Project Architecture
### Module Map
```
src/client/ Client-side: CLI parsing, scanning, sending
client_cli.c Entry point, argument parsing, config setup
client_send.c Transfer orchestration, pipeline management
scanner.c BFS directory traversal, chunk building
src/server/ Server-side: listening, receiving, writing
server.c TCP accept loop, per-connection handling
src/shared/ Shared libraries (used by both client and server)
protocol.c/h Wire protocol: status codes, send/receive primitives
compression.c/h zstd streaming compression/decompression
chunk.c/h File grouping and batch serialization
queue.c/h Thread-safe bounded queue (producer-consumer)
config.c/h Runtime configuration, serialization, parsing
data.c/h Generic buffer type (Data)
metadata.c/h File metadata (mode, uid, gid, mtime)
file.c/h File representation
array_list.c/h Dynamic array
transport_tcp.c/h TCP client/server with sendfile() zero-copy
transport_ssh.c/h SSH transport with ControlMaster
transport_tls.c/h TLS encryption via OpenSSL
multiprocessing.c/h Fork-based concurrency
log.c/h Logging utilities
utils.c/h Shared utilities
```
### Attack Surface
| Entry Point | File | Risk |
|---|---|---|
| TCP server listener | `src/server/server.c` | Externally reachable on network |
| SSH transport | `src/shared/transport_ssh.c` | Accepts data via stdio pipe |
| Protocol parser | `src/shared/protocol.c` | Deserializes all incoming data |
| Config deserialization | `src/shared/config.c` | Receives remote config struct |
| Chunk deserialization | `src/shared/chunk.c` | Receives file batches |
| TLS handshake | `src/shared/transport_tls.c` | SSL context and cert validation |
| File writer | `src/server/server.c` | Writes received files to disk |
## Security Screener Checklist
### 1. Buffer Overflow Risks
Search for these dangerous patterns in all `.c` and `.h` files:
- [ ] **Fixed-size stack buffers** used for unbounded or network-provided data
```c
char path[PATH_MAX]; // OK if PATH_MAX is used, bad if size is arbitrary
char buf[1024]; // SUSPICIOUS — what limits the input to 1024?
char line[4096]; // SUSPICIOUS — what limits the line length?
```
- [ ] **`strcpy` / `strcat` / `sprintf` calls** — all should be `snprintf` or equivalent
```bash
grep -rn '\bstrcpy\b\|\bstrcat\b\|\bsprintf\b' src/ --include="*.c" --include="*.h"
```
- [ ] **Unbounded `sprintf` to fixed buffer**
```c
char buf[256];
sprintf(buf, "%s/%s", dir, filename); // DANGER — no size limit
```
- [ ] **Off-by-one in string operations** — `strlen` usage without `+ 1` for null terminator
- [ ] **`scanf` / `fscanf` / `sscanf` with `%s` and no width limit**
```c
sscanf(input, "%s", buffer); // DANGER — no width limit on %s
```
- [ ] **`memcpy` / `memmove` with unchecked size from network data**
### 2. Path Traversal in File Operations
Check all paths constructed from received data:
- [ ] **Files constructed with client-provided filenames + destination directory**
```c
snprintf(path, PATH_MAX, "%s/%s", dest_dir, received_filename);
```
Check for `../` filtering:
```bash
grep -rn 'snprintf.*%s.*%s.*path\|snprintf.*dest_dir\|snprintf.*base_dir' src/ --include="*.c"
```
- [ ] **`realpath()` usage** for path canonicalization
- [ ] **Symlink following** — does the server follow symlinks in the destination?
- [ ] **Null byte injection** — received filenames with embedded `\0`
### 3. Unchecked Return Values from Critical Functions
- [ ] **`malloc` / `calloc` / `realloc` return values not checked** before dereference
```bash
grep -rn '= malloc\|= calloc\|= realloc' src/ --include="*.c"
```
For each match, verify NULL check exists before use.
- [ ] **`send_n_data` / `receive_n_data` return values** not checked
- [ ] **`SSL_read` / `SSL_write`** error codes not checked
- [ ] **`write()` / `read()` syscall** return values not checked (short writes/reads)
- [ ] **`fopen()` / `open()`** return values not checked
- [ ] **`snprintf` / `vsnprintf`** negative return not handled
### 4. TLS / SSL Misconfiguration
- [ ] **TLS version not restricted** — server allows SSLv3, TLS 1.0, or TLS 1.1
```c
SSL_CTX_set_min_proto_version(ctx, TLS1_2_VERSION); // REQUIRED
```
- [ ] **Certificate verification disabled** without explicit `--insecure` flag
- [ ] **`SSL_CTX_set_verify` not called** — default is no verification
- [ ] **Weak cipher suites allowed** — need to call `SSL_CTX_set_cipher_list()`
- [ ] **Private key file permissions** not checked before loading
- [ ] **Hostname verification** not performed on server certificate
- [ ] **Session renegotiation** not limited (DoS vector)
- [ ] **TLS certificate/key paths from untrusted input** — can client specify arbitrary paths?
### 5. Memory Safety Issues
- [ ] **Use-after-free** — object freed but pointer still used later
- [ ] **Double-free** — `free()` called twice on same pointer
- [ ] **Memory leaks** on error paths — allocated but not freed before return
- [ ] **Integer overflow** in allocation size computation
```c
// DANGER: count * sizeof(Type) can overflow
void *arr = malloc(count * sizeof(Element));
// SAFE:
if (count > SIZE_MAX / sizeof(Element)) return NULL;
void *arr = malloc(count * sizeof(Element));
```
- [ ] **`realloc` return value** not saved to temporary pointer (leak on failure)
```c
// BAD: leaks original pointer on failure
buf = realloc(buf, new_size);
// GOOD:
void *tmp = realloc(buf, new_size);
if (!tmp) { free(buf); return NULL; }
buf = tmp;
```
### 6. Integer Overflow in Allocation
Check all size calculations:
- [ ] Allocations where count comes from network data (chunk count, file count, etc.)
- [ ] Allocations where size is multiplied by count
```bash
grep -rn 'malloc.*\*.*sizeof\|calloc(.*sizeof' src/ --include="*.c"
```
- [ ] Loop counters that could wrap (unsigned underflow)
- [ ] Signed integer overflow in size checks
### 7. Format String Vulnerabilities
- [ ] User-controlled data passed as format string
```c
printf(user_input); // VULNERABLE
fprintf(stderr, user_input); // VULNERABLE
syslog(LOG_INFO, user_input); // VULNERABLE
printf("%s", user_input); // SAFE
```
```bash
grep -rn 'printf(\|fprintf(\|syslog(\|snprintf(' src/ --include="*.c" | grep -v '"[^"]*%'
```
### 8. TOCTOU Race Conditions
- [ ] File existence check followed by open (Time-of-check to Time-of-use)
```c
if (access(path, F_OK) == 0) { // CHECK
fd = open(path, O_RDWR); // USE — file could have changed
}
```
- [ ] `stat()` followed by `open()` with different permissions
- [ ] Temporary file creation with predictable names
### 9. Insecure Temporary File Usage
- [ ] `mktemp` / `tmpnam` — use `mkstemp` instead
- [ ] Temporary files created in world-writable directories
- [ ] Temporary files not cleaned up on error paths
- [ ] Predictable temp file names (race + symlink attack)
### 10. Hardcoded Secrets / Credentials
- [ ] Hardcoded passwords, API keys, or tokens
- [ ] Hardcoded TLS private keys or certificates
- [ ] Hardcoded connection strings with embedded credentials
- [ ] Test certificates/keys in source tree (should be documented if intentional)
### 11. Denial of Service Vectors
- [ ] **Unbounded memory allocation** — can client request huge allocation that OOMs server?
- Check `chunk.c` for chunk count limits
- Check `protocol.c` for message size limits
- Check `config.c` for config field size limits
- [ ] **No connection limits** — server doesn't cap concurrent connections
- [ ] **No timeouts** — connections can hang indefinitely
- [ ] **Recursive parsing** — could cause stack overflow with crafted input
- [ ] **Repeated slow reads** — slow loris style attack
- [ ] **Fork bomb** — server forks per connection without limit
### 12. Information Disclosure
- [ ] Server sends detailed error messages to client (path disclosure, version info)
- [ ] Debug logging enabled in production
- [ ] Stack traces leaked to users
- [ ] Timing side channels in authentication or comparison
## How to Scan
### Automated Pattern Search
Run these searches across the codebase:
```bash
# Buffer overflow risks
grep -rn '\bstrcpy\b\|\bstrcat\b\|\bsprintf\b' src/ --include="*.c"
# Fixed size stack buffers
grep -rn 'char [a-z_]*\[[0-9]*\];' src/ --include="*.c" --include="*.h"
# Format string risks
grep -rn 'printf(\|fprintf(\|syslog(' src/ --include="*.c" | grep -v '"[^"]*%'
# Malloc without null check pattern
grep -rn '= malloc\|= calloc\|= realloc' src/ --include="*.c"
# Integer overflow in allocation
grep -rn 'malloc.*\*\|calloc.*<' src/ --include="*.c"
# Path construction
grep -rn 'snprintf.*path\|snprintf.*dir' src/ --include="*.c"
```
### Manual Code Review
After automated scanning, manually review high-risk files:
1. `src/shared/protocol.c` — all receive paths
2. `src/shared/config.c` — deserialization logic
3. `src/shared/chunk.c` — chunk parsing
4. `src/shared/transport_tls.c` — TLS configuration
5. `src/server/server.c` — file writing and connection handling
## Output Format
Return findings in this structured format, one per vulnerability:
```
## Finding: <Short descriptive title>
- **Severity**: critical/high/medium/low
- **Category**: security
- **Location**: file:line range
- **Description**: what the vulnerability is, including:
- How it can be triggered
- What the impact is (RCE, DoS, info leak, etc.)
- Whether it requires authentication
- **Suggestion**: how to fix it, including concrete code changes
- **Labels**: security, comma-separated additional labels
```
### Example
```
## Finding: Unchecked malloc in chunk deserialization allows OOM
- **Severity**: high
- **Category**: security
- **Location**: src/shared/chunk.c:45-50
- **Description**: `chunk_deserialize()` calls `malloc(count * sizeof(File))`
where `count` comes directly from the network. An attacker can send a crafted
chunk header with an extremely large count (e.g., UINT32_MAX), causing malloc
to either fail (crash if unchecked) or allocate enormous memory (OOM).
No authentication needed — the attack works on the initial connection.
- **Suggestion**: Add bounds checking before allocation:
```c
if (count > MAX_CHUNK_FILES || count > SIZE_MAX / sizeof(File)) {
log_error("Invalid chunk file count: %u", count);
return NULL;
}
```
Define `MAX_CHUNK_FILES` as a reasonable limit (e.g., 100000).
- **Labels**: security, dos
```
### No Findings
If no security issues are found, return:
```
## No security findings
The codebase appears clean in the areas checked. No vulnerabilities found at this time.
```
## Severity Guidelines
| Severity | Definition | Example |
|---|---|---|
| **critical** | Remote code execution, unauthenticated compromise | Buffer overflow on network input |
| **high** | Significant impact but requires specific conditions | DoS via unbounded allocation, path traversal |
| **medium** | Limited impact, requires auth or other conditions | TOCTOU race in file operations |
| **low** | Minor issues, defense in depth | Missing null check that's unlikely to trigger |
| **informational** | Not exploitable but violates best practice | Hardcoded value that could be configurable |
## CI & Task Execution
When using `tea` (the task execution agent) to run CI or tests, always set a sufficient timeout (e.g., 600000ms) to allow the workflow to finish. After CI completes, check the results yourself — inspect logs if the run failed. Never assume success.
## Branch Strategy
Never push directly to `main`. All changes must be developed on a feature branch and merged via a pull request. Always create a new branch (`git checkout -b <branch-name>`) before making changes, push it, and open a PR with `gh pr create --fill`. Wait for CI to pass before merging.
## Dependency Installation
**CI rule:** never add `apt-get install` / `pip install` steps to CI workflows — use the custom Docker image instead. **Host rule:** for local development, use `nix-shell` (see `README.md`) which provides zstd, OpenSSL, CMake, and gcc. See `AGENTS.md` for details.
+3 -5
View File
@@ -138,11 +138,9 @@ int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size) {
Build for fuzzing:
```bash
cmake -B build-fuzz -S . \
-DCMAKE_C_FLAGS="-fsanitize=fuzzer,address,undefined -g" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=fuzzer,address,undefined"
CC=clang CXX=clang++ cmake -B build-fuzz -S . -DENABLE_FUZZ=ON
cmake --build build-fuzz -j$(nproc)
./build-fuzz/tests/fuzz_chunk_deserialize corpus/ -max_len=1048576
./build-fuzz/fuzz_chunk_deserialize corpus/ -max_len=1048576
```
### AFL++ Harness
@@ -216,7 +214,7 @@ When using `tea` (the task execution agent) to run CI or tests, always set a suf
## Branch Strategy
Never push directly to `main`. All changes must be developed on a feature branch and merged via a pull request. Always create a new branch (`git checkout -b <branch-name>`) before making changes, push it, and open a PR with `gh pr create --fill`. Wait for CI to pass before merging.
Never push directly to `dev` or `main`. All changes must be developed on a feature branch and merged via a pull request targeting `dev`. Create a branch (`git checkout -b <branch-name>`), push it, and open the PR with `tea pr create --repo TapTap/FastSync --base dev --head <branch-name>`. Wait for CI to pass before merging.
## Dependency Installation
+29 -16
View File
@@ -38,16 +38,20 @@ dd if=/dev/urandom of=/tmp/fastsync_bench/src/large.bin bs=1M count=10 2>/dev/nu
Test each configuration 3 times, record median:
```bash
# Real FastSync flags: -z=compression, -j=multithreading,
# --chunk-serialization, --sendfile (long form only). The old rsync-style
# spellings -c/-m/-s/-f are NOT the same options (-c=--checksum,
# -m=--prune-empty-dirs, -s=--secluded-args, -f=--filter) and must not be used.
CONFIGS=(
"Standard|"
"Compression|-c"
"Multithreading|-m"
"MT+Compression|-m -c"
"Chunk Serialization|-s"
"MT+Compression+Chunk|-m -c -s"
"Sendfile|-f"
"Compression|-z"
"Multithreading|-j"
"MT+Compression|-j -z"
"Chunk Serialization|-j -z --chunk-serialization"
"Sendfile|--sendfile"
)
PORT=18080
for config in "${CONFIGS[@]}"; do
IFS='|' read -r name flags <<< "$config"
echo "=== $name ==="
@@ -55,13 +59,14 @@ for config in "${CONFIGS[@]}"; do
rm -rf /tmp/fastsync_bench/dst
mkdir -p /tmp/fastsync_bench/dst
./build/server &
./build/server -p "$PORT" --allow-unauthenticated &
SERVER_PID=$!
sleep 0.5
START=$(date +%s%N)
./build/client --source-dir /tmp/fastsync_bench/src \
--dest-dir /tmp/fastsync_bench/dst \
--server-port "$PORT" \
--save-to-disk $flags
END=$(date +%s%N)
@@ -74,14 +79,21 @@ for config in "${CONFIGS[@]}"; do
done
```
### Step 4: Full Integration Benchmark (Optional)
### Step 4: Full Benchmark Tool (Preferred)
The maintained benchmark tool is `benchmark/bench.py`. It handles building,
data generation, network shaping (LAN/WAN profiles or custom `--delay`/`--jitter`/
`--throughput`/`--loss`), rsync comparison, and JSON/table reporting:
For comprehensive benchmarking with network shaping:
```bash
python3 test.py --full
python3 benchmark/bench.py --help
python3 benchmark/bench.py --runs 5 --profiles unlimited
python3 benchmark/bench.py --size-mb 100 --random-ratio 0.5 --output json
python3 benchmark/bench.py --delay 50ms --jitter 10ms --throughput 100mbit
```
This tests LAN/WAN profiles, SSH, TLS, and compares against rsync.
Network shaping needs root (`tc`/`netem` on `lo`). SSH and TLS coverage lives in
the pytest integration suite, not the benchmark tool.
### Step 5: Report Results
@@ -93,12 +105,13 @@ Platform: <OS, CPU, network>
Configuration | Run 1 | Run 2 | Run 3 | Median
-----------------------|---------|---------|---------|--------
Standard | 0.12s | 0.11s | 0.12s | 0.12s
Compression (-c) | 0.09s | 0.08s | 0.09s | 0.09s
Multithreading (-m) | 0.07s | 0.07s | 0.08s | 0.07s
MT+Compression (-m -c) | 0.05s | 0.05s | 0.06s | 0.05s
Sendfile (-f) | 0.04s | 0.04s | 0.04s | 0.04s
Compression (-z) | 0.09s | 0.08s | 0.09s | 0.09s
Multithreading (-j) | 0.07s | 0.07s | 0.08s | 0.07s
MT+Compression (-j -z) | 0.05s | 0.05s | 0.06s | 0.05s
Chunk Serialization (--chunk-serialization) | 0.05s | 0.04s | 0.05s | 0.05s
Sendfile (--sendfile) | 0.04s | 0.04s | 0.04s | 0.04s
Best configuration: MT+Compression (-m -c)
Best configuration: Sendfile (--sendfile)
Throughput: <X> MB/s
```
+12 -17
View File
@@ -32,23 +32,19 @@ Try to reproduce the issue with the exact command the user provides.
**Memory errors (first priority):**
```bash
rm -rf build
cmake -B build -S . \
-DCMAKE_C_FLAGS="-fsanitize=address -fno-omit-frame-pointer -g" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address"
cmake --build build -j$(nproc)
./build/tests
rm -rf build-asan
cmake -B build-asan -S . -DSANITIZER=address
cmake --build build-asan -j$(nproc)
./build-asan/tests
# or run the failing command
```
**Thread errors:**
```bash
rm -rf build
cmake -B build -S . \
-DCMAKE_C_FLAGS="-fsanitize=thread -g" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=thread"
cmake --build build -j$(nproc)
./build/tests
rm -rf build-tsan
cmake -B build-tsan -S . -DSANITIZER=thread
cmake --build build-tsan -j$(nproc)
./build-tsan/tests
```
**Valgrind (if ASan doesn't find it):**
@@ -108,13 +104,12 @@ cmake -B build -S . && cmake --build build -j$(nproc)
./build/tests
# If integration test needed
python3 test.py
python3 -m pytest tests/integration/ -n 4 --dist=load -m "not setpriv"
# Re-run under sanitizer to confirm fix
rm -rf build
cmake -B build -S . -DCMAKE_C_FLAGS="-fsanitize=address -fno-omit-frame-pointer" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address"
cmake --build build -j$(nproc)
rm -rf build-asan
cmake -B build-asan -S . -DSANITIZER=address
cmake --build build-asan -j$(nproc)
# reproduce the original failing command
```
+5 -9
View File
@@ -19,7 +19,7 @@ tea pr checkout <number>
If already on a PR branch, verify with:
```bash
git branch --show-current
git log main..HEAD --oneline
git log dev..HEAD --oneline
```
### Step 2: Clean build
@@ -39,17 +39,13 @@ If the PR touches threading, memory management, or network code, also build with
```bash
# AddressSanitizer
rm -rf build-asan
cmake -B build-asan -S . \
-DCMAKE_C_FLAGS="-fsanitize=address -fno-omit-frame-pointer -g" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address"
cmake -B build-asan -S . -DSANITIZER=address
cmake --build build-asan -j$(nproc)
./build-asan/tests
# ThreadSanitizer (if threading changes)
rm -rf build-tsan
cmake -B build-tsan -S . \
-DCMAKE_C_FLAGS="-fsanitize=thread -g" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=thread"
cmake -B build-tsan -S . -DSANITIZER=thread
cmake --build build-tsan -j$(nproc)
./build-tsan/tests
```
@@ -91,10 +87,10 @@ If tests fail:
### Step 6: Run integration tests (optional)
```bash
python3 test.py
python3 -m pytest tests/integration/ -n 4 --dist=load -m "not setpriv"
```
This runs the integration + benchmark suite. It takes longer — only run if the user asks or if unit tests pass.
This runs the integration suite (benchmarking is `benchmark/bench.py`). It takes longer — only run if the user asks or if unit tests pass.
### Step 7: Fix and commit
+3 -3
View File
@@ -19,13 +19,13 @@ tea pr checkout <number>
If already on a PR branch, verify with:
```bash
git branch --show-current
git log main..HEAD --oneline
git log dev..HEAD --oneline
```
### Step 2: Get changed files
```bash
git diff main --name-only -- '*.c' '*.h'
git diff dev --name-only -- '*.c' '*.h'
```
This gives the list of C source and header files changed in the PR.
@@ -125,7 +125,7 @@ STYLE: <count>
If the user wants to post the review as a PR comment:
```bash
tea pr comment <number> --comment "<review report>"
tea comment --repo TapTap/FastSync <number> "<review report>"
```
## Rules
+19 -10
View File
@@ -16,7 +16,7 @@ Ask the user or determine from context:
- **Minor** (x.Y.0) — new features, backward compatible
- **Patch** (x.y.Z) — bug fixes, no protocol changes
Current version: `PROTOCOL_VERSION "1.1.0"` in `src/shared/config.h`
Current version: `PROTOCOL_VERSION "2.26.0"` in `src/shared/config.h`
### Step 2: Check Protocol Version
@@ -37,7 +37,7 @@ rm -rf build
cmake -B build -S .
cmake --build build -j$(nproc)
./build/tests
python3 test.py
python3 -m pytest tests/integration/ -n 4 --dist=load -m "not setpriv"
```
ALL tests must pass before release.
@@ -46,12 +46,10 @@ ALL tests must pass before release.
```bash
# ASan
rm -rf build
cmake -B build -S . \
-DCMAKE_C_FLAGS="-fsanitize=address -fno-omit-frame-pointer" \
-DCMAKE_EXE_LINKER_FLAGS="-fsanitize=address"
cmake --build build -j$(nproc)
./build/tests
rm -rf build-asan
cmake -B build-asan -S . -DSANITIZER=address
cmake --build build-asan -j$(nproc)
./build-asan/tests
```
### Step 5: Update README (If Needed)
@@ -79,12 +77,23 @@ git commit -m "Release vX.Y.Z
git tag -a vX.Y.Z -m "Release vX.Y.Z"
```
### Step 8: Push
### Step 8: Push and Open dev → main PR
`main` is protected and only receives changes via `dev` → `main` PRs (see AGENTS.md). Never push directly to `main`.
```bash
git push origin main --tags
# Push the release commit and tag to dev
git push origin dev
git push origin vX.Y.Z
# Open the dev → main release PR for review + CI
tea pr create --repo TapTap/FastSync --head dev --base main \
--title "Release vX.Y.Z" \
--description "Release vX.Y.Z"
```
Then wait for the full CI to pass and request review before the PR is merged to `main`.
### Step 9: Report
```
+2 -2
View File
@@ -102,9 +102,9 @@ Informational: <count>
...
=== VERDICT ===
[PASS] No critical/high issues found
[PASS] No critical/high-severity issues found
— or —
[FAIL] <N> critical/high issues must be fixed
[FAIL] <N> critical/high-severity issues must be fixed
```
## Rules
+12 -11
View File
@@ -4,18 +4,19 @@ FastSync is a high-performance file synchronization system written in C11. It su
## Dependency installation
**CI rule:** never add `apt-get install` / `pip install` steps to CI workflows — use the custom Docker image instead. The image is built from the repo-root `Dockerfile` and is the same image CI uses: `gitea.tap-tap.win/taptap/fastsync-ci:v10`. It contains the full toolchain: gcc/g++, CMake, libzstd-dev, libssl-dev, make, git, cppcheck, clang-format, python3 + pytest + pytest-xdist, openssh-client, and Node.js.
**CI rule:** never add `apt-get install` / `pip install` steps to CI workflows — use the custom Docker image instead. The image is built from the repo-root `Dockerfile` and is the same image CI uses: `gitea.tap-tap.win/taptap/fastsync-ci:v11`. It contains the full toolchain: gcc/g++, CMake, libzstd-dev, libssl-dev, make, git, cppcheck, clang-format, python3 + pytest + pytest-xdist, openssh-client, Node.js, plus `rsync` 3.4.1 (with zstd/xxhash/lz4), `acl` and `attr` (setfacl/getfacl, setfattr/getfattr) for drop-in parity tests.
**Host rule:** for local development, use `nix-shell` (see `README.md`) which provides zstd, OpenSSL, CMake, and gcc. The Docker image can also be used locally for CI parity.
```bash
# Use the prebuilt CI image directly (faster, guaranteed CI parity)
docker pull gitea.tap-tap.win/taptap/fastsync-ci:v10
docker tag gitea.tap-tap.win/taptap/fastsync-ci:v10 fastsync-ci:local
docker pull gitea.tap-tap.win/taptap/fastsync-ci:v11
docker tag gitea.tap-tap.win/taptap/fastsync-ci:v11 fastsync-ci:local
# Or build the image from the repo-root Dockerfile
# (Note: the prebuilt :v10 image reflects the previous Dockerfile state;
# rebuild from source to pick up any newly added packages like lcov/valgrind.)
# (Note: the prebuilt :v11 image is built from the current Dockerfile and
# includes rsync 3.4.1 plus acl/attr; rebuild from source after changing
# the Dockerfile.)
docker build -t fastsync-ci:local .
# Build, run unit tests, and run integration tests inside the container
@@ -28,7 +29,7 @@ docker run --rm --user "$(id -u):$(id -g)" -v "$PWD:/workspace" \
sh -c 'cmake -B build -S . && cmake --build build -j$(nproc) && ./build/tests && python3 -m pytest tests/integration/ -n 4 --dist=load'
```
> **Note:** The first `cmake configure` (`cmake -B build -S .`) fetches xxHash from GitHub via `FetchContent` — network access is required. Subsequent reconfigures reuse the cached source.
> **Note:** The first `cmake configure` (`cmake -B build -S .`) fetches xxHash via `FetchContent` — network access is required. Subsequent reconfigures reuse the cached source.
If a dependency is missing from the CI image, add it to the `Dockerfile` (and rebuild) rather than adding an install step to the CI workflow.
@@ -59,28 +60,28 @@ python3 -m pytest tests/integration/ -n 4 --dist=load -m ci # PR-gate subset o
## CI Workflow — Waiting for Results
When running the CI workflow via `tea` (the task execution agent), always set a sufficient timeout (e.g., 600000ms) to allow CI to finish. After CI completes, check the results yourself — do not assume success. Use `gh run watch` or similar to monitor CI status, then inspect logs on failure.
When running the CI workflow via `tea` (the task execution agent), always set a sufficient timeout (e.g., 600000ms) to allow CI to finish. After CI completes, check the results yourself — do not assume success. Monitor CI status via the Gitea API (see below) or `tea actions`, then inspect logs on failure.
## CI Troubleshooting
### If lint (clang-format) fails
Run clang-format in the CI Docker image to match the exact CI version:
```bash
docker run --rm -v "$PWD:/workspace" -w /workspace gitea.tap-tap.win/taptap/fastsync-ci:v10 \
docker run --rm -v "$PWD:/workspace" -w /workspace gitea.tap-tap.win/taptap/fastsync-ci:v11 \
sh -c 'find src/ tests/ -name "*.c" -o -name "*.h" | xargs clang-format -i'
```
### If cppcheck fails
Fix reported issues locally, then verify with:
```bash
docker run --rm -v "$PWD:/workspace" -w /workspace gitea.tap-tap.win/taptap/fastsync-ci:v10 \
docker run --rm -v "$PWD:/workspace" -w /workspace gitea.tap-tap.win/taptap/fastsync-ci:v11 \
sh -c 'cppcheck --enable=warning,style,performance,portability --suppress=missingIncludeSystem --error-exitcode=1 --inline-suppr src/ tests/'
```
### If integration tests fail
Run locally before pushing:
```bash
python3 -m pytest tests/ -v --tb=short
python3 -m pytest tests/integration/ -n 4 --dist=load -m "not setpriv"
```
## Branch Strategy
@@ -165,7 +166,7 @@ This can be cron'd locally if desired (e.g., `crontab -e` with `opencode run`).
## Is opencode a good option?
**Yes, for FastSync's needs.** The hybrid model works well:
- opencode's 17 specialized agents handle deep code analysis, fixes, tests, and reviews
- opencode's 16 specialized agents handle deep code analysis, fixes, tests, and reviews
- The assistant orchestrates subagents, merges branches, and iterates on CI
- You only review the final output
+380
View File
@@ -0,0 +1,380 @@
# Changelog
All notable changes to FastSync are documented here. Versions match
`PROTOCOL_VERSION` (printed by `fastsync --version`); the client and server must
run the same version because the handshake is strict.
## [2.26.0] - 2026-09-17
### Added
- **Parity-completion wave.** Closed the remaining rsync-parity gaps against
rsync 3.4.1 and reclassified the inherently non-rsync rows. It moved the wire
protocol three times (`2.23.0 → 2.24.0 → 2.25.0 → 2.26.0`).
- **Delete timing (2.24.0):** per-directory delete plans
(`STATUS_DELETE_PLAN`) for `--delete-during`/`--delete-delay`. An interrupted
during-transfer has already removed the reached directories' extras, while a
delayed transfer commits per directory only after the whole transfer
succeeds (a late-created extra survives `--delete-delay` but not
`--delete-after`). `-R --delete` is scoped to the transferred prefix; empty
in-scope source directories survive; dry-run never deletes.
- **Wire stats (2.25.0):** `STATUS_STATS` carries the receiver counters
(matched data, deleted files) and the dry-run would-delete list. `--stats`
prints rsync's protocol-independent lines; `--progress`/`-P` print per-file
blocks; `--out-format` gains `%b` (wire bytes), `%c` (block-sum bytes) and
`%C` (whole-file digest); `-n --delete` prints escaped `*deleting` lines in
the sequential and `--threads` paths.
- **Codecs (2.26.0):** `lz4`/`zlib`/`zlibx` compression and `md4`/`sha1`/
`none` checksums, with rsync-style `auto` negotiation (default `xxh128` +
`zstd`) and exit-4 rejection of unknown names; the resolved `compression_algo`
crosses the wire.
- General `-R`/`--relative` (including the `/./` cut) and `--no-implied-dirs`;
one-level `-d`/`--dirs` listing for `dir`, `dir/` and `.`; the full filter
grammar (`merge`/`dir-merge`/`hide`/`show`/`protect`/`risk`/`clear` and
modifiers) with `-f` bound to `--filter`; a single `-F` transfers
`.rsync-filter` and `-FF` excludes it.
- Receiver-side `--chown`/`--usermap`/`--groupmap` TO-name resolution; absolute
basis directories and a `--link-dest` relink of an up-to-date destination;
a receiver-side `--ignore-existing` short-circuit before any payload;
`--preallocate` now wins over `--sparse` via `fallocate(2)`.
- Client quick wins: `--iconv=.`/`-`/`--no-iconv`, a lone `-h` prints help, an
empty `--files-from` succeeds (exit 0), a broken referent under
`-L`/`--copy-unsafe-links` exits 23, the full `--info`/`--debug`
vocabularies, and the aliases `--ignore-non-existing`, `--protect-args`,
`--msgs2stderr`.
### Changed
- `PROTOCOL_VERSION` bumped `2.23.0 → 2.24.0` (delete plans),
`2.24.0 → 2.25.0` (`STATUS_STATS` + `report_stats`), and
`2.25.0 → 2.26.0` (codec negotiation + `md4`/`sha1`/`none`).
- `--checksum-choice`/`--cc` now accepts `md4`, `sha1`, `none` and the two-name
form; the negotiated whole-file default is `xxh128`.
- `--compress-choice`/`--zc` now accepts `lz4`, `zlib`, `zlibx`.
- `RSYNC_COMPAT.md` reclassifies the matrix: 9 already-parity rows to ✅, 17
inherently non-rsync rows to ❌ (native daemon config/auth, batch, privileged
xattr namespaces, and the safe-subset device/privilege flags), and the genuine
fixes to ✅; new rows cover `--bwlimit`, `--partial`, `--partial-dir`,
`--no-whole-file`, `--inc-recursive`/`--no-inc-recursive`, `--protect-args`
and `--msgs2stderr`.
- The client `--help` `--max-delete` text now describes the implemented partial
semantics (delete up to N, skip the rest, exit 25).
### Notes
- Remaining documented divergences include the `--stats` per-type file-count
breakdown, `%b`/`%c` being FastSync wire counts, `-n --delete` line ordering,
the default `--delete` timing (delete-after, not rsync's delete-during),
destination-only exclude protection (still sender-derived), `--temp-dir`
absolute paths, basis-dir attribute re-application and the 256 MiB whole-file
cap, `--fuzzy` tie-breaking, `--bwlimit=0`/decimal rates, `zlibx`==`zlib`, and
recursive empty-directory creation.
- Build: adds zlib and lz4 as link dependencies.
## [2.23.0] - 2026-09-16
### Added
- **Rsync-parity wave.** Closed the remaining CLI, filesystem, ownership,
deletion, and output gaps against rsync 3.4.1.
- Short options `-r` (`--recursive`), `-b` (`--backup`), `-L`
(`--copy-links`), and `-B` (`--block-size`/`--delta-block`); rsync
short-option clustering (`-av`, `-aAX`, `-rlpt`) and attached/inline values
(`--opt=value`, `-B1000`, `-essh`, `-MOPT`). A value that starts with `-`
is not mistaken for a cluster.
- `-c`/`--checksum` now implies the incremental checksum quick-check (and,
like rsync, does not imply `-t`).
- `--checksum-choice`/`--cc` accepts `xxh64`/`xxhash`/`xxh3`/`xxh128`/`md5`/
`auto` and rejects `md4`/`sha1`/`none` and the two-name form by name;
`--checksum-seed=0` (the default) is randomized per transfer and the chosen
seed is sent to the receiver.
- `--compress-choice`/`--zc` accepts `zstd`/`none`/`auto` and rejects
`lz4`/`zlib`/`zlibx` by name; `--skip-compress` defaults to rsync 3.4.1's
built-in suffix list; `--no-whole-file` is accepted.
- `--timeout` defaults to 0 (disabled) and `--contimeout` to 60 s (both `0`
disables), matching rsync; `--max-alloc=0` means no local limit.
- `--temp-dir` is confined to the receive root (absolute/`..` rejected by the
receiver) and an `EXDEV` install falls back to a non-atomic copy.
- `--numeric-ids` is documented as a mapping modifier only;
`--usermap`/`--groupmap` support inclusive `LOW-HIGH` ranges, `*`,
empty-`FROM` (unnamed ids), and receiver-resolved `TO` names; `--chown`
conflicts with a map on the same side are rejected.
- `--fake-super` records the *resolved* owner (never a real chown) and replays
mode/time; directory ownership and directory xattrs/ACLs are preserved.
- `-l`/`--links` stores symlink targets verbatim (absolute and `..`-bearing
included), matching rsync; `--safe-links`/`--copy-unsafe-links` are applied
sender-side and `--munge-links` uses rsync's `/rsyncd-munged/` marker;
`--trust-sender` no longer affects symlink targets.
- `--specials` recreates unix sockets with `mknod(S_IFSOCK)` (so `-D` covers
the full rsync node set).
- Deletion: the manifest carries a synchronized-directory section so
`--files-from` subsets no longer delete untransmitted paths;
`--delete-excluded` leaves size-pruned mirrors protected; extraneous
destination symlinks are unlinked (never followed); `--max-delete=N` is
partial (delete up to N, skip the rest, exit 25) and `--delete-missing-args`
removals draw from the same budget; `--force` is honored during
`--delay-updates` publication.
- `-x`/`--one-file-system` emits the mount-point directory entry; the
`--include`/`--exclude` layers are an ordered first-match rule list.
- `--chmod` is a faithful port of rsync 3.4.1 (numeric/symbolic, `D`/`F`/`X`,
`s`/`t`, append semantics, no `-p` implication, no sanitization).
### Changed
- `PROTOCOL_VERSION` bumped `2.22.0 → 2.23.0`: the delete manifest gains a
synchronized-directory section and the terminal status gains
`STATUS_DELETE_LIMIT` (client exit 25 on a `--max-delete`-capped commit).
- **The 2.22.0 mode-masking divergence is removed.** Under `-p` the source mode
is copied exactly, including `S_IWGRP`/`S_IWOTH` and setuid/setgid/sticky;
`--chmod` no longer implies `-p`. New files without `-p` still use
`source_mode & ~umask` when metadata is present (else `0644`), and new
directories without `-p` still use the `0755` creation default.
- `--protocol=NUM` accepts only the current `2.23.0` version string.
### Notes
- The rsync-compatibility matrix (`RSYNC_COMPAT.md`) now classifies every row
as **parity**, **caveat** (works with a documented divergence), or
**divergent** (not supported/no-op/impossible), replacing the previous
misleading "N implemented / 0 divergence" summary. Durable documented
divergences remain: receiver-side symlink target containment is not enforced
by default (verbatim storage is rsync parity; use `--safe-links`),
`--temp-dir` rejects absolute/foreign-filesystem paths, `--copy-devices`
reads a bounded `st_size`, a broken referent under `--copy-links` exits 0,
new directories without `-p` use `0755`, `--stats` receiver-only counters are
0, and `--password-file`/`--early-input`/`--hash-credentials`/`--iterations`
and the batch format are FastSync-native.
## [2.22.0] - 2026-09-15
### Added
- **Per-attribute metadata preservation (protocol 2.22.0).** The former single
metadata bundle is split into four independent, rsync-compatible flags:
`-p/--perms`, `-t/--times`, `-o/--owner`, and `-g/--group`, each applied
independently on the receiver, with negations `--no-perms`/`--no-times`/
`--no-owner`/`--no-group` (short `--no-p`/`--no-t`/`--no-o`/`--no-g`) and
`--no-preserve` clearing all four. `-a/--archive` is now full rsync
`-rlptgoD` (owner and group included; their application stays
privilege-gated). `-A/--acls` and `--chmod` imply `-p`, `-X/--xattrs` does
not, `-E/--executability` sets only executability, and `-U`/`-N` do not imply
`-t`. `--incremental`/`--delta` still auto-preserve perms+times unless the
user explicitly negated them.
- Receiver applies directory modes under `-p` (at the end of the transfer,
alongside the deferred directory times) and symlink mode under `-p`; `-O`
suppresses directory times only.
### Changed
- `PROTOCOL_VERSION` bumped `2.21.0 → 2.22.0`: the binary config frame gains
four appended booleans (`preserve_perms`/`preserve_times`/`preserve_owner`/
`preserve_group`) after `omit_link_times`. The fixed-width `FileMetadata`
layout is unchanged; the receiver derives the metadata-frame gate
(`use_metadata`) from the four attributes.
### Notes
- Documented divergences from rsync: a client-supplied mode never grants
group/other write (`S_IWGRP|S_IWOTH` are stripped for files, directories,
symlinks, and specials; rsync's `-p` preserves them exactly); a brand-new file
without `-p` gets `source_mode & ~umask` (sanitized) when metadata is present,
else the historical fixed `0644`; `--chmod` implies `-p` (rsync does not);
`-o`/`-g` map by name on the receiver with a raw-numeric fallback (only
numeric ids cross the wire); and a daemon module without `client owner = yes`
does not refuse a plain `-a`/`-o`/`-g` but forces super-user activities off,
applies no ownership, and logs a warning (explicit `--chown`/`--usermap`/
`--groupmap`/`--numeric-ids`/`--copy-as`/`--super` are still refused).
## [2.21.0] - 2026-09-14
### Added
- Optional server→client rejection detail (protocol 2.21.0). A rejected
operation may now carry a bounded human-readable reason via
`STATUS_ERROR_DETAIL` instead of a bare `STATUS_ERROR`, so the client can
report *why* the server refused (daemon module gate, config validation,
receiver-side path/node validation). `receive_status()` transparently maps the
new status back to `STATUS_ERROR` for every existing call site and captures
the reason into a thread-local buffer exposed by `protocol_last_error()`. The
detail body is always consumed, so the stream cannot desynchronize, and
messages are sliced to `MAX_ERROR_DETAIL_BYTES` (4096) on send.
- **Server-contacting `--dry-run` (protocol 2.21.0).** `--dry-run` now performs
a real handshake with a remote/daemon receiver and reports exactly what WOULD
change based on receiver state (existing destination files, mtimes, checksums,
basis dirs). The wire config carries the dry-run intent (`Config.dry_run`) and
the receiver answers each per-file check with `STATUS_DRY_RUN_TRANSFER` (would
transfer) or `STATUS_OK` (already up to date); the sender prints the
would-transfer set and its trailer without sending any file data. The receiver
performs the normal read-only incremental decision but mutates nothing: no temp
files, writes, renames, deletes, metadata/xattr/chown, or directory creation.
A plain local destination (no explicit `--server-port`/remote) keeps the
original client-side dry-run. Would-delete reporting for `--delete*` is
deferred to a follow-up; dry-run never deletes.
- Daemon `max connections per host` (per-source-IP concurrent cap, default 0 =
unlimited), `auth lockout threshold` (default 10; 0 disables) and
`auth lockout duration` (default 300 s) config keys.
- `fastsync-server --allow-super` opt-in for a privileged standalone TCP server;
without it a root standalone receiver forces super-user activities off (device
nodes, `--write-devices`, ownership). The `--stdio` SSH argv is client-composed,
so super activities always stay off there.
### Changed
- Config wire fields are now declared once in an X-macro table
(`CONFIG_WIRE_FIELDS` in `src/shared/config.h`) that generates the struct
members, defaults, and the send/receive sequence, removing the manual
six-site field sync. Wire bytes and `PROTOCOL_VERSION` are unchanged.
- `receive_incremental_check()` (the per-file `STATUS_CHECK` fast path) is split
into small static helpers with a short linear orchestrator. Pure refactor: the
wire byte stream and all cleanup are unchanged.
- `authorized_root` state has a single owner (`utils.c`) with read accessors; the
duplicated statics in `file.c` and the server were removed.
- `Data` records its owning `ProtocolSession` so its memory charge is returned to
the session that reserved it, regardless of the destroying thread.
- The receiver pipeline moved out of `shared` into `server/receiver_pipeline.[ch]`;
the build now uses explicit `fastsync_shared` / `fastsync_client_core` /
`fastsync_server_core` targets instead of a GLOB, and the client no longer links
server code.
- The benchmark tool generates the requested random/compressible data mix
accurately, verifies each transfer before recording it, computes correct
percentiles, adds a MB/s column, handles `tc`/netem without requiring `sudo`
when already root, builds into a dedicated `build-bench/` directory, and adds a
`--warm` incremental-transfer mode.
- The `nix-shell` dev environment provides the full toolchain (clang-format,
cppcheck, pytest-xdist, OpenSSH, rsync, iproute2, valgrind, lcov) and no longer
builds on entry.
### Security
- Enforce the daemon's per-module `max connections` cap (0 = unlimited) and add
the shared per-source `max connections per host` cap plus a cross-process
`auth lockout`. Because the listener forks one child per connection, the
counters live in an anonymous shared mapping created before the accept loop and
reclaimed by the parent's `SIGCHLD` handler, so the per-module, per-source and
auth-failure state is shared across every child (including after `SIGKILL`). The
per-source table has a bounded lifetime (expired/idle entries are reclaimed,
with a rate-limited warning when genuinely full), and the occupancy counters are
re-derived from the shared slot table on every child exit. Trusted loopback
peers are exempt (they share one address); clients behind a shared NAT/proxy
share a single per-host budget and lockout, which is documented.
- Hardening from a full security audit:
- Fail a truncated zstd frame instead of spinning forever (remote DoS).
- Open receiver destination/basis/hard-link entries `O_NONBLOCK` so a
client-planted FIFO cannot block a worker indefinitely.
- Require a regular file before `--inplace` writes, closing a FIFO-hang and a
raw-device write that bypassed the `--write-devices` gate.
- Reject SSH destinations whose user/host begins with `-` and insert `--` before
the host token, closing `-o ProxyCommand=…` argument injection (RCE).
- Gate client `--force` recursive removal behind the server `--allow-delete`
policy.
- Reject empty `hosts allow`/`hosts deny`/`auth users` values instead of
silently meaning "unrestricted".
- Restrict TLS 1.2 to AEAD suites and set server cipher preference; load the
private key TOCTOU-safely from an `O_NOFOLLOW` fd; verify IP literals against
IP SANs; guard client-cert CN truncation.
- Make `--dry-run` content-blind: it neither reads destination files nor
hashes basis files, removing a 1-bit content oracle against `read only`
modules.
- Bound glob matching (iterative DP, no exponential backtracking) and bound
line reads for filter/`--files-from`/pattern files.
- Gate `system.posix_acl_*` xattrs on `--acls` and charge decompression/chunk
allocations against the per-connection memory budget.
### Fixed
- Pre-auth NULL dereference in `config_delete()` when an over-long
`basis_count` (and the analogous count fields) was received and then failed
validation; received counts are now validated before being published.
- Leaked inherited `Data` in the forked compression-truncation unit test
(valgrind definite leak).
- `receive_status()` no longer loses a captured rejection reason when owed
keepalives are drained.
## [2.20.0] - 2026-09-13
### Security
- Cap cumulative `DirTimeList` growth and bound pre-auth config-string memory
(remote memory-exhaustion DoS).
- Daemon host access control (`hosts allow`/`hosts deny`, IPv4/IPv6/CIDR),
configurable global `max connections`, connection audit logging, and a
bounded `auth failure delay` throttle. IPv4-mapped peers are normalized and
invalid patterns are rejected at parse time (no silent fail-open).
- Honor `--timeout` for protocol I/O and bound idle/session time to defeat
keepalive slowloris; child-safe signal handling in the forked daemon.
- Compiler/linker hardening (`_FORTIFY_SOURCE`, stack protector, PIE, RELRO)
and pinned build dependencies.
### Fixed
- Use-after-free in the basis-dir oversize preflight.
- Placeholder `Data` leaks, `missing_args` leak, scanner chunk leak.
- Thread-safe logging; single fd owner and cleanup epilogue in the server
handler.
### Performance
- Metadata now crosses the wire as one packed frame (protocol 2.20.0).
- Delete keep-set and `--files-from` lookups indexed (O(n*m) → O(n)).
- Reused per-thread zstd contexts; `TCP_NODELAY` by default.
- Byte-bounded sender queues; removed a redundant scanner `stat()`.
## [2.19.0] - 2026-09-12
### Security
- **Daemon authentication rewritten as SCRAM-SHA-256 challenge/response**
(`STATUS_AUTH_CHALLENGE` → `STATUS_AUTH_RESPONSE` → `STATUS_AUTH_OK`/`STATUS_AUTH_FAILED`),
replacing the old replayable static `SHA-256(password)` bearer credential.
Each proof is bound to a fresh per-connection server nonce plus a client
nonce, so a captured response can never be reused.
- **Salted verifier store.** `--password-file`/`--early-input` now hold
`user:$fastsync$1$pbkdf2-sha256$<iters>$<salt>$<stored_key>$<server_key>`
(PBKDF2-HMAC-SHA256, default 600000 iterations, range 100000–10000000). The
legacy `user:SHA256HEX` form is hard-rejected; there is no auto-upgrade.
Generate stores offline with `fastsync-server --hash-credentials FILE
[--iterations N]`.
- **Username-enumeration hardening.** Unknown/off-list users are answered with a
dummy verifier whose salt is a deterministic per-username value
(`HMAC-SHA256(dummy_key, username)`), using the store-wide uniform iteration
count and a constant-time full-length membership scan. The dummy key is
persisted in an owner-only `<store>.dummykey` sidecar (atomic publish, exact
mode 0600) so challenges are stable across restarts.
- **Verified transport for auth-required modules.** A module with `auth users`
accepts credentials only over verified TLS whose client certificate matches
`--client-cn`, or — when `--allow-unauthenticated` is explicitly set —
plaintext from a loopback peer. Remote plaintext is refused before any
challenge. Clients must use `--tls` to send `--password-file` credentials to a
non-loopback daemon; `--client-cn` is mandatory with `--tls`.
- **Secret hygiene.** The plaintext password, derived keys, nonces/proofs and
the dummy key are wiped from memory on every path and never logged.
- Carried-over hardening: `-K` TOCTOU-safe directory walk
(`openat(O_NOFOLLOW)` per component), always shell-quoted SSH remote path,
TLS compression/renegotiation disabled, race-free (open-then-`fstat`)
`--password-file`/`--early-input` checks, log-injection escaping, and lazy
protocol debug escaping.
### Added
- `fastsync-server --hash-credentials FILE [--iterations N]` offline tool.
- `<store>.dummykey` sidecar (auto-created, owner-only, 0600).
- Integration tests for auth replay rejection, malformed frames, legacy-store
refusal, and the loopback/TLS transport policy; fuzz targets for config
receive and daemon-auth parsing.
### Changed
- **Protocol version 2.18.0 → 2.19.0 (breaking).** The config-frame auth block
is now `[present][username]` (digest removed) and the auth challenge/response
frames are interleaved between the config frame and its `STATUS_OK`. A 2.19.0
client and a 2.18.0 server (or vice versa) fail cleanly at the handshake.
- Daemon modules declaring `auth users` require a configured credential store at
startup (fail closed); operators regenerate stores from plaintext with
`--hash-credentials`.
### Notes
- First tagged release. FastSync implements rsync-compatible file
synchronization over TCP and SSH with TLS (OpenSSL), streaming zstd
compression, multithreaded transfers, and incremental sync. See
[RSYNC_COMPAT.md](RSYNC_COMPAT.md) for the flag-parity matrix.
+212 -26
View File
@@ -1,6 +1,6 @@
cmake_minimum_required(VERSION 3.22)
project(FastFileTransfer)
project(FastFileTransfer VERSION 2.26.0)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
set(CMAKE_C_STANDARD 11)
@@ -38,11 +38,24 @@ if(ENABLE_COVERAGE)
add_link_options(--coverage)
endif()
# --- Build hardening option ---
# Production hardening is applied to the shipping server/client binaries only,
# and only when no sanitizer or coverage instrumentation is active: sanitizers
# carry their own instrumentation, and _FORTIFY_SOURCE requires an optimising
# build (never the -O0 used for coverage).
option(ENABLE_HARDENING "Enable compiler/linker hardening for production targets" ON)
set(HARDENING_ACTIVE OFF)
if(ENABLE_HARDENING AND SANITIZER STREQUAL "none" AND NOT ENABLE_COVERAGE)
set(HARDENING_ACTIVE ON)
endif()
include(FetchContent)
FetchContent_Declare(
xxhash
GIT_REPOSITORY https://github.com/Cyan4973/xxHash
GIT_TAG v0.8.3
# v0.8.3 is a lightweight tag pointing at this exact commit (no ^{} peel
# entry); pin the commit SHA instead of the mutable tag.
GIT_TAG e626a72bc2321cd320e953a0ccf1584cad60f363 # v0.8.3
SOURCE_SUBDIR cmake_unofficial
)
FetchContent_MakeAvailable(xxhash)
@@ -55,37 +68,194 @@ if(NOT ZSTD_LIBRARY)
message(FATAL_ERROR "zstd library not found. Ensure it is in your nix-shell!")
endif()
find_library(ZLIB_LIBRARY z)
if(NOT ZLIB_LIBRARY)
message(FATAL_ERROR "zlib library not found. Ensure zlib1g-dev / nix zlib is available!")
endif()
find_library(LZ4_LIBRARY lz4)
if(NOT LZ4_LIBRARY)
message(FATAL_ERROR "lz4 library not found. Ensure liblz4-dev / nix lz4 is available!")
endif()
find_package(OpenSSL REQUIRED)
file(GLOB SHARED_SRCS "src/shared/*.c")
set(FILE_STORE_SRCS "${CMAKE_CURRENT_SOURCE_DIR}/src/shared/file_store.c")
list(REMOVE_ITEM SHARED_SRCS ${FILE_STORE_SRCS})
file(GLOB SERVER_SRCS "src/server/*.c")
set(SERVER_RECEIVER_SRCS src/server/receiver.c)
file(GLOB CLIENT_SRCS "src/client/*.c")
# --- Explicit source lists ---
# The shared library is self-contained: it must never depend on the client or
# server modules. In particular, the receiver pipeline (receive_thread /
# write_thread) lives under src/server, not here, so the client executable can
# link the shared library without pulling in any server code.
set(SHARED_SRCS
src/shared/array_list.c
src/shared/batch.c
src/shared/charset.c
src/shared/checksum.c
src/shared/chmod.c
src/shared/chunk.c
src/shared/compression.c
src/shared/config.c
src/shared/credentials.c
src/shared/daemon_conf.c
src/shared/daemon_limits.c
src/shared/data.c
src/shared/delay_updates.c
src/shared/delete_plan.c
src/shared/delta.c
src/shared/file.c
src/shared/file_list.c
src/shared/file_receive.c
src/shared/file_send.c
src/shared/file_store.c
src/shared/filter.c
src/shared/format.c
src/shared/hardlink.c
src/shared/identity.c
src/shared/log.c
src/shared/metadata.c
src/shared/motd.c
src/shared/multiprocessing.c
src/shared/protocol.c
src/shared/queue.c
src/shared/stop_condition.c
src/shared/transport_ssh.c
src/shared/transport_tcp.c
src/shared/transport_tls.c
src/shared/utils.c
src/shared/xattr.c
)
# Server implementation (no main): the receiver read/write pipeline plus the
# CLI parser. The server executable adds its own main (server.c).
set(SERVER_CORE_SRCS
src/server/receiver.c
src/server/receiver_pipeline.c
src/server/server_cli.c
)
set(SERVER_MAIN_SRCS src/server/server.c)
# Client implementation (no main): everything except the CLI entry point.
set(CLIENT_CORE_SRCS
src/client/change_list.c
src/client/client_send.c
src/client/client_validation.c
src/client/scanner.c
src/client/usage.c
)
set(CLIENT_MAIN_SRCS src/client/client_cli.c)
# --- Library targets ---
add_library(fastsync_shared STATIC ${SHARED_SRCS})
target_include_directories(fastsync_shared PUBLIC src/shared)
target_link_libraries(fastsync_shared PUBLIC Threads::Threads ${ZSTD_LIBRARY} ${ZLIB_LIBRARY}
${LZ4_LIBRARY} OpenSSL::SSL OpenSSL::Crypto xxhash)
add_library(fastsync_client_core STATIC ${CLIENT_CORE_SRCS})
target_include_directories(fastsync_client_core PUBLIC src/client)
target_link_libraries(fastsync_client_core PUBLIC fastsync_shared)
add_library(fastsync_server_core STATIC ${SERVER_CORE_SRCS})
target_include_directories(fastsync_server_core PUBLIC src/server)
target_link_libraries(fastsync_server_core PUBLIC fastsync_shared)
# --- Main executables ---
add_executable(server ${SERVER_SRCS} ${SHARED_SRCS} ${FILE_STORE_SRCS})
target_include_directories(server PRIVATE src/shared src/server src/client)
target_link_libraries(server PRIVATE Threads::Threads ${ZSTD_LIBRARY} OpenSSL::SSL OpenSSL::Crypto xxhash)
# The client links only the shared library and its own core; it deliberately
# does NOT get src/server on its include path nor compile receiver.c.
add_executable(server ${SERVER_MAIN_SRCS})
target_link_libraries(server PRIVATE fastsync_server_core)
add_executable(client ${CLIENT_SRCS} ${SHARED_SRCS} ${FILE_STORE_SRCS} ${SERVER_RECEIVER_SRCS})
target_include_directories(client PRIVATE src/shared src/server src/client)
target_link_libraries(client PRIVATE Threads::Threads ${ZSTD_LIBRARY} OpenSSL::SSL OpenSSL::Crypto xxhash)
add_executable(client ${CLIENT_MAIN_SRCS})
target_link_libraries(client PRIVATE fastsync_client_core)
# --- Production hardening ---
# Each compile flag is probed so a compiler/architecture that lacks it still
# configures cleanly. _FORTIFY_SOURCE is guarded separately because it only
# works in an optimising build. xxHash is a static archive built by
# FetchContent, so it must be position-independent for the -pie link; the same
# applies to the first-party static libraries linked into the -pie binaries.
if(HARDENING_ACTIVE)
set_target_properties(xxhash fastsync_shared fastsync_server_core fastsync_client_core
PROPERTIES POSITION_INDEPENDENT_CODE ON)
include(CheckCCompilerFlag)
foreach(flag -fstack-protector-strong -fstack-clash-protection -fPIE)
string(MAKE_C_IDENTIFIER "HARDEN_${flag}" _harden_var)
check_c_compiler_flag("${flag}" ${_harden_var})
endforeach()
check_c_compiler_flag("-D_FORTIFY_SOURCE=2" HARDEN_FORTIFY_SOURCE)
foreach(target fastsync_shared fastsync_server_core fastsync_client_core server client)
foreach(flag -fstack-protector-strong -fstack-clash-protection -fPIE)
string(MAKE_C_IDENTIFIER "HARDEN_${flag}" _harden_var)
if(${_harden_var})
target_compile_options(${target} PRIVATE ${flag})
endif()
endforeach()
if(HARDEN_FORTIFY_SOURCE)
target_compile_options(${target} PRIVATE -D_FORTIFY_SOURCE=2)
endif()
endforeach()
foreach(target server client)
target_link_options(${target} PRIVATE -pie -Wl,-z,relro -Wl,-z,now -Wl,-z,noexecstack)
endforeach()
endif()
# --- Testing ---
enable_testing()
# Common test libraries
set(TEST_LIBS Threads::Threads ${ZSTD_LIBRARY} OpenSSL::SSL OpenSSL::Crypto xxhash)
set(TEST_INCLUDES tests src/shared src/server src/client)
# --- Unit tests ---
# The monolithic test binary exercises both client and server code, so it is
# the one place that legitimately sees both include directories and links both
# core libraries. client_cli.c is compiled here directly (with the test build
# define) rather than linked from fastsync_client_core so its test-only shims
# and the absence of main() are preserved.
set(TEST_SRCS
tests/runner.c
tests/test_array_list.c
tests/test_batch.c
tests/test_change_list.c
tests/test_checksum.c
tests/test_chunk.c
tests/test_client_cli.c
tests/test_compression.c
tests/test_config.c
tests/test_credentials.c
tests/test_daemon_conf.c
tests/test_daemon_limits.c
tests/test_data.c
tests/test_delay_updates.c
tests/test_delta.c
tests/test_file.c
tests/test_file_list.c
tests/test_file_sendfile.c
tests/test_format.c
tests/test_fuzz_smoke.c
tests/test_glob.c
tests/test_hardlink.c
tests/test_iconv.c
tests/test_log.c
tests/test_metadata.c
tests/test_motd.c
tests/test_multiprocessing.c
tests/test_property.c
tests/test_protocol.c
tests/test_protocol_error.c
tests/test_queue.c
tests/test_receiver_timeout.c
tests/test_robustness.c
tests/test_scanner.c
tests/test_server.c
tests/test_server_cli.c
tests/test_shared_utils.c
tests/test_stop.c
tests/test_stress.c
tests/test_transport_ssh.c
tests/test_transport_tcp.c
tests/test_transport_tls.c
tests/test_xattr.c
)
# Monolithic test binary (backward compatible)
file(GLOB TEST_SRCS "tests/test_*.c" "tests/runner.c")
add_executable(tests ${TEST_SRCS} ${SHARED_SRCS} ${FILE_STORE_SRCS} ${SERVER_RECEIVER_SRCS} src/client/scanner.c src/client/change_list.c src/client/client_cli.c src/client/client_validation.c src/client/usage.c)
target_include_directories(tests PRIVATE ${TEST_INCLUDES})
add_executable(tests ${TEST_SRCS} src/client/client_cli.c)
target_include_directories(tests PRIVATE tests)
target_compile_definitions(tests PRIVATE FASTSYNC_TEST_BUILD)
target_link_libraries(tests PRIVATE ${TEST_LIBS})
target_link_libraries(tests PRIVATE fastsync_server_core fastsync_client_core)
add_test(NAME unit_all COMMAND tests)
# --- Fuzz targets (requires clang) ---
@@ -94,13 +264,29 @@ if(ENABLE_FUZZ)
if(NOT CMAKE_C_COMPILER_ID MATCHES "Clang")
message(FATAL_ERROR "ENABLE_FUZZ requires Clang (compiler is ${CMAKE_C_COMPILER_ID})")
endif()
file(GLOB FUZZ_SRCS "tests/fuzz/*.c")
set(FUZZ_SRCS
tests/fuzz/fuzz_chunk_deserialize.c
tests/fuzz/fuzz_compress_decompress.c
tests/fuzz/fuzz_config_receive.c
tests/fuzz/fuzz_delta_deserialize.c
tests/fuzz/fuzz_delta_signature_deserialize.c
tests/fuzz/fuzz_glob_match.c
tests/fuzz/fuzz_identity_parse.c
tests/fuzz/fuzz_manifest.c
tests/fuzz/fuzz_metadata_from_buf.c
tests/fuzz/fuzz_protocol_framing.c
tests/fuzz/fuzz_xattr_block.c
)
# Compile the sources under test directly so libFuzzer's coverage
# instrumentation sees them (static libraries would be uninstrumented).
set(FUZZ_CORE_SRCS ${SHARED_SRCS} src/server/receiver.c src/server/receiver_pipeline.c)
foreach(FUZZ_SRC ${FUZZ_SRCS})
get_filename_component(FUZZ_NAME ${FUZZ_SRC} NAME_WE)
add_executable(${FUZZ_NAME} ${FUZZ_SRC} ${SHARED_SRCS} ${FILE_STORE_SRCS} ${SERVER_RECEIVER_SRCS})
target_include_directories(${FUZZ_NAME} PRIVATE ${TEST_INCLUDES})
add_executable(${FUZZ_NAME} ${FUZZ_SRC} ${FUZZ_CORE_SRCS})
target_include_directories(${FUZZ_NAME} PRIVATE tests src/shared src/server)
target_compile_options(${FUZZ_NAME} PRIVATE -fsanitize=fuzzer,address,undefined -fno-omit-frame-pointer)
target_link_options(${FUZZ_NAME} PRIVATE -fsanitize=fuzzer,address,undefined)
target_link_libraries(${FUZZ_NAME} PRIVATE ${TEST_LIBS})
target_link_libraries(${FUZZ_NAME} PRIVATE Threads::Threads ${ZSTD_LIBRARY} ${ZLIB_LIBRARY}
${LZ4_LIBRARY} OpenSSL::SSL OpenSSL::Crypto xxhash)
endforeach()
endif()
+15 -1
View File
@@ -2,8 +2,22 @@ FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
gcc g++ make libc6-dev cmake libzstd-dev libssl-dev git ca-certificates curl cppcheck clang-format \
python3 python3-pip python3-venv openssl openssh-client \
lcov valgrind clang libclang-rt-18-dev && \
lcov valgrind clang libclang-rt-18-dev \
acl attr zlib1g-dev liblz4-dev libxxhash-dev && \
pip3 install --break-system-packages pytest pytest-xdist && \
curl -fsSL https://deb.nodesource.com/setup_20.x | bash - && \
apt-get install -y --no-install-recommends nodejs && \
rm -rf /var/lib/apt/lists/*
# rsync is used as the reference implementation for drop-in parity tests.
# Ubuntu 24.04 ships 3.2.7, so build the pinned 3.4.1 reference from source.
ARG RSYNC_VERSION=3.4.1
ARG RSYNC_SHA256=2924bcb3a1ed8b551fc101f740b9f0fe0a202b115027647cf69850d65fd88c52
RUN curl -fsSL "https://download.samba.org/pub/rsync/src/rsync-${RSYNC_VERSION}.tar.gz" -o /tmp/rsync.tar.gz && \
echo "${RSYNC_SHA256} /tmp/rsync.tar.gz" | sha256sum -c - && \
tar -xzf /tmp/rsync.tar.gz -C /tmp && \
cd "/tmp/rsync-${RSYNC_VERSION}" && \
./configure --enable-zstd --enable-xxhash --enable-lz4 && \
make -j"$(nproc)" && \
make install && \
rm -rf "/tmp/rsync-${RSYNC_VERSION}" /tmp/rsync.tar.gz
+67
View File
@@ -0,0 +1,67 @@
# FastSync — Session Handoff (2026-09-17)
## Current status
- **Release `v2.21.0`** tagged (`919a729`, "Release v2.21.0"); full CI green
(run 552: lint, build-and-test, ASan, UBSan, fuzz-build, coverage, valgrind).
`dev` has the release commit plus later doc-only merges (a README refresh and
this handoff).
- **Release PR #284 (`dev` -> `main`)** open, CI green (run 553).
`main` is protected: it needs review/approval to merge.
https://gitea.tap-tap.win/TapTap/FastSync/pulls/284
- **`PROTOCOL_VERSION` = `"2.26.0"`** (`src/shared/config.h`); CMake
`project(FastFileTransfer VERSION 2.26.0)`.
- Working tree clean; no wave worktrees remain.
## What landed this session
1. **Wave 8 (refactors):** Config X-macro wire table; single-owner `authorized_root`;
daemon per-module/per-host caps + cross-process auth lockout (`daemon_limits.[ch]`);
`Data` charge returns to its owning `ProtocolSession`.
2. **Wave 9 (protocol 2.21.0):** optional `STATUS_ERROR_DETAIL` rejection reasons;
server-contacting `--dry-run` (`STATUS_DRY_RUN_TRANSFER`, receiver mutates nothing).
3. **Security wave:** ran 5 parallel audits (wire parsing; daemon/transport/TLS/auth;
receiver confinement; client/CLI/SSH; crypto/memory/limits). Fixed all HIGH and the
confirmed MEDIUMs:
- SSH `-o ProxyCommand=…` argument injection (RCE) — reject leading `-`, insert `--`.
- Truncated zstd frame infinite CPU loop (remote DoS).
- FIFO receiver opens lacked `O_NONBLOCK` (indefinite hang).
- `--inplace` could write a FIFO/device (bypass of `--write-devices` gate).
- `--force` not gated by server `--allow-delete`.
- Privileged standalone server defaulted super activities on; added `--allow-super`
(never honored with `--stdio`).
- `--dry-run` content/hash oracle on `read only`/basis files removed.
- Empty `hosts allow`/`deny`/`auth users` now rejected.
- TLS: AEAD-only 1.2 + server preference, TOCTOU-safe key load, IP-SAN verify,
CN-truncation guard. Glob backtracking bounded; line reads bounded; ACL xattrs
gated on `--acls`; decompression/chunk memory charged; pre-auth `basis_count`
NULL-deref fixed.
4. **Tooling:** benchmark accuracy (data mix, verification, percentiles, `tc`,
`build-bench/`, `--warm` mode); `shell.nix` full toolchain and no build-on-entry;
docs state push-only / remote-source unsupported.
5. **Preserve-attribute split (protocol 2.22.0)** landed on `feat/preserve-attr-split`: per-attribute `-p/-t/-o/-g` + `--no-*` negations, `-a` = `-rlptgoD`, and the 2.21.0 → 2.22.0 wire bump.
6. **Rsync-parity wave (protocol 2.23.0)** on `feat/rsync-parity`: rsync short options/clustering/attached values (`-r`/`-b`/`-L`/`-B`, `-av`, `-aAX`, `-B1000`, `-essh`, `-MOPT`), `-c` checksum quick-check, `--checksum-choice`/`--compress-choice` validation and seed randomization, rsync timeout/max-alloc defaults, temp-dir confinement + `EXDEV` fallback, ownership/mapping parity (numeric-ids modifier, map ranges/`*`/empty-FROM, `--chown`+map conflicts, fake-super resolved-owner record), verbatim symlink storage with rsync `--safe-links`/`--munge-links`, socket recreation under `--specials`, `--chmod` 3.4.1 semantics, and delete scoping + `--max-delete` partial/exit-25. Wire: appended delete-manifest synchronized-directory section and `STATUS_DELETE_LIMIT`.
7. **Parity-completion wave (protocol 2.24.0 → 2.26.0)** on `feat/parity-completion`: per-directory delete plans (`STATUS_DELETE_PLAN`) for `--delete-during`/`--delete-delay`; receiver `STATUS_STATS` counters feeding `--stats`/`--progress` and `--out-format %b/%c/%C`, plus `-n --delete` lines; `lz4`/`zlib`/`zlibx` compression and `md4`/`sha1`/`none` checksums with `auto` negotiation (default `xxh128`/`zstd`); general `-R`/`--no-implied-dirs`/`-d`; the full filter grammar (`merge`/`dir-merge`/`hide`/`show`/`protect`/`risk`/`clear` + modifiers) and corrected `-F`/`-FF`; receiver-side `--chown`/map TO-name resolution; absolute basis dirs + `--link-dest` relink; receiver-side `--ignore-existing` short-circuit; `--preallocate` over `--sparse` via `fallocate(2)`; `--iconv=.`/`-`/`--no-iconv`; lone `-h` help; aliases `--ignore-non-existing`/`--protect-args`/`--msgs2stderr`; and the full `--info`/`--debug` vocabulary. `RSYNC_COMPAT.md` reclassifies the matrix to 106 ✅ / 27 ⚠️ / 23 ❌.
## Next steps
1. **Merge PR #284** (`dev` -> `main`) once reviewed (protected branch).
2. **Deferred security items** (documented, not implemented):
- Pre-auth config/daemon-auth handshake has no aggregate wall-clock deadline
(per-message timeout only) — slowloris holds connection slots.
- Per-source registry fails open when the shared table is full (per-module/global
caps and host ACLs still apply); consider fail-closed or larger/evicting table.
- SCRAM-like daemon auth has no TLS channel binding (and is not RFC 5802).
- `cleanup()` signal handler calls non-async-signal-safe teardown; daemon `umask(0)`.
- Wire protocol assumes homogeneous word size/endianness (lengths are native
`size_t`) — document or move to fixed-width framing.
3. **Out of scope / intentional:** pull (remote source) mode is **not** planned —
FastSync is push-only; see `RSYNC_COMPAT.md#direction`.
## Key facts / commands
- CI image: `gitea.tap-tap.win/taptap/fastsync-ci:v11` (alias `fastsync-ci:local`).
- Build/test: `cmake -B build -S . -DSTRICT_WARNINGS=ON && cmake --build build -j$(nproc) && ./build/tests`
then `python3 -m pytest tests/integration/ -n 4 --dist=load -m "not setpriv"`.
- Dev shell: `nix-shell` (provides clang-format, cppcheck, pytest-xdist, openssh,
rsync, iproute2, valgrind, lcov; does not build on entry).
- Gitea API token: supplied out-of-band via the `TOKEN` environment variable; it is
intentionally **not** recorded in this file.
- CI polling: `GET /api/v1/repos/TapTap/FastSync/actions/runs?limit=N`, match `head_sha`,
then `/actions/runs/<id>/jobs`.
+539 -186
View File
@@ -1,4 +1,4 @@
#FastSync
# FastSync
FastSync is a high-performance file synchronization tool designed to become a
drop-in replacement for common `rsync` workflows. It keeps the familiar
@@ -6,6 +6,10 @@ source/destination model and rsync-style options while adding optional
multithreading, streaming zstd compression, chunking, zero-copy TCP transfers,
and native TCP/TLS transports.
The release version is FastSync's client/server protocol version (printed by
`./build/client --version`); client and server must match. See
[CHANGELOG.md](CHANGELOG.md) for the history.
The compatibility target is straightforward:
- Existing rsync commands should keep the same meaning.
@@ -24,7 +28,7 @@ FastSync uses a producer-consumer transfer pipeline and can combine several
optimizations for large or high-latency transfers:
- Multithreaded scanning, loading, and sending.
- Streaming zstd compression with levels 1 through 22.
- Streaming compression (zstd by default, plus lz4/zlib/zlibx) with levels 1 through 22.
- Configurable file chunking and compact chunk serialization.
- `sendfile()` zero-copy transfers over TCP.
- Batched incremental checks to reduce round trips.
@@ -47,94 +51,221 @@ replacement for every rsync feature or protocol mode.
- Rsync-style source and destination arguments.
- SSH transport using `user@host:destination` paths below the remote authorized root.
- TCP client/server transfers.
- Dry runs, excludes, includes, size filters, backups, statistics, and
bandwidth limiting.
- Incremental size/mtime checks and optional xxHash64 content checks.
- Dry runs (server-contacting since protocol 2.21.0 for server-routed targets),
excludes, includes, size filters, backups, statistics, and bandwidth
limiting.
- Incremental size/mtime checks and optional content checks (`xxh128` by
default, selectable with `--checksum-choice`).
- FastSync-native delta transfer for changed files.
- Optional mode and timestamp preservation.
- Delete manifests with server-side delete authorization.
- Temporary-file writes with atomic rename by default.
- Path traversal checks and destination-root confinement.
### Not yet equivalent to rsync
### Boundaries and documented divergences
The items below summarize FastSync's rsync compatibility status — recently
closed gaps and the remaining known divergences. Each row of the detailed
matrix is classified as parity, caveat, or divergent in
[`RSYNC_COMPAT.md`](RSYNC_COMPAT.md).
- The FastSync wire protocol is not the rsync wire protocol.
- SSH mode requires `fastsync-server` on the remote host.
- Archive mode does not yet provide all of rsync's `-rlptgoD` behavior.
- Symlink transfer is incomplete; link targets are not yet recreated in all
modes.
- Owner/group, ACL, xattr, hard-link, device, and special-file handling is
incomplete or unavailable.
- Sparse-file handling does not yet preserve all holes correctly.
- `--partial`, `--partial-dir`, `-P`, `--append`, and `--append-verify` are not
yet full rsync-style resumable transfers. Interrupted files are not retained
for resumption.
- `--dirs` is not implemented. Its compatibility aliases `--old-dirs` and
`--old-d` are recognized but rejected explicitly rather than silently using
FastSync's recursive directory behavior.
- Several rsync short options currently have FastSync-specific meanings. Do
not assume every short option is interchangeable yet.
- Archive mode covers rsync's `-rlptgoD` behavior — links, permissions, times,
owner, group, devices, and special files — and does not imply compression or
multithreading (see [Client](#client)). Ownership application is still
privilege-gated: a receiver that cannot `chown` logs a warning and skips it.
Under `-p` the source mode is copied exactly, including setuid/setgid/sticky
and group/other-write bits (strict rsync parity; see
[`RSYNC_COMPAT.md`](RSYNC_COMPAT.md)).
- Symlink transfer stores targets **verbatim** (`-l`/`--links`), including
absolute and `..`-bearing targets, matching rsync. The receiver does not
enforce a containment predicate by default; `--safe-links` drops unsafe
targets on the sender, and `--munge-links` rewrites them with rsync's
`/rsyncd-munged/` marker. `--trust-sender` does not affect symlink targets.
A destination later consumed by a link-following tool can therefore follow a
link outside the receive root — use `--safe-links` for untrusted sources.
- Hard links (`-H`/`--hard-links`), extended attributes (`-X`/`--xattrs`), and
POSIX ACLs (`-A`/`--acls`) are preserved; owner/group is applied through
`-o`/`-g` (or an `-a`/`--archive` transfer), through the opt-in identity flags
(`--chown`/`--usermap`/`--groupmap`/`--numeric-ids`/`--copy-as`), and only when
the receiver has permission. See
[`RSYNC_COMPAT.md`](RSYNC_COMPAT.md) for the exact semantics and documented
divergences.
- Device and special-file preservation is implemented with documented
divergences: recreated device nodes require `CAP_MKNOD` on the receiver (a
non-root receiver skips the entry), while FIFOs **and unix sockets** are
recreated (`--specials`).
- Sparse-file hole preservation (`-S`, `--sparse`) is implemented receiver-side:
long all-zero runs are written as holes (no wire change; the full file image
is already in memory).
- `--partial`, `--partial-dir`, `-P`, `--append`, and `--append-verify` keep
the write atomic (temp + rename). With `--partial`, a failed/interrupted write
now retains the already-written temp at the destination path (best-effort) so
a later `--append`/`--append-verify` run can resume it.
- `-d`/`--dirs` and its aliases `--old-dirs`/`--old-d` transfer the named
directory entries without recursing into their contents.
- Short-option names are now rsync-parity (Phase 7 Wave A): FastSync's former
collisions were renamed (`-j`/`--threads`, `--preserve`, `--sendfile`,
`--chunk-serialization`, `--timeout`, `--ssh-port`), so `-m`, `-M`, `-f`,
`-s`, `-T`, `-p`, `-c`, `-a`, and `-z` follow rsync.
- Short-option clustering (`-av`, `-aAX`, `-rlpt`) and attached values
(`-B1000`, `-essh`, `-MOPT`, `--opt=value`) are accepted, matching rsync.
- `-r`, `-b`, `-L`, and `-B` are parsed with the rsync short names.
- `--stats` prints the counters FastSync can observe plus the receiver-only
counters (`Matched data`, deleted files) reported over the wire; rsync's
per-type `Number of files` breakdown is not reproduced. `--progress` prints
rsync-style per-file blocks (without rsync's leading `./` line).
- Codecs match rsync 3.4.1: `zstd`/`lz4`/`zlib`/`zlibx` compression and
`xxh128`/`xxh3`/`xxh64`/`md5`/`md4`/`sha1`/`none` checksums, negotiated with
`auto`; `zlibx` behaves as `zlib`, and the transfer checksum is not separately
selectable.
The detailed flag matrix is maintained in
[`RSYNC_COMPAT.md`](RSYNC_COMPAT.md). It distinguishes implemented,
partial, alternate, and planned behavior.
[`RSYNC_COMPAT.md`](RSYNC_COMPAT.md). It reports each row as **parity**,
**caveat** (works with a documented divergence), or **divergent** (not
supported), rather than treating "parsed" as parity.
## Quick Start
### Build
```bash
cmake -B build -S .
cmake --build build -j$(nproc)
```
This produces `./build/client` and `./build/server`. `compile_commands.json` is a symlink to `build/compile_commands.json` and is used by clangd/editor tooling; its target is generated by the build, so it dangles until the first build.
### Client
| Argument | Description |
|----------|-------------|
| Positional | `<source> <dest>` — automatic SSH detection if dest contains `:` |
| `-c [level]` | Compression with optional level (1–22, default 5) |
| `-z [level]` | Alias for `-c` |
| `-a, --archive` | Archive mode: enables `-c -m -M` (no `-s`) |
| `-m` | Multithreading mode |
| `-s` | Chunk serialization (batch all files per chunk) |
| `--secluded-args` | Accepted as an rsync compatibility option with no effect; `-s` remains chunk serialization. |
| `-f, --sendfile` | Sendfile zero-copy. Incompatible with `-c` / `-s`. TCP only. |
| `-M, --preserve` | Preserve supported file metadata (mode and mtime; ownership and atime are unsupported) |
| `-n, --dry-run` | Scan and print what would be transferred |
| `-p <port>` | SSH port (default: 22) |
| `-v, --verbose` | Enable debug logging |
| `-q, --quiet` | Suppress non-error output |
| `--progress` | Show real-time transfer speed |
| `-P` | Enables partial-transfer mode and progress output (partial retention is incomplete) |
| `--delete` | Delete files on receiver not present in source (default timing: delete-after, i.e. only after the whole transfer succeeded) |
| `-c, --checksum` | Verify content by checksum instead of size+mtime (implies the incremental checksum quick-check) |
| `--checksum-choice <alg>` | Whole-file checksum algorithm: `xxh128` (default), `xxh3`, `xxh64`/`xxhash`, `md5`, `md4`, `sha1`, `none`, or `auto` (plus rsync's two-name `transfer,pre-transfer` form) |
| `-z, --compress [level]` | Enable streaming compression (default `zstd`; level 1–22, default 5) |
| `--compress-choice <alg>` | Compression algorithm: `zstd` (default), `lz4`, `zlib`, `zlibx`, `none`, or `auto` |
| `--skip-compress <list>` | Skip compression for suffixes (`/`- or `,`-separated); defaults to rsync 3.4.1's built-in suffix list |
| `-a, --archive` | rsync archive mode (`-rlptgoD`): links, perms, times, owner, group, devices and specials; ownership application stays privilege-gated (not compression/multithreading) |
| `-j, --threads[=N]` | Multithreading mode; `N` (1–256) sets the parallel scanner worker count, bare `-j`/`--threads` uses the default |
| `-m` | rsync `--prune-empty-dirs` (short form now rsync-parity) |
| `-r, --recursive` | Recurse into directories (FastSync is always recursive; accepted for rsync compatibility) |
| `-d, --dirs` | Transfer the named directory entries without recursing into their contents; aliases `--old-dirs`/`--old-d` |
| `-R, --relative` | Use rsync's relative path semantics (including the `/./` cut); with `--files-from`, preserve each listed entry's relative path below the destination root |
| `--chunk-serialization` | Chunk serialization (batch all files per chunk; long form only) |
| `-s` | rsync `--secluded-args` compatibility no-op (remote SSH argv is already injection-safe) |
| `--sendfile` | Sendfile zero-copy. Incompatible with compression / chunk serialization. TCP only. Long form only. |
| `--preallocate` | Allocate destination file space up front (fail-fast on a full disk) |
| `--append` | Resume a shorter destination by appending only its tail (prefix not verified; requires `--incremental`) |
| `--append-verify` | Like `--append`, but verifies the retained prefix checksum first (falls back to a full transfer on mismatch) |
| `-W, --whole-file` | Transfer changed files without delta processing; `--no-whole-file` clears it |
| `-B <n>, --block-size <n>` | Delta block size in bytes (alias `--delta-block`) |
| `--checksum-seed <n>` | Seed for the whole-file xxHash digest; an unset/`0` seed is randomized per transfer, matching rsync |
| `-I, --ignore-times` | Transfer files even when size and mtime match |
| `--size-only` | Skip incremental files matching in size, ignoring mtime |
| `--preserve` | Preserve mode and mtime (`-p` + `-t`; add `-o`/`-g` for owner/group or `-U`/`--atimes` for atime; `-N`/`--crtimes` captures birth time but cannot apply it) |
| `-U, --atimes` | Preserve access times. Captured with the metadata payload; does not enable ownership. |
| `-N, --crtimes` | Capture birth time; cannot be applied (documented divergence) |
| `-p, --perms` | Preserve permission bits. Strict rsync parity: the source mode is copied exactly, including setuid/setgid/sticky and group/other-write bits |
| `-t, --times` | Preserve modification times |
| `-o, --owner` | Preserve the source owner (privilege-gated; mapped by name on the receiver with a numeric fallback) |
| `-g, --group` | Preserve the source group (privilege-gated; mapped by name on the receiver with a numeric fallback) |
| `--no-perms`, `--no-times`, `--no-owner`, `--no-group`, `--no-preserve` | Negate the per-attribute flags (short `--no-p`/`--no-t`/`--no-o`/`--no-g`; `--no-preserve` clears all four) |
| `-E, --executability` | Preserve executable permission bits |
| `-X, --xattrs` | Preserve user `user.*` extended attributes |
| `-A, --acls` | Preserve POSIX ACLs |
| `--chmod <changes>` | Modify transferred permissions (rsync syntax) |
| `--chown=USER:GROUP` | Override the ownership of transferred files |
| `--usermap=MAP` | Map usernames when applying ownership |
| `--groupmap=MAP` | Map group names when applying ownership |
| `--numeric-ids` | Apply source numeric uid/gid directly instead of mapping by name |
| `--copy-as=USER[:GROUP]` | Force every written entry to USER[:GROUP] (requires a privileged receiver) |
| `--fake-super` | Record the resolved owner plus mode/time in a reserved `user.fastsync.stat` xattr and replay mode/time; never performs a real chown |
| `--super` | Permit the receiver to attempt confined super-user activities (device nodes) |
| `-D` | Preserve device and special files (implies `--devices --specials`) |
| `--devices` | Recreate device nodes on the destination (privileged; skipped without `CAP_MKNOD`) |
| `--specials` | Recreate special files: FIFOs and unix sockets |
| `--remove-source-files` | Remove regular source files after a successful transfer |
| `--exclude <pattern>` | Exclude files matching glob pattern (repeatable) |
| `--exclude-from <file>` | Read exclude patterns from a file (one per line) |
| `--include <pattern>` | Only transfer files matching glob pattern (repeatable, whitelist) |
| `--include-from <file>` | Read include patterns from a file |
| `--files-from <file>` | Read the source file list from FILE (paths relative to the source root) |
| `--max-size <n>` | Skip files larger than n bytes |
| `--min-size <n>` | Skip files smaller than n bytes |
| `-x, --one-file-system` | Do not cross filesystem boundaries; the mount-point directory entry is emitted (empty at the destination) without descending |
| `--max-alloc <SIZE>` | Maximum single allocation (binary units: B, K, M, G, T, P, E; default 1G; `0` = no local limit, matching rsync) |
| `-u, --update` | Skip files newer than the source on the receiver |
| `--incremental` | Skip files unchanged since last transfer (size + mtime). Auto-enables `--preserve`. Incompatible with `--chunk-serialization`. |
| `--existing` | Skip files not already present at the destination; update existing files normally. |
| `--compare-dest <dir>` | Extra comparison basis: unchanged files are not transferred (requires/implies `--incremental`) |
| `--copy-dest <dir>` | Like `--compare-dest`, but copies the unchanged file from DIR into the destination |
| `--link-dest <dir>` | Like `--copy-dest`, but hard-links the unchanged file from DIR (repeatable; earlier DIRs win) |
| `--delete` | Delete files on receiver not present in source (default timing: delete-after, i.e. only after the whole transfer succeeded). Scoped to the synchronized directories, so `--files-from` subsets are safe |
| `--delete-before` | Delete extras before the transfer starts (implies `--delete`) |
| `--delete-during`, `--del` | Delete extras once the keep-set is known, before data is applied (implies `--delete`) |
| `--delete-delay` | Delete extras only after a successful transfer (implies `--delete`) |
| `--delete-after` | Explicit delete-after timing (implies `--delete`) |
| `--exclude <pattern>` | Exclude files matching glob pattern (repeatable) |
| `--exclude-from <file>` | Read exclude patterns from a file (one per line) |
| `--include <pattern>` | Only transfer files matching glob pattern (repeatable, whitelist) |
| `--max-size <n>` | Skip files larger than n bytes |
| `--min-size <n>` | Skip files smaller than n bytes |
| `--max-alloc <SIZE>` | Maximum single allocation (binary units: B, K, M, G, T, P, E; default 1G) |
| `--incremental` | Skip files unchanged since last transfer (size + mtime). Auto-enables `--preserve`. Incompatible with `-s`. |
| `--existing` | Skip files not already present at the destination; update existing files normally. |
| `--bwlimit <KB/s>` | Bandwidth limit in kilobytes per second |
| `--chunk-size <n>` | Chunk size in bytes (default: 10485760) |
| `--timeout <sec>` | I/O timeout in seconds (default: 30) |
| `--contimeout <sec>` | Connection timeout in seconds (default: 10) |
| `--backup` | Backup existing destination files before overwriting |
| `--backup-dir <dir>` | Target directory for backups (requires `--backup`) |
| `--stats` | Print transfer statistics at end (bytes, files, timing) |
| `-h, --human-readable` | Format transfer byte sizes with binary units |
| `--delete-excluded` | Also delete filter-excluded destination mirrors (size-pruned mirrors stay protected) |
| `--max-delete <n>` | Delete at most n destination entries; the rest are skipped and the run exits 25 (partial), matching rsync |
| `--delay-updates` | Put updated files into place only at the end of the transfer (`--force` is honored at publication) |
| `-T, --temp-dir <dir>` | Scratch directory for temp files before the atomic install; confined to the receive root (relative only), with an `EXDEV` non-atomic copy fallback |
| `-n, --dry-run` | Report what would be transferred without mutating the destination. Since protocol 2.21.0 a server-routed target contacts the receiver and reports would-transfer based on receiver state; a plain local destination keeps the client-side scan. Never mutates or deletes. |
| `-v, --verbose` | Enable debug logging |
| `-q, --quiet` | Suppress non-error output |
| `--progress` | Show rsync-style per-file progress blocks from the receiver's wire counters (FastSync does not print rsync's leading `./` line) |
| `-P` | Enables partial-transfer mode + progress output; interrupted writes retain the already-written temp for resumption |
| `--stats` | Print transfer statistics at end (bytes, files, timing), including the receiver-only counters reported over the wire; rsync's per-type `Number of files` breakdown is not reproduced |
| `-i, --itemize-changes` | Print an rsync-style per-file change line |
| `--out-format=FORMAT` | Output format for changed files (`%f %n %l %b %M %%`) |
| `--list-only` | List source files instead of transferring |
| `--fsync` | Fsync every written file before publication |
| `-h, --human-readable` | Format transfer byte/rate counts with rsync's decimal (base-1000) units |
| `--max-depth <n>` | Maximum directory depth to recurse (0 = unlimited, default: 0) |
| `--log-file <path>` | Write log messages to file instead of stderr |
| `--write-batch=FILE` | Run the normal live transfer and also emit a self-contained batch file of the source tree |
| `--only-write-batch=FILE` | Emit the batch file only (no destination, no server) |
| `--read-batch=FILE` | Apply a batch file to the destination (no source, no server) |
| `--source-dir <path>` | Source directory (overrides `FASTSYNC_SOURCE_DIR`) |
| `--dest-dir <path>` | Server destination directory (overrides `FASTSYNC_DEST_DIR`) |
| `--save-to-disk` | Write received files to disk |
| `--server-host <ip>` | Server IP address (default: `127.0.0.1`) |
| `--server-port <n>` | Server port (default: `8080`) |
| `--ssh-port <port>` | SSH port (default: 22) |
| `-e, --rsh <command>` | Remote shell to launch for the SSH transport (default: `ssh`; may include arguments, e.g. `-e "ssh -p 2222"`) |
| `-M, --remote-option=OPT` | Append OPT to the remote server invocation over SSH (repeatable) |
| `--address <ip>` | Bind the outgoing client socket to this source address |
| `-4, --ipv4` | Force IPv4 for destination resolution |
| `-6, --ipv6` | Force IPv6 for destination resolution |
| `--sockopts=OPTS` | Comma-separated OPT=VAL socket options applied before connect (`TCP_NODELAY`, `SO_KEEPALIVE`, `SO_RCVBUF`, `SO_SNDBUF`, `SO_REUSEADDR`) |
| `--bwlimit <KB/s>` | Bandwidth limit in kilobytes per second |
| `--chunk-size <n>` | Chunk size in bytes (default: 10485760) |
| `--timeout <sec>` | I/O timeout in seconds, applied to both the socket (`SO_RCVTIMEO`/`SO_SNDTIMEO`) and the per-message protocol poll deadline. Default `0` = disabled (matching rsync); `0` disables it. `--no-timeout` is the negation. The value is not sent on the wire; the server side keeps its own safe floor. |
| `--contimeout <sec>` | Connection timeout in seconds (default: 60, matching rsync); `0` disables it (`--no-contimeout` is the negation) |
| `--stop-after=MINS` | Stop the transfer after MINS minutes (a positive integer); whatever was already transferred is kept |
| `--stop-at=TIME` | Stop at an absolute time (`HH:MM`, `HH:MM:SS`, or `now+N[smhd]`); an early stop skips the late `--delete` keep-set |
| `-b, --backup` | Backup existing destination files before overwriting |
| `--backup-dir <dir>` | Target directory for backups (requires `--backup`) |
| `--tls` | Enable TLS encryption |
| `--cert <path>` | TLS certificate file (PEM) |
| `--key <path>` | TLS private key file (PEM) |
| `--ca <path>` | TLS CA certificate file for verification (PEM) |
| `--client-cn <name>` | Required TLS client certificate common name |
The exhaustive rsync flag matrix is in [`RSYNC_COMPAT.md`](RSYNC_COMPAT.md).
**Per-message vs. connection timeouts.** `--timeout` bounds each individual protocol
send/receive (the `poll()` deadline), so a peer that stops mid-frame is dropped. It
does not, by itself, stop a peer that keeps sending well-formed frames forever. The
receiver therefore also enforces two wall-clock (`CLOCK_MONOTONIC`) bounds on a
connection: a **1 hour** idle limit and a **24 hour** overall session cap. Only
frames that move real work (not `STATUS_KEEPALIVE`/`STATUS_ABORT` and not an
empty `STATUS_CHECK_BATCH`/`STATUS_DIR_TIMES`) refresh the idle timestamp, so a
peer cannot hold a connection slot by emitting cheap empty frames; a peer that
fabricates minimal non-empty frames can still occupy a slot until the 24 hour
cap, since no bound can require actual payload without risking a legitimate
long operation. Both are deliberately generous so a legitimate long-running
transfer is never aborted.
### Server
@@ -148,7 +279,8 @@ partial, alternate, and planned behavior.
| `--ca <path>` | TLS CA certificate file for verification (PEM) |
| `--destination-root <path>` | Authorized destination root (default: `.`) |
| `--allow-delete` | Permit manifest deletion |
| `--allow-unauthenticated` | Permit plaintext TCP clients |
| `--allow-super` | Standalone TCP listener only: keep super-user activities enabled for a **root** receiver. Without it a root standalone server forces `SUPER_MODE_OFF`, so client `--devices`/`--write-devices`/`--super` and client-chosen ownership requests are skipped/refused. **Rejected with `--stdio`** (the SSH remote argv is client-composed, so a client could otherwise pass it and defeat the secure default; operators exposing `fastsync-server --stdio` over SSH must use a forced command if the default must hold). No effect when not root. |
| `--allow-unauthenticated` | Permit plaintext TCP clients. For an `auth users` module this opts in **loopback plaintext only**; remote auth still requires verified TLS, so the flag never permits remote plaintext auth. |
| `-v, --verbose` | Enable debug logging |
| `--help` | Show help |
@@ -159,60 +291,64 @@ partial, alternate, and planned behavior.
| `FASTSYNC_SOURCE_DIR` | — | Source directory fallback |
| `FASTSYNC_DEST_DIR` | — | Destination directory fallback |
| `FASTSYNC_SAVE_TO_DISK` | `false` | Disk persistence fallback |
| `FASTSYNC_SSH_PORT` | `22` | Default SSH port |
| `FASTSYNC_SERVER_HOST` | `127.0.0.1` | Default server host |
| `FASTSYNC_SERVER_PORT` | `8080` | Default server port |
| `FASTSYNC_TLS_CERT` | — | Default TLS certificate path |
| `FASTSYNC_TLS_KEY` | — | Default TLS private key path |
| `FASTSYNC_TLS_CA` | — | Default TLS CA certificate path |
## Implementation Details
### Data Structures
1. **Chunk** — collection of files (~10 MB total by default)
2. **File** — path, content (`Data`), optional `FileMetadata` pointer
3. **FileMetadata** — `mode`, `uid`, `gid`, `mtime_sec`, `mtime_nsec`;
uid / gid are advisory wire fields and are never applied by the receiver;
atime is unsupported
4. **Config** — runtime parameters (transported over wire, TLS settings excluded). Includes `timeout`, `contimeout`, `quiet`, `backup`, `backup_dir`, `stats`, `max_depth`, `log_file`, `queue_size`.
5. **Queue** — thread-safe bounded queue with condition variables
6. **DirectoryScanner** — recursive BFS traversal with exclude and include pattern support, max-depth enforcement
1. **Chunk** — collection of files (~10 MB total by default).
2. **File** — path, content (`Data`), optional `FileMetadata` pointer.
3. **FileMetadata** — `mode`, `uid`, `gid`, `mtime_sec`, `mtime_nsec` (plus
atime/crtime fields). `uid`/`gid` are applied only through the opt-in
identity path; atime is preserved with `-U`/`--atimes`; crtime is captured
but cannot be set on the destination.
4. **Config** — runtime parameters. Most cross the wire (TLS settings
excluded); `backup` and `backup_dir` are in the serialized wire table, while
`timeout`, `contimeout`, `quiet`, `stats`, `max_depth`, and `log_file` are
client-only.
5. **Queue** — thread-safe bounded queue with condition variables.
6. **DirectoryScanner** — recursive BFS traversal with exclude and include
pattern support, max-depth enforcement.
### Key Algorithms
1. **File scanning** — BFS directory traversal;
entries matched against exclude and include patterns,
max - depth enforced 2. * *Chunking ** — files accumulated until `chunk_size` threshold,
then flushed 3. *
*Compression ** — streaming zstd
via `ZSTD_compressStream2` / `ZSTD_decompressStream` 4. *
*Network protocol ** — status -
code - driven exchange with metadata packing,
keep - alive,
and abort support 5. * *Incremental check ** — client sends `STATUS_CHECK` + path + size +
mtime and,
with `--checksum`, XXH64 content checksum; server compares against destination. Can be batched via `STATUS_CHECK_BATCH` for reduced round-trips.
6. **Bandwidth limiting** — token-bucket algorithm with `nanosleep` throttling on 64 KB write chunks
7. **Metadata restoration** — `chmod()`, `chown()`, `utimensat()` on the receiving side
8. **`--delete`** — sender tracks all sent paths;
receiver walks destination tree and removes unlisted files / directories 9. *
*SSH transport *
* — `socketpair()` + `fork()` + `execvp("ssh",
...)` with `ControlMaster` and port support
10. *
*TLS transport ** — OpenSSL `SSL_CTX` with TLS
1.2 minimum,
mutual CA verification,
transparent `SSL_read`/`SSL_write` via `io_set_ssl()` 11. *
*Path traversal protection ** — `has_path_traversal()` rejects any file path
containing `..` components,
preventing directory escape attacks 12. *
*Connection limiting ** — server tracks active connections and rejects
new ones beyond `max_connections` (default 100)13. *
*Keep
- alive ** — idle connections receive periodic `STATUS_KEEPALIVE` to detect half
- open TCP connections 14. * *Abort handling ** — `SIGINT` sets an abort flag; the next protocol operation sends `STATUS_ABORT` for clean server cleanup
15. **Atomic writes** — files are written to a `.tmp` suffix then atomically renamed via `rename()`, preventing partial files
16. **Backup** — before overwriting, existing files are moved to `--backup-dir` (or same directory with `~` suffix) preserving the original
1. **File scanning** — BFS directory traversal; entries matched against exclude
and include patterns, with max-depth enforced.
2. **Chunking** — files accumulated until the `chunk_size` threshold (default
10 MiB) is reached, then flushed.
3. **Compression** — streaming zstd via `ZSTD_compressStream2()` /
`ZSTD_decompressStream()`.
4. **Network protocol** — status-code-driven exchange with metadata packing,
keep-alive, and abort support.
5. **Incremental check** — the client sends `STATUS_CHECK` + path + size +
mtime and, with `--checksum`, a whole-file content checksum (`xxh128` by
default; selectable via `--checksum-choice`/`--cc`, seeded by
`--checksum-seed`); the server compares against the destination. Can be
batched via `STATUS_CHECK_BATCH` for reduced round-trips.
6. **Bandwidth limiting** — token-bucket algorithm with sleep throttling on
64 KiB write chunks.
7. **Metadata restoration** — mode via `chmod()`/`fchmod()`, times via
`utimensat()`/`futimens()`, and ownership only with an identity flag via
fd-relative `fchown()`/`fchownat()`.
8. **`--delete`** — the sender tracks all sent paths; the receiver walks the
destination tree and removes unlisted files and directories.
9. **SSH transport** — `socketpair()` + `fork()` + `execvp("ssh", ...)` with
`ControlMaster` and port support.
10. **TLS transport** — OpenSSL `SSL_CTX` with TLS 1.2 minimum, mutual CA
verification, and transparent `SSL_read()`/`SSL_write()` via
`io_set_ssl()`.
11. **Path traversal protection** — `has_path_traversal()` rejects any file
path containing `..` components, preventing directory escape attacks.
12. **Connection limiting** — the server tracks active connections and rejects
new ones beyond `max_connections` (default 100).
13. **Keep-alive** — idle connections receive periodic `STATUS_KEEPALIVE` to
detect half-open TCP connections.
14. **Abort handling** — `SIGINT` sets an abort flag; the next protocol
operation sends `STATUS_ABORT` for clean server cleanup.
15. **Atomic writes** — files are written to a `.tmp` suffix then atomically
renamed via `rename()`, preventing partial files.
16. **Backup** — before overwriting, existing files are moved to `--backup-dir`
(or the same directory with a `~` suffix), preserving the original.
## Security Features
@@ -269,22 +405,34 @@ cmake --build build -j$(nproc)
### SSH transfer
The remote host must have `fastsync-server` available in `PATH`, or use
The remote host must have `fastsync-server` available in `PATH` (install or
copy the built `./build/server` there as `fastsync-server`), or use
`--fastsync-server-path`. SSH starts `fastsync-server --stdio` in its remote
working directory, so use a destination below that directory unless the
remote server is otherwise configured with a matching authorized root.
The remote `--stdio` server argv is composed by the client, so it must never
be trusted to opt a root receiver into super-user activities: `--allow-super`
is rejected with `--stdio` and super stays off on that path. Operators
exposing `fastsync-server --stdio` over SSH must use a forced command (e.g. an
`authorized_keys` `command=` entry) if the default must hold.
```bash
ssh user@host 'mkdir -p destination'
./build/client /path/to/source user@host:destination
```
FastSync is **push-only**: the source (first argument) is always a local
directory and only the destination may be remote. A remote source such as
`client user@host:src ./local` (a "pull") is intentionally not supported; see
[RSYNC_COMPAT.md](RSYNC_COMPAT.md#direction).
### TCP transfer
Start the FastSync server:
```bash
./build/server --destination-root /path/to -p 8080
./build/server --destination-root /path/to -p 8080 --allow-unauthenticated
```
Then run the client:
@@ -299,8 +447,13 @@ Plain TCP requires the explicit `--allow-unauthenticated` server option. Use TLS
authenticated network connections.
### TLS transfer
Server TLS requires `--cert`, `--key`, `--ca`, and `--client-cn`; the client
requires `--cert`, `--key`, and `--ca`.
```bash
./build/server --destination-root /path/to --tls --cert server.pem --key server-key.pem -p 8443
./build/server --destination-root /path/to --tls --cert server.pem --key server-key.pem \
--ca ca.pem --client-cn client -p 8443
./build/client --tls --cert client.pem --key client-key.pem --ca ca.pem \
--server-host example.com --server-port 8443 \
--source-dir /path/to/source --dest-dir /path/to/destination \
@@ -313,32 +466,32 @@ These examples show the intended rsync-style workflow. Options marked as
FastSync-native are optional performance or transport extensions.
```bash
#Basic synchronization
# Basic synchronization
./build/client /source/ /destination/
#Archive - style synchronization(current FastSync archive behavior)
# Archive-style synchronization (current FastSync archive behavior)
./build/client -a /source/ user@host:destination/
#Preview a transfer without changing the destination
# Preview a transfer without changing the destination
./build/client -n /source/ /destination/
#Exclude temporary and object files
# Exclude temporary and object files
./build/client --exclude '*.tmp' --exclude '*.o' \
/source/ user@host:destination/
#Remove destination entries not present in the source
# Remove destination entries not present in the source
./build/client --delete /source/ user@host:destination/
#Skip unchanged files using size and modification time
# Skip unchanged files using size and modification time
./build/client --incremental /source/ user@host:destination/
#Verify content when size and time are not sufficient
# Verify content when size and time are not sufficient
./build/client --incremental --checksum /source/ user@host:destination/
#Preserve supported mode and timestamp metadata
./build/client -M /source/ user@host:destination/
# Preserve supported mode and timestamp metadata
./build/client --preserve /source/ user@host:destination/
#Keep backups of overwritten destination files
# Keep backups of overwritten destination files
./build/client --backup --backup-dir backups \
/source/ user@host:destination/
```
@@ -350,37 +503,41 @@ features without changing the meaning of ordinary compatibility options.
| Option | Purpose |
|---|---|
| `-m` | Enable the multithreaded scanner/loader/sender pipeline. |
| `-c [level]`, `-z [level]` | Enable streaming zstd compression, levels 1-22. |
| `--compress-level <n>` | Set the zstd compression level. |
| `--zc <alg>` | Alias for `--compress-choice`. FastSync supports `zstd` and `none`. |
| `-j`, `--threads[=N]` | Enable the multithreaded scanner/loader/sender pipeline. `N` (1–256) sets the parallel scanner worker count; bare `-j`/`--threads` uses the default. |
| `-z [level]`, `--compress [level]` | Enable streaming compression (default `zstd`), levels 1-22. |
| `--compress-level <n>` | Set the compression level. |
| `--zc <alg>` | Alias for `--compress-choice`. FastSync supports `zstd` (default), `lz4`, `zlib`, `zlibx`, `none`, and `auto`; `zlibx` behaves as `zlib`. |
| `--zl <n>` | Alias for `--compress-level`. |
| `--skip-compress <list>` | Skip compression for comma-separated suffixes; incompatible with `-s`. |
| `--skip-compress <list>` | Skip compression for `/`- or `,`-separated suffixes; defaults to rsync 3.4.1's built-in list. Incompatible with `--chunk-serialization`. |
| `--compress-threads <n>` | Use `n` zstd compression workers. Requires compression and a zstd build with threaded support; the setting affects sender CPU work only. |
| `--chunk-size <bytes>` | Set the transfer chunk size. |
| `-s` | Enable FastSync chunk serialization. |
| `-f`, `--sendfile` | Use TCP `sendfile()` zero-copy transfer. Incompatible with compression and chunk serialization. |
| `--chunk-serialization` | Enable FastSync chunk serialization (long form only; `-s` is rsync's `--secluded-args`). |
| `--sendfile` | Use TCP `sendfile()` zero-copy transfer. Incompatible with compression and chunk serialization. Long form only. |
| `--delta` | Use FastSync-native block delta transfer. Requires `--incremental`. |
| `--delta-block <bytes>` | Set the FastSync delta block size. |
| `--delta-block <bytes>` | Set the FastSync delta block size (`--block-size` is an alias). |
| `--delta-max <bytes>` | Limit files eligible for FastSync delta transfer. |
| `--server-host <host>` | Select the TCP server host. |
| `--server-port <port>` | Select the TCP server port. |
| `--server-port <port>` | Select the TCP server port (`--port <port>` and `--port=<port>` are rsync-friendly aliases). |
| `--tls` | Enable TLS for TCP transport. |
| `--bwlimit <KB/s>` | Apply token-bucket bandwidth limiting. |
| `--progress` | Show transfer progress and throughput. |
| `--stats` | Print transfer statistics. |
| `--timeout <seconds>` | Set I/O timeout. |
| `--contimeout <seconds>` | Set connection timeout. |
| `--progress` | Show rsync-style per-file progress blocks from the receiver's wire counters (FastSync omits rsync's leading `./` line). |
| `--stats` | Print transfer statistics, including the receiver-only counters reported over the wire; rsync's per-type `Number of files` breakdown is not reproduced. |
| `--timeout <seconds>` | Set the socket **and** per-message protocol I/O timeout. Default `0` = disabled (matching rsync); `0` disables it. |
| `--contimeout <seconds>` | Connection timeout (default 60, matching rsync); `0` disables it. |
Current short-option conflicts are tracked as compatibility work. In
particular, FastSync currently uses `-p` for SSH port, `-s` for chunk
serialization, and `-S` for sparse handling. These meanings must be reconciled
before FastSync can claim full rsync CLI compatibility.
Short-option conflicts with rsync have been resolved for the CLI namespace
(Phase 7): `-c` is now rsync's `--checksum`, `-m` is `--prune-empty-dirs`, `-M`
is `--remote-option`, `-f` is `--filter`, `-s` is `--secluded-args`, `-p` is
`--perms`, and `-T` is `--temp-dir`. FastSync's own flags were renamed to
long-form-only or new shorts: multithreading is `-j`/`--threads`, metadata
is `--preserve`, sendfile is `--sendfile`, chunk serialization is
`--chunk-serialization`, timeout is `--timeout`, and SSH port is `--ssh-port`.
`-a`/`--archive` is now rsync archive `-rlptgoD` (owner/group implied, but the
receiver still needs privilege to apply them).
`--secluded-args` is accepted as a long-form compatibility no-op. It does not
change FastSync's transport or protocol behavior. The rsync short form `-s` is
intentionally not aliased because it remains FastSync's chunk-serialization
option.
`--secluded-args` (and its short form `-s`) is accepted as a compatibility
no-op. It does not change FastSync's transport or protocol behavior, because
remote SSH argv is already built injection-safe.
## Client Options
@@ -388,50 +545,107 @@ option.
| Option | Description |
|---|---|
| `-a`, `--archive` | Enable current archive preset. Full rsync archive semantics are planned. |
| `-n`, `--dry-run` | Scan and report without writing files. |
| `--delete` | Request removal of destination entries absent from the source. The server must allow deletion. Default timing is delete-after: extras are removed only after the whole transfer succeeded. |
| `-a`, `--archive` | rsync archive mode (`-rlptgoD`): links, perms, times, owner, group, devices and specials; ownership application stays privilege-gated. |
| `-n`, `--dry-run` | Report what would be transferred without mutating the destination. Since protocol 2.21.0 a server-routed target contacts the receiver and reports would-transfer based on receiver state; a plain local destination keeps the client-side scan. Never mutates or deletes. |
| `--remove-source-files` | Remove regular source files after a successful transfer. |
| `--incremental` | Skip files matching destination size and mtime. Auto-enables `--preserve`. Incompatible with `--chunk-serialization`. |
| `-c, --checksum` | Verify content by checksum (implies the incremental quick-check). Algorithm selectable with `--checksum-choice`. |
| `--checksum-choice <alg>` | Whole-file checksum algorithm: `xxh64`/`xxhash` (default), `xxh3`, `xxh128`, `md5`, or `auto`. |
| `--checksum-seed <n>` | Seed for the whole-file xxHash digest; an unset/`0` seed is randomized per transfer, matching rsync. |
| `--size-only` | Skip incremental files matching in size, ignoring mtime. |
| `-I, --ignore-times` | Transfer files even when size and mtime match. |
| `-u, --update` | Skip files newer than the source on the receiver. |
| `-W, --whole-file` | Transfer changed files without delta processing (`--no-whole-file` clears it). |
| `-B <n>, --block-size <n>` | Delta block size in bytes (alias `--delta-block`). |
| `-d, --dirs` | Transfer the named directory entries without recursing into their contents (aliases `--old-dirs`/`--old-d`). |
| `-R, --relative` | Use rsync's relative path semantics (including the `/./` cut); with `--files-from`, preserve each listed entry's relative path below the destination root. |
| `--files-from <file>` | Read the source file list from FILE (paths relative to the source root). |
| `--delay-updates` | Put updated files into place only at the end of the transfer. |
| `--compare-dest <dir>` | Extra comparison basis: unchanged files are not transferred (requires/implies `--incremental`). |
| `--copy-dest <dir>` | Like `--compare-dest`, but copies the unchanged file from DIR into the destination. |
| `--link-dest <dir>` | Like `--copy-dest`, but hard-links the unchanged file from DIR (repeatable; earlier DIRs win). |
| `--preallocate` | Allocate destination file space up front (fail-fast on a full disk). |
| `--append` | Resume a shorter destination by appending only its tail (prefix not verified; requires `--incremental`). |
| `--append-verify` | Like `--append`, but verifies the retained prefix checksum first (falls back to a full transfer on mismatch). |
| `--delete` | Request removal of destination entries absent from the source. The server must allow deletion. Default timing is delete-after: extras are removed only after the whole transfer succeeded. Scoped to the synchronized directories, so `--files-from` subsets are safe. |
| `--delete-before` | Delete extras before the transfer starts (implies `--delete`). |
| `--delete-during`, `--del` | Delete extras once the keep-set manifest is known, before data is applied (implies `--delete`; early mode, same engine behaviour as `--delete-before`). |
| `--delete-delay` | Delete extras only after a successful transfer (implies `--delete`; commit mode, same behaviour as `--delete-after`). |
| `--delete-after` | Explicit delete-after timing: delete only after the transfer succeeded (implies `--delete`). |
| `--delete-excluded` | Also delete filter-excluded destination mirrors (size-pruned mirrors stay protected). |
| `--max-delete <n>` | Delete at most n destination entries; the rest are skipped and the run exits 25 (partial), matching rsync. |
| `--force` | Allow an incoming file/symlink to replace a destination directory (also during `--delay-updates` publication). |
| `--exclude <pattern>` | Exclude matching paths. Repeatable. |
| `--include <pattern>` | Include matching paths. Repeatable. |
| `--exclude-from <file>` | Read exclude patterns from a file. |
| `--include-from <file>` | Read include patterns from a file. |
| `-f, --filter=RULE` | Add an rsync-style filter rule (`+`/`-`, `include`/`exclude`, `merge`/`.`, `dir-merge`/`:`, `hide`/`H`, `show`/`S`, `protect`/`P`, `risk`/`R`, `clear`/`!`, and modifiers; repeatable). |
| `--max-size <bytes>` | Skip files larger than the limit. |
| `--min-size <bytes>` | Skip files smaller than the limit. |
| `--max-depth <n>` | Limit recursive scanning depth;
zero means unlimited.| | `--incremental` | Skip files matching destination size and mtime.|
| `--checksum` | Include xxHash64 content checks in incremental comparisons.| | `--backup` |
Back up overwritten files.| | `--backup - dir<dir>` | Store backups under a separate directory.|
| `--suffix<suffix>` | Set the backup filename suffix.| | `--partial` |
Select partial - transfer handling.With `--partial - dir`,
completed files are written there;
resumable transfers are not implemented.| | `--partial - dir<dir>` |
Set a relative partial - transfer directory below the server destination root;
use with `--partial`. |
| `--max-alloc <SIZE>` | Maximum single allocation (binary units; default 1G; `0` = no local limit). |
| `--max-depth <n>` | Limit recursive scanning depth; zero means unlimited. |
| `-b, --backup` | Back up overwritten files. |
| `-T, --temp-dir <dir>` | Scratch directory for temp files before the atomic install (confined to the receive root; `EXDEV` falls back to a non-atomic copy). |
| `--backup-dir <dir>` | Store backups under a separate directory (requires `--backup`). |
| `--suffix <suffix>` | Set the backup filename suffix (default: `~`). |
| `--partial` | Select partial-transfer handling. On failed/interrupted writes the already-written temp file is retained (best-effort) for resumption. With `--partial --partial-dir <dir>`, completed files are written under the partial directory and installed atomically. |
| `--partial-dir <dir>` | Set a relative partial-transfer directory below the server destination root. Use with `--partial`. |
| `--inplace` | Write directly to the destination instead of using a temporary file. |
| `--fsync` | Fsync every written file before publication. |
| `--write-batch=FILE` | Run the normal live transfer and also emit a self-contained batch file of the source tree. |
| `--only-write-batch=FILE` | Emit the batch file only (no destination, no server). |
| `--read-batch=FILE` | Apply a batch file to the destination (no source, no server). |
| `--stop-after=MINS` | Stop the transfer after MINS minutes; whatever was already transferred is kept. |
| `--stop-at=TIME` | Stop at an absolute time (`HH:MM`, `HH:MM:SS`, or `now+N[smhd]`). An early stop skips the late `--delete` keep-set. |
### Metadata and links
| Option | Description |
|---|---|
| `-M`, `--preserve` | Preserve supported file metadata, currently mode and modification time. |
| `-l`, `--links` | Request symlink preservation;
link-target transfer remains incomplete. |
| `--copy-links` | Copy symlink referents. |
| `--safe-links` | Skip symlinks that point outside the transfer tree. |
| `--preserve` | Preserve mode and mtime (long form only; equivalent to `-p` + `-t`). Add `-o`/`-g` for owner/group, `-U`/`--atimes` for atime, or an identity flag (`--chown`/`--usermap`/`--groupmap`/`--numeric-ids`/`--copy-as`) for mapped ownership. |
| `-U`, `--atimes` | Preserve access times. Captured with the metadata payload; does not enable ownership. |
| `-N`, `--crtimes` | Capture birth time and transmit it; it cannot be applied because no portable filesystem call can set a birth time (documented divergence). |
| `-p`, `--perms` | Preserve permission bits. One of the four per-attribute preserve flags (with `-t`/`-o`/`-g`); under `-p` the source mode is copied exactly (setuid/setgid/sticky and group/other-write included), matching rsync. |
| `-t`, `--times` | Preserve modification times. Independent of the other attributes; `-O`/`--omit-dir-times` suppresses directories only. |
| `-o`, `--owner` | Preserve the source owner (uid). Mapped by name on the receiver with a raw-numeric fallback (only numeric ids cross the wire); application is privilege-gated. |
| `-g`, `--group` | Preserve the source group (gid). Same name-mapping/numeric-fallback and privilege gating as `-o`. |
| `--no-perms`, `--no-times`, `--no-owner`, `--no-group` | Negate each per-attribute flag (also `--no-p`/`--no-t`/`--no-o`/`--no-g`); `--no-preserve` clears all four. |
| `-E`, `--executability` | Preserve executable permission bits. |
| `-X`, `--xattrs` | Preserve user `user.*` extended attributes. |
| `-A`, `--acls` | Preserve POSIX ACLs. |
| `--chmod <changes>` | Modify transferred permissions (rsync syntax, including `D`/`F`/`X` selectors and `s`/`t`); does not imply `-p`. |
| `--chown=USER:GROUP` | Override the ownership of transferred files (`USER:GROUP`, `USER`, or `:GROUP`); conflicts with `--usermap`/`--groupmap` on the same side. |
| `--usermap=MAP` | Map usernames when applying ownership (`FROM:TO` rules; names, ids, `LOW-HIGH` ranges, `*`, empty-`FROM`). |
| `--groupmap=MAP` | Map group names when applying ownership (same syntax as `--usermap`). |
| `--numeric-ids` | Mapping modifier: apply the source numeric uid/gid directly instead of mapping by name (combine with `-o`/`-g`, `-a`, or a map). |
| `--copy-as=USER[:GROUP]` | Force every written entry to USER[:GROUP]; requires a privileged receiver. |
| `--fake-super` | Record the resolved owner plus mode/time in a reserved `user.fastsync.stat` xattr and replay mode/time; never performs a real chown. |
| `--super` | Permit the receiver to attempt confined super-user activities (device nodes). |
| `--no-super` | Forbid those super-user activities even when the receiver is root. |
| `-l`, `--links` | Copy symlinks as symlinks; the target is stored verbatim (absolute and `..`-bearing targets included), matching rsync. |
| `-L`, `--copy-links` | Copy symlink referents (a broken referent makes the run exit 23, matching rsync). |
| `--safe-links` | Skip symlinks whose target points outside the transfer tree (applied on the sender). |
| `--copy-unsafe-links` | Copy unsafe symlink referents. |
| `-S`, `--sparse` | Request sparse-file handling; full hole preservation is planned. |
| `--munge-links` | Rewrite stored symlink targets with rsync's `/rsyncd-munged/` marker. |
| `-k`, `--copy-dirlinks` | Treat a symlink to a directory as a real directory on the sender. |
| `-K`, `--keep-dirlinks` | Follow an existing destination symlink-to-directory (confined to the receive root). |
| `-H`, `--hard-links` | Preserve hard-link relationships across the transfer. |
| `-D` | Preserve device and special files (implies `--devices --specials`). |
| `--devices` | Recreate device nodes on the destination (privileged; skipped without `CAP_MKNOD`). |
| `--specials` | Recreate special files: FIFOs and unix sockets. |
| `-S`, `--sparse` | Sparse-file handling: receiver preserves holes (zero runs are written as holes; no wire change). |
### Output and logging
| Option | Description |
|---|---|
| `-v`, `--verbose` | Enable debug logging. |
| `--progress` | Show live transfer progress. |
| `--stats` | Print transfer statistics. |
| `-q`, `--quiet` | Suppress non-error output. |
| `--progress` | Show rsync-style per-file progress blocks (not rsync's leading `./` line). |
| `--stats` | Print transfer statistics, including the receiver-only counters reported over the wire. |
| `-i`, `--itemize-changes` | Print an rsync-style per-file change line. |
| `--out-format=FORMAT` | Output format for changed files (`%f %n %l %b %M %%`). |
| `--list-only` | List source files instead of transferring. |
| `--log-file <path>` | Write log output to a file. |
| `-V`, `--version` | Print the FastSync protocol version. |
| `--help` | Print command usage. |
@@ -440,34 +654,117 @@ link-target transfer remains incomplete. |
| Option | Description |
|---|---|
| `-p <port>` | SSH port in the current CLI. This conflicts with rsync's `-p` permissions option and is planned for correction. |
| `--fastsync-server-path <path>` | Remote FastSync server path for SSH mode. |
| `--ssh-port <port>` | SSH port for the SSH transport (default: 22). Note the short `-p` is now rsync's `--perms`. |
| `-e`, `--rsh <command>` | Remote shell to launch for the SSH transport (default: `ssh`; may include arguments). |
| `--fastsync-server-path <path>` | Remote FastSync server path for SSH mode (client-only; never crosses the wire). |
| `--rsync-path <path>` | Alias for `--fastsync-server-path`. |
| `-M`, `--remote-option=OPT` | Append OPT to the remote server invocation over SSH (repeatable; rejected for daemon/TCP destinations). |
| `--trust-sender` | Receiver-local: trust the remote sender's file list and skip path re-validation (does not affect symlink targets). |
| `--timeout <sec>` | Socket + per-message I/O timeout; default `0` = disabled. |
| `--contimeout <sec>` | Connection timeout; default 60; `0` disables. |
| `--source-dir <path>` | Set the source directory explicitly. |
| `--dest-dir <path>` | Set the destination directory explicitly. |
| `--save-to-disk` | Enable server-side disk persistence. |
| `--server-host <host>` | TCP server address. |
| `--server-port <port>` | TCP server port. |
| `--tls` | Enable TLS. Requires `--cert` and `--key`. |
| `--server-port <port>` | TCP server port. `--port <port>` / `--port=<port>` is an alias. |
| `--address <ip>` | Bind the outgoing client socket to this source address. |
| `-4`, `--ipv4` | Force IPv4 for destination resolution. |
| `-6`, `--ipv6` | Force IPv6 for destination resolution. |
| `--sockopts=OPTS` | Comma-separated OPT=VAL socket options applied before connect. |
| `--tls` | Enable TLS. Requires `--cert`, `--key`, and `--ca`. |
| `--cert <path>` | TLS certificate file. |
| `--key <path>` | TLS private key file. |
| `--ca <path>` | CA file for peer verification. |
| `--ca <path>` | CA file for peer verification (always required with `--tls`). |
## Server Options
| Option | Description |
|---|---|
| `--stdio` | Serve one SSH connection over standard input/output. |
| `-p <port>` | TCP listen port. |
| `--daemon` | Run as a persistent daemon listener using a module config file; the daemon default port is 873 (unlike `-p`, which defaults to 8080). |
| `--config=FILE` | Daemon config file (default: `~/.config/fastsync/fastsyncd.conf`, else `/etc/fastsyncd.conf`). Requires `--daemon`. |
| `--dparam=KEY=VALUE` | Override one global config key on the command line. Requires `--daemon`. |
| `--no-detach` | Stay in the foreground (default detaches to the background when running `--daemon`). |
| `-p, --port <port>` | TCP listen port (default: 8080, range: 1–65535). |
| `--tls` | Enable TLS. |
| `--cert <path>` | TLS certificate file. |
| `--key <path>` | TLS private key file. |
| `--ca <path>` | CA file for peer verification. |
| `--destination-root <path>` | Confine received files to this server-side root;
defaults to the current directory. |
| `--allow-delete` | Permit client delete manifests. Deletion is refused by default. |
| `--cert <path>` | TLS certificate file (PEM). |
| `--key <path>` | TLS private key file (PEM). |
| `--ca <path>` | CA file for peer verification (PEM). |
| `--client-cn <name>` | TLS client certificate CN; mandatory with `--tls` (the server verifies the client CN). |
| `--destination-root <path>` | Confine received files to this server-side root; defaults to the current directory. |
| `--address <addr>` | Bind the listening socket to this address. |
| `-4`, `--ipv4` | Bind an IPv4 socket (default). |
| `-6`, `--ipv6` | Bind an IPv6 socket. |
| `--allow-delete` | Permit client delete manifests. Deletion is refused by default. This also gates `--force` (which can recursively replace/remove a destination directory tree). |
| `--allow-super` | Standalone TCP listener only: keep super-user activities enabled for a **root** receiver. Without it a root standalone server forces `SUPER_MODE_OFF`, so client `--devices`/`--write-devices`/`--super` and client-chosen ownership requests are skipped/refused. Rejected with `--stdio` (the SSH remote argv is client-composed; use a forced command if the default must hold). No effect when not root. Daemon modules opt in per module with `client owner = yes`. |
| `--trust-sender` | Trust the remote sender's file list: skip the receiver's up-front path-traversal re-validation (fewer checks, faster, potentially unsafe; off by default). It does not affect symlink targets, which are stored verbatim either way. |
| `--no-super` | Operator veto: never attempt super-user activities (ownership, device nodes) even as root, and refuse any client `--copy-as`/`--super` request. |
| `--allow-unauthenticated` | Permit plaintext/anonymous network clients; an auth-required module still accepts only opted-in loopback plaintext. |
| `--iconv=LOCAL[,REMOTE]` | Declare this server's LOCAL charset for file-name conversion. |
| `--password-file=FILE` | Credential store for modules that declare `auth users`. Requires `--daemon`. |
| `--early-input=FILE` | Second credential store layered over `--password-file`. Requires `--daemon`. |
| `--hash-credentials <file>` | Read `<file>`'s `user:password` lines and print PBKDF2 credential-store lines to stdout, then exit. Cannot be combined with `--daemon` or `--stdio`. |
| `--iterations N` | PBKDF2 iteration count for `--hash-credentials` (default 600000, range 100000–10000000). Requires `--hash-credentials`. |
| `-v`, `--verbose` | Enable debug logging. |
| `--help` | Print server usage. |
### Daemon configuration
`fastsync-server --daemon --config FILE` reads a line-based module config (an
implicit global section, then `[module]` sections). Besides `port`, `motd file`,
and `address`, the global section accepts:
- `max connections = N` — global cap on concurrent connections, default 100. The
listener enforces it; `0`, negative, and non-numeric values are parse errors.
- `max connections per host = N` — cap on concurrent connections from a single
source IP, default 0 (unlimited). Enforced across all forked connection
children through a shared registry.
- `auth failure delay = MS` — milliseconds to sleep after a failed
authentication, default 500. `0` disables it and the value is capped at 5000,
so online password guessing is rate-limited per connection. Successful auths
are never delayed.
- `auth lockout threshold = N` — number of failed authentications from one source
IP before that source is locked out, default 10; `0` disables the lockout. The
failure counter is shared across every connection child, so the lockout holds
even when the next attempt is handled by a different forked child.
- `auth lockout duration = SECONDS` — how long a locked-out source is refused
(default 300). A locked-out client is refused before any SCRAM challenge is
sent; a successful authentication clears the counter.
- `hosts allow` / `hosts deny` — comma- and/or whitespace-separated host access
patterns.
A `[module]` requires `path`, and may also set `read only`, `client owner`,
`auth users`, `max connections` (0 = unlimited; enforced per module across all
connection children), and its own `hosts allow`/`hosts deny`.
The per-host cap and the shared auth lockout identify a source by its numeric
peer IP. **Loopback peers (127.0.0.0/8, IPv6 `::1`) are exempt**: every local
client shares that one address, so counting or locking them out would let one
local process deny service to all the others. The per-module and global
`max connections` caps still apply to loopback. Because the key is the peer IP,
`max connections per host` and `auth lockout` also cannot distinguish clients
behind the same NAT, proxy, or reverse-proxy address — they share one budget and
one lockout counter, so an over-aggressive lockout can affect unrelated users
behind that address. Prefer TLS client certificates (`--client-cn`) plus
`hosts allow`/`hosts deny` for per-client policy when clients share an address,
and size `auth lockout threshold` accordingly.
The shared per-source table has a bounded lifetime: an entry with no live
connection is reclaimed once its lockout has expired, or after it has been idle
(300 s). If every entry is still live or locked, a new source is admitted without
per-host accounting (fail open) and a rate-limited warning is logged; the
per-module cap and host ACLs still apply. The occupancy counters are re-derived
from the shared slot table after every child exit, so a child killed mid-transfer
(or mid-registration) cannot leak a slot or an occupancy count.
Host patterns are `*` (match all), IPv4/IPv6 literals, or IPv4/IPv6 CIDR
(`10.0.0.0/8`, `2001:db8::/32`). Hostnames are not resolved, so hostname globs
are rejected at parse time rather than silently never matching. A matching
`hosts deny` rejects; if any `hosts allow` entries exist, a peer matching none of
them is rejected; deny takes precedence over allow. The global list is checked
before the module list, before authentication, and the connecting peer address
(IPv4 or IPv6) appears in the connection and authentication audit log lines.
## Architecture
### Client
@@ -492,17 +789,61 @@ defaults to the current directory. |
## Protocol and Security
FastSync protocol version `2.5.0` is shared by the client and server. The
FastSync protocol version `2.26.0` is shared by the client and server. The
current protocol is sender-driven and includes configuration negotiation,
including the maximum allocation limit, incremental checks, checksums,
manifests, keep-alives, abort handling, per-file remove-source results, and
FastSync-native delta messages.
Client and server versions must currently match exactly.
TLS provides encrypted TCP transport. Supplying `--ca` enables certificate
verification; without it, traffic is encrypted but peer identity is not
verified. Use certificate verification for deployments where authentication
matters. The default TCP transport is not encrypted.
Daemon modules that declare `auth users` authenticate with a SCRAM-SHA-256-style
challenge/response against a salted PBKDF2 verifier store: no password and no
replayable bearer credential crosses the wire or is stored on the daemon. All
store entries share one iteration count, and an unknown user is answered with a
deterministic per-username dummy challenge, so probing the daemon cannot
enumerate users. Store lines are generated with
`fastsync-server --hash-credentials <plaintext-file>` (see `RSYNC_COMPAT.md`);
redirect that output to an owner-only (mode 0600) file, and note that legacy
`user:SHA256HEX` stores are rejected. FastSync also maintains an owner-only
(mode 0600) `<store>.dummykey` sidecar next to the store: it holds the store-wide
dummy key, is auto-created on first load, and must be preserved across daemon
restarts so the dummy challenge for an unknown user stays stable (the key is
never regenerated while the sidecar exists). The sidecar is secret material and
must be protected like the credential store: keep it owner-only (mode 0600) and
include it with the store in backups and credential rotation. If the sidecar
cannot be created (a process-substitution/FIFO store path such as `/dev/fd/N`, a
read-only filesystem, a missing directory, or a create, write, fsync, link, or
fchmod failure), the daemon logs a warning and uses a transient key, so the
cross-restart guarantee does not hold for those deployments. One residual is
accepted: the store
iteration count is observable pre-auth by design, since the miss path must match
a hit.
An `auth users` module accepts credentials only when one of two conditions
holds: (a) the connection is an encrypted, verified TLS connection whose client
certificate matches the server's `--client-cn`, or (b) the connection is
plaintext from a loopback peer **and** the operator explicitly passed
`--allow-unauthenticated`. A remote plaintext peer is refused before any
challenge is sent, and `--allow-unauthenticated` never permits remote plaintext
auth: remote peers still require verified TLS regardless of the flag. Clients
sending daemon credentials with `--password-file` to a non-loopback daemon must
therefore use `--tls`; the client rejects a non-local plaintext credential
destination before any network I/O. Daemon modules are a `--daemon`-only
feature: the SSH `--stdio` path never loads a daemon config and is not an auth
transport for them.
Because the loopback allowance trusts whichever peer the kernel reports as
`127.0.0.1`, it assumes nothing relays remote connections to the daemon. A local
TCP forwarder or a TLS-terminating proxy in front of an auth-module listener
makes remote clients appear as loopback and bypasses the mutual-TLS identity
check, so do not front an auth-module listener with such a relay. `--tls` always
mandates `--client-cn`, so a TLS connection to an auth-required module always
has its client CN verified (`--client-cn` matches the certificate's CN only, not
a subjectAltName, which is acceptable for a private CA).
TLS provides encrypted TCP transport. Both the client and the server require
`--ca` together with `--tls`, so peer certificates are always verified
(`SSL_VERIFY_PEER`, depth 4). The default TCP transport is not encrypted.
The receiver protects its destination root with path validation, `openat()`
directory traversal, `O_NOFOLLOW`, temporary files, and atomic renames. Delete
@@ -513,18 +854,27 @@ operations require the server's explicit `--allow-delete` policy.
The project will reach the drop-in replacement goal in stages:
1. Correct rsync option meanings, including short options, combined options,
and `--option=value` syntax.
and `--option=value` syntax — **done** in the rsync-parity wave: `-r`/`-b`/
`-L`/`-B`, short-option clustering (`-av`, `-aAX`, `-rlpt`), and attached
values (`-B1000`, `-essh`, `-MOPT`) all parse.
2. Add differential tests that compare FastSync and rsync contents, metadata,
links, deletes, filters, dry runs, and exit codes.
3. Make `-a` implement the expected recursive, links, permissions, times,
owner/group, and supported special-file behavior.
4. Complete symlink, sparse-file, metadata, delete-policy, and resumable-write
semantics.
links, deletes, filters, dry runs, and exit codes — **done** for the
completion wave's scope; the tests live in `tests/integration/` and skip
cleanly when rsync is unavailable.
3. `-a` implements full rsync `-rlptgoD`; under `-p` the source mode is copied
exactly (no masking). Ownership application stays privilege-gated, as in
rsync.
4. Symlink (verbatim storage), sparse-file, metadata, delete-policy (including
`--max-delete` partial + exit 25, per-directory `--delete-during`/
`--delete-delay`), codecs, and resumable-write semantics are implemented;
remaining work is the documented edge cases, which the **Parity Completion
Wave** section of `RSYNC_COMPAT.md` enumerates honestly.
5. Add rsync remote-shell and daemon protocol interoperability.
6. Keep FastSync performance options as negotiated, optional extensions.
The exhaustive implementation matrix and compatibility notes are in
[`RSYNC_COMPAT.md`](RSYNC_COMPAT.md).
[`RSYNC_COMPAT.md`](RSYNC_COMPAT.md); each row is classified as parity, caveat,
or divergent.
## Testing
@@ -537,7 +887,7 @@ Run the unit test binary:
Run the Python integration suite:
```bash
python3 -m pytest tests/
python3 -m pytest tests/integration/ -n 4 --dist=load -m "not setpriv"
```
For stricter local validation:
@@ -561,10 +911,13 @@ rsync protocol or filesystem-semantic compatibility.
## Performance Guidance
- Use `-m` for workloads with many files or enough CPU parallelism.
- Use `-c` or `-z` when network bandwidth is more constrained than CPU.
- Use `-j`/`--threads` for workloads with many files or enough CPU parallelism
(`-m` is `--prune-empty-dirs`).
- Use `-z` when network bandwidth is more constrained than CPU (`-c` is
`--checksum`, not a bandwidth option).
- Tune `--chunk-size` for file sizes, memory limits, and network latency.
- Use `-f` for large uncompressed TCP transfers where zero-copy I/O helps.
- Use `--sendfile` for large uncompressed TCP transfers where zero-copy I/O
helps (`-f` is `--filter`).
- Use `--incremental` to avoid retransmitting unchanged files.
- Use `--delta` for changed files when both endpoints are FastSync peers.
- Use `--bwlimit` when sharing a link with other traffic.
+930 -221
View File
File diff suppressed because it is too large. Load diff
+412 -67
View File
@@ -3,19 +3,24 @@
Compares FastSync configs against rsync (no compression) and rsync+zstd.
Data is ~75% random/incompressible and ~25% structured/compressible by default,
controllable via --random-ratio.
controllable via --random-ratio. Transfers are verified by default (source and
destination must match) so a fast-but-broken copy is never counted.
Usage:
python3 benchmark/bench.py
python3 benchmark/bench.py --runs 5 --profiles lan wan
python3 benchmark/bench.py --random-ratio 0.5 --size-mb 50
python3 benchmark/bench.py --delay 50ms --jitter 10ms --throughput 100mbit
python3 benchmark/bench.py --warm --runs 3
python3 benchmark/bench.py --output json
"""
import argparse
import filecmp
import json
import math
import os
import random
import shlex
import shutil
import socket
import statistics
@@ -25,8 +30,11 @@ import tempfile
import time
PROJECT_ROOT = os.path.abspath(os.path.join(os.path.dirname(__file__), ".."))
BUILD_DIR = os.path.join(PROJECT_ROOT, "build")
SERVER_CMD = [os.path.join(BUILD_DIR, "server")]
DEFAULT_BUILD_DIR = "build-bench"
# Populated by configure_build_dirs(); default to the dedicated bench dir so
# importing this module never depends on the user's existing build/ tree.
BUILD_DIR = os.path.join(PROJECT_ROOT, DEFAULT_BUILD_DIR)
SERVER_CMD = [os.path.join(BUILD_DIR, "server"), "--allow-unauthenticated"]
CLIENT_CMD = [os.path.join(BUILD_DIR, "client")]
BENCH_DIR = os.path.join(PROJECT_ROOT, "bench_data")
@@ -44,10 +52,11 @@ NETWORK_PROFILES = {
FASTSYNC_CONFIGS = [
{"name": "fastsync", "flags": [], "tool": "fastsync"},
{"name": "fastsync -c", "flags": ["-c"], "tool": "fastsync"},
{"name": "fastsync -m", "flags": ["-m"], "tool": "fastsync"},
{"name": "fastsync -m -c", "flags": ["-m", "-c"], "tool": "fastsync"},
{"name": "fastsync -m -c -s", "flags": ["-m", "-c", "-s"], "tool": "fastsync"},
{"name": "fastsync -z", "flags": ["-z"], "tool": "fastsync"},
{"name": "fastsync -j", "flags": ["-j"], "tool": "fastsync"},
{"name": "fastsync -j -z", "flags": ["-j", "-z"], "tool": "fastsync"},
{"name": "fastsync -j -z --chunk-serialization", "flags": ["-j", "-z", "--chunk-serialization"], "tool": "fastsync"},
{"name": "fastsync --sendfile", "flags": ["--sendfile"], "tool": "fastsync"},
]
RSYNC_CONFIGS = [
@@ -56,6 +65,7 @@ RSYNC_CONFIGS = [
{"name": "rsync -z --zstd", "flags": ["-z", "--zc", "zstd"],"tool": "rsync"},
]
class RsyncDaemon:
"""Manages an rsync daemon for network-fair benchmarking."""
@@ -117,6 +127,12 @@ STRUCTURED_FILES = {
"nested/another.txt": b"another nested file\n" * 50,
}
# Repeated text used to synthesize genuinely compressible filler of any size.
COMPRESSIBLE_TEXT = (
b"FastSync benchmark payload: the quick brown fox jumps over the lazy dog. "
b"0123456789 ABCDEFGHIJKLMNOPQRSTUVWXYZ abcdefghijklmnopqrstuvwxyz\n"
)
class Progress:
"""Simple progress bar with ETA."""
@@ -151,35 +167,127 @@ class Progress:
sys.stderr.flush()
def write_compressible(path, nbytes):
"""Write exactly nbytes of highly compressible, repeated text content."""
if nbytes <= 0:
return
block = COMPRESSIBLE_TEXT * (max(1, 8192 // len(COMPRESSIBLE_TEXT)) + 1)
remaining = nbytes
with open(path, "wb") as f:
while remaining > 0:
piece = block if remaining >= len(block) else block[:remaining]
f.write(piece)
remaining -= len(piece)
def generate_bench_data(source_dir, size_mb=25, random_ratio=0.75):
"""Generate test data. ~random_ratio is incompressible, rest is structured."""
"""Generate test data honouring the requested random/compressible split.
Exactly ``random_ratio * target`` bytes are incompressible random data and
the remainder is genuinely compressible structured/repeated content. The
measured byte counts are returned so callers can report the real mix.
"""
if os.path.exists(source_dir):
shutil.rmtree(source_dir)
os.makedirs(source_dir)
target = size_mb * 1024 * 1024
structured_budget = int(target * (1 - random_ratio))
written = 0
random_budget = int(target * random_ratio)
compressible_budget = target - random_budget
compressible_written = 0
random_written = 0
files = 0
# A handful of fixed, human-meaningful files (directories, small files, a
# binary blob) as long as they fit inside the compressible budget.
for rel_path, content in STRUCTURED_FILES.items():
if written >= structured_budget:
if compressible_written + len(content) > compressible_budget:
break
full_path = os.path.join(source_dir, rel_path)
os.makedirs(os.path.dirname(full_path), exist_ok=True)
with open(full_path, "wb") as f:
f.write(content)
written += len(content)
compressible_written += len(content)
files += 1
os.makedirs(os.path.join(source_dir, "bulk"), exist_ok=True)
# Fill the rest of the compressible share with generated repeated content.
if compressible_written < compressible_budget:
os.makedirs(os.path.join(source_dir, "compressible"), exist_ok=True)
i = 0
while written < target:
chunk_size = min(5 * 1024 * 1024, target - written)
with open(os.path.join(source_dir, f"bulk/file_{i}.dat"), "wb") as f:
f.write(random.randbytes(chunk_size))
written += chunk_size
while compressible_written < compressible_budget:
chunk = min(1024 * 1024, compressible_budget - compressible_written)
write_compressible(os.path.join(source_dir, "compressible", f"text_{i}.dat"), chunk)
compressible_written += chunk
files += 1
i += 1
return written
# Incompressible share.
if random_written < random_budget:
os.makedirs(os.path.join(source_dir, "bulk"), exist_ok=True)
i = 0
while random_written < random_budget:
chunk = min(5 * 1024 * 1024, random_budget - random_written)
with open(os.path.join(source_dir, "bulk", f"file_{i}.dat"), "wb") as f:
f.write(random.randbytes(chunk))
random_written += chunk
files += 1
i += 1
return {
"total_bytes": compressible_written + random_written,
"compressible_bytes": compressible_written,
"random_bytes": random_written,
"files": files,
}
def list_relative_files(root):
"""Return the set of file paths (relative to root) under a directory."""
found = set()
for dirpath, _dirnames, filenames in os.walk(root):
for name in filenames:
full = os.path.join(dirpath, name)
found.add(os.path.relpath(full, root))
return found
def verify_transfer(source_dir, dest_dir):
"""Recursively check dest matches source (paths, sizes, content).
Returns (ok, detail). Content is compared byte-for-byte, never hashed, so
collisions are impossible. This is intentionally not part of the timing.
"""
if not os.path.isdir(dest_dir):
return False, "destination directory missing"
src_files = list_relative_files(source_dir)
dst_files = list_relative_files(dest_dir)
if src_files != dst_files:
missing = src_files - dst_files
extra = dst_files - src_files
return False, f"path set mismatch (missing {len(missing)}, extra {len(extra)})"
for rel in sorted(src_files):
src = os.path.join(source_dir, rel)
dst = os.path.join(dest_dir, rel)
if os.path.getsize(src) != os.path.getsize(dst):
return False, f"size mismatch: {rel}"
if not filecmp.cmp(src, dst, shallow=False):
return False, f"content mismatch: {rel}"
return True, ""
def percentile(values, pct):
"""Linear-interpolation percentile (matches numpy's default method)."""
if not values:
return None
ordered = sorted(values)
if len(ordered) == 1:
return ordered[0]
rank = (len(ordered) - 1) * (pct / 100.0)
low = math.floor(rank)
high = math.ceil(rank)
if low == high:
return ordered[int(rank)]
return ordered[low] + (ordered[high] - ordered[low]) * (rank - low)
def find_free_port():
@@ -207,18 +315,45 @@ def wait_proc(proc, timeout=5):
proc.wait()
def _tc_base_cmd():
"""Return the command prefix for tc, honouring root vs sudo."""
tc = shutil.which("tc")
if not tc:
raise RuntimeError(
"tc (iproute2) not found in PATH; install iproute2 to use network profiles")
if os.geteuid() == 0:
return [tc]
sudo = shutil.which("sudo")
if sudo:
return [sudo, tc]
raise RuntimeError(
"applying network limits requires root or sudo; "
"re-run as root or install sudo")
def _run_tc(args, check=True):
return subprocess.run(_tc_base_cmd() + args, check=check, capture_output=True)
def netem_apply(delay=None, jitter=None, throughput=None, loss=None):
"""Apply tc/netem rules to loopback. Pass None to skip a parameter."""
netem_reset()
cmd = ["sudo", "tc", "qdisc", "add", "dev", "lo", "root", "netem"]
params = []
if throughput:
cmd += ["rate", throughput]
params += ["rate", throughput]
if delay:
cmd += ["delay", delay, jitter or "0ms"]
params += ["delay", delay, jitter or "0ms"]
if loss:
cmd += ["loss", loss]
if len(cmd) > 6:
subprocess.run(cmd, check=True, capture_output=True)
params += ["loss", loss]
if not params:
return
try:
_run_tc(["qdisc", "add", "dev", "lo", "root", "netem"] + params)
except subprocess.CalledProcessError as exc:
detail = exc.stderr.decode(errors="replace").strip() if exc.stderr else str(exc)
raise RuntimeError(f"failed to apply network profile via tc/netem: {detail}") from exc
except RuntimeError:
raise
def netem_apply_profile(profile_name):
@@ -235,7 +370,11 @@ def netem_apply_profile(profile_name):
def netem_reset():
subprocess.run("sudo tc qdisc del dev lo root".split(), capture_output=True)
"""Best-effort removal of any loopback qdisc. Always safe to call."""
try:
_run_tc(["qdisc", "del", "dev", "lo", "root"], check=False)
except Exception:
pass
def run_fastsync(source_dir, dest_dir, flags, port):
@@ -252,8 +391,10 @@ def run_fastsync(source_dir, dest_dir, flags, port):
duration = time.monotonic() - start
if result.returncode == 0:
return duration
sys.stderr.write(f" fastsync failed (exit {result.returncode}): "
f"{result.stderr.strip()[:500]}\n")
except subprocess.TimeoutExpired:
pass
sys.stderr.write(" fastsync timed out after 120s\n")
return None
@@ -269,8 +410,10 @@ def run_rsync(source_dir, dest_dir, flags, rsync_daemon=None):
duration = time.monotonic() - start
if result.returncode == 0:
return duration
sys.stderr.write(f" rsync failed (exit {result.returncode}): "
f"{result.stderr.strip()[:500]}\n")
except subprocess.TimeoutExpired:
pass
sys.stderr.write(" rsync timed out after 120s\n")
return None
@@ -282,7 +425,81 @@ def run_transfer(config, source_dir, dest_dir, port=None, rsync_daemon=None):
return run_fastsync(source_dir, dest_dir, config["flags"], port)
def run_benchmark(source_dir, dest_dir, configs, runs, profile_name, progress=None):
def apply_incremental_changes(source_dir, target_bytes):
"""Add and modify a few files so a warm transfer has real work to do.
Returns a mutation record (changed byte count plus enough data to revert
and re-apply it) so every warm run can start from a pristine source.
"""
modified_n = 3
added_n = 2
per_file = max(4096, target_bytes // (modified_n + added_n))
modified = {}
added = {}
changed = 0
existing = sorted(list_relative_files(source_dir))
if existing:
step = max(1, len(existing) // modified_n)
for rel in existing[::step][:modified_n]:
path = os.path.join(source_dir, rel)
original_size = os.path.getsize(path)
with open(path, "ab") as f:
f.write(random.randbytes(per_file))
modified[rel] = (original_size, per_file)
changed += per_file
for i in range(added_n):
os.makedirs(os.path.join(source_dir, "incremental"), exist_ok=True)
rel = os.path.join("incremental", f"new_{i}.dat")
write_compressible(os.path.join(source_dir, rel), per_file)
added[rel] = per_file
changed += per_file
return {"changed": changed, "modified": modified, "added": added}
def revert_incremental_changes(source_dir, mutation):
"""Undo apply_incremental_changes so the source is pristine again."""
if not mutation:
return
for rel, (original_size, _appended) in mutation["modified"].items():
path = os.path.join(source_dir, rel)
if os.path.exists(path):
with open(path, "r+b") as f:
f.truncate(original_size)
for rel in mutation["added"]:
path = os.path.join(source_dir, rel)
if os.path.exists(path):
os.remove(path)
def reapply_incremental_changes(source_dir, mutation):
"""Re-apply a mutation after an untimed pristine seed transfer."""
if not mutation:
return
for rel, (_original_size, appended) in mutation["modified"].items():
with open(os.path.join(source_dir, rel), "ab") as f:
f.write(random.randbytes(appended))
for rel, size in mutation["added"].items():
write_compressible(os.path.join(source_dir, rel), size)
def expected_received_root(dest_dir, source_dir, tool):
"""Where a tool places transferred files inside dest_dir.
FastSync mirrors the absolute source path under dest_dir (see the
integration suite's get_dest_received_dir); rsync copies the source tree
contents directly into dest_dir.
"""
if tool == "rsync":
return dest_dir
return os.path.join(dest_dir, os.path.abspath(source_dir).lstrip(os.sep))
def run_benchmark(source_dir, dest_dir, configs, runs, profile_name,
measure_bytes, verify=True, warm=False, mutation=None,
progress=None):
"""Run benchmark for all configs, returns list of results."""
is_limited = profile_name != "unlimited"
has_rsync = any(c["tool"] == "rsync" for c in configs)
@@ -298,7 +515,10 @@ def run_benchmark(source_dir, dest_dir, configs, runs, profile_name, progress=No
results = []
for config in configs:
times = []
invalid = 0
for run_idx in range(runs):
if warm:
revert_incremental_changes(source_dir, mutation)
if os.path.exists(dest_dir):
shutil.rmtree(dest_dir)
os.makedirs(dest_dir, exist_ok=True)
@@ -306,15 +526,33 @@ def run_benchmark(source_dir, dest_dir, configs, runs, profile_name, progress=No
port = find_free_port()
server = None
try:
if config["tool"] == "fastsync":
if config["tool"] == "fastsync" or warm:
server = subprocess.Popen(
SERVER_CMD + ["-p", str(port)],
stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL,
)
wait_for_port(port)
if warm:
seed = run_transfer(config, source_dir, dest_dir, port, rsync_daemon)
if seed is None:
invalid += 1
sys.stderr.write(" warm-mode seeding failed; run not counted\n")
continue
reapply_incremental_changes(source_dir, mutation)
t = run_transfer(config, source_dir, dest_dir, port, rsync_daemon)
if t is not None:
if t is None:
invalid += 1
elif verify:
root = expected_received_root(dest_dir, source_dir, config["tool"])
ok, detail = verify_transfer(source_dir, root)
if ok:
times.append(t)
else:
invalid += 1
sys.stderr.write(f" verification FAILED ({detail}); run not counted\n")
else:
times.append(t)
finally:
if server:
@@ -327,15 +565,22 @@ def run_benchmark(source_dir, dest_dir, configs, runs, profile_name, progress=No
"config": config["name"],
"tool": config["tool"],
"profile": profile_name,
"warm": warm,
"runs": len(times),
"invalid": invalid,
"times": [round(t, 4) for t in times],
}
if times:
entry["p50"] = round(statistics.median(times), 4)
entry["p95"] = round(sorted(times)[int(len(times) * 0.95)], 4) if len(times) > 1 else entry["p50"]
p50 = percentile(times, 50)
p95 = percentile(times, 95)
entry["p50"] = round(p50, 4)
entry["p95"] = round(p95, 4)
entry["min"] = round(min(times), 4)
entry["max"] = round(max(times), 4)
entry["stdev"] = round(statistics.stdev(times), 4) if len(times) > 1 else 0.0
if measure_bytes:
entry["throughput_mbps"] = round(
(measure_bytes / (1024 * 1024)) / p50, 3)
results.append(entry)
return results
finally:
@@ -345,44 +590,59 @@ def run_benchmark(source_dir, dest_dir, configs, runs, profile_name, progress=No
netem_reset()
def print_table(results, total_bytes, random_ratio):
def print_table(results, measure_bytes, stats, warm):
"""Print results as a human-readable table grouped by profile."""
profiles = {}
for r in results:
profiles.setdefault(r["profile"], []).append(r)
total = stats["total_bytes"]
comp_pct = stats["compressible_bytes"] / total * 100 if total else 0
rand_pct = stats["random_bytes"] / total * 100 if total else 0
for profile, entries in profiles.items():
params = NETWORK_PROFILES.get(profile, {})
print(f"\n{'=' * 85}")
print(f"\n{'=' * 95}")
print(f" Profile: {profile.upper()}")
if params.get("rate"):
print(f" Network: {params['rate']}, {params['delay']} +/- {params['jitter']}, loss {params['loss']}")
else:
print(f" Network: unlimited")
print(f" Data: {total_bytes / (1024*1024):.1f} MB ({random_ratio*100:.0f}% random, {(1-random_ratio)*100:.0f}% compressible)")
print(f"{'=' * 85}")
print(f" Data: {total / (1024*1024):.1f} MB "
f"({rand_pct:.0f}% random, {comp_pct:.0f}% compressible actual)")
if warm:
print(f" Mode: warm (incremental) — measured {measure_bytes / (1024*1024):.2f} MB "
f"changed after an untimed full seed")
else:
print(" Mode: cold (full copy)")
print(f"{'=' * 95}")
fs_entries = [e for e in entries if e.get("tool") == "fastsync"]
rsync_entries = [e for e in entries if e.get("tool") == "rsync"]
header = (f" {'Config':<38} {'p50':>8} {'p95':>8} {'min':>8} {'max':>8} "
f"{'stdev':>8} {'MB/s':>9} {'runs':>5} {'bad':>4}")
rule = (f" {'-' * 38} {'-' * 8} {'-' * 8} {'-' * 8} {'-' * 8} "
f"{'-' * 8} {'-' * 9} {'-' * 5} {'-' * 4}")
if fs_entries:
print(f"\n FastSync:")
print(f" {'Config':<25} {'p50':>8} {'p95':>8} {'min':>8} {'max':>8} {'stdev':>8} {'runs':>5}")
print(f" {'-' * 25} {'-' * 8} {'-' * 8} {'-' * 8} {'-' * 8} {'-' * 8} {'-' * 5}")
print(header)
print(rule)
for e in sorted(fs_entries, key=lambda x: x.get("p50", 999)):
_print_entry(e)
if rsync_entries:
print(f"\n rsync:")
print(f" {'Config':<25} {'p50':>8} {'p95':>8} {'min':>8} {'max':>8} {'stdev':>8} {'runs':>5}")
print(f" {'-' * 25} {'-' * 8} {'-' * 8} {'-' * 8} {'-' * 8} {'-' * 8} {'-' * 5}")
print(header)
print(rule)
for e in sorted(rsync_entries, key=lambda x: x.get("p50", 999)):
_print_entry(e)
if params.get("rate_bps") and fs_entries and rsync_entries:
fs_best = min((e["p50"] for e in fs_entries if "p50" in e), default=None)
rsync_best = min((e["p50"] for e in rsync_entries if "p50" in e), default=None)
theoretical = total_bytes / params["rate_bps"]
theoretical = measure_bytes / params["rate_bps"]
if fs_best and rsync_best:
print(f"\n Theoretical max (line rate): {theoretical:.4f}s")
print(f" FastSync best: {fs_best:.4f}s ({theoretical/fs_best:.2f}x vs line rate)")
@@ -392,10 +652,43 @@ def print_table(results, total_bytes, random_ratio):
def _print_entry(e):
if "p50" in e:
print(f" {e['config']:<25} {e['p50']:>7.4f}s {e['p95']:>7.4f}s "
f"{e['min']:>7.4f}s {e['max']:>7.4f}s {e['stdev']:>7.4f} {e['runs']:>5}")
tp = f"{e['throughput_mbps']:.2f}" if "throughput_mbps" in e else "N/A"
print(f" {e['config']:<38} {e['p50']:>7.4f}s {e['p95']:>7.4f}s "
f"{e['min']:>7.4f}s {e['max']:>7.4f}s {e['stdev']:>7.4f} "
f"{tp:>9} {e['runs']:>5} {e.get('invalid', 0):>4}")
else:
print(f" {e['config']:<25} {'N/A':>8} {'N/A':>8} {'N/A':>8} {'N/A':>8} {'N/A':>8} {e['runs']:>5}")
print(f" {e['config']:<38} {'N/A':>8} {'N/A':>8} {'N/A':>8} {'N/A':>8} "
f"{'N/A':>8} {'N/A':>9} {e['runs']:>5} {e.get('invalid', 0):>4}")
def configure_build_dirs(build_dir):
"""Install the selected build directory and derived binary paths."""
global BUILD_DIR, SERVER_CMD, CLIENT_CMD
if not os.path.isabs(build_dir):
build_dir = os.path.join(PROJECT_ROOT, build_dir)
BUILD_DIR = os.path.abspath(build_dir)
SERVER_CMD = [os.path.join(BUILD_DIR, "server"), "--allow-unauthenticated"]
CLIENT_CMD = [os.path.join(BUILD_DIR, "client")]
def build_project():
"""Configure (Release) and build into the dedicated bench build dir."""
if shutil.which("cmake") is None:
sys.stderr.write("cmake not found in PATH; cannot build\n")
sys.exit(1)
os.makedirs(BUILD_DIR, exist_ok=True)
configure = ["cmake", "-B", BUILD_DIR, "-S", PROJECT_ROOT,
"-DCMAKE_BUILD_TYPE=Release"]
result = subprocess.run(configure, capture_output=True, text=True)
if result.returncode != 0:
sys.stderr.write("CMake configure failed:\n" + result.stdout + result.stderr + "\n")
sys.exit(1)
jobs = str(os.cpu_count() or 1)
result = subprocess.run(["cmake", "--build", BUILD_DIR, "-j", jobs],
capture_output=True, text=True)
if result.returncode != 0:
sys.stderr.write("Build failed:\n" + result.stdout + result.stderr + "\n")
sys.exit(1)
def main():
@@ -412,12 +705,20 @@ Custom network limits (--delay/--jitter/--throughput) override profiles.
Data mix:
Default is ~75%% random/incompressible + ~25%% structured/compressible,
reflecting typical real-world file sets.
reflecting typical real-world file sets. The actual mix is measured and
reported. Transfers are verified (destination must match source) unless
--no-verify is given.
Warm mode:
--warm seeds the destination with an untimed full copy of a pristine base,
then measures only the incremental transfer after modifying a few files.
Examples:
%(prog)s --profiles wan --runs 5
%(prog)s --throughput 50mbit --delay 30ms --jitter 5ms
%(prog)s --random-ratio 0.5 --size-mb 100
%(prog)s --warm --runs 3 --no-rsync
%(prog)s --dry-run --size-mb 4 --random-ratio 0.25
""")
parser.add_argument("--runs", type=int, default=3,
help="Number of runs per config (default: 3)")
@@ -425,7 +726,8 @@ Examples:
choices=list(NETWORK_PROFILES.keys()),
help="Predefined network profiles (default: unlimited)")
parser.add_argument("--configs", nargs="+", default=None,
help="Custom FastSync config flags")
help="Custom FastSync config flags (shell-quoted, e.g. "
"\"-j -z --chunk-serialization\")")
parser.add_argument("--size-mb", type=int, default=25,
help="Test data size in MB (default: 25)")
parser.add_argument("--random-ratio", type=float, default=0.75,
@@ -440,6 +742,14 @@ Examples:
help="Custom packet loss (e.g. 1%%)")
parser.add_argument("--no-rsync", action="store_true",
help="Skip rsync comparison")
parser.add_argument("--no-verify", action="store_true",
help="Skip source/destination verification after each run")
parser.add_argument("--warm", action="store_true",
help="Incremental mode: seed dest first, measure only changes")
parser.add_argument("--build-dir", default=DEFAULT_BUILD_DIR,
help=f"Build directory (default: {DEFAULT_BUILD_DIR})")
parser.add_argument("--dry-run", action="store_true",
help="Only generate data and report its composition, then exit")
parser.add_argument("--progress", action="store_true",
help="Show progress bar with ETA")
parser.add_argument("--output", choices=["table", "json"], default="table",
@@ -448,12 +758,48 @@ Examples:
help="Don't clean up test data")
args = parser.parse_args()
# Build
print("Building...")
if os.system(f"cmake -B {BUILD_DIR} -S {PROJECT_ROOT} > /dev/null 2>&1") != 0:
print("CMake configure failed"); sys.exit(1)
if os.system(f"cmake --build {BUILD_DIR} -j$(nproc) > /dev/null 2>&1") != 0:
print("Build failed"); sys.exit(1)
if not 0.0 <= args.random_ratio <= 1.0:
parser.error("--random-ratio must be between 0.0 and 1.0")
if args.size_mb <= 0:
parser.error("--size-mb must be positive")
configure_build_dirs(args.build_dir)
# Generate data
source_dir = os.path.join(BENCH_DIR, "source")
dest_dir = os.path.join(BENCH_DIR, "dest")
stats = generate_bench_data(source_dir, args.size_mb, args.random_ratio)
total_bytes = stats["total_bytes"]
comp_pct = stats["compressible_bytes"] / total_bytes * 100 if total_bytes else 0
rand_pct = stats["random_bytes"] / total_bytes * 100 if total_bytes else 0
print(f"Generated {total_bytes / (1024*1024):.1f} MB in {stats['files']} files "
f"({rand_pct:.0f}% random, {comp_pct:.0f}% compressible actual)",
file=sys.stderr)
if args.dry_run:
print(f"size_mb={args.size_mb} random_ratio={args.random_ratio:.4f} "
f"total_bytes={stats['total_bytes']} "
f"compressible_bytes={stats['compressible_bytes']} "
f"random_bytes={stats['random_bytes']} files={stats['files']}")
if not args.keep_data:
shutil.rmtree(BENCH_DIR, ignore_errors=True)
return
# Warm mode: keep a pristine base copy, then mutate the live source.
base_dir = None
measure_bytes = total_bytes
mutation = None
if args.warm:
change_target = max(64 * 1024, min(int(total_bytes * 0.01), 4 * 1024 * 1024))
mutation = apply_incremental_changes(source_dir, change_target)
measure_bytes = mutation["changed"]
revert_incremental_changes(source_dir, mutation)
print(f"Warm mode: each run seeds a full copy, then measures "
f"{measure_bytes / (1024*1024):.3f} MB of add/change deltas", file=sys.stderr)
# Build (Release: benchmarking a debug build is meaningless)
print(f"Building (Release) into {BUILD_DIR}...", file=sys.stderr)
build_project()
# Determine active profile for display
has_custom_net = args.delay or args.jitter or args.throughput or args.loss
@@ -473,18 +819,10 @@ Examples:
else:
profiles_to_run = args.profiles or ["unlimited"]
# Generate data
source_dir = os.path.join(BENCH_DIR, "source")
dest_dir = os.path.join(BENCH_DIR, "dest")
total_bytes = generate_bench_data(source_dir, args.size_mb, args.random_ratio)
compressible_pct = (1 - args.random_ratio) * 100
random_pct = args.random_ratio * 100
print(f"Generated {total_bytes / (1024*1024):.1f} MB "
f"({random_pct:.0f}% random, {compressible_pct:.0f}% compressible)")
# Build config list
# Build config list (shlex so quoted/space-separated flags survive)
if args.configs:
fastsync_configs = [{"name": c, "flags": c.split(), "tool": "fastsync"} for c in args.configs]
fastsync_configs = [{"name": c, "flags": shlex.split(c), "tool": "fastsync"}
for c in args.configs]
else:
fastsync_configs = list(FASTSYNC_CONFIGS)
@@ -496,14 +834,21 @@ Examples:
total_runs = len(configs) * args.runs * len(profiles_to_run)
progress = Progress(total_runs, "Benchmarking") if args.progress else None
if progress:
print(f"Running {total_runs} transfers...")
print(f"Running {total_runs} transfers...", file=sys.stderr)
all_results = []
try:
for profile in profiles_to_run:
results = run_benchmark(source_dir, dest_dir, configs, args.runs, profile, progress)
results = run_benchmark(source_dir, dest_dir, configs, args.runs, profile,
measure_bytes, verify=not args.no_verify,
warm=args.warm, mutation=mutation,
progress=progress)
all_results.extend(results)
except RuntimeError as exc:
sys.stderr.write(f"error: {exc}\n")
sys.exit(1)
finally:
netem_reset()
if not args.keep_data:
shutil.rmtree(BENCH_DIR, ignore_errors=True)
@@ -511,7 +856,7 @@ Examples:
if args.output == "json":
print(json.dumps(all_results, indent=2))
else:
print_table(all_results, total_bytes, args.random_ratio)
print_table(all_results, measure_bytes, stats, args.warm)
print()
+3
View File
@@ -5,3 +5,6 @@ markers =
setpriv: privilege-dependent tests (drop to an unprivileged user); excluded
from CI because their result depends on the runner/container uid and the
host mount permissions, but run locally as root
daemon_detach: real double-fork backgrounding path (--daemon without
--no-detach); slower/fragile, so it runs in the full suite but not the
fast PR gate
+35 -2
View File
@@ -3,26 +3,59 @@
}:
pkgs.mkShell {
# Development shell for FastSync. Provides the host-side toolchain needed to
# build, lint, unit-test, integration-test and benchmark the project.
# It deliberately does NOT build on entry: run the CMake commands in README.md
# (or use the CI Docker image for exact CI parity).
nativeBuildInputs = with pkgs; [
# build
gcc
cmake
gnumake
pkg-config
# lint / static analysis (matches CI)
clang-tools # clang-format
cppcheck
# tests
(python3.withPackages (ps: with ps; [ pytest pytest-xdist psutil ]))
openssh # SSH transport integration tests
# debugging
gdb
valgrind
# coverage
lcov
# benchmark tooling
rsync
iproute2 # tc/netem for network shaping
# misc
git
curl
nodejs
nixpkgs-fmt
docker
tea
];
buildInputs = with pkgs; [
zstd
zlib
lz4
openssl
(python3.withPackages (ps: with ps; [ pytest ]))
];
# The CMake configure step fetches xxHash via FetchContent, which needs
# network access; NIX_ENFORCE_PURITY must be off so the sandbox does not block.
NIX_ENFORCE_PURITY = 0;
shellHook = ''
export NIX_ENFORCE_PURITY=0
cmake -B build
# Make an existing build tree available on PATH, but never build here.
if [ -d "$PWD/build" ]; then
export PATH="$PWD/build:$PATH"
fi
echo "FastSync dev shell ready."
echo " Build: cmake -B build -S . && cmake --build build -j\$(nproc)"
echo " Unit: ./build/tests"
echo " CI parity: docker run --rm --user \"\$(id -u):\$(id -g)\" -v \"\$PWD:/workspace\" -w /workspace gitea.tap-tap.win/taptap/fastsync-ci:v11 ..."
'';
}
+455 -154
View File
@@ -1,5 +1,7 @@
#include "change_list.h"
#include "checksum.h"
#include "utils.h"
#include <fcntl.h>
#include <limits.h>
#include <stdint.h>
#include <stdio.h>
@@ -7,21 +9,7 @@
#include <string.h>
#include <sys/stat.h>
#include <time.h>
/* Itemize code emitted for a transferred regular file.
*
* Layout (rsync-compatible 11-char item): `>f` marks a regular file that was
* transferred to the remote host; the trailing nine markers are, in order,
* c(hecksum) s(ize) t(ime) p(erms) o(wner) g(roup) u(ser/acl) a(ttrs) x(attrs).
* Every marker is `+` (FastSync does not compare each attribute on the
* receiving side, so a sent file is reported as fully updated). Files that
* are already up to date print no line at all, matching rsync's single -i
* which only itemizes changes.
*
* Because the scanner only yields regular-file transfer candidates, `>d`
* (directory) lines are never produced; directories are not transferred as
* items by FastSync. */
#define ITEMIZE_SENT_FILE ">f+++++++++"
#include <unistd.h>
typedef struct {
char* data;
@@ -80,103 +68,14 @@ static bool strbuf_append(StrBuf* buf, const char* text) {
return true;
}
static bool strbuf_append_ull(StrBuf* buf, unsigned long long value) {
char digits[32];
int written = snprintf(digits, sizeof(digits), "%llu", value);
if (written < 0 || (size_t)written >= sizeof(digits))
return false;
return strbuf_append(buf, digits);
}
static bool strbuf_append_longlong(StrBuf* buf, long long value) {
char digits[32];
int written = snprintf(digits, sizeof(digits), "%lld", value);
if (written < 0 || (size_t)written >= sizeof(digits))
return false;
return strbuf_append(buf, digits);
}
bool change_list_enabled(const Config* config) {
return config != NULL && (config->itemize_changes || config->out_format != NULL ||
(config->log_file != NULL && config->log_file_format != NULL));
}
char* change_render_itemize(const ChangeEvent* event) {
if (event == NULL || event->decision != CHANGE_SENT)
return str_dup("");
const char* code = event->is_directory ? ">d+++++++++" : ITEMIZE_SENT_FILE;
StrBuf line = {0};
bool ok = strbuf_append(&line, code) && strbuf_append(&line, " ") &&
strbuf_append(&line, event->path != NULL ? event->path : "");
if (!ok) {
strbuf_free(&line);
return NULL;
}
return line.data;
}
/* ---- Itemize code ---- */
static const char* leaf_name(const char* path) {
if (path == NULL)
return "";
const char* slash = strrchr(path, '/');
return slash != NULL && slash[1] != '\0' ? slash + 1 : path;
}
char* change_render_format(const char* format, const ChangeEvent* event) {
if (format == NULL)
return NULL;
StrBuf line = {0};
bool ok = true;
for (const char* p = format; *p != '\0' && ok;) {
if (*p != '%') {
ok = strbuf_append_char(&line, *p);
p++;
continue;
}
char token = p[1];
if (token == '\0') {
ok = strbuf_append_char(&line, '%');
break;
}
switch (token) {
case '%':
ok = strbuf_append_char(&line, '%');
break;
case 'f':
ok = strbuf_append(&line, event->path != NULL ? event->path : "");
break;
case 'n':
ok = strbuf_append(&line, leaf_name(event->path));
break;
case 'l':
ok = strbuf_append_ull(&line, event->size);
break;
case 'b':
ok = strbuf_append_ull(&line, event->bytes_sent);
break;
case 'M':
ok = strbuf_append_longlong(&line, (long long)event->mtime_sec);
break;
default:
/* Unknown escape sequences are preserved verbatim. */
ok = strbuf_append_char(&line, '%') && strbuf_append_char(&line, token);
break;
}
p += 2;
}
if (!ok) {
strbuf_free(&line);
return NULL;
}
if (line.data == NULL) {
line.data = str_dup("");
if (!line.data)
return NULL;
}
return line.data;
}
/* Format a mode as an `ls -l` permission string, e.g. `-rw-r--r--`. */
/* Format the permission bits as an `ls -l` string, e.g. `-rw-r--r--`. */
static void mode_to_ls_string(mode_t mode, char out[11]) {
out[0] = S_ISDIR(mode) ? 'd'
: S_ISLNK(mode) ? 'l'
@@ -198,29 +97,94 @@ static void mode_to_ls_string(mode_t mode, char out[11]) {
out[10] = '\0';
}
char* change_render_list_line(mode_t mode, unsigned long long size, time_t mtime,
const char* path) {
char permission[11];
mode_to_ls_string(mode, permission);
char date[32];
struct tm broken_down;
if (localtime_r(&mtime, &broken_down) != NULL) {
if (strftime(date, sizeof(date), "%Y/%m/%d %H:%M:%S", &broken_down) == 0)
snprintf(date, sizeof(date), "?");
} else {
snprintf(date, sizeof(date), "?");
static char itemize_type_char(const ChangeEvent* event) {
if (event->is_directory)
return 'd';
if (event->is_symlink)
return 'L';
if (event->is_special) {
if (S_ISCHR(event->mode) || S_ISBLK(event->mode))
return 'D';
return 'S';
}
return 'f';
}
static bool times_match(const Config* config, const ChangeEvent* event) {
if (!event->dest.known || !event->dest.existed)
return false;
if (event->mtime_sec == event->dest.mtime_sec)
return event->mtime_nsec == event->dest.mtime_nsec;
long long delta = (long long)event->mtime_sec - (long long)event->dest.mtime_sec;
if (delta < 0)
delta = -delta;
return delta <= (long long)config->modify_window;
}
/* Fill the 11-character itemize code (10 chars + NUL). `created` means the
* destination entry did not exist, so every attribute marker is `+`. */
static void itemize_code(const Config* config, const ChangeEvent* event, char code[12]) {
bool known = event->dest.known;
bool created = !known || !event->dest.existed;
char update;
if (event->is_hardlink)
update = 'h';
else if (created)
update = (event->is_directory || event->is_symlink || event->is_special) ? 'c' : '>';
else
update = '>';
code[0] = update;
code[1] = itemize_type_char(event);
if (created) {
for (int i = 0; i < 9; i++)
code[2 + i] = '+';
code[11] = '\0';
return;
}
bool size_diff = event->size != event->dest.size;
bool time_diff = !times_match(config, event);
bool perms_diff = (event->mode & 07777) != (event->dest.mode & 07777);
bool owner_diff = event->uid != (uid_t)event->dest.uid;
bool group_diff = event->gid != (gid_t)event->dest.gid;
code[2] = '.'; /* checksum: no destination digest available */
code[3] = size_diff ? 's' : '.';
code[4] = time_diff ? 't' : '.';
code[5] = (config->preserve_perms && perms_diff) ? 'p' : '.';
code[6] = (config->preserve_owner && owner_diff) ? 'o' : '.';
code[7] = (config->preserve_group && group_diff) ? 'g' : '.';
code[8] = '.'; /* reserved */
code[9] = '.'; /* acl: not compared */
code[10] = '.';
code[11] = '\0';
}
/* rsync %n: the transfer-relative name, with a trailing slash for directories. */
static bool append_name(StrBuf* buf, const ChangeEvent* event) {
if (!strbuf_append(buf, event->name != NULL ? event->name : ""))
return false;
if (event->is_directory && (event->name == NULL || event->name[0] == '\0' ||
event->name[strlen(event->name) - 1] != '/'))
return strbuf_append_char(buf, '/');
return true;
}
/* rsync %L: " -> target" for a symlink, " => target" for a hard link, else "". */
static bool append_link_suffix(StrBuf* buf, const ChangeEvent* event) {
if (event->is_symlink && event->symlink_target != NULL)
return strbuf_append(buf, " -> ") && strbuf_append(buf, event->symlink_target);
if (event->is_hardlink && event->hardlink_target != NULL)
return strbuf_append(buf, " => ") && strbuf_append(buf, event->hardlink_target);
return true;
}
char* change_render_itemize(const Config* config, const ChangeEvent* event) {
if (event == NULL || event->decision != CHANGE_SENT)
return str_dup("");
char code[12];
itemize_code(config, event, code);
StrBuf line = {0};
char size_field[32];
int written = snprintf(size_field, sizeof(size_field), "%llu", size);
if (written < 0 || (size_t)written >= sizeof(size_field)) {
strbuf_free(&line);
return NULL;
}
bool ok = strbuf_append(&line, permission) && strbuf_append_char(&line, ' ') &&
strbuf_append(&line, size_field) && strbuf_append_char(&line, ' ') &&
strbuf_append(&line, date) && strbuf_append_char(&line, ' ') &&
strbuf_append(&line, path != NULL ? path : "");
bool ok = strbuf_append(&line, code) && strbuf_append_char(&line, ' ') &&
append_name(&line, event) && append_link_suffix(&line, event);
if (!ok) {
strbuf_free(&line);
return NULL;
@@ -228,6 +192,238 @@ char* change_render_list_line(mode_t mode, unsigned long long size, time_t mtime
return line.data;
}
/* ---- --out-format / --log-file-format ---- */
/* rsync 3.4.1's `%C` uses the negotiated transfer checksum; with the default
* "auto" choice on both ends that is xxh128. FastSync's internal XXH64 default
* is not an rsync algorithm, so map it to xxh128 for parity. */
static ChecksumAlgo out_format_checksum_algo(const Config* config) {
switch ((ChecksumAlgo)config->checksum_algo) {
case CHECKSUM_ALGO_MD5:
return CHECKSUM_ALGO_MD5;
case CHECKSUM_ALGO_XXH3:
return CHECKSUM_ALGO_XXH3;
case CHECKSUM_ALGO_XXH128:
return CHECKSUM_ALGO_XXH128;
case CHECKSUM_ALGO_XXH64:
default:
return CHECKSUM_ALGO_XXH128;
}
}
/* Render a digest as rsync's sum_as_hex: for xxh128 the HIGH 64-bit half is
* printed before the low half; every other algorithm prints its bytes in order. */
static void digest_to_hex(ChecksumAlgo algo, const uint8_t* digest, size_t len, char* out) {
if (algo == CHECKSUM_ALGO_XXH128 && len == 16) {
uint64_t low = 0;
uint64_t high = 0;
memcpy(&low, digest, sizeof(low));
memcpy(&high, digest + 8, sizeof(high));
snprintf(out, len * 2 + 1, "%016llx%016llx", (unsigned long long)high, (unsigned long long)low);
return;
}
static const char hex[] = "0123456789abcdef";
for (size_t i = 0; i < len; i++) {
out[i * 2] = hex[(digest[i] >> 4) & 0xf];
out[i * 2 + 1] = hex[digest[i] & 0xf];
}
out[len * 2] = '\0';
}
static bool format_uses_checksum(const char* format) {
if (format == NULL)
return false;
for (const char* p = format; *p != '\0';) {
if (*p != '%') {
p++;
continue;
}
char token = p[1];
if (token == '\0')
break;
if (token == 'C')
return true;
p += 2;
}
return false;
}
/* Fill event->checksum/checksum_known for a transferred regular file. A
* non-regular entry (or a hard-link sibling) leaves checksum_known false, which
* renders as spaces like rsync. */
static void fill_event_checksum(const Config* config, const File* file, ChangeEvent* event) {
if (file == NULL || file->is_dir || file->is_symlink || file->is_special ||
(file->link_group != 0 && !file->link_first))
return;
if (!format_uses_checksum(config->out_format) && !format_uses_checksum(config->log_file_format))
return;
if (file->path == NULL)
return;
ChecksumAlgo algo = out_format_checksum_algo(config);
uint8_t digest[CHECKSUM_MAX_DIGEST_LEN];
size_t len = 0;
/* rsync's %C is the transfer checksum, which is always seeded with 0 (it is
* independent of --checksum-seed, as rsync 3.4.1 demonstrates). */
if (!checksum_digest_file(algo, 0, file->path, digest, sizeof(digest), &len))
return;
digest_to_hex(algo, digest, len, event->checksum);
event->checksum_known = true;
}
char* change_render_format(const char* format, const Config* config, const ChangeEvent* event) {
if (format == NULL || event == NULL)
return NULL;
StrBuf line = {0};
bool ok = true;
for (const char* p = format; *p != '\0' && ok;) {
if (*p != '%') {
ok = strbuf_append_char(&line, *p);
p++;
continue;
}
char token = p[1];
if (token == '\0') {
ok = strbuf_append_char(&line, '%');
break;
}
switch (token) {
case '%':
ok = strbuf_append_char(&line, '%');
break;
case 'i': {
if (event->deleted) {
/* rsync's ITEM_DELETED itemize code: `*deleting ` (11 chars). */
ok = strbuf_append(&line, "*deleting ");
break;
}
char code[12];
itemize_code(config, event, code);
ok = strbuf_append(&line, code);
break;
}
case 'f':
ok = strbuf_append(&line, event->path != NULL ? event->path : "");
break;
case 'n':
ok = append_name(&line, event);
break;
case 'L':
ok = append_link_suffix(&line, event);
break;
case 'l': {
char digits[32];
int written = snprintf(digits, sizeof(digits), "%llu", event->size);
ok = written >= 0 && (size_t)written < sizeof(digits) && strbuf_append(&line, digits);
} break;
case 'b': {
char digits[32];
int written = snprintf(digits, sizeof(digits), "%llu", event->bytes_sent);
ok = written >= 0 && (size_t)written < sizeof(digits) && strbuf_append(&line, digits);
} break;
case 'c': {
char digits[32];
int written = snprintf(digits, sizeof(digits), "%llu", event->bytes_read);
ok = written >= 0 && (size_t)written < sizeof(digits) && strbuf_append(&line, digits);
} break;
case 'C': {
if (event->checksum_known) {
ok = strbuf_append(&line, event->checksum);
} else {
/* rsync pads a non-regular / untransferred entry with spaces. */
ChecksumAlgo algo = out_format_checksum_algo(config);
int width = checksum_digest_len(algo) * 2;
for (int i = 0; i < width && ok; i++)
ok = strbuf_append_char(&line, ' ');
}
} break;
case 'M': {
char when[32];
if (format_rsync_datetime(event->mtime_sec, true, when, sizeof(when)))
ok = strbuf_append(&line, when);
} break;
case 't': {
char when[32];
if (format_rsync_datetime(time(NULL), false, when, sizeof(when)))
ok = strbuf_append(&line, when);
} break;
case 'o':
ok = strbuf_append(&line, "send");
break;
case 'p': {
char digits[32];
int written = snprintf(digits, sizeof(digits), "%ld", (long)getpid());
ok = written >= 0 && (size_t)written < sizeof(digits) && strbuf_append(&line, digits);
} break;
case 'B': {
char permission[11];
mode_to_ls_string(event->mode, permission);
ok = strbuf_append(&line, permission + 1);
} break;
case 'U': {
char digits[32];
int written = snprintf(digits, sizeof(digits), "%u", (unsigned)event->uid);
ok = written >= 0 && (size_t)written < sizeof(digits) && strbuf_append(&line, digits);
} break;
case 'G': {
char digits[32];
int written = snprintf(digits, sizeof(digits), "%u", (unsigned)event->gid);
ok = written >= 0 && (size_t)written < sizeof(digits) && strbuf_append(&line, digits);
} break;
default:
/* Unknown escape sequences are preserved verbatim. */
ok = strbuf_append_char(&line, '%') && strbuf_append_char(&line, token);
break;
}
p += 2;
}
if (!ok) {
strbuf_free(&line);
return NULL;
}
if (line.data == NULL) {
line.data = str_dup("");
if (!line.data)
return NULL;
}
return line.data;
}
/* ---- --list-only ---- */
char* change_render_list_line(const Config* config, const ChangeEvent* event) {
(void)config;
if (event == NULL)
return NULL;
char permission[11];
mode_to_ls_string(event->mode, permission);
char date[32];
if (!format_rsync_datetime(event->mtime_sec, false, date, sizeof(date)))
snprintf(date, sizeof(date), "?");
StrBuf line = {0};
char size_field[40];
char grouped[32];
if (!format_big_num(event->size, false, grouped, sizeof(grouped))) {
strbuf_free(&line);
return NULL;
}
int written = snprintf(size_field, sizeof(size_field), "%15s", grouped);
if (written < 0 || (size_t)written >= sizeof(size_field)) {
strbuf_free(&line);
return NULL;
}
const char* name = event->name != NULL && event->name[0] != '\0' ? event->name : ".";
bool ok = strbuf_append(&line, permission) && strbuf_append(&line, size_field) &&
strbuf_append_char(&line, ' ') && strbuf_append(&line, date) &&
strbuf_append_char(&line, ' ') && strbuf_append(&line, name);
if (!ok) {
strbuf_free(&line);
return NULL;
}
return line.data;
}
/* ---- Event emission ---- */
static void print_escaped_line(FILE* stream, const char* line, bool eight_bit_output) {
char* escaped = output_escape(line, eight_bit_output);
if (escaped != NULL) {
@@ -247,15 +443,16 @@ void change_emit(const Config* config, const ChangeEvent* event) {
bool to_stdout = config->itemize_changes || config->out_format != NULL;
bool to_log = config->log_file != NULL && config->log_file_format != NULL;
if (to_stdout) {
char* line = config->out_format != NULL ? change_render_format(config->out_format, event)
: change_render_itemize(event);
char* line = config->out_format != NULL
? change_render_format(config->out_format, config, event)
: change_render_itemize(config, event);
if (line != NULL) {
print_escaped_line(stdout, line, config->eight_bit_output);
free(line);
}
}
if (to_log) {
char* line = change_render_format(config->log_file_format, event);
char* line = change_render_format(config->log_file_format, config, event);
if (line != NULL) {
print_escaped_line(config->log_file, line, config->eight_bit_output);
free(line);
@@ -266,9 +463,6 @@ void change_emit(const Config* config, const ChangeEvent* event) {
static bool format_uses_mtime(const char* format) {
if (format == NULL)
return false;
/* Mirror change_render_format's tokenizer: "%%" is a literal percent (so
* "%%M" does NOT expand %M) and unknown "%X" escapes consume both chars.
* This keeps the optional stat() fallback below in step with the renderer. */
for (const char* p = format; *p != '\0';) {
if (*p != '%') {
p++;
@@ -284,46 +478,153 @@ static bool format_uses_mtime(const char* format) {
return false;
}
void change_emit_file_sent(const Config* config, const File* file) {
/* Relative path of an entry below the transfer root (no leading slash). Uses
* the sender-side send_path override when present (bare-relative -R layout). */
static char* relative_name(const Config* config, const File* file) {
const char* full = file_wire_path(file);
if (file->send_path != NULL)
return str_dup(full != NULL ? full : "");
const char* root = config->send_directory;
if (root == NULL || full == NULL)
return str_dup(full != NULL ? full : "");
size_t root_len = strlen(root);
while (root_len > 1 && root[root_len - 1] == '/')
root_len--;
if (strncmp(root, full, root_len) == 0) {
if (full[root_len] == '\0')
return str_dup("");
if (full[root_len] == '/')
return str_dup(full + root_len + 1);
}
return str_dup(full);
}
/* rsync %f long form: the source argument as typed (leading '/' removed,
* trailing '/' removed, leading "./" removed) joined to the relative name. */
static char* display_name(const Config* config, const char* name) {
const char* root = config->send_directory;
if (root == NULL)
return str_dup(name != NULL ? name : "");
const char* p = root;
while (*p == '/')
p++;
if (p[0] == '.' && p[1] == '/')
p += 2;
size_t root_len = strlen(p);
while (root_len > 0 && p[root_len - 1] == '/')
root_len--;
size_t name_len = name != NULL ? strlen(name) : 0;
if (root_len == 0 && name_len == 0)
return str_dup("");
char* out = malloc(root_len + (root_len > 0 && name_len > 0 ? 1 : 0) + name_len + 1);
if (!out)
return NULL;
size_t offset = 0;
if (root_len > 0) {
memcpy(out, p, root_len);
offset = root_len;
}
if (root_len > 0 && name_len > 0)
out[offset++] = '/';
if (name_len > 0)
memcpy(out + offset, name, name_len);
out[offset + name_len] = '\0';
return out;
}
static void fill_event_from_file(const Config* config, const File* file, ChangeEvent* event,
char** name_out, char** path_out) {
char* name = relative_name(config, file);
char* path = display_name(config, name);
event->name = name;
event->path = path;
*name_out = name;
*path_out = path;
if (file->metadata != NULL) {
event->mtime_sec = file->metadata->mtime_sec;
event->mtime_nsec = file->metadata->mtime_nsec;
event->mode = file->metadata->mode;
event->uid = file->metadata->uid;
event->gid = file->metadata->gid;
} else if (format_uses_mtime(config->out_format) || format_uses_mtime(config->log_file_format)) {
struct stat st;
if (file->path != NULL && stat(file->path, &st) == 0) {
event->mtime_sec = st.st_mtime;
event->mtime_nsec = st.st_mtim.tv_nsec;
}
}
}
void change_emit_file_sent_bytes(const Config* config, const File* file,
unsigned long long bytes_sent, unsigned long long bytes_read) {
if (file == NULL || !change_list_enabled(config))
return;
ChangeEvent event;
memset(&event, 0, sizeof(event));
/* The displayed path is the one transmitted (with -R + --files-from this is
the bare relative destination path); the metadata fallback below still
stats the local absolute path. */
event.path = file_wire_path(file);
event.decision = CHANGE_SENT;
event.is_directory = false;
event.is_symlink = false;
event.is_special = false;
event.is_hardlink = false;
event.size = file->data != NULL ? file->data->size : 0;
/* FastSync has no wire-byte counter yet, so %b reports the source length
* that had to be delivered (always equal to %l); the actual bytes written
* to the socket (compressed/delta) are not measured. */
event.bytes_sent = event.size;
if (file->metadata != NULL) {
event.mtime_sec = file->metadata->mtime_sec;
} else if (format_uses_mtime(config->out_format) || format_uses_mtime(config->log_file_format)) {
/* Best-effort fallback for %M when no metadata was captured (no -M): the
* path is stat()ed just to fill the field, and any failure leaves 0. */
struct stat st;
if (file->path != NULL && stat(file->path, &st) == 0)
event.mtime_sec = st.st_mtime;
event.dest = file->dest_state;
if (file->is_symlink) {
event.is_symlink = true;
event.symlink_target = file->symlink_target;
event.size = file->symlink_target != NULL ? strlen(file->symlink_target) : 0;
event.bytes_sent = 0;
} else if (file->is_special) {
event.is_special = true;
event.bytes_sent = 0;
} else if (file->link_group != 0 && !file->link_first) {
event.is_hardlink = true;
event.hardlink_target = file->hardlink_target;
event.bytes_sent = 0;
} else {
event.bytes_sent = bytes_sent;
/* rsync's %c is the block-checksum bytes received for the file. Even a
* whole-file transfer (no basis; --append/--inplace included) receives
* rsync's 16-byte sum header, so rsync reports 16; a dry run transfers
* nothing and reports 0. FastSync's whole-file path has no sum header, so
* report rsync's value for parity. With delta enabled the real received
* bytes are kept, but FastSync's signature framing differs from rsync's so
* those stay numerically divergent. */
bool delta_active = config->use_delta && !config->whole_file;
event.bytes_read = (!config->dry_run && !delta_active) ? 16 : bytes_read;
}
char* name = NULL;
char* path = NULL;
fill_event_from_file(config, file, &event, &name, &path);
if (name != NULL && path != NULL) {
fill_event_checksum(config, file, &event);
change_emit(config, &event);
}
free(name);
free(path);
}
void change_emit_file_sent(const Config* config, const File* file) {
if (file == NULL)
return;
unsigned long long payload = file->data != NULL ? file->data->size : 0;
change_emit_file_sent_bytes(config, file, payload, 0);
}
/* Build and emit a CHANGE_SENT event for an explicit directory entry (-d). */
void change_emit_dir_sent(const Config* config, const File* file) {
if (file == NULL || !change_list_enabled(config))
return;
ChangeEvent event;
memset(&event, 0, sizeof(event));
event.path = file_wire_path(file);
event.decision = CHANGE_SENT;
event.is_directory = true;
event.size = 0;
event.bytes_sent = 0;
if (file->metadata != NULL)
event.mtime_sec = file->metadata->mtime_sec;
event.dest = file->dest_state;
char* name = NULL;
char* path = NULL;
fill_event_from_file(config, file, &event, &name, &path);
if (name != NULL && path != NULL)
change_emit(config, &event);
free(name);
free(path);
}
+51 -24
View File
@@ -2,7 +2,9 @@
#define CHANGE_LIST_H
#include "config.h"
#include "checksum.h"
#include "file_types.h"
#include "format.h"
#include <stdbool.h>
#include <sys/stat.h>
#include <time.h>
@@ -26,42 +28,59 @@ typedef enum {
} ChangeDecision;
typedef struct {
const char* path; /* full source path */
const char* path; /* long-form display path (rsync %f) */
const char* name; /* transfer-relative path (rsync %n), no trailing slash */
ChangeDecision decision;
bool is_directory;
bool is_symlink;
bool is_special;
bool is_hardlink; /* a hard-link sibling (linked, no data sent) */
bool deleted; /* a would-delete report (-n --delete); no source file */
const char* symlink_target;
const char* hardlink_target;
unsigned long long size; /* source file length in bytes */
/* The number of bytes reported for a sent file. FastSync has no wire-byte
* counter, so this is always the source length (== size / %l); actual
* post-compression/delta bytes on the wire are not counted. */
unsigned long long bytes_sent;
time_t mtime_sec; /* 0 when unknown */
unsigned long long bytes_sent; /* wire bytes actually transferred (rsync %b) */
unsigned long long bytes_read; /* wire bytes read back for this file (rsync %c) */
/* rsync %C: whole-file checksum hex for a transferred regular file. Only
* filled when the active format uses %C (checksum_known == false otherwise,
* which renders as spaces like rsync for non-regular entries). */
bool checksum_known;
char checksum[CHECKSUM_MAX_DIGEST_LEN * 2 + 1];
time_t mtime_sec;
long mtime_nsec;
mode_t mode;
uid_t uid;
gid_t gid;
/* Receiver-reported pre-transfer destination state (OutputDestState.known is
* false when no report was requested/received). */
OutputDestState dest;
} ChangeEvent;
/* True when any output mode is active and per-file events matter. */
bool change_list_enabled(const Config* config);
/* Render the rsync-style itemize line for a transferred file:
* `>f+++++++++ <path>`
* The 11-char code is `>f` (regular file transferred to the remote host)
* followed by c/s/t/p/o/g/u/a/x markers that are all `+` (value will be set
* / differs) because FastSync does not separately compare checksums, size,
* mtime, perms, owner, group, uid, acl, or xattr on the receiving side, so a
* sent file is reported as fully updated. Up-to-date files print no line
* (rsync single `-i` only shows changes). Caller frees the result. */
char* change_render_itemize(const ChangeEvent* event);
/* Render the rsync-style itemize line for a transferred item
* (`%i %n%L`): `>f+++++++++ sub/b.txt`. Caller frees the result. */
char* change_render_itemize(const Config* config, const ChangeEvent* event);
/* Expand an --out-format/--log-file-format template. Tokens:
* %f full source path %b "bytes sent" == the source length (%l);
* %n leaf (base) name actual post-compression/delta wire bytes
* %l file length in bytes are not counted
* %M mtime in whole seconds %% a literal percent sign
/* Expand an --out-format/--log-file-format template. Supported tokens:
* %i itemize code %n transfer-relative name (dir: trailing /)
* %f long display path %l file length in bytes
* %b wire bytes transferred %c block-checksum bytes received (rsync: 16
* for a whole-file transfer, 0 for a dry run)
* %C whole-file checksum hex (xxh128 by default; spaces for non-regular)
* %M mtime (YYYY/MM/DD-HH:MM:SS)
* %t current time %o operation ("send"/"del.")
* %p pid %B permission bits without the type char
* %U uid %G gid
* %L " -> target" / " => target" %% a literal percent sign
* Unknown %X sequences are preserved verbatim. Caller frees the result. */
char* change_render_format(const char* format, const ChangeEvent* event);
char* change_render_format(const char* format, const Config* config, const ChangeEvent* event);
/* Render one --list-only long-listing entry:
* `-rw-r--r-- 12 2026/09/06 10:00:00 <path>`
* `-rw-r--r-- 12 2026/09/06 10:00:00 sub/b.txt`
* (ls -l style columns; mtime in the local time zone). Caller frees it. */
char* change_render_list_line(mode_t mode, unsigned long long size, time_t mtime, const char* path);
char* change_render_list_line(const Config* config, const ChangeEvent* event);
/* Emit an event to every active destination:
* stdout: --itemize-changes line, or the --out-format expansion when set;
@@ -69,7 +88,15 @@ char* change_render_list_line(mode_t mode, unsigned long long size, time_t mtime
* CHANGE_UP_TO_DATE events produce no output. */
void change_emit(const Config* config, const ChangeEvent* event);
/* Build and emit a CHANGE_SENT event for a file the client just sent. */
/* Build and emit a CHANGE_SENT event for a file the client just sent. `bytes_sent`
* is the process-wide wire-byte delta for this file (rsync's %b) and `bytes_read`
* the received bytes used for the delta handshake; pass 0 when unknown. For a
* whole-file transfer %c is pinned to rsync's 16-byte sum header regardless. */
void change_emit_file_sent_bytes(const Config* config, const File* file,
unsigned long long bytes_sent, unsigned long long bytes_read);
/* Build and emit a CHANGE_SENT event for a file the client just sent, deriving
* the wire byte counts from the source payload length. */
void change_emit_file_sent(const Config* config, const File* file);
/* Build and emit a CHANGE_SENT event for an explicit directory entry (-d). */
+2115 -394
View File
File diff suppressed because it is too large. Load diff
+1751 -294
View File
File diff suppressed because it is too large. Load diff
+19 -2
View File
@@ -4,10 +4,27 @@
#include "chunk.h"
#include "config.h"
#include "transport_tcp.h"
#include <signal.h>
#include <stdbool.h>
int send_chunk(Client* client, Chunk* chunk, Config* config);
/* Set ONLY by the client's SIGINT/SIGTERM handler (async-signal-safe: the
* handler stores 1 and does nothing else). The send loops poll it via
* client_abort_pending() and, when set, best-effort send STATUS_ABORT so the
* receiver can clean up before the client exits. */
extern volatile sig_atomic_t client_abort_requested;
bool client_abort_pending(void);
/* Arm/disarm abort handling around the network phase. While disarmed, a
* SIGINT/SIGTERM takes the default action (immediate termination) so local-only
* modes are not left unresponsive. Defined in client_cli.c. */
void client_set_abort_armed(bool armed);
/* Both sender entry points BORROW `config` for the duration of the call; they
* never free it, and the caller retains ownership (freeing it with
* config_delete() once the call returns). */
int send_files(Config* config);
/* Takes ownership only when *config is set to NULL on return. */
int send_files_multithreaded(Config** config);
/* Phase 6 residual-batch (client-only). See client_send.c. */
int write_batch_from_source(const Config* config, const char* batch_path);
int apply_batch_to_dest(const Config* config, const char* batch_path, const char* dest_root);
#endif
+70 -53
View File
@@ -1,25 +1,55 @@
#include "client_validation.h"
#include "delay_updates.h"
#include "log.h"
#include "usage.h"
#include "utils.h"
#include <string.h>
#include <stdio.h>
/* Validate config after parsing. Returns true if valid. */
bool validate_config(const Config* config) {
if (!config->send_directory || !config->receive_root_directory) {
log_message(LOG_LEVEL_ERROR, "source and destination directories are required");
/* Phase 6 residual-batch modes relax the normal source+destination pair: the
batch driver is local and needs only what it consumes. --only-write-batch
emits a batch from the source (no destination, no server);
--read-batch applies a batch to the destination (no source, no server);
--write-batch runs the live transfer AND emits a batch, so it keeps the
full pair. */
bool write_batch = config->write_batch != NULL;
bool only_write_batch = config->only_write_batch != NULL;
bool read_batch = config->read_batch != NULL;
if ((write_batch && only_write_batch) || (write_batch && read_batch) ||
(only_write_batch && read_batch)) {
log_message(LOG_LEVEL_ERROR,
"--write-batch, --only-write-batch, and --read-batch are mutually exclusive");
return false;
}
/* A dry-run of a local batch apply is not meaningful: --read-batch bypasses
the client-side scan/server decision entirely, so dry-run would have no
wire state to report (and must not be used as a mutation escape hatch).
--only-write-batch likewise never contacts a receiver. --write-batch DOES
run a live transfer but additionally mutates the filesystem by emitting the
batch file, so a dry-run must not write it either. Reject all three up
front instead of silently ignoring --dry-run. */
if (config->dry_run && (read_batch || only_write_batch || write_batch)) {
log_message(LOG_LEVEL_ERROR,
"--dry-run cannot be combined with --read-batch, --only-write-batch, or "
"--write-batch; a dry-run must not mutate anything, including batch files");
return false;
}
if (read_batch) {
if (!config->receive_root_directory) {
log_message(LOG_LEVEL_ERROR, "--read-batch requires a destination directory");
print_usage();
return false;
}
if (config_has_basis(config) && config->use_chunk_serialization) {
log_message(LOG_LEVEL_ERROR,
"--compare-dest/--copy-dest/--link-dest require per-file incremental checks and "
"cannot be combined with -s (chunk serialization)");
} else if (only_write_batch) {
if (!config->send_directory) {
log_message(LOG_LEVEL_ERROR, "--only-write-batch requires a source directory");
print_usage();
return false;
}
if (config->use_sendfile && (config->use_chunk_serialization || config->use_compression)) {
log_message(LOG_LEVEL_ERROR, "-f/--sendfile cannot be combined with -c (compression) or -s "
"(chunk serialization)");
} else if (!config->send_directory || !config->receive_root_directory) {
log_message(LOG_LEVEL_ERROR, "source and destination directories are required");
print_usage();
return false;
}
if (config->compression_threads > 0 && !config->use_compression) {
@@ -30,43 +60,19 @@ bool validate_config(const Config* config) {
log_message(LOG_LEVEL_ERROR, "-f/--sendfile is not supported with SSH transport");
return false;
}
if (config->use_incremental && config->use_chunk_serialization) {
log_message(LOG_LEVEL_ERROR, "--incremental is not supported with -s (chunk serialization)");
return false;
}
if (config->skip_compress_set && config->use_chunk_serialization) {
/* -M/--remote-option appends an option to the REMOTE server's argv, which
* only exists on the SSH (user@host:path) transport. A daemon
* (host::module/path) or local TCP destination has no remote command line,
* so the option would be silently ignored; reject it by name instead. */
if (config->remote_option_count > 0 && config->transport != TRANSPORT_SSH) {
log_message(LOG_LEVEL_ERROR,
"--skip-compress cannot be combined with -s (chunk serialization)");
"-M/--remote-option is only valid with the SSH transport (user@host:path); it "
"cannot be used with a daemon (host::module/path) or local TCP destination");
return false;
}
if (config->use_delta && !config->whole_file && !config->use_incremental) {
log_message(LOG_LEVEL_ERROR, "--delta requires --incremental");
return false;
}
if (config->use_delta && !config->whole_file && config->use_chunk_serialization) {
log_message(LOG_LEVEL_ERROR, "--delta cannot be combined with -s (chunk serialization)");
return false;
}
if (config->use_delta && !config->whole_file && config->use_sendfile) {
log_message(LOG_LEVEL_ERROR, "--delta cannot be combined with -f (sendfile)");
return false;
}
/* --append / --append-verify resume a shorter existing destination by
transmitting only the tail. The resume needs the per-file STATUS_CHECK
handshake (so the dest length is learned), which chunk serialization -s
disables; and whole-file is the opposite intent (send everything), so the
two would silently make the resume pointless. Both are rejected up front
rather than silently degrading to a full transfer. */
if ((config->append || config->append_verify) && config->use_chunk_serialization) {
log_message(LOG_LEVEL_ERROR,
"--append/--append-verify require the per-file incremental check and cannot be "
"combined with -s (chunk serialization)");
return false;
}
if ((config->append || config->append_verify) && config->whole_file) {
log_message(LOG_LEVEL_ERROR,
"--append/--append-verify are incompatible with --whole-file (which forces a "
"full transfer)");
/* -4 and -6 are mutually exclusive: a socket address family cannot be both. */
if (config->ipv4 && config->ipv6) {
log_message(LOG_LEVEL_ERROR, "-4/--ipv4 and -6/--ipv6 are mutually exclusive");
return false;
}
if (config->log_file_format && !config->log_file) {
@@ -79,20 +85,31 @@ bool validate_config(const Config* config) {
return false;
}
}
if (config->delay_updates && config->inplace) {
log_message(LOG_LEVEL_ERROR, "--delay-updates does not work with --inplace");
/* Daemon credentials (A7, protocol 2.19.0): a --password-file would send the
username in the clear and derive a SCRAM proof a network sniffer could
attack offline, so it is only allowed over TLS (which itself mandates a
verified --cert/--key/--ca set above) or to a loopback destination. A
remote plaintext daemon is refused here, before any network I/O. */
if (config->password_file && !config->use_tls && !utils_host_is_loopback(config->server_host)) {
log_message(LOG_LEVEL_ERROR, "sending daemon credentials to a non-local server requires --tls");
return false;
}
if (config->delay_updates && delay_updates_staging_name_conflict(config->backup_dir)) {
log_message(LOG_LEVEL_ERROR,
"--backup-dir is reserved when --delay-updates is active (used for the internal "
"staging directory)");
/* Every cross-field invariant the receiver enforces lives in one shared
predicate so the client and the server can never disagree. The client
reports the specific reason here, before any network I/O. */
const char* invariants_error = config_invariants_error(config);
if (invariants_error) {
log_message(LOG_LEVEL_ERROR, "%s", invariants_error);
return false;
}
if (!config_has_valid_delete_timing(config)) {
/* --protocol: FastSync has exactly one wire format, so the forced version
must equal the current PROTOCOL_VERSION exactly. Rejected here, before any
network I/O, rather than letting the server hit its own mismatch check. */
if (strcmp(config->version, PROTOCOL_VERSION) != 0) {
log_message(LOG_LEVEL_ERROR,
"--delete-before/--delete-during/--delete-delay/--delete-after select the delete "
"timing; at most one may be given and each implies --delete");
"--protocol must be %s (FastSync supports only its current wire "
"protocol version and cannot speak an older or virtual one)",
PROTOCOL_VERSION);
return false;
}
return true;
+1026 -206
View File
File diff suppressed because it is too large. Load diff
+108 -39
View File
@@ -4,16 +4,30 @@
#include "chunk.h"
#include "file_list.h"
#include "filter.h"
#include "hardlink.h"
#include "protocol.h"
#include "queue.h"
#include "stop_condition.h"
#include <dirent.h>
#include <stdbool.h>
#include <stdatomic.h>
#include <sys/types.h>
#include <threads.h>
/* Upper bound on the configurable parallel scanner worker count (--threads=N):
* keeps one transfer from spawning an unbounded pool on a very large machine. */
#define MAX_SCANNER_THREADS 256
typedef struct {
bool use_metadata;
/* Phase 4 metadata capture: -U/--atimes and -N/--crtimes tell the scanner to
* capture the source access / birth time into each entry's FileMetadata. */
bool preserve_atimes;
bool preserve_crtimes;
/* Phase 4 xattrs: when preserve_xattrs || preserve_acls is set the scanner
* captures each regular file's whitelisted xattr set onto the File. */
bool preserve_xattrs;
bool preserve_acls;
unsigned long long chunk_size;
char** exclude_patterns;
int exclude_count;
@@ -27,16 +41,45 @@ typedef struct {
bool copy_links;
bool safe_links;
bool copy_unsafe_links;
/* Phase 4 symlink-trust sender options: -k/--copy-dirlinks (dereference a
* symlink to a directory as a directory, keeping symlinks-to-files as
* symlinks) and --munge-links (rewrite each transmitted symlink target with a
* marker; escaping targets are never transmitted). Both are client/sender
* side only and never serialized to the wire (keep_dirlinks is the
* receiver-side counterpart). */
bool copy_dirlinks;
bool munge_links;
bool checksum;
bool one_file_system;
/* Phase 4 special/devices: whether device nodes (--devices) and special files
* (--specials) are preserved via recreation, and whether --copy-devices
* copies a device's content as an ordinary regular file. */
bool preserve_devices;
bool preserve_specials;
bool copy_devices;
/* Phase 2 (files-from / filter layer). All pointers are shared read-only
* across scanner instances and worker threads; ownership stays with the
* caller (client_send). */
const FileListSet* file_list; /* --files-from allow-set, or NULL */
const FilterRuleList* base_filters; /* command-line + -C rules, or NULL */
bool per_dir_filters; /* -F: read .rsync-filter per directory */
/* --delete-excluded: per-directory plain rules become sender-only, so they no
longer protect the receiver from deletion. */
bool delete_excluded;
/* -FF: also exclude the per-directory filter files themselves from the
transfer (single -F transfers them). */
bool exclude_per_dir_filter_files;
bool dirs; /* -d/--dirs: transfer dir entries, no recursion */
bool relative; /* -R/--relative (dest rel paths, with --files-from) */
/* -R/--relative outside --files-from: the destination-relative path prefix
* reconstructed from the source spec (rsync's '/./' cut point), or NULL when
* -R is off or --files-from is in use (the bare-relative path then comes from
* the listed entry). Borrowed read-only; owned by client_send. */
const char* relative_prefix;
/* --list-only: emit an is_dir File for every traversed directory (the listing
* includes directory entries, matching rsync). Client-only; never set on a
* real transfer, which relies on implicit parent creation. */
bool list_dirs;
/* --prune-empty-dirs (long only): in --dirs mode an empty source directory's
explicit entry is omitted from the transfer file list (so nothing is
created at the destination and it can be pruned by --delete); explicitly
@@ -45,17 +88,39 @@ typedef struct {
bool prune_empty_dirs;
/* Delete-excluded protection sink (optional): when non-NULL the scanner
* appends the destination-relative path of every entry it prunes because a
* USER SELECTION rule excluded it (--filter/-C/per-dir rules, the legacy
* --exclude/--include layer, and --max-size/--min-size). The sender turns
* this list into the manifest's protected prefixes so `--delete` leaves the
* destination mirror of excluded source paths alone (rsync's default), and
* empties it when --delete-excluded opts back into deleting them. NOT
* recorded for --files-from subset pruning (whose delete semantics stay
* keep-set-only) or for -R/--files-from relative wire paths. When
* `excluded_mutex` is non-NULL it is taken around every append (the parallel
* scanner shares one list across its worker threads). */
* USER SELECTION rule excluded it (--filter/-C/per-dir rules and the legacy
* --exclude/--include layer). The sender turns this list into the manifest's
* protected prefixes so `--delete` leaves the destination mirror of excluded
* source paths alone (rsync's default), and drops it when --delete-excluded
* opts back into deleting them. NOT recorded for --files-from subset pruning
* (whose delete semantics derive from the synchronized-directory set) or for
* -R/--files-from relative wire paths. When `excluded_mutex` is non-NULL it
* is taken around every append (the parallel scanner shares one list across
* its worker threads). */
ArrayList* excluded_paths;
mtx_t* excluded_mutex;
/* Size-prune protection sink (optional): when non-NULL the scanner appends
* the destination-relative path of every entry it skipped because of
* --max-size/--min-size. rsync never deletes a size-skipped source mirror,
* even under --delete-excluded, so the sender always transmits this list as
* protected prefixes (unlike excluded_paths, which --delete-excluded drops).
* Guarded by `excluded_mutex` like excluded_paths. */
ArrayList* size_skipped_paths;
/* Synchronized-directory sink (optional): when non-NULL the scanner appends
* the destination-relative path of every directory it is about to traverse
* that lies inside a --files-from listed directory (or of every traversed
* directory when there is no list). The sender sends this set with the delete
* manifest so the receiver confines its extras walk to synchronized
* directories, exactly like rsync; the receive root is the "." sentinel.
* Guarded by `excluded_mutex`. */
ArrayList* synced_dirs;
/* Delete-plan directory sink (optional): when non-NULL the scanner appends
* the destination-relative path of every directory it traverses (except the
* receive root). The per-directory --delete-during/--delete-delay plan
* builder uses this to keep an empty in-scope source directory (rsync keeps
* it) and to emit its plan after the data stream, when no file frame would
* otherwise trigger it. Guarded by `excluded_mutex`. */
ArrayList* plan_dirs;
/* --ignore-errors: an unreadable directory during the scan is recorded as an
* I/O error and skipped instead of aborting the scan. Client-only. */
bool ignore_io_errors;
@@ -64,6 +129,28 @@ typedef struct {
* instead of failing (the --dirs generator is the only scanner path that
* observes a listed-but-missing entry). */
bool ignore_missing_args;
/* --hard-links (-H): shared, mutable (mutex-guarded) link-group detection
* table, NULL when -H is off. Owned by the caller (client_send), shared
* read-only here; the parallel scanner passes it unchanged to every worker so
* one table detects every group across all subdirectories. */
HardLinkTable* hardlinks;
/* Phase 6: optional sender stop deadline. When non-NULL the scanner checks
* it at natural loop boundaries and stops emitting chunks once reached
* (without marking the scan as failed), so a busy scan itself stops early.
* Client-only, never serialized to the wire. */
const StopCondition* stop_condition;
/* P7 Wave D (protocol 2.17.0): directory-time capture sink. When
* `capture_dir_times` is true the recursive scan appends one is_dir File
* (with metadata, no payload) per source directory it traverses to
* `dir_entries`, so the sender can transmit trailing STATUS_DIR_TIMES
* frame(s) and the receiver can apply directory mtimes AFTER all children
* are written. `dir_entries_mutex` (optional) guards the list
* for the parallel scanner's shared worker threads; the caller owns both.
* The --dirs generator does not use this (its directory entries carry their
* metadata inline through STATUS_MKDIR). */
bool capture_dir_times;
ArrayList* dir_entries;
mtx_t* dir_entries_mutex;
} ScannerOptions;
/* Internal per-scanner filter state. FilterNode chains represent the ordered
@@ -71,25 +158,14 @@ typedef struct {
typedef struct FilterNode FilterNode;
typedef struct {
/* Scan inputs, copied once at create time. Everything that is also a
ScannerOptions field lives here (with the normalized chunk_size); only
scanner-owned bookkeeping stays as direct members below. */
ScannerOptions options;
Queue* directories;
DIR* current_dir;
char* current_path;
bool use_metadata;
unsigned long long chunk_size;
char** exclude_patterns;
int exclude_count;
char** include_patterns;
int include_count;
unsigned long long max_size;
unsigned long long min_size;
int max_depth;
int current_depth;
bool follow_symlinks;
bool copy_links;
bool safe_links;
bool copy_unsafe_links;
bool checksum;
bool one_file_system;
dev_t root_dev;
bool failed;
/* Phase 2 (files-from / filter layer). */
@@ -99,27 +175,13 @@ typedef struct {
FilterNode* seed_node; /* inherited context of the seed dir, or NULL */
FilterNode* current_node; /* filter context of the open directory */
ArrayList* filter_nodes; /* owned FilterNode arena (may be NULL) */
const FileListSet* file_list;
const FilterRuleList* base_filters;
bool per_dir_filters;
/* --dirs / -R state for the directory-entry generator (dirs_mode replaces
/* --dirs / -R state for the directory-entry generator (options.dirs replaces
the recursive scan). */
bool dirs_mode;
bool relative_mode; /* file_list && relative: send bare relative wire paths */
bool prune_empty_dirs;
bool dirs_root_emitted;
int list_index;
ArrayList* dirs_batch; /* owned when non-NULL */
unsigned long long dirs_batch_size;
/* Excluded-path sink (see ScannerOptions). `excluded_mutex` is shared across
parallel worker threads. */
ArrayList* excluded_paths;
mtx_t* excluded_mutex;
/* --ignore-errors: continue past unreadable directories (records io_error). */
bool ignore_io_errors;
/* --ignore-missing-args: --dirs listed-but-missing entries are skipped, not
fatal (see ScannerOptions.ignore_missing_args). */
bool ignore_missing_args;
/* A directory could not be opened (I/O error, e.g. EACCES). With
--ignore-errors the scan continues past it and the caller decides what to
do; `failed` is reserved for fatal errors that always abort the scan. */
@@ -169,6 +231,13 @@ bool scanner_same_filesystem(bool one_file_system, dev_t root_device, dev_t entr
* "/". Exposed so tests can exercise the mapping directly. */
char* scanner_path_relative(const char* root, const char* fs_path);
/* -R/--relative destination-relative prefix reconstructed from a source spec:
* the path after rsync's first '.' path component (the '/./' cut point), with
* leading/trailing slashes removed, or the whole spec (normalized) when there
* is no cut. Returns "" for the receive root, or NULL when `spec` is NULL or
* allocation fails. Exposed so tests can exercise the mapping directly. */
char* scanner_relative_prefix(const char* spec);
ParallelScanner* parallel_scanner_create_with_options(const char* root_directory,
const ScannerOptions* options,
ProtocolSession* allocation_session);
+224 -54
View File
@@ -2,6 +2,7 @@
#include <stdio.h>
#include <delta.h>
#include <chunk.h>
#include "scanner.h"
void print_usage(void) {
printf("Usage:\n");
@@ -11,36 +12,73 @@ void print_usage(void) {
printf("Destination formats:\n");
printf(" user@host:/path SSH transport (rsync-style)\n");
printf(" host:/path SSH transport (current user)\n");
printf(" host::module/path Daemon TCP transport (fastsync-server --daemon);\n");
printf(" module names a server-side module, path is relative\n");
printf(" within it (connect with --server-port)\n");
printf(" /local/path TCP transport (requires server on localhost:8080)\n");
printf("\n");
printf("Options:\n");
printf(" -c [level] Enable compression (level 1-22, default 5)\n");
printf(" -z [level] Alias for -c\n");
printf(" -a, --archive Archive mode (-c -m -M)\n");
printf(" -c, --checksum Verify content by checksum instead of size+mtime\n");
printf(" -z, --compress [level] Enable compression (level 1-22, default 5)\n");
printf(" -a, --archive rsync archive mode (-rlptgoD): links, perms, times,\n");
printf(" owner, group, devices and specials; not\n");
printf(" compression/multithreading\n");
printf(" -r, --recursive Recurse into directories (FastSync is always recursive)\n");
printf(" -n, --dry-run Show what would be transferred\n");
printf(" --remove-source-files Remove regular source files after successful transfer\n");
printf(" -p <port> SSH port (default: 22)\n");
printf(" -p, --perms Preserve permission bits\n");
printf(" -t, --times Preserve modification times\n");
printf(" -o, --owner Preserve owner (uid)\n");
printf(" -g, --group Preserve group (gid)\n");
printf(" --ssh-port <port> SSH port (default: 22)\n");
printf(" -e, --rsh <command> Remote shell to launch on the client for the SSH\n");
printf(" transport (default: ssh). The command may include\n");
printf(" arguments, e.g. -e \"ssh -p 2222\"\n");
printf(" --rsync-path <path> Alias for --fastsync-server-path (path to the\n");
printf(" fastsync server binary on the remote side)\n");
printf(" --blocking-io Leave the SSH transport socket without read/write\n");
printf(" timeouts so it blocks naturally\n");
printf(" --outbuf=MODE stdout/stderr buffering: N (none/unbuffered),\n");
printf(" L (line-buffered), or B (block-buffered, default)\n");
printf(" --progress Show transfer progress\n");
printf(" -P Partial mode with progress (retention incomplete)\n");
printf(" -8, --8-bit-output Leave high-bit characters unescaped in output\n");
printf(" --iconv=LOCAL[,REMOTE] Convert file-NAME charsets at the wire boundary:\n");
printf(" LOCAL is the charset of our file names, REMOTE is the\n");
printf(" remote side's charset (defaults to LOCAL). Names are\n");
printf(" converted before transmission and back on receipt; a\n");
printf(" name that cannot be represented in the target charset\n");
printf(" fails that transfer cleanly (rsync-compatible)\n");
printf(" --protocol=NUM Force the wire protocol version (must equal the current\n");
printf(" PROTOCOL_VERSION; FastSync cannot speak older/virtual\n");
printf(" wire formats)\n");
printf(" --write-batch=FILE Run the normal live transfer AND also emit a\n");
printf(" self-contained batch file of the whole source tree\n");
printf(" (implies the single-threaded transfer path)\n");
printf(" --only-write-batch=FILE\n");
printf(" Emit the batch file only (no destination, no server)\n");
printf(" --read-batch=FILE Apply the batch file to the destination (no source, no\n");
printf(" server); takes only the destination as an argument\n");
printf(" NOTE: the FastSync batch format is NOT interoperable with rsync's batch\n");
printf(" files (different container format); do not mix the two tools.\n");
printf(" --delete Delete files on receiver not in source\n");
printf(" (default timing: delete only after the whole\n");
printf(" transfer has succeeded)\n");
printf(" --delete-before Delete extras before the transfer starts\n");
printf(" (implies --delete)\n");
printf(" --delete-during Delete extras once the keep-set manifest is known,\n");
printf(" before the data is applied (implies --delete)\n");
printf(" --delete-during Delete a directory's extras as that directory is\n");
printf(" processed (implies --delete)\n");
printf(" --del Alias for --delete-during\n");
printf(" --delete-delay Delete extras only after a successful transfer\n");
printf(" (implies --delete)\n");
printf(" --delete-delay Record the extras during the scan but remove them\n");
printf(" only after a successful transfer (implies --delete)\n");
printf(" --delete-after Delete only after the whole transfer succeeded\n");
printf(" (the default --delete timing; implies --delete)\n");
printf(" --delete-excluded Also delete destination files that were excluded on\n");
printf(" the source (default protects them, matching rsync)\n");
printf(" --max-delete=NUM Never delete more than NUM destination entries per run;\n");
printf(" if the extras would exceed NUM, nothing is deleted and\n");
printf(" the run fails with a clear error (implies --delete only\n");
printf(" when used with it)\n");
printf(" --max-delete=NUM Delete at most NUM destination entries per run; if the\n");
printf(" extras exceed NUM, the rest are skipped and the run is\n");
printf(" reported as partial (exit 25, matching rsync). Only\n");
printf(" applies together with --delete\n");
printf(" --ignore-errors Continue (and still delete) when a source directory is\n");
printf(" unreadable during the scan, instead of aborting with no\n");
printf(" deletion\n");
@@ -52,9 +90,8 @@ void print_usage(void) {
printf(" entry's destination mirror receiver-side. Independent of\n");
printf(" --delete (it does not imply --delete; a non-empty directory\n");
printf(" mirror is removed only with --force or --delete)\n");
printf(" --prune-empty-dirs Do not transfer empty directory entries (--dirs mode);\n");
printf(" recursive transfers never send empty dirs. rsync's -m\n");
printf(" short form stays FastSync multithreading\n");
printf(" -m, --prune-empty-dirs Do not transfer empty directory entries (--dirs mode);\n");
printf(" recursive transfers never send empty dirs\n");
printf(" Note: each timing flag implies --delete. Combining a timing flag with\n");
printf(" --no-delete (in either order) is rejected as a config error.\n");
printf(" --ignore-existing Skip files that already exist on receiver\n");
@@ -70,20 +107,23 @@ void print_usage(void) {
printf(" parent directory is not itself listed\n");
printf(" --mkpath Create the destination root directory on the server when it\n");
printf(" does not exist yet\n");
printf(" --exclude <pattern> Exclude files matching pattern\n");
printf(" --include <pattern> Only include files matching pattern\n");
printf(" --exclude-from <file> Read exclude patterns from file\n");
printf(" --include-from <file> Read include patterns from file\n");
printf(" --exclude <pattern>, --exclude=<pattern> Exclude files matching pattern\n");
printf(" --include <pattern>, --include=<pattern> Only include files matching pattern\n");
printf(" --exclude-from <file>, --exclude-from=<file> Read exclude patterns from file\n");
printf(" --include-from <file>, --include-from=<file> Read include patterns from file\n");
printf(" --files-from <file> Read the source file list from FILE (paths relative to the "
"source root)\n");
printf(" -0, --from0 Entries in --files-from are NUL-delimited\n");
printf(" --filter=RULE rsync-style filter rule (+/- include/exclude; repeatable; the\n");
printf(" rsync -f short form conflicts with FastSync sendfile -f)\n");
printf(" -f, --filter=RULE rsync-style filter rule: exclude/- include/+ hide/H show/S\n");
printf(" protect/P risk/R merge/. dir-merge/: clear/! with modifiers\n");
printf(" (repeatable; --filter=RULE and -f RULE / -f=RULE both work)\n");
printf(" -C, --cvs-exclude Auto-ignore common CVS/SCM files (.git/, .svn/, *.o, *~, ...)\n");
printf(" -F Apply per-directory .rsync-filter files during the scan\n");
printf(" -F Apply per-directory .rsync-filter files; repeated -FF also\n");
printf(" excludes the .rsync-filter files themselves\n");
printf(" --max-size <n> Skip files larger than n bytes\n");
printf(" --min-size <n> Skip files smaller than n bytes\n");
printf(" --max-alloc <SIZE> Maximum single allocation (default: 1G)\n");
printf(" --max-alloc <SIZE> Maximum single allocation (default: 1G; 0 = no limit,\n");
printf(" matching rsync)\n");
printf(" --incremental Skip files unchanged since last transfer\n");
printf(" --size-only Skip incremental files matching in size, ignoring mtime\n");
printf(" -I, --ignore-times Transfer files even when size and mtime match\n");
@@ -98,77 +138,192 @@ void print_usage(void) {
printf(" --link-dest <dir> Like --copy-dest, but hard-links the unchanged file from DIR\n");
printf(" into the destination (repeatable; earlier DIRs win)\n");
printf(" --checksum-choice, --cc <alg> Whole-file checksum algorithm for --incremental/\n");
printf(" --checksum compares (xxh64/xxhash or md5; default xxh64 with\n");
printf(" seed 0). The seed comes from --checksum-seed\n");
printf(" --checksum-seed <num> Seed for the whole-file xxHash64 digest (and the delta\n");
printf(" block strong hash, low 32 bits); md5 ignores the seed. The\n");
printf(" digest algorithm and seed must match on sender and receiver\n");
printf(" --checksum compares. Accepted: xxh128 (default), xxh3, xxh64\n");
printf(" (aka xxhash), md5, md4, sha1, or none. A two-name\n");
printf(" 'transfer,pre-transfer' form is accepted like rsync; 'none' as\n");
printf(" the pre-transfer algorithm is rejected with --checksum\n");
printf(" --checksum-seed <num> Seed for the whole-file xxHash digest (and the delta\n");
printf(" block strong hash, low 32 bits); md5 ignores the seed. A seed\n");
printf(" of 0 (the default) is randomized per transfer, exactly like\n");
printf(" rsync, and the chosen seed is sent to the receiver\n");
printf(" --delta Delta transfer for changed files (requires --incremental)\n");
printf(" -W, --whole-file Transfer changed files without delta processing\n");
printf(" --no-whole-file rsync spelling that clears -W/--whole-file\n");
printf(" -y, --fuzzy Use a similar-named file already in the destination\n");
printf(" directory as the delta basis when the destination has no\n");
printf(" usable file at the exact path (saves bandwidth; implies\n");
printf(" --incremental and --delta; inert with --whole-file,\n");
printf(" --no-delta, or --no-incremental)\n");
printf(" --no-fuzzy Disable --fuzzy\n");
printf(" --delta-block <n> Delta block size in bytes (default: %d)\n",
DELTA_BLOCK_SIZE_DEFAULT);
printf(" -B <n>, --block-size <n>, --delta-block <n>\n");
printf(" Delta block size in bytes (default: %d)\n", DELTA_BLOCK_SIZE_DEFAULT);
printf(" --delta-max <n> Max file size for delta transfer (default: %llu)\n",
DELTA_MAX_FILE_SIZE);
printf(" -m Enable multithreading\n");
printf(" -s Enable chunk serialization\n");
printf(" --secluded-args Accept rsync compatibility option (no effect)\n");
printf(" -f Enable sendfile (TCP only, not with -c or -s)\n");
printf(" --compress-choice <alg> Compression algorithm (default: zstd)\n");
printf(" -j, --threads[=N] Enable the multithreaded scanner/loader/sender\n");
printf(" pipeline; N (1-%d) sets the parallel scanner worker\n",
MAX_SCANNER_THREADS);
printf(" count (bare -j/--threads uses the default)\n");
printf(" --chunk-serialization Enable chunk serialization (long form only)\n");
printf(" -s, --secluded-args Protect-args compatibility option (no effect; remote\n");
printf(" SSH argv is already built injection-safe)\n");
printf(" --sendfile Enable sendfile zero-copy (TCP only; long form only;\n");
printf(" -f is bound to --filter, not --sendfile)\n");
printf(" --compress-choice <alg> Compression algorithm: zstd (default), lz4, zlib,\n");
printf(" zlibx, none, or auto\n");
printf(" --zc <alg> Alias for --compress-choice\n");
printf(" -v, --verbose Enable debug logging\n");
printf(" -q, --quiet Suppress non-error output\n");
printf(" --debug=FLAGS Fine-grained debug logging (use --debug=help for flags)\n");
printf(" --info=FLAGS Fine-grained info: copy,misc,skip,stats,all,none\n");
printf(" none suppresses info even with --verbose\n");
printf(" -M, --preserve Preserve file metadata\n");
printf(" --info=FLAGS Fine-grained info: copy,name,misc,skip,stats,all,none\n");
printf(" (use --info=help for flags; none suppresses --verbose)\n");
printf(" --preserve Preserve permissions and times (= -pt; long form only)\n");
printf(" --no-perms Negate -p/--perms\n");
printf(" --no-times Negate -t/--times\n");
printf(" --no-owner Negate -o/--owner\n");
printf(" --no-group Negate -g/--group\n");
printf(" --no-preserve Disable metadata preservation (negates --preserve)\n");
printf(" -E, --executability Preserve executable permission bits\n");
printf(" --chmod <changes> Modify transferred permissions (rsync syntax)\n");
printf(" -U, --atimes Preserve access times\n");
printf(" -N, --crtimes Capture birth time; cannot be applied (documented\n");
printf(" divergence)\n");
printf(" -X, --xattrs Preserve user extended attributes (user.* only;\n");
printf(" privileged security.*/trusted.* namespaces are\n");
printf(" never captured or applied)\n");
printf(" -A, --acls Preserve POSIX ACLs (the system.posix_acl_* xattrs;\n");
printf(" setting an ACL the receiver is not permitted to\n");
printf(" set is warned and skipped, never fatal)\n");
printf(" --fake-super Store the source uid/gid/mode/mtime in a reserved\n");
printf(" user.fastsync.stat xattr on each written file and\n");
printf(" re-apply it (fd-relative) on a privileged run; the\n");
printf(" recording format diverges from rsync's user.rsync.%%stat%%\n");
printf(" --super Permit the receiver to attempt super-user activities\n");
printf(" (char/block device-node creation, --write-devices)\n");
printf(" within the confined receive root. Never elevates\n");
printf(" privileges and never bypasses confinement; ownership\n");
printf(" is still applied only with -o/--owner, -g/--group, or an\n");
printf(" explicit identity flag (--chown/--usermap/--groupmap/\n");
printf(" --copy-as); --numeric-ids only changes how ids map\n");
printf(" --no-super Forbid those super-user activities even when the\n");
printf(" receiver is running as root\n");
printf(
" --chmod <changes> Modify new/transferred permissions (rsync syntax; implies no -p)\n");
printf(" --numeric-ids Map uid/gid by id instead of by name (a modifier, not\n");
printf(" an ownership request: combine with -o/-g or a map)\n");
printf(" --usermap=MAP Map usernames when applying ownership: comma-separated\n");
printf(" FROM:TO rules, first match wins. FROM is a name (from\n");
printf(" the source), an id, an inclusive LOW-HIGH range, *\n");
printf(" (any id), or empty (ids with no name). TO is an id, *\n");
printf(" (current user), or a name resolved on the receiver.\n");
printf(" e.g. 0-99:nobody,*:normal (cannot mix with --chown)\n");
printf(" --groupmap=MAP Map group names when applying ownership (same syntax)\n");
printf(" --chown=USER:GROUP Override the ownership of transferred files. Forms:\n");
printf(" USER:GROUP, USER (owner only), :GROUP (group only); a\n");
printf(" value of * means the current/root user as appropriate.\n");
printf(" Names resolve on the source machine; @N for numerics.\n");
printf(" (Implies owner/group metadata; -M now means rsync's\n");
printf(" --remote-option.)\n");
printf(" --copy-as=USER[:GROUP] Force every written entry (files, dirs, symlinks\n");
printf(" and special nodes) to USER[:GROUP], resolved on the\n");
printf(" source machine like --chown. Requires a privileged\n");
printf(" (root) receiver and implies owner/group metadata; an\n");
printf(" unprivileged receiver refuses the transfer. Never\n");
printf(" switches process credentials (safe-subset; see\n");
printf(" RSYNC_COMPAT.md). A daemon refuses it.\n");
printf(" --chunk-size <n> Chunk size in bytes (default: %d)\n", DEFAULT_CHUNK_SIZE);
printf(" --source-dir <path> Source directory\n");
printf(" --dest-dir <path> Destination directory\n");
printf(" --save-to-disk Write received files to disk\n");
printf(" --server-host <ip> Server IP address (default: 127.0.0.1)\n");
printf(" --server-port <n> Server port (default: 8080)\n");
printf(" --port <n> Alias for --server-port\n");
printf(" --password-file <f> Authenticate a host::module/path daemon destination.\n");
printf(" FastSync-native SCRAM/PBKDF2 credential scheme (NOT\n");
printf(" rsync's --password-file): the file's first user:password\n");
printf(" line supplies the username and password; no password or\n");
printf(" reusable digest is sent (keep the file mode 0600)\n");
printf(" --no-motd Suppress display of the daemon's MOTD (the server\n");
printf(" still sends it; the client just does not show it)\n");
printf(" --bwlimit <KB/s> Bandwidth limit in kilobytes per second\n");
printf(" --tls Enable TLS encryption\n");
printf(" --cert <path> TLS certificate file (PEM)\n");
printf(" --key <path> TLS private key file (PEM)\n");
printf(" --ca <path> TLS CA certificate file (PEM)\n");
printf(" --timeout <sec> I/O timeout in seconds (default: 30)\n");
printf(" -T <sec> Alias for --timeout\n");
printf(" --contimeout <sec> Connection timeout in seconds (default: 10)\n");
printf(" --backup Backup existing files before overwriting\n");
printf(" --timeout <sec> I/O timeout in seconds (default: 0 = disabled, matching\n");
printf(" rsync). 0 disables it; --no-timeout is the same\n");
printf(" --contimeout <sec> Connection timeout in seconds (default: 60, matching\n");
printf(" rsync); 0 disables it (--no-contimeout)\n");
printf(" --stop-after=MINS Stop the transfer after MINS minutes (a positive\n");
printf(" integer); whatever was already transferred is kept\n");
printf(" --stop-at=TIME Stop at an absolute time. Accepts rsync's date form\n");
printf(" (Y-M-DTh:m, Y/M/DTh:m, abbreviable fields such as 12-31,\n");
printf(" 14:00, :59, 1) plus FastSync's HH:MM[:SS] and now+N[smhd]\n");
printf(" (a time already in the past stops the transfer\n");
printf(" immediately; client-only). An early stop skips the late\n");
printf(" --delete keep-set so it cannot delete source mirrors that\n");
printf(" were not yet scanned\n");
printf(" --address <ip> Bind the outgoing client socket to this source address\n");
printf(" -4, --ipv4 Force IPv4 for destination resolution\n");
printf(" -6, --ipv6 Force IPv6 for destination resolution\n");
printf(" --sockopts=OPTS Comma-separated OPT=VAL socket options applied before connect:\n");
printf(" TCP_NODELAY, SO_KEEPALIVE, SO_RCVBUF, SO_SNDBUF, SO_REUSEADDR\n");
printf(" -b, --backup Backup existing files before overwriting\n");
printf(" --backup-dir <dir> Directory for backups (requires --backup)\n");
printf(" --suffix <str> Backup suffix (default: ~)\n");
printf(" --stats Print transfer statistics at end\n");
printf(" -i, --itemize-changes Print an rsync-style per-file change line\n");
printf(" --out-format=FORMAT Output format for changed files (%%f %%n %%l %%b %%M %%%%)\n");
printf(" --out-format=FORMAT Output format (%%f %%n %%l %%b %%c %%C %%i %%M %%%%)\n");
printf(" --list-only List source files instead of transferring\n");
printf(" --log-file-format=FORMAT Per-file log line format (needs --log-file)\n");
printf(" -h, --human-readable Print byte sizes in human-readable form\n");
printf(" --max-depth <n> Maximum directory depth (0=unlimited)\n");
printf(" -x, --one-file-system Do not cross filesystem boundaries\n");
printf(" --log-file <path> Write log messages to file\n");
printf(" --log-file <path>, --log-file=<path> Write log messages to file\n");
printf(" --stderr=MODE Route logging to stderr: errors or all\n");
printf(" --partial Keep partial files on interrupted transfer\n");
printf(" --partial-dir <dir> Directory for partial files\n");
printf(" --temp-dir <dir> Scratch dir for temp files before atomic install\n");
printf(" -T, --temp-dir <dir> Scratch dir for temp files before atomic install.\n");
printf(" Confined to the receive root: a relative dir resolves below\n");
printf(" it and an absolute/traversal dir is rejected. The dir must\n");
printf(" already exist; a different filesystem falls back to a\n");
printf(" non-atomic copy instead of aborting\n");
printf(" --fastsync-server-path <path>\n");
printf(" Path to fastsync-server on remote (default: fastsync-server)\n");
printf(
" --old-args Disable safe SSH command argument quoting (legacy compatibility)\n");
printf(" --old-args Accepted for rsync CLI compatibility; no effect (the\n");
printf(" remote server path is always safely quoted now)\n");
printf(" -M, --remote-option=OPT Append OPT to the REMOTE server invocation. SSH\n");
printf(" transport ONLY (user@host:path): a daemon (host::module) or\n");
printf(" local TCP destination rejects it (no remote command line to\n");
printf(" append to). Repeatable; each value is single-quote-escaped on\n");
printf(" the remote command line; empty values and values with control\n");
printf(" characters are rejected; -M OPT, -M=OPT and\n");
printf(" --remote-option=OPT work\n");
printf(" --trust-sender RECEIVER-LOCAL policy: trust the remote sender's file list\n");
printf(" and skip the receiver's own up-front path-traversal/\n");
printf(" containment re-validation of the incoming list (fewer checks,\n");
printf(" faster, potentially unsafe). It is never sent to the peer, so\n");
printf(" for a push it must be enabled on the receiving SERVER\n");
printf(" (fastsync-server --trust-sender) or forwarded with\n");
printf(" -M--trust-sender; the client flag alone has no effect\n");
printf(" -l, --links Copy symlinks as symlinks\n");
printf(" --copy-links Transform symlinks into referent files\n");
printf(" --safe-links Skip symlinks that point outside transfer tree\n");
printf(" --copy-unsafe-links Only transform unsafe symlinks into referent files\n");
printf(" -L, --copy-links Transform symlinks into referent files\n");
printf(" --safe-links Skip symlinks whose target points outside the tree\n");
printf(" --copy-unsafe-links Copy unsafe symlinks (outside tree) as referent files\n");
printf(" -k, --copy-dirlinks Transform symlinks to directories into real dirs\n");
printf(" -K, --keep-dirlinks Keep an existing symlink-to-dir as that dir\n");
printf(" --munge-links Munge stored symlink targets (/rsyncd-munged/) on the receiver\n");
printf(" -H, --hard-links Preserve hard-link relationships across the transfer\n");
printf(" -S, --sparse Handle sparse files efficiently\n");
printf(
" -D Preserve device and special files (implies --devices --specials)\n");
printf(
" --devices Recreate device nodes on the destination (privileged; skipped when\n");
printf(" the receiver lacks CAP_MKNOD)\n");
printf(" --specials Recreate special files (FIFOs, sockets) on the destination\n");
printf(" --copy-devices Copy a source device's content as a regular file instead\n");
printf(" --write-devices Write received data into an existing destination device node\n");
printf(" --inplace Update files in-place (no temp+rename)\n");
printf(
" --preallocate Allocate destination file space up front (fail-fast on full disk)\n");
printf(" --append Resume a shorter destination by appending only its tail\n");
printf(" (prefix is not verified; requires --incremental)\n");
printf(" --append-verify Like --append, but verifies the retained prefix checksum\n");
@@ -176,7 +331,9 @@ void print_usage(void) {
printf(" --fsync Fsync every written file before publication\n");
printf(" --compress-level <n> Compression level (default: 5)\n");
printf(" --zl <n> Alias for --compress-level\n");
printf(" --skip-compress=LIST Skip compression for comma-separated suffixes\n");
printf(" --skip-compress=LIST Skip compression for suffixes in LIST (separated by\n");
printf(" '/' as in rsync, or ','); a leading dot is optional. The\n");
printf(" default is rsync 3.4.1's built-in skip-compress list\n");
printf(" --compress-threads <n> Compression worker threads (requires zstd threaded support)\n");
printf(" --no-OPTION Disable a supported boolean option\n");
printf(" --help Show this help\n");
@@ -184,7 +341,20 @@ void print_usage(void) {
}
void print_debug_usage(void) {
printf("Supported debug flags: IO,PROTO,PACK,UTIL,ALL,NONE\n");
printf("Emitting debug flags: IO,PROTO,PACK,UTIL,ALL,NONE\n");
printf("Also accepted for rsync CLI parity (silent): ACL,BACKUP,BIND,CHDIR,\n");
printf("CONNECT,CMD,DEL,DELTASUM,DUP,EXIT,FILTER,FLIST,FUZZY,GENR,HASH,HLINK,\n");
printf("ICONV,NSTR,OWN,RECV,SEND,TIME.\n");
printf("Flags may be comma-separated, for example: --debug=io,proto\n");
printf("Other rsync debug flags are unsupported and rejected.\n");
printf("An optional level suffix is accepted (e.g. --debug=io2); level 0\n");
printf("silences that item. Unknown names are rejected.\n");
}
void print_info_usage(void) {
printf("Emitting info flags: COPY,NAME,MISC,SKIP,STATS,ALL,NONE\n");
printf("Also accepted for rsync CLI parity (silent): BACKUP,DEL,FLIST,MOUNT,\n");
printf("NONREG,PROGRESS,REMOVE,SYMSAFE.\n");
printf("Flags may be comma-separated, for example: --info=name,stats\n");
printf("An optional level suffix is accepted (e.g. --info=stats2); level 0\n");
printf("silences that item. Unknown names are rejected.\n");
}
+1
View File
@@ -3,5 +3,6 @@
void print_usage(void);
void print_debug_usage(void);
void print_info_usage(void);
#endif
+365 -34
View File
@@ -1,15 +1,20 @@
#include "receiver.h"
#include "charset.h"
#include "chunk.h"
#include "config.h"
#include "delete_plan.h"
#include "delay_updates.h"
#include "file.h"
#include "file_receive.h"
#include "log.h"
#include "metadata.h"
#include "protocol.h"
#include "utils.h"
#include <stdlib.h>
#include <string.h>
#include <sys/stat.h>
#include <time.h>
bool receiver_outcomes_append(ReceiverOutcomes* outcomes, unsigned char code) {
if (!outcomes)
@@ -40,17 +45,49 @@ void receiver_outcomes_destroy(ReceiverOutcomes* outcomes) {
/* End-of-transfer success frame. When --remove-source-files was negotiated
each processed data file is acknowledged first (STATUS_NEXT = written,
STATUS_OK = skipped) so the sender never removes a source the receiver did
not actually store. The frame always ends with a plain STATUS_OK. */
bool receiver_send_final_success(int fd, const Config* config, const ReceiverOutcomes* outcomes) {
not actually store. The frame ends with `final_status` (STATUS_OK, or
STATUS_DELETE_LIMIT when a --max-delete commit was capped). */
bool receiver_send_final_success(int fd, const Config* config, const ReceiverOutcomes* outcomes,
Status final_status) {
if (!config->remove_source_files)
return send_status(fd, STATUS_OK);
return send_status(fd, final_status);
size_t count = outcomes ? outcomes->count : 0;
for (size_t i = 0; i < count; i++) {
Status per_file = outcomes->entries[i] == FILE_SAVE_WRITTEN ? STATUS_NEXT : STATUS_OK;
if (!send_status(fd, per_file))
return false;
}
return send_status(fd, STATUS_OK);
return send_status(fd, final_status);
}
bool receiver_send_stats_frame(int fd, const Config* config, const ReceiverStats* stats,
const struct ArrayList* would_delete) {
if (!config->report_stats)
return true;
ReceiverStats local;
memset(&local, 0, sizeof(local));
const ReceiverStats* out = stats ? stats : &local;
size_t count = would_delete ? (size_t)would_delete->size : 0;
if (count > (size_t)MAX_MANIFEST_ENTRIES)
count = MAX_MANIFEST_ENTRIES;
ReceiverStats record = *out;
record.would_delete_count = count;
if (!send_status(fd, STATUS_STATS) || !format_stats_send(fd, &record) ||
!send_int(fd, (int)count))
return false;
for (size_t i = 0; i < count; i++) {
const char* path = (const char*)would_delete->items[i];
if (!send_wire_str(fd, path ? path : ""))
return false;
}
return true;
}
/* Add a delete commit's tally to the sink's end-of-transfer wire counters (when
the sink reports them). Runs on the receiving thread, so no locking. */
static void receiver_tally_deleted(const ReceiverSink* sink, size_t deleted) {
if (sink && sink->stats && deleted > 0)
sink->stats->deleted_files += deleted;
}
static bool receiver_process_chunk(Chunk* chunk, const ReceiverSink* sink) {
@@ -72,13 +109,32 @@ static bool receiver_process_chunk(Chunk* chunk, const ReceiverSink* sink) {
return true;
}
/* P7 Wave D: read one STATUS_DIR_TIMES frame (a count followed by that many
* (path, metadata) directory entries) and route every entry through the regular
* store_file sink. A dir-time entry is RECORD-ONLY (file->dir_time_only): the
* sink accumulates its metadata for end-of-transfer application but creates
* nothing, so an empty/pruned source directory is never resurrected. A large
* tree arrives as repeated frames, each bounded by MAX_MANIFEST_ENTRIES; a
* malformed count or entry is a hard error. */
static bool receiver_process_dir_times(int fd, const Config* config, const ReceiverSink* sink) {
int count;
if (!receive_int(fd, &count) || count < 0 || count > MAX_MANIFEST_ENTRIES)
return false;
for (int i = 0; i < count; i++) {
File* dir = file_receive_dir_time(fd, config);
if (!dir || !sink->store_file(dir, sink->context))
return false;
}
return true;
}
static bool receiver_process_batch(Config* config, int file_descriptor) {
int count;
if (config->checksum || !receive_int(file_descriptor, &count) || count < 0 ||
count > MAX_MANIFEST_ENTRIES)
return false;
for (int i = 0; i < count; i++) {
char* check_path = receive_str(file_descriptor);
char* check_path = receive_wire_str(file_descriptor);
if (!check_path)
return false;
unsigned long long check_size;
@@ -92,7 +148,11 @@ static bool receiver_process_batch(Config* config, int file_descriptor) {
send_status(file_descriptor, STATUS_ERROR);
return false;
}
if (!utils_valid_batch_path(check_path)) {
/* --trust-sender: accept a ``..``/absolute check path (a trusted sender's
odd-but-legit entry) and defer containment to the secure stat below;
an empty path is still always rejected. */
if (check_path[0] == '\0' ||
(!file_get_trust_sender() && !utils_valid_batch_path(check_path))) {
free(check_path);
send_status(file_descriptor, STATUS_ERROR);
return false;
@@ -128,8 +188,95 @@ static bool receiver_process_batch(Config* config, int file_descriptor) {
return true;
}
/* ---- Anti-slowloris connection bounds ----
* A legitimate transfer either streams data frames continuously or, when it
* must pause, sends STATUS_KEEPALIVE so the peer sees the connection is alive.
* An attacker can therefore squat on a connection slot indefinitely by sending
* only keepalives under the per-message timeout. Two CLOCK_MONOTONIC bounds
* defeat that without ever punishing a real transfer:
*
* MAX_SESSION_IDLE_SEC (1 h): the longest a stream may make no forward
* progress. Data/status frames count as progress and refresh the timer;
* keepalives do not. One hour is far longer than any real pause between
* data frames, yet small enough to reap a slowloris well before the 24 h
* session cap.
*
* MAX_SESSION_WALL_SEC (24 h): an absolute ceiling on one connection's
* lifetime as defense-in-depth against a trickle of progress frames that
* resets the idle timer just below its limit. Larger than any plausible
* single transfer while still bounding resource occupancy.
*
* Both are wall-clock deltas, so the per-message poll timeout (60 s by default,
* or --timeout) can never fool them, and both the single-threaded and the -m
* receiver paths (receiver_process_pending) share the same logic. */
#define MAX_SESSION_IDLE_SEC 3600u
#define MAX_SESSION_WALL_SEC 86400u
static unsigned int g_max_session_idle_sec = MAX_SESSION_IDLE_SEC;
static unsigned int g_max_session_wall_sec = MAX_SESSION_WALL_SEC;
void receiver_set_time_limits(unsigned int idle_sec, unsigned int wall_sec) {
g_max_session_idle_sec = idle_sec;
g_max_session_wall_sec = wall_sec;
}
void receiver_reset_time_limits(void) {
g_max_session_idle_sec = MAX_SESSION_IDLE_SEC;
g_max_session_wall_sec = MAX_SESSION_WALL_SEC;
}
bool receiver_time_limit_exceeded(const struct timespec* session_start,
const struct timespec* last_progress,
const struct timespec* now) {
if (!session_start || !last_progress || !now)
return false;
if (now->tv_sec - session_start->tv_sec >= (time_t)g_max_session_wall_sec)
return true;
if (now->tv_sec - last_progress->tv_sec >= (time_t)g_max_session_idle_sec)
return true;
return false;
}
/* A frame proves forward progress only when it cannot be fabricated for free.
* KEEPALIVE/ABORT are pure liveness, and CHECK_BATCH/DIR_TIMES may carry zero
* entries, so a peer must not be able to hold a connection slot forever by
* merely emitting empty frames. */
static bool status_counts_as_progress(Status status) {
switch (status) {
case STATUS_KEEPALIVE:
case STATUS_ABORT:
case STATUS_CHECK_BATCH:
case STATUS_DIR_TIMES:
return false;
default:
return true;
}
}
/* Refresh the progress timestamp for a forward-moving frame and enforce the
* bounds above. Returns false when the connection must be dropped; the
* terminal STATUS_ERROR is sent only when the sink owns error reporting (the
* -m sink sets send_error=false so the main thread emits exactly one). */
static bool receiver_note_status(const struct timespec* session_start,
struct timespec* last_progress, Status status, int file_descriptor,
const ReceiverSink* sink) {
struct timespec now;
if (clock_gettime(CLOCK_MONOTONIC, &now) != 0)
now = *last_progress;
if (status_counts_as_progress(status))
*last_progress = now;
if (!receiver_time_limit_exceeded(session_start, last_progress, &now))
return true;
log_message(LOG_LEVEL_ERROR,
"Receive session exceeded its time bound (idle %us / total %us); aborting connection",
g_max_session_idle_sec, g_max_session_wall_sec);
if (!sink || sink->send_error)
send_status(file_descriptor, STATUS_ERROR);
return false;
}
int receiver_process(Config* config, int file_descriptor, const ReceiverSink* sink) {
return receiver_process_pending(config, file_descriptor, sink, NULL);
return receiver_process_pending(config, file_descriptor, sink, NULL, NULL);
}
/* Runs the whole receive loop. The delete manifest may legitimately arrive
@@ -143,18 +290,36 @@ int receiver_process(Config* config, int file_descriptor, const ReceiverSink* si
the whole transfer succeeded. See receiver_process_pending() for how the -m
receiver defers that commit until its disk writer has drained. */
int receiver_process_pending(Config* config, int file_descriptor, const ReceiverSink* sink,
DeleteManifest** pending_manifest) {
DeleteManifest** pending_manifest, DeletePlanSession** pending_plans) {
Status status;
if (!receive_status(file_descriptor, &status))
return -1;
/* Wall-clock (=CLOCK_MONOTONIC) anti-slowloris bookkeeping. session_start is
* fixed for the whole connection; last_progress is refreshed by every frame
* that is not a keepalive/abort. */
struct timespec session_start;
struct timespec last_progress;
clock_gettime(CLOCK_MONOTONIC, &session_start);
last_progress = session_start;
if (!receiver_note_status(&session_start, &last_progress, status, file_descriptor, sink))
return -1;
bool early_delete = config_delete_timing_early(config);
bool per_dir_delete = config_delete_timing_per_dir(config);
/* Parked keep-set for the late/commit timing. Every exit path below frees it
exactly once; the only exception is the successful FINISHED handoff, which
transfers ownership to *pending_manifest (used by the -m receiver). */
DeleteManifest* deferred_manifest = NULL;
/* Per-directory delete session for --delete-during/--delete-delay. During the
loop it applies plans inline (during) or snapshots their extras (delay); on
a successful FINISHED it is either committed here or handed to
*pending_plans so the -m caller commits after its disk writer drained. */
DeletePlanSession* plan_session = NULL;
bool delete_limit_noted = false;
while (status == STATUS_NEXT || status == STATUS_CHUNK || status == STATUS_CHECK ||
status == STATUS_KEEPALIVE || status == STATUS_ABORT || status == STATUS_CHECK_BATCH ||
status == STATUS_MKDIR || status == STATUS_MANIFEST) {
status == STATUS_MKDIR || status == STATUS_MANIFEST || status == STATUS_HARDLINK ||
status == STATUS_SYMLINK || status == STATUS_SPECIAL || status == STATUS_DIR_TIMES ||
status == STATUS_DELETE_PLAN) {
if (status == STATUS_KEEPALIVE) {
if (!send_status(file_descriptor, STATUS_KEEPALIVE))
goto fail;
@@ -165,10 +330,19 @@ int receiver_process_pending(Config* config, int file_descriptor, const Receiver
goto fail;
}
if (status == STATUS_CHECK) {
bool skipped;
File* file = receive_incremental_check(file_descriptor, config, &skipped);
if (!skipped && (!file || !sink->store_file(file, sink->context)))
bool skipped = false;
bool would_transfer = false;
File* file = receive_incremental_check_ex(file_descriptor, config, &skipped, &would_transfer);
if (config->dry_run) {
/* Server-contacting --dry-run: the reply has already been sent
(STATUS_OK = up to date, STATUS_DRY_RUN_TRANSFER = would transfer) and
nothing may be stored. Both flags false means a genuine protocol
error (STATUS_ERROR already sent or sent by receive_error below). */
if (!skipped && !would_transfer)
goto receive_error;
} else if (!skipped && (!file || !sink->store_file(file, sink->context))) {
goto receive_error;
}
} else if (status == STATUS_CHUNK) {
Chunk* chunk = receive_chunk_data(file_descriptor, config);
if (!chunk || !receiver_process_chunk(chunk, sink))
@@ -178,33 +352,69 @@ int receiver_process_pending(Config* config, int file_descriptor, const Receiver
goto fail;
goto next_status;
} else if (status == STATUS_MKDIR) {
File* dir = file_receive_directory(file_descriptor);
File* dir = file_receive_directory(file_descriptor, config);
if (!dir || !sink->store_file(dir, sink->context))
goto receive_error;
} else if (status == STATUS_DIR_TIMES) {
if (!receiver_process_dir_times(file_descriptor, config, sink))
goto receive_error;
} else if (status == STATUS_HARDLINK) {
File* file = file_receive_hardlink(file_descriptor);
if (!file || !sink->store_file(file, sink->context))
goto receive_error;
} else if (status == STATUS_SYMLINK) {
File* sym = file_receive_symlink(file_descriptor, config);
if (!sym || !sink->store_file(sym, sink->context))
goto receive_error;
} else if (status == STATUS_SPECIAL) {
File* file = file_receive_special(file_descriptor);
if (!file || !sink->store_file(file, sink->context))
goto receive_error;
} else if (status == STATUS_MANIFEST) {
DeleteManifest* manifest = receive_manifest_entries(file_descriptor);
if (!manifest)
goto fail; /* receive_manifest_entries already sent STATUS_ERROR */
if (early_delete) {
/* --delete-before / --delete-during: the manifest is authoritative the
moment it arrives, before any file data. Delete now and acknowledge
so the sender only starts streaming once the deletion committed (or
failed). This is the rsync delete-before/delete-during window: a
later transfer failure does not restore these deletions. */
bool deletion_ok = (config->use_delete || config->delete_missing_args)
? manifest_delete_all(config, manifest)
: true;
if (config->dry_run) {
/* Server-contacting --dry-run mutates nothing, so a keep-set manifest
is consumed and discarded. The early-delete mode still needs its ACK
so a sender blocked on the delete handshake is not left hanging.
When would-delete reporting is armed, enumerate (read-only) the
destination extras so the terminal STATUS_STATS frame can list them. */
if (config->use_delete && sink->would_delete) {
size_t count = 0;
if (!manifest_would_delete_list(config, manifest, sink->would_delete, &count))
log_message(LOG_LEVEL_WARNING, "dry-run: could not enumerate would-delete paths");
}
delete_manifest_free(manifest);
if (!deletion_ok) {
if (early_delete && !send_status(file_descriptor, STATUS_OK))
goto fail;
goto next_status;
}
if (early_delete) {
/* --delete-before: the whole-tree manifest is authoritative the moment
it arrives, before any file data. Delete now and acknowledge so the
sender only starts streaming once the deletion committed (or failed).
A later transfer failure does not restore these deletions. A
--max-delete-capped commit still succeeds and the transfer proceeds;
the terminal success frame reports the cap. */
size_t deleted = 0;
DeleteCommitResult deletion = (config->use_delete || config->delete_missing_args)
? manifest_delete_all_counted(config, manifest, &deleted)
: DELETE_COMMIT_OK;
receiver_tally_deleted(sink, deleted);
delete_manifest_free(manifest);
if (deletion == DELETE_COMMIT_ERROR) {
send_status(file_descriptor, STATUS_ERROR);
goto fail;
}
if (deletion == DELETE_COMMIT_LIMIT_REACHED && sink->note_delete_limit)
sink->note_delete_limit(sink->context);
if (!send_status(file_descriptor, STATUS_OK))
goto fail;
} else if (config->use_delete || config->delete_missing_args) {
/* Plain --delete / --delete-after / --delete-delay and the
--delete-missing-args exact-path deletions: hold the manifest and
commit it only after STATUS_FINISHED. */
/* Plain --delete / --delete-after and the --delete-missing-args
exact-path deletions: hold the manifest and commit it only after
STATUS_FINISHED. The per-directory modes never send this frame. */
if (deferred_manifest) {
log_message(LOG_LEVEL_ERROR, "Received a second delete manifest");
delete_manifest_free(deferred_manifest);
@@ -218,6 +428,23 @@ int receiver_process_pending(Config* config, int file_descriptor, const Receiver
delete_manifest_free(manifest);
}
goto next_status;
} else if (status == STATUS_DELETE_PLAN) {
if (!per_dir_delete) {
log_message(LOG_LEVEL_ERROR, "Received a per-directory delete plan without a per-dir "
"delete timing");
send_status(file_descriptor, STATUS_ERROR);
goto fail;
}
if (!plan_session)
plan_session = delete_plan_session_create(config);
if (!plan_session || delete_plan_session_receive(plan_session, config, file_descriptor) != 0)
goto fail;
if (delete_plan_session_limit_reached(plan_session) && !delete_limit_noted &&
sink->note_delete_limit) {
sink->note_delete_limit(sink->context);
delete_limit_noted = true;
}
goto next_status;
} else {
File* file = file_receive(config, file_descriptor);
if (!file) {
@@ -230,6 +457,8 @@ int receiver_process_pending(Config* config, int file_descriptor, const Receiver
next_status:
if (!receive_status(file_descriptor, &status))
goto receive_error;
if (!receiver_note_status(&session_start, &last_progress, status, file_descriptor, sink))
goto fail;
}
if (status != STATUS_FINISHED) {
log_message(LOG_LEVEL_ERROR, "Did not receive FINISHED Status");
@@ -249,13 +478,45 @@ int receiver_process_pending(Config* config, int file_descriptor, const Receiver
*pending_manifest = deferred_manifest;
deferred_manifest = NULL;
} else {
bool deletion_ok = manifest_delete_all(config, deferred_manifest);
size_t deleted = 0;
DeleteCommitResult deletion =
manifest_delete_all_counted(config, deferred_manifest, &deleted);
receiver_tally_deleted(sink, deleted);
delete_manifest_free(deferred_manifest);
deferred_manifest = NULL;
if (!deletion_ok) {
if (deletion == DELETE_COMMIT_ERROR) {
send_status(file_descriptor, STATUS_ERROR);
goto fail;
}
if (deletion == DELETE_COMMIT_LIMIT_REACHED && sink->note_delete_limit)
sink->note_delete_limit(sink->context);
}
}
/* Per-directory deletion: --delete-during already applied each plan inline, so
this only finishes the missing-args deletions; --delete-delay committed
nothing yet and applies its decompressed snapshot here. The -m receiver
hands the session to its caller instead, which commits after the disk
writer drained. */
if (plan_session) {
if (pending_plans) {
*pending_plans = plan_session;
plan_session = NULL;
} else if (config->dry_run) {
/* Central dry-run no-op: never commit a deletion for a -n run. */
delete_plan_session_destroy(plan_session);
plan_session = NULL;
} else {
DeleteCommitResult deletion = delete_plan_session_commit(plan_session, config);
bool limit = delete_plan_session_limit_reached(plan_session);
receiver_tally_deleted(sink, delete_plan_session_deleted(plan_session));
delete_plan_session_destroy(plan_session);
plan_session = NULL;
if (deletion == DELETE_COMMIT_ERROR) {
send_status(file_descriptor, STATUS_ERROR);
goto fail;
}
if (limit && !delete_limit_noted && sink->note_delete_limit)
sink->note_delete_limit(sink->context);
}
}
if (sink->send_success) {
@@ -270,11 +531,14 @@ int receiver_process_pending(Config* config, int file_descriptor, const Receiver
fail:
/* Failure exits that must not (or already did) report a STATUS_ERROR. The
parked keep-set is dropped: never commit a deletion for a failed stream. */
parked keep-set/session is dropped: never commit a deletion for a failed
stream. */
if (deferred_manifest) {
delete_manifest_free(deferred_manifest);
deferred_manifest = NULL;
}
if (plan_session)
delete_plan_session_destroy(plan_session);
return -1;
receive_error:
@@ -282,6 +546,8 @@ receive_error:
delete_manifest_free(deferred_manifest);
deferred_manifest = NULL;
}
if (plan_session)
delete_plan_session_destroy(plan_session);
if (sink->send_error)
send_status(file_descriptor, STATUS_ERROR);
return -1;
@@ -292,20 +558,54 @@ receive_error:
typedef struct {
Config* config;
ReceiverOutcomes outcomes;
/* P7 Wave D: directory metadata accumulated during the stream, applied only
after the whole transfer (and its delete/publication phases) has run so a
child write never clobbers a directory mtime. */
DirTimeList dir_times;
/* Set when a --max-delete commit was capped; the terminal frame then carries
STATUS_DELETE_LIMIT so the sender exits 25 like rsync. */
bool delete_limit_reached;
/* End-of-transfer wire counters (protocol 2.25.0) and the -n/--dry-run
--delete would-delete path list collected while processing the manifest. */
ReceiverStats stats;
ArrayList* would_delete;
} ReceiverSaveContext;
static bool receiver_save_file(File* file, void* context_pointer) {
ReceiverSaveContext* context = context_pointer;
FileSaveResult result = FILE_SAVE_ERROR;
if (!context->config->save_to_disk) {
if (context->config->dry_run) {
/* Defense in depth: a dry-run receiver mutates nothing even if a data
frame reaches the sink (the sender is not supposed to send one). */
result = FILE_SAVE_SKIPPED;
} else if (!context->config->save_to_disk) {
/* Nothing is stored; report the file as not-written so a
--remove-source-files sender keeps its source. */
result = FILE_SAVE_SKIPPED;
} else {
result = file_save_to_disk_full(context->config->receive_root_directory, file, context->config);
}
if (result != FILE_SAVE_ERROR && context->config->remove_source_files && !file->is_dir &&
!file->skip && !receiver_outcomes_append(&context->outcomes, (unsigned char)result)) {
/* Wire-stats tally: bytes reconstructed from the basis file (delta matches)
count as matched data in the end-of-transfer report. */
if (result != FILE_SAVE_ERROR && file->matched_bytes > 0)
context->stats.matched_data += file->matched_bytes;
/* A directory's metadata is deferred, never applied inline: collect it now
and apply it at the end. -O/--omit-dir-times and --preserve_perms/-times
are honored by dir_metadata_list_apply's caller (see
receiver_send_success_frame). */
if (result != FILE_SAVE_ERROR && file->is_dir && file->metadata &&
dir_metadata_should_capture(context->config) &&
!dir_time_list_add(&context->dir_times, file->path, file->metadata, file->xattrs)) {
file_destroy(file);
return false;
}
/* A dry-run receiver mutates nothing AND records no per-file outcomes: a
hostile dry-run client that streamed data frames anyway must not be able to
grow `outcomes` without bound (receiver_outcomes_append reallocs uncharged)
or force a per-frame ack. */
if (!context->config->dry_run && result != FILE_SAVE_ERROR &&
context->config->remove_source_files && !file->is_dir && !file->is_special && !file->skip &&
!receiver_outcomes_append(&context->outcomes, (unsigned char)result)) {
file_destroy(file);
return false;
}
@@ -313,8 +613,20 @@ static bool receiver_save_file(File* file, void* context_pointer) {
return result != FILE_SAVE_ERROR;
}
static void receiver_note_delete_limit(void* context_pointer) {
ReceiverSaveContext* context = context_pointer;
context->delete_limit_reached = true;
}
static bool receiver_send_success_frame(int fd, void* context_pointer) {
ReceiverSaveContext* context = context_pointer;
Status final_status = context->delete_limit_reached ? STATUS_DELETE_LIMIT : STATUS_OK;
if (!receiver_send_stats_frame(fd, context->config, &context->stats, context->would_delete))
return false;
/* Server-contacting --dry-run: nothing was staged or written, so there is
nothing to publish and no directory times to stamp. */
if (context->config->dry_run)
return receiver_send_final_success(fd, context->config, &context->outcomes, final_status);
/* --delay-updates: the whole protocol stream (including manifest/delete
handling, which ran inside receiver_process) has succeeded and every
staged file was fully written. Publish them atomically now, before the
@@ -326,15 +638,34 @@ static bool receiver_send_success_frame(int fd, void* context_pointer) {
return false;
}
}
return receiver_send_final_success(fd, context->config, &context->outcomes);
/* P7 Wave D: every child is now written and the delete / --delay-updates
phases have committed, so it is finally safe to stamp directory times.
This runs after the deferred deletion because receiver_process commits it
before calling this success frame. */
dir_metadata_list_apply(&context->dir_times, context->config->receive_root_directory,
context->config);
return receiver_send_final_success(fd, context->config, &context->outcomes, final_status);
}
int receiver_receive_files(Config* config, int file_descriptor) {
ReceiverSaveContext context = {.config = config, .outcomes = {0}};
ReceiverSink sink = {receiver_save_file, &context, true, true, receiver_send_success_frame};
dir_time_list_init(&context.dir_times);
context.would_delete = array_list_create(free);
if (!context.would_delete)
return -1;
ReceiverSink sink = {receiver_save_file,
&context,
true,
true,
receiver_send_success_frame,
receiver_note_delete_limit,
&context.stats,
context.would_delete};
int ret = receiver_process(config, file_descriptor, &sink);
if (ret != 0 && config->delay_updates && config->delay_context)
delay_updates_cleanup(config->delay_context);
receiver_outcomes_destroy(&context.outcomes);
dir_time_list_free(&context.dir_times);
array_list_delete(context.would_delete);
return ret;
}
+53 -5
View File
@@ -2,8 +2,12 @@
#define RECEIVER_H
#include "config.h"
#include "delete_plan.h"
#include "file.h"
#include "file_receive.h"
#include "protocol.h"
#include <stdbool.h>
#include <time.h>
typedef bool (*ReceiverFileSink)(File* file, void* context);
@@ -19,6 +23,12 @@ typedef struct {
typedef bool (*ReceiverSuccessFrame)(int fd, void* context);
/* Records that a --max-delete commit stopped with extras left over, so the
caller's terminal success frame can carry STATUS_DELETE_LIMIT instead of
STATUS_OK. The commit runs on the receiver thread, so the flag is stored in
the sink's own context rather than in a shared global. */
typedef void (*ReceiverNoteDeleteLimit)(void* context);
typedef struct {
ReceiverFileSink store_file;
void* context;
@@ -26,23 +36,61 @@ typedef struct {
bool send_success;
/* Emits the end-of-transfer success frame. When the sender requested
--remove-source-files this includes one per-file status per processed
data file followed by the final STATUS_OK; otherwise just STATUS_OK. */
data file followed by the final status; otherwise just the final status. */
ReceiverSuccessFrame send_success_frame;
/* Optional; may be NULL when the sink has no --max-delete handling. */
ReceiverNoteDeleteLimit note_delete_limit;
/* Optional end-of-transfer wire counters (protocol 2.25.0). When non-NULL
and the wire config carries report_stats, the success frame is preceded by
a STATUS_STATS record; `would_delete` (optional, receiver-owned strings)
carries the -n/--dry-run --delete path list. */
ReceiverStats* stats;
struct ArrayList* would_delete;
} ReceiverSink;
bool receiver_outcomes_append(ReceiverOutcomes* outcomes, unsigned char code);
void receiver_outcomes_destroy(ReceiverOutcomes* outcomes);
bool receiver_send_final_success(int fd, const Config* config, const ReceiverOutcomes* outcomes);
/* Send the terminal success frame. `final_status` is usually STATUS_OK, or
STATUS_DELETE_LIMIT when a --max-delete commit was capped. */
bool receiver_send_final_success(int fd, const Config* config, const ReceiverOutcomes* outcomes,
Status final_status);
/* Emit STATUS_STATS (a fixed ReceiverStats record plus, when `would_delete` is
non-NULL, a count and that many wire strings) when the wire config requested
report_stats. A no-op otherwise. */
bool receiver_send_stats_frame(int fd, const Config* config, const ReceiverStats* stats,
const struct ArrayList* would_delete);
int receiver_process(Config* config, int file_descriptor, const ReceiverSink* sink);
/* receiver_process with an escape hatch for the commit-style (late) deletion:
when `pending_manifest` is non-NULL the receiver does NOT delete at
STATUS_FINISHED itself; instead it stores the owned keep-set manifest there
(leaving *pending_manifest untouched on early modes/errors) so the caller can
commit the deletion only after its disk writer has fully drained. Pass NULL
to keep the default behaviour (delete before the success frame). */
commit the deletion only after its disk writer has fully drained. Likewise,
when `pending_plans` is non-NULL the --delete-delay per-directory session is
handed to the caller instead of being committed at STATUS_FINISHED. Pass NULL
for either to keep the default behaviour (delete before the success frame). */
int receiver_process_pending(Config* config, int file_descriptor, const ReceiverSink* sink,
DeleteManifest** pending_manifest);
DeleteManifest** pending_manifest, DeletePlanSession** pending_plans);
int receiver_receive_files(Config* config, int file_descriptor);
/* ---- Connection time bounds (anti-slowloris) ----
* receiver_process_pending() aborts a connection that makes no forward progress
* (only STATUS_KEEPALIVE/STATUS_ABORT frames) beyond a wall-clock idle limit,
* and enforces a hard cap on the whole session. Both are CLOCK_MONOTONIC
* deltas, independent of the per-message poll deadline, so a 60 s (or
* --timeout) receive window can never reset them. Defaults are deliberately
* generous (see MAX_SESSION_IDLE_SEC / MAX_SESSION_WALL_SEC in receiver.c). */
/* Test seam: override the idle/session wall-clock limits (0 = abort on the
* next status). Always restore with receiver_reset_time_limits(). */
void receiver_set_time_limits(unsigned int idle_sec, unsigned int wall_sec);
void receiver_reset_time_limits(void);
/* Pure predicate over explicit monotonic timestamps, exposed so the bound is
* unit-testable without sleeping. True when either the idle or the overall
* session limit has elapsed. */
bool receiver_time_limit_exceeded(const struct timespec* session_start,
const struct timespec* last_progress, const struct timespec* now);
#endif
+290
View File
@@ -0,0 +1,290 @@
#include "receiver_pipeline.h"
#include "log.h"
#include "protocol.h"
#include "queue.h"
#include "utils.h"
#include <stdlib.h>
#include <string.h>
#include <threads.h>
PipelineContextReceiver* pipeline_context_receiver_create(Config* config, Queue* queue,
int file_descriptor, SSL* ssl) {
PipelineContextReceiver* context = malloc(sizeof(PipelineContextReceiver));
if (context == NULL)
return NULL;
context->config = config;
context->queue = queue;
context->file_descriptor = file_descriptor;
context->ssl = ssl;
context->outcomes.entries = NULL;
context->outcomes.count = 0;
context->outcomes.capacity = 0;
dir_time_list_init(&context->dir_times);
protocol_session_init(&context->session, file_descriptor, file_descriptor);
protocol_session_set_ssl(&context->session, ssl);
context->receiver_done = false;
context->queued_bytes = 0;
context->max_queue_bytes = 0;
context->deferred_manifest = NULL;
context->deferred_plans = NULL;
context->delete_limit_reached = false;
memset(&context->stats, 0, sizeof(context->stats));
context->would_delete = NULL;
atomic_init(&context->cancelled, false);
int init = 0;
if (mtx_init(&context->mutex, mtx_plain) != thrd_success)
goto fail;
init++;
if (cnd_init(&context->condition_not_full) != thrd_success)
goto fail;
init++;
if (cnd_init(&context->condition_not_empty) != thrd_success)
goto fail;
// cppcheck-suppress unreadVariable
init++;
context->would_delete = array_list_create(free);
if (!context->would_delete)
goto fail;
return context;
fail:
log_perror("Error initializing synchronization objects");
if (init >= 3)
cnd_destroy(&context->condition_not_empty);
if (init >= 2)
cnd_destroy(&context->condition_not_full);
if (init >= 1)
mtx_destroy(&context->mutex);
free(context);
return NULL;
}
void pipeline_context_receiver_destroy(PipelineContextReceiver* context) {
config_delete(context->config);
if (context->deferred_manifest)
delete_manifest_free(context->deferred_manifest);
if (context->deferred_plans)
delete_plan_session_destroy(context->deferred_plans);
queue_destroy(context->queue);
receiver_outcomes_destroy(&context->outcomes);
dir_time_list_free(&context->dir_times);
if (context->would_delete)
array_list_delete(context->would_delete);
mtx_destroy(&context->mutex);
cnd_destroy(&context->condition_not_full);
cnd_destroy(&context->condition_not_empty);
free(context);
}
void pipeline_context_receiver_set_queue_byte_limit(PipelineContextReceiver* context,
size_t max_bytes) {
if (context == NULL)
return;
mtx_lock(&context->mutex);
context->max_queue_bytes = max_bytes;
context->queued_bytes = 0;
cnd_broadcast(&context->condition_not_full);
mtx_unlock(&context->mutex);
}
void pipeline_context_receiver_note_bytes_released(PipelineContextReceiver* context,
size_t released_bytes) {
if (context == NULL || context->max_queue_bytes == 0 || released_bytes == 0)
return;
mtx_lock(&context->mutex);
if (released_bytes >= context->queued_bytes)
context->queued_bytes = 0;
else
context->queued_bytes -= released_bytes;
cnd_signal(&context->condition_not_full);
mtx_unlock(&context->mutex);
}
bool pipeline_context_receiver_enqueue_file(PipelineContextReceiver* context, File* file) {
if (context == NULL || file == NULL)
return false;
size_t file_bytes = file->data ? file->data->size : 0;
mtx_lock(&context->mutex);
while (!atomic_load(&context->cancelled)) {
bool blocked_by_count = queue_is_full(context->queue);
bool blocked_by_budget = false;
if (context->max_queue_bytes > 0) {
size_t budget = context->max_queue_bytes;
size_t used = context->queued_bytes;
if (used >= budget) {
blocked_by_budget = true;
} else if (file_bytes > budget - used) {
/* A single payload larger than the whole budget (not possible with
the per-file receive cap) is only admitted to an empty pipeline so
the wait can never deadlock. */
blocked_by_budget = used != 0;
}
}
if (!blocked_by_count && !blocked_by_budget)
break;
cnd_wait(&context->condition_not_full, &context->mutex);
}
if (atomic_load(&context->cancelled)) {
mtx_unlock(&context->mutex);
file_destroy(file);
return false;
}
if (!queue_enqueue(context->queue, file)) {
mtx_unlock(&context->mutex);
file_destroy(file);
return false;
}
context->queued_bytes += file_bytes;
cnd_signal(&context->condition_not_empty);
mtx_unlock(&context->mutex);
return true;
}
static bool receiver_enqueue_file(File* file, void* context_pointer) {
PipelineContextReceiver* context = (PipelineContextReceiver*)context_pointer;
if (file && file->matched_bytes > 0) {
mtx_lock(&context->mutex);
context->stats.matched_data += file->matched_bytes;
mtx_unlock(&context->mutex);
}
return pipeline_context_receiver_enqueue_file(context, file);
}
/* Early delete modes (--delete-before/--delete-during) commit the manifest
inside receiver_process_pending on this thread; record a capped commit so
server.c's terminal frame can report STATUS_DELETE_LIMIT. The plain bool is
safe: receive_thread writes it before the main thread joins the thread. */
static void receiver_pipeline_note_delete_limit(void* context_pointer) {
PipelineContextReceiver* context = (PipelineContextReceiver*)context_pointer;
context->delete_limit_reached = true;
}
static void receiver_thread_fail(PipelineContextReceiver* context) {
mtx_lock(&context->mutex);
atomic_store(&context->cancelled, true);
context->receiver_done = true;
cnd_broadcast(&context->condition_not_empty);
cnd_broadcast(&context->condition_not_full);
mtx_unlock(&context->mutex);
}
int receive_thread(void* pipeline_context) {
PipelineContextReceiver* context = (PipelineContextReceiver*)pipeline_context;
protocol_session_bind(&context->session);
mtx_lock(&context->mutex);
int file_descriptor = context->file_descriptor;
const Config* config = context->config;
mtx_unlock(&context->mutex);
ReceiverSink sink = {receiver_enqueue_file,
context,
false,
false,
NULL,
receiver_pipeline_note_delete_limit,
&context->stats,
context->would_delete};
if (receiver_process_pending((Config*)config, file_descriptor, &sink, &context->deferred_manifest,
&context->deferred_plans) != 0) {
receiver_thread_fail(context);
protocol_session_unbind();
return thrd_error;
}
mtx_lock(&context->mutex);
context->receiver_done = true;
cnd_signal(&context->condition_not_empty);
mtx_unlock(&context->mutex);
protocol_session_unbind();
return thrd_success;
}
int write_thread(void* pipeline_context) {
PipelineContextReceiver* context = (PipelineContextReceiver*)pipeline_context;
protocol_session_bind(&context->session);
mtx_lock(&context->mutex);
bool save_to_disk = context->config->save_to_disk;
char* root_directory = str_dup(context->config->receive_root_directory);
mtx_unlock(&context->mutex);
if (save_to_disk && !root_directory) {
mtx_lock(&context->mutex);
atomic_store(&context->cancelled, true);
context->receiver_done = true;
cnd_broadcast(&context->condition_not_full);
cnd_broadcast(&context->condition_not_empty);
mtx_unlock(&context->mutex);
protocol_session_unbind();
return thrd_error;
}
while (true) {
File* file =
queue_dequeue_multithreaded(context->queue, &context->mutex, &context->condition_not_empty,
&context->condition_not_full, &context->receiver_done);
if (file == NULL) {
free(root_directory);
protocol_session_unbind();
return thrd_success;
}
size_t file_bytes = file->data ? file->data->size : 0;
FileSaveResult result = FILE_SAVE_SKIPPED;
/* Server-contacting --dry-run: never write. The receiver thread does not
enqueue anything on the dry-run path, but this keeps the writer thread
provably mutation-free if a data frame ever reached it. */
bool dry_run = context->config->dry_run;
if (save_to_disk && !dry_run) {
result = file_save_to_disk_full(root_directory, file, context->config);
if (result == FILE_SAVE_ERROR) {
file_destroy(file);
pipeline_context_receiver_note_bytes_released(context, file_bytes);
mtx_lock(&context->mutex);
atomic_store(&context->cancelled, true);
context->receiver_done = true;
cnd_broadcast(&context->condition_not_full);
cnd_broadcast(&context->condition_not_empty);
mtx_unlock(&context->mutex);
free(root_directory);
protocol_session_unbind();
return thrd_error;
}
}
/* P7 Wave D: a directory's times are never applied inline (a later child
write would clobber them); accumulate the metadata here and let the
caller apply it once every writer has drained. */
if (!dry_run && result != FILE_SAVE_ERROR && file->is_dir && file->metadata &&
dir_metadata_should_capture(context->config) &&
!dir_time_list_add(&context->dir_times, file->path, file->metadata, file->xattrs)) {
file_destroy(file);
pipeline_context_receiver_note_bytes_released(context, file_bytes);
mtx_lock(&context->mutex);
atomic_store(&context->cancelled, true);
context->receiver_done = true;
cnd_broadcast(&context->condition_not_full);
cnd_broadcast(&context->condition_not_empty);
mtx_unlock(&context->mutex);
free(root_directory);
protocol_session_unbind();
return thrd_error;
}
/* Record the per-file outcome so a --remove-source-files sender learns
which sources were actually written versus skipped on the receiver.
Explicit directory entries and recreated device/special nodes have no
source and are never acknowledged (mirrors receiver.c). */
if (!dry_run && context->config->remove_source_files && !file->is_dir && !file->is_special &&
!file->skip && !receiver_outcomes_append(&context->outcomes, (unsigned char)result)) {
file_destroy(file);
pipeline_context_receiver_note_bytes_released(context, file_bytes);
mtx_lock(&context->mutex);
atomic_store(&context->cancelled, true);
context->receiver_done = true;
cnd_broadcast(&context->condition_not_full);
cnd_broadcast(&context->condition_not_empty);
mtx_unlock(&context->mutex);
free(root_directory);
protocol_session_unbind();
return thrd_error;
}
file_destroy(file);
pipeline_context_receiver_note_bytes_released(context, file_bytes);
}
}
+84
View File
@@ -0,0 +1,84 @@
#ifndef RECEIVER_PIPELINE_H
#define RECEIVER_PIPELINE_H
#include <stdatomic.h>
#include <stdbool.h>
#include <threads.h>
#include "config.h"
#include "file.h"
#include "file_receive.h"
#include "protocol.h"
#include "queue.h"
#include "receiver.h"
#include <openssl/ssl.h>
typedef struct PipelineContextReceiver {
Queue* queue;
Config* config;
int file_descriptor;
SSL* ssl;
ProtocolSession session;
ReceiverOutcomes outcomes;
mtx_t mutex;
cnd_t condition_not_full;
cnd_t condition_not_empty;
bool receiver_done;
atomic_bool cancelled;
/* Aggregate payload bytes that have been received but not yet released by
the disk writer (queued or in the writer's hand). Guarded by `mutex`.
When `max_queue_bytes` is non-zero the receiver blocks before enqueuing
once this total would exceed it, so decompressed/copied file payloads
buffered ahead of a slow disk writer respect the per-connection memory
budget instead of growing without bound. */
size_t queued_bytes;
size_t max_queue_bytes;
/* Keep-set manifest for the commit-style (late) deletion
(--delete/--delete-after/--delete-delay). receive_thread parses the whole
protocol stream but hands the manifest here instead of deleting while the
disk writer may still be draining; the caller (server.c) commits the
deletion after both threads have joined, so no extra is removed unless the
transfer truly succeeded. NULL in the early delete modes (which delete at
the manifest). */
DeleteManifest* deferred_manifest;
/* Per-directory delete session for --delete-delay: receive_thread snapshots
each plan's extras as it arrives and hands the session here instead of
committing while the disk writer may still be draining; server.c commits it
after both threads joined. NULL for every other timing. */
DeletePlanSession* deferred_plans;
/* Set by server.c when the deferred delete commit hit the --max-delete
budget; the terminal success frame then carries STATUS_DELETE_LIMIT
(rsync exit 25) while the transfer itself still succeeds. */
bool delete_limit_reached;
/* P7 Wave D: directory metadata collected by write_thread from received
directory entries. Only write_thread mutates it (before it joins); the
caller (server.c) applies it after the delete/delay-updates phase. */
DirTimeList dir_times;
/* End-of-transfer wire counters (protocol 2.25.0). receive_thread accumulates
matched_data under `mutex`; server.c adds the delete-commit tallies after
both threads join and emits the STATUS_STATS frame. */
ReceiverStats stats;
/* -n/--dry-run --delete would-delete path list, collected by receive_thread
and reported in the STATUS_STATS frame. */
struct ArrayList* would_delete;
} PipelineContextReceiver;
PipelineContextReceiver* pipeline_context_receiver_create(Config* config, Queue* queue_receiver,
int file_descriptor, SSL* ssl);
void pipeline_context_receiver_destroy(PipelineContextReceiver* context);
/* Bound the bytes buffered ahead of the disk writer (see max_queue_bytes). */
void pipeline_context_receiver_set_queue_byte_limit(PipelineContextReceiver* context,
size_t max_bytes);
/* Blocking enqueue used by the receive pipeline sink. Blocks while the queue
is full by element count or when adding `file` would push queued_bytes over
the configured byte limit; waits until the disk writer releases bytes.
Takes ownership of `file` on success and destroys it on failure/cancel. */
bool pipeline_context_receiver_enqueue_file(PipelineContextReceiver* context, File* file);
/* Account for `released_bytes` of payload memory that has been freed by the
disk writer, unblocking a receiver that is waiting on the byte limit. */
void pipeline_context_receiver_note_bytes_released(PipelineContextReceiver* context,
size_t released_bytes);
int receive_thread(void* pipeline_context);
int write_thread(void* pipeline_context);
#endif
+1187 -169
View File
File diff suppressed because it is too large. Load diff
+311
View File
@@ -0,0 +1,311 @@
#include "server_cli.h"
#include "charset.h"
#include "credentials.h"
#include "utils.h"
#include <limits.h>
#include <stdarg.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/socket.h>
static void set_error(char* err, size_t err_size, const char* fmt, ...) {
if (!err || err_size == 0)
return;
va_list args;
va_start(args, fmt);
vsnprintf(err, err_size, fmt, args);
va_end(args);
}
void server_cli_options_default(ServerCliOptions* opts) {
if (!opts)
return;
memset(opts, 0, sizeof(*opts));
opts->destination_root = ".";
opts->port = 8080;
opts->bind_family = AF_UNSPEC;
}
static bool arg_is(const char* arg, const char* name) {
return strcmp(arg, name) == 0;
}
/* Match "--opt" against "--opt=value" / separate-value forms; on the "=" form
* *value receives the inline value. Returns true when the argument is the
* named option in either form. */
static bool arg_has_value(const char* arg, const char* name, const char** value) {
if (strcmp(arg, name) == 0)
return true; /* separate form; caller takes the next argv slot */
size_t name_len = strlen(name);
if (strncmp(arg, name, name_len) == 0 && arg[name_len] == '=') {
*value = arg + name_len + 1;
return true;
}
return false;
}
static int parse_port_arg(const char* value, int* port, char* err, size_t err_size) {
char* end;
long p = strtol(value, &end, 10);
if (*end != '\0' || p <= 0 || p > 65535) {
char* escaped = output_escape(value, false);
set_error(err, err_size, "invalid port '%s' (must be 1-65535)",
escaped ? escaped : "<allocation failed>");
free(escaped);
return -1;
}
*port = (int)p;
return 0;
}
int server_cli_parse(int argc, char* argv[], ServerCliOptions* opts, char* err, size_t err_size) {
if (err && err_size)
err[0] = '\0';
server_cli_options_default(opts);
for (int i = 1; i < argc; i++) {
const char* inline_value = NULL;
if (arg_is(argv[i], "--help")) {
opts->show_help = true;
return 1;
} else if (arg_is(argv[i], "--stdio")) {
opts->stdio_mode = true;
} else if (arg_is(argv[i], "--daemon")) {
opts->daemon_mode = true;
} else if (arg_is(argv[i], "--no-detach")) {
opts->no_detach = true;
} else if (arg_is(argv[i], "-v") || arg_is(argv[i], "--verbose")) {
opts->verbose = true;
} else if (arg_is(argv[i], "--tls")) {
opts->use_tls = true;
} else if (arg_is(argv[i], "--cert")) {
if (i + 1 >= argc) {
set_error(err, err_size, "missing argument for --cert");
return -1;
}
opts->tls_cert = argv[++i];
} else if (arg_is(argv[i], "--key")) {
if (i + 1 >= argc) {
set_error(err, err_size, "missing argument for --key");
return -1;
}
opts->tls_key = argv[++i];
} else if (arg_is(argv[i], "--ca")) {
if (i + 1 >= argc) {
set_error(err, err_size, "missing argument for --ca");
return -1;
}
opts->tls_ca = argv[++i];
} else if (arg_is(argv[i], "--client-cn")) {
if (i + 1 >= argc) {
set_error(err, err_size, "missing argument for --client-cn");
return -1;
}
opts->client_cn = argv[++i];
} else if (arg_is(argv[i], "--destination-root")) {
if (i + 1 >= argc) {
set_error(err, err_size, "missing argument for --destination-root");
return -1;
}
opts->destination_root = argv[++i];
opts->destination_root_set = true;
} else if (arg_has_value(argv[i], "--password-file", &inline_value)) {
if (!inline_value) {
if (i + 1 >= argc) {
set_error(err, err_size, "missing argument for --password-file");
return -1;
}
inline_value = argv[++i];
}
opts->password_file = inline_value;
} else if (arg_has_value(argv[i], "--early-input", &inline_value)) {
if (!inline_value) {
if (i + 1 >= argc) {
set_error(err, err_size, "missing argument for --early-input");
return -1;
}
inline_value = argv[++i];
}
opts->early_input_file = inline_value;
} else if (arg_has_value(argv[i], "--hash-credentials", &inline_value)) {
if (!inline_value) {
if (i + 1 >= argc) {
set_error(err, err_size, "missing argument for --hash-credentials");
return -1;
}
inline_value = argv[++i];
}
opts->hash_credentials_file = inline_value;
} else if (arg_has_value(argv[i], "--iterations", &inline_value)) {
if (!inline_value) {
if (i + 1 >= argc) {
set_error(err, err_size, "missing argument for --iterations");
return -1;
}
inline_value = argv[++i];
}
char* end = NULL;
long n = strtol(inline_value, &end, 10);
if (!end || *end != '\0' || n < (long)CREDENTIAL_MIN_ITERS ||
n > (long)CREDENTIAL_MAX_ITERS) {
set_error(err, err_size, "--iterations must be in [%u,%u], got '%s'", CREDENTIAL_MIN_ITERS,
CREDENTIAL_MAX_ITERS, inline_value);
return -1;
}
opts->hash_iterations = (uint32_t)n;
opts->hash_iterations_set = true;
} else if (arg_is(argv[i], "--address")) {
if (i + 1 >= argc) {
set_error(err, err_size, "missing argument for --address");
return -1;
}
opts->bind_address = argv[++i];
} else if (arg_is(argv[i], "-4") || arg_is(argv[i], "--ipv4")) {
if (opts->bind_family == AF_INET6) {
set_error(err, err_size, "--ipv4 and --ipv6 are mutually exclusive");
return -1;
}
opts->bind_family = AF_INET;
} else if (arg_is(argv[i], "-6") || arg_is(argv[i], "--ipv6")) {
if (opts->bind_family == AF_INET) {
set_error(err, err_size, "--ipv4 and --ipv6 are mutually exclusive");
return -1;
}
opts->bind_family = AF_INET6;
} else if (arg_is(argv[i], "--allow-delete")) {
opts->allow_delete = true;
} else if (arg_is(argv[i], "--trust-sender")) {
opts->trust_sender = true;
} else if (arg_is(argv[i], "--no-super")) {
opts->no_super = true;
} else if (arg_is(argv[i], "--allow-super")) {
opts->allow_super = true;
} else if (arg_is(argv[i], "--allow-unauthenticated")) {
opts->allow_unauthenticated = true;
} else if (arg_has_value(argv[i], "--iconv", &inline_value)) {
if (!inline_value) {
if (i + 1 >= argc) {
set_error(err, err_size, "missing argument for --iconv");
return -1;
}
inline_value = argv[++i];
}
opts->iconv_spec = inline_value;
} else if (arg_is(argv[i], "-p") || arg_has_value(argv[i], "--port", &inline_value)) {
if (inline_value) {
opts->port_set = true;
if (parse_port_arg(inline_value, &opts->port, err, err_size) != 0)
return -1;
} else {
if (i + 1 >= argc) {
set_error(err, err_size, "missing argument for %s", argv[i]);
return -1;
}
opts->port_set = true;
if (parse_port_arg(argv[++i], &opts->port, err, err_size) != 0)
return -1;
}
} else {
if (arg_has_value(argv[i], "--config", &inline_value)) {
if (!inline_value) {
if (i + 1 >= argc) {
set_error(err, err_size, "missing argument for --config");
return -1;
}
inline_value = argv[++i];
}
opts->config_path = inline_value;
} else if (arg_has_value(argv[i], "--dparam", &inline_value)) {
if (!inline_value) {
if (i + 1 >= argc) {
set_error(err, err_size, "missing argument for --dparam");
return -1;
}
inline_value = argv[++i];
}
const char** grown =
realloc((char**)opts->dparams, (size_t)(opts->dparam_count + 1) * sizeof(const char*));
if (!grown) {
set_error(err, err_size, "out of memory parsing --dparam");
return -1;
}
opts->dparams = grown;
opts->dparams[opts->dparam_count++] = inline_value;
} else if (argv[i][0] == '-') {
char* escaped = output_escape(argv[i], false);
set_error(err, err_size, "unknown option: %s", escaped ? escaped : "<allocation failed>");
free(escaped);
return -1;
} else {
set_error(err, err_size, "unexpected argument '%s'", argv[i]);
return -1;
}
}
}
/* Cross-mode validation. */
if (opts->stdio_mode && opts->daemon_mode) {
set_error(err, err_size, "--stdio and --daemon are mutually exclusive");
return -1;
}
if (opts->daemon_mode && opts->destination_root_set) {
set_error(err, err_size,
"--destination-root cannot be combined with --daemon (module paths "
"replace it)");
return -1;
}
if (!opts->daemon_mode &&
(opts->config_path != NULL || opts->dparam_count > 0 || opts->no_detach ||
opts->password_file != NULL || opts->early_input_file != NULL)) {
set_error(err, err_size,
"--config, --dparam, --no-detach, --password-file, and --early-input require "
"--daemon");
return -1;
}
if (opts->hash_credentials_file != NULL && (opts->daemon_mode || opts->stdio_mode)) {
set_error(err, err_size, "--hash-credentials cannot be combined with --daemon or --stdio");
return -1;
}
if (opts->allow_super && opts->no_super) {
set_error(err, err_size, "--allow-super and --no-super are mutually exclusive");
return -1;
}
if (opts->allow_super && opts->daemon_mode) {
set_error(err, err_size,
"--allow-super is for a locally-launched standalone TCP server; daemon modules opt "
"in per module with 'client owner = yes'");
return -1;
}
/* --stdio is the SSH transport: the remote server argv is composed by the
* CLIENT (directly and via --remote-option), so a client could otherwise pass
* --allow-super to a root --stdio receiver and defeat the C3 secure default.
* Never honor it there; the super mode stays forced OFF. An operator who
* must keep the historical permissive behavior over SSH has to launch the
* receiver through a forced command, not via client-composed argv. */
if (opts->allow_super && opts->stdio_mode) {
set_error(err, err_size,
"--allow-super is not accepted with --stdio (the remote argv is client-composed; "
"use a forced command if the default must hold)");
return -1;
}
if (opts->hash_iterations_set && opts->hash_credentials_file == NULL) {
set_error(err, err_size, "--iterations requires --hash-credentials");
return -1;
}
/* --iconv: reject a malformed CONVERT_SPEC or an unsupported charset name at
startup (a probe iconv_open is attempted). */
if (opts->iconv_spec != NULL && !charset_spec_valid(opts->iconv_spec)) {
set_error(err, err_size, "--iconv requires LOCAL[,REMOTE] charset names supported by iconv");
return -1;
}
return 0;
}
void server_cli_options_free(ServerCliOptions* opts) {
if (!opts)
return;
free((char**)opts->dparams);
opts->dparams = NULL;
opts->dparam_count = 0;
}
+79
View File
@@ -0,0 +1,79 @@
#ifndef SERVER_CLI_H
#define SERVER_CLI_H
#include <stdbool.h>
#include <stddef.h>
#include <stdint.h>
/* Parsed fastsync-server command line. All string members are borrowed
* pointers into the original argv (valid for the life of the argv array the
* caller passed to server_cli_parse); dparams points at the raw --dparam
* argument strings. No member owns heap memory. */
typedef struct ServerCliOptions {
bool stdio_mode; /* --stdio */
bool daemon_mode; /* --daemon */
bool no_detach; /* --no-detach */
bool verbose; /* -v / --verbose */
bool show_help; /* --help */
bool use_tls; /* --tls */
const char* tls_cert; /* --cert */
const char* tls_key; /* --key */
const char* tls_ca; /* --ca */
const char* client_cn; /* --client-cn */
bool destination_root_set; /* an explicit --destination-root was given */
const char* destination_root; /* --destination-root value ("." if unset) */
bool port_set; /* an explicit -p was given */
int port; /* -p value (default 8080 when unset) */
const char* config_path; /* --config value, or NULL */
const char* password_file; /* --password-file value, or NULL (daemon) */
const char* early_input_file; /* --early-input value, or NULL (daemon) */
/* --hash-credentials=FILE: read `user:password` lines from FILE and print
* new-format credential-store lines to stdout, then exit. Standalone mode
* (mutually exclusive with --daemon/--stdio). */
const char* hash_credentials_file;
bool hash_iterations_set; /* an explicit --iterations was given */
uint32_t hash_iterations; /* --iterations value (default CREDENTIAL_DEFAULT_ITERS) */
const char** dparams; /* raw --dparam override strings */
int dparam_count;
const char* bind_address; /* --address */
int bind_family; /* AF_UNSPEC / AF_INET / AF_INET6 */
bool allow_delete; /* --allow-delete */
bool trust_sender; /* --trust-sender */
bool allow_unauthenticated; /* --allow-unauthenticated */
/* --no-super: operator veto forcing SUPER_MODE_OFF for every connection, so
* the receiver never attempts super-user activities (ownership application,
* device-node creation) even when running as root. Applies to --stdio and
* --daemon alike; also makes the server refuse any client --copy-as. */
bool no_super; /* --no-super */
/* --allow-super: locally-launched standalone TCP listener opt-in that keeps
* the historical permissive behavior for a PRIVILEGED (root) receiver.
* Without it a root standalone server forces SUPER_MODE_OFF, so a client
* --devices / --write-devices / --super / ownership request cannot make it
* create device nodes, write raw devices, or apply client-chosen ownership.
* It is rejected for --stdio: that path's remote argv is composed by the
* client (directly and via --remote-option), so it must never opt a root
* receiver back into super mode. Non-root receivers are unaffected (the
* kernel refuses the confined attempts). The daemon path instead uses the
* per-module `client owner = yes` opt-in. */
bool allow_super; /* --allow-super */
/* --iconv=CONVERT_SPEC: the server's own LOCAL charset declaration. The
* client's full spec rides the wire config frame anyway; when the server is
* started with its own --iconv, its LOCAL half overrides the local charset
* the client assumed so the server converts received names to ITS charset.
* Borrowed pointer into argv (never owns heap). */
const char* iconv_spec; /* --iconv value, or NULL */
} ServerCliOptions;
/* Parse argc/argv into *opts. Zero-initialize *opts before calling (or use
* server_cli_options_default). Returns:
* 1 -- --help was requested (opts->show_help set; caller prints usage).
* 0 -- parsed successfully.
* -1 -- invalid arguments (err is filled with the reason).
*/
void server_cli_options_default(ServerCliOptions* opts);
int server_cli_parse(int argc, char* argv[], ServerCliOptions* opts, char* err, size_t err_size);
/* Release the only heap the parsed options own (the dparams pointer array; the
* strings it points at are borrowed from argv and are not freed). Safe to
* call on a zero-initialized/defaulted struct. */
void server_cli_options_free(ServerCliOptions* opts);
#endif
+3
View File
@@ -1,6 +1,7 @@
#include "log.h"
#include "array_list.h"
#include "protocol.h"
#include <limits.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
@@ -39,6 +40,8 @@ void array_list_delete(ArrayList* array_list) {
static bool array_list_extend(ArrayList* array_list) {
if (array_list == NULL)
return false;
if (array_list->capacity > INT_MAX / 2)
return false;
int new_capacity = array_list->capacity * 2;
if (new_capacity == 0)
new_capacity = INITIAL_ARRAY_SIZE;
+195
View File
@@ -0,0 +1,195 @@
#include "batch.h"
#include "data.h"
#include "file.h"
#include "file_receive.h"
#include "identity.h"
#include "log.h"
#include <errno.h>
#include <stdlib.h>
#include <string.h>
#include <unistd.h>
/* Serialization metadata mode for the batch stream, captured from the config at
* batch_write_header time. The header persists it into the file so a batch is
* self-describing about whether per-entry metadata was CAPTURED in the stream:
* batch_read_apply re-reads it from the file (not from the reading config) to
* decode the chunk records correctly. Which attributes are actually APPLIED,
* however, comes from the INVOKING process's per-attribute config (the
* FileAttrPolicy and the dir-metadata gate), so a batch written with -M is NOT
* automatically applied identically by an invoking process with a different
* -p/-t/-o/-g: --read-batch must be invoked with the same -p/-t/-o/-g as the
* write side (rsync requires the same options). The batch driver is a single
* sequential scan pass within one thread, so this module-level flag is safe. */
static bool batch_metadata_mode = false;
static bool write_all_bytes(int fd, const void* data, size_t size) {
const unsigned char* p = (const unsigned char*)data;
size_t done = 0;
while (done < size) {
ssize_t n = write(fd, p + done, size - done);
if (n < 0 && errno == EINTR)
continue;
if (n <= 0)
return false;
done += (size_t)n;
}
return true;
}
bool batch_write_header(int fd, const Config* config) {
if (fd < 0)
return false;
batch_metadata_mode = config != NULL && config->use_metadata;
if (!write_all_bytes(fd, BATCH_MAGIC, BATCH_MAGIC_LEN))
return false;
unsigned char version = BATCH_FORMAT_VERSION;
if (!write_all_bytes(fd, &version, 1))
return false;
unsigned char mode = batch_metadata_mode ? 1 : 0;
return write_all_bytes(fd, &mode, 1);
}
bool batch_write_chunk(int fd, Chunk* chunk) {
if (fd < 0 || chunk == NULL)
return false;
Data* serialized = chunk_serialize(chunk, batch_metadata_mode);
if (serialized == NULL)
return false;
bool ok = false;
unsigned long long length = (unsigned long long)serialized->size;
if (length > BATCH_MAX_RECORD) {
log_message(LOG_LEVEL_ERROR, "batch: record size %llu exceeds the %llu-byte cap", length,
(unsigned long long)BATCH_MAX_RECORD);
} else if (write_all_bytes(fd, &length, sizeof(length)) &&
(length == 0 || write_all_bytes(fd, serialized->data, (size_t)length))) {
ok = true;
}
data_destroy(serialized);
return ok;
}
/* Read exactly `size` bytes. Returns true on success. On reaching EOF, sets
* *clean_eof only when no bytes had been read yet (a clean boundary) and returns
* that value, so a truncated record (EOF mid-read) yields false. */
static bool read_exact(int fd, void* data, size_t size, bool* clean_eof) {
unsigned char* p = (unsigned char*)data;
size_t done = 0;
while (done < size) {
ssize_t n = read(fd, p + done, size - done);
if (n < 0 && errno == EINTR)
continue;
if (n == 0) {
if (clean_eof)
*clean_eof = done == 0;
return done == 0;
}
if (n < 0)
return false;
done += (size_t)n;
}
if (clean_eof)
*clean_eof = false;
return true;
}
int batch_read_apply(int fd, const Config* config, const char* dest_root) {
if (fd < 0 || dest_root == NULL || dest_root[0] == '\0')
return -1;
/* Directory metadata is deferred to the end of the apply (a child write would
* otherwise clobber its parent's mtime/mode). The batch header's single
* metadata bit only says whether metadata is present in the stream; which
* attributes are APPLIED comes from the invoking process's config, so
* --read-batch must be invoked with the same -p/-t/-o/-g as the write side
* (rsync requires the same options). The identity snapshot is activated so
* -o/-g and the explicit ownership flags can apply. */
DirTimeList dir_times;
dir_time_list_init(&dir_times);
int result = -1;
if (!identity_set_active(config)) {
log_message(LOG_LEVEL_ERROR, "batch: could not activate the identity policy");
goto done;
}
char magic[BATCH_MAGIC_LEN];
bool eof = false;
if (!read_exact(fd, magic, BATCH_MAGIC_LEN, &eof) || eof ||
memcmp(magic, BATCH_MAGIC, BATCH_MAGIC_LEN) != 0) {
log_message(LOG_LEVEL_ERROR, "batch: malformed header (bad magic)");
goto done;
}
unsigned char version;
if (!read_exact(fd, &version, 1, &eof) || eof || version != BATCH_FORMAT_VERSION) {
log_message(LOG_LEVEL_ERROR, "batch: malformed header (bad or missing format version)");
goto done;
}
unsigned char mode;
if (!read_exact(fd, &mode, 1, &eof) || eof || (mode != 0 && mode != 1)) {
log_message(LOG_LEVEL_ERROR, "batch: malformed header (bad metadata flag)");
goto done;
}
bool use_metadata = mode == 1;
while (1) {
unsigned long long length;
if (!read_exact(fd, &length, sizeof(length), &eof)) {
log_message(LOG_LEVEL_ERROR, "batch: truncated length prefix");
goto done;
}
if (eof)
break; /* clean end of stream */
if (length == 0 || length > BATCH_MAX_RECORD) {
log_message(LOG_LEVEL_ERROR, "batch: rejected record length %llu (valid range 1..%llu)",
length, (unsigned long long)BATCH_MAX_RECORD);
goto done;
}
char* record = (char*)malloc((size_t)length);
if (record == NULL) {
log_message(LOG_LEVEL_ERROR, "batch: could not allocate a %llu-byte record", length);
goto done;
}
if (!read_exact(fd, record, (size_t)length, &eof) || eof) {
log_message(LOG_LEVEL_ERROR, "batch: truncated chunk record");
free(record);
goto done;
}
Data* data = data_create(record, (size_t)length);
if (data == NULL)
goto done; /* data_create frees `record` on failure */
Chunk* chunk = chunk_deserialize(data, use_metadata);
data_destroy(data);
if (chunk == NULL) {
log_message(LOG_LEVEL_ERROR, "batch: rejected malformed chunk record");
goto done;
}
for (int i = 0; i < chunk->element_count; i++) {
File* file = chunk->items[i];
chunk->items[i] = NULL;
if (file == NULL)
continue;
FileSaveResult save = file_save_to_disk_full(dest_root, file, config);
/* Accumulate directory metadata (when it applies) before the File is
* destroyed; applied once the whole stream has been consumed. */
if (save != FILE_SAVE_ERROR && file->is_dir && file->metadata &&
dir_metadata_should_capture(config) &&
!dir_time_list_add(&dir_times, file->path, file->metadata, file->xattrs)) {
file_destroy(file);
chunk_destroy(chunk);
goto done;
}
file_destroy(file);
if (save == FILE_SAVE_ERROR) {
chunk_destroy(chunk);
goto done;
}
}
chunk_destroy(chunk);
}
dir_metadata_list_apply(&dir_times, dest_root, config);
result = 0;
done:
identity_clear_active();
dir_time_list_free(&dir_times);
return result;
}
+28
View File
@@ -0,0 +1,28 @@
#ifndef BATCH_H
#define BATCH_H
#include "chunk.h"
#include "config.h"
/* Phase 6 residual-batch codec. A residual batch is a self-contained
* single-file record of a whole source tree: a magic+format-version header
* followed by length-prefixed chunk blobs (each built with chunk_serialize),
* byte-identical by construction. The batch is a client-only driver feature:
* it never crosses the wire, so there is no PROTOCOL_VERSION bump and no server
* change. */
#define BATCH_MAGIC "FSTRESBATCH"
#define BATCH_MAGIC_LEN 11
#define BATCH_FORMAT_VERSION 1
/* Max size of a single length-prefixed record (a whole serialized chunk,
* which can span several files). A single source file near the 64 MB wire
* limit plus per-file headers can produce a record slightly over 64 MB, so a
* large file just under the wire cap may be refused by the batch writer; this
* is documented upstream and the failure is clean (the partial batch is
* unlinked), never a truncated/corrupt batch. */
#define BATCH_MAX_RECORD (64ULL * 1024 * 1024)
bool batch_write_header(int fd, const Config* config);
bool batch_write_chunk(int fd, Chunk* chunk);
int batch_read_apply(int fd, const Config* config, const char* dest_root);
#endif
+384
View File
@@ -0,0 +1,384 @@
#include "charset.h"
#include "log.h"
#include "protocol.h"
#include "utils.h"
#include <errno.h>
#include <iconv.h>
#include <stdlib.h>
#include <string.h>
typedef struct {
iconv_t cd;
} CharsetConversion;
/* Process-wide wire conversion descriptor (one direction per process: a client
* only sends, a server only receives). CONCURRENCY CONTRACT: iconv_t is not
* guaranteed thread-safe, so every conversion MUST run on a single thread at a
* time. This holds today -- on the client the conversions run on the sender
* thread (in the -m pipeline chunk_serialize/send happen on the sender thread
* only), on the server on the receive-loop thread; the descriptor is
* initialized on one thread before any transfer thread spawns and torn down
* (charset_wire_free) only after all threads have joined. Do not add a
* concurrent conversion path (e.g. parallel chunk serialization) without
* guarding access with a mutex. */
static CharsetConversion* g_wire_conv;
/* Grow *buf to double capacity, freeing it on failure. realloc preserves the
* already-written prefix, so the caller only tracks its write offset. */
static bool grow_charset_buffer(char** buf, size_t* cap) {
size_t new_cap = *cap * 2;
if (new_cap <= *cap) {
free(*buf);
*buf = NULL;
return false;
}
char* grown = realloc(*buf, new_cap);
if (!grown) {
free(*buf);
*buf = NULL;
return false;
}
*buf = grown;
*cap = new_cap;
return true;
}
/* Throw away any pending shift state so a subsequent conversion starts clean.
* The flush output is discarded; for the stateless single-byte/UTF charsets
* this feature targets it is a no-op. */
static void charset_conversion_reset(const CharsetConversion* conv) {
char scratch[64];
char* sp = scratch;
size_t sl = sizeof(scratch);
(void)iconv(conv->cd, NULL, NULL, &sp, &sl);
}
int charset_spec_parse(const char* spec, char** local_out, char** remote_out) {
if (!local_out || !remote_out)
return -1;
*local_out = NULL;
*remote_out = NULL;
if (!spec || spec[0] == '\0')
return -1;
char* dup = str_dup(spec);
if (!dup)
return -1;
char* comma = strchr(dup, ',');
if (comma) {
if (comma == dup || comma[1] == '\0') {
free(dup);
return -1;
}
*comma = '\0';
*local_out = str_dup(dup);
*remote_out = str_dup(comma + 1);
free(dup);
} else {
*local_out = str_dup(dup);
*remote_out = str_dup(dup);
free(dup);
}
if (!*local_out || !*remote_out) {
free(*local_out);
free(*remote_out);
*local_out = NULL;
*remote_out = NULL;
return -1;
}
return 0;
}
void* charset_conversion_open(const char* from_charset, const char* to_charset) {
if (!from_charset || !to_charset)
return NULL;
iconv_t cd = iconv_open(to_charset, from_charset);
if (cd == (iconv_t)-1)
return NULL;
CharsetConversion* conv = malloc(sizeof(CharsetConversion));
if (!conv) {
iconv_close(cd);
return NULL;
}
conv->cd = cd;
return conv;
}
void charset_conversion_close(void* conversion) {
if (!conversion)
return;
CharsetConversion* conv = (CharsetConversion*)conversion;
iconv_close(conv->cd);
free(conv);
}
/* Probe a single conversion direction: the from/to charsets both open AND a
* representative ASCII name converts to a byte string containing no embedded
* NUL (so a target charset like UTF-16 that emits NUL bytes for ordinary ASCII
* names is rejected up front -- such an output would be silently truncated by
* the C-string wire helpers). */
static bool direction_probe_valid(const char* from, const char* to) {
if (!from || !to)
return false;
void* conv = charset_conversion_open(from, to);
if (!conv)
return false;
bool ok = true;
char input = 'a';
char* in_ptr = &input;
size_t in_left = 1;
char out_buf[64];
char* out_ptr = out_buf;
size_t out_left = sizeof(out_buf);
if (iconv(((CharsetConversion*)conv)->cd, &in_ptr, &in_left, &out_ptr, &out_left) == (size_t)-1)
ok = false;
char flush_buf[64];
char* flush_ptr = flush_buf;
size_t flush_left = sizeof(flush_buf);
if (ok &&
iconv(((CharsetConversion*)conv)->cd, NULL, NULL, &flush_ptr, &flush_left) == (size_t)-1)
ok = false;
size_t produced = (size_t)(out_ptr - out_buf);
if (ok && produced > 0 && memchr(out_buf, '\0', produced) != NULL)
ok = false;
charset_conversion_close(conv);
return ok;
}
bool charset_pair_valid(const char* local, const char* remote) {
/* Both ends convert in opposite directions with the same two charsets, so a
* valid spec must open (and be NUL-free) in BOTH directions: the sender
* opens local->remote, the receiver opens remote->local. */
return direction_probe_valid(local, remote) && direction_probe_valid(remote, local);
}
bool charset_spec_valid(const char* spec) {
if (!spec)
return true;
char* local;
char* remote;
if (charset_spec_parse(spec, &local, &remote) != 0)
return false;
bool ok = charset_pair_valid(local, remote);
free(local);
free(remote);
return ok;
}
bool charset_spec_valid_direction(const char* from_charset, const char* to_charset) {
return direction_probe_valid(from_charset, to_charset);
}
/* The receiver's real conversion is wire(client REMOTE) -> server-local (the
* server's own --iconv LOCAL half, or the client's LOCAL half when the server
* has no --iconv). A dedicated pre-ack check so an impossible direction is
* rejected before the connection instead of refusing mid-transfer. */
bool charset_wire_receiver_spec_valid(const char* spec, const char* server_spec) {
if (!spec)
return true;
char* local;
char* remote;
if (charset_spec_parse(spec, &local, &remote) != 0)
return false;
const char* wire = remote;
const char* target_local = local;
char* server_local = NULL;
char* server_remote = NULL;
if (server_spec) {
if (charset_spec_parse(server_spec, &server_local, &server_remote) != 0) {
free(local);
free(remote);
return false;
}
target_local = server_local;
}
bool ok = charset_spec_valid_direction(wire, target_local);
free(server_local);
free(server_remote);
free(local);
free(remote);
return ok;
}
char* charset_convert(const void* conversion, const char* in, int* err_out) {
if (!conversion || !in)
return NULL;
const CharsetConversion* conv = (const CharsetConversion*)conversion;
size_t in_len = strlen(in);
size_t cap = in_len + 16;
char* out = malloc(cap);
if (!out)
return NULL;
size_t in_left = in_len;
char* in_ptr = (char*)in;
size_t out_used = 0;
while (in_left > 0) {
char* out_ptr = out + out_used;
size_t out_left = cap - out_used;
if (iconv(conv->cd, &in_ptr, &in_left, &out_ptr, &out_left) == (size_t)-1) {
if (errno != E2BIG) {
if (err_out)
*err_out = errno;
charset_conversion_reset(conv);
free(out);
return NULL;
}
/* Output exhausted but input remains. E2BIG does not roll the output
pointer back: the bytes iconv already emitted before the failure must
be preserved, so advance out_used before growing. */
out_used = (size_t)(out_ptr - out);
if (!grow_charset_buffer(&out, &cap))
return NULL;
continue;
}
out_used = (size_t)(out_ptr - out);
}
/* Flush any pending shift state (a no-op for the stateless single-byte and
UTF charsets this feature targets, but keeps the descriptor clean). */
for (;;) {
char* out_ptr = out + out_used;
size_t out_left = cap - out_used;
if (iconv(conv->cd, NULL, NULL, &out_ptr, &out_left) == (size_t)-1) {
if (errno != E2BIG) {
if (err_out)
*err_out = errno;
charset_conversion_reset(conv);
free(out);
return NULL;
}
out_used = (size_t)(out_ptr - out);
if (!grow_charset_buffer(&out, &cap))
return NULL;
continue;
}
out_used = (size_t)(out_ptr - out);
break;
}
/* A successful iconv call may legitimately consume the whole buffer (output
exactly fills cap), leaving no room for the terminator: guarantee headroom
before the final write. */
if (out_used >= cap && !grow_charset_buffer(&out, &cap))
return NULL;
/* Defense in depth: a target charset that emits embedded NUL bytes would
truncate at the first NUL in the C-string wire helpers; fail cleanly
(validation already rejects such charsets up front). */
if (memchr(out, '\0', out_used) != NULL) {
if (err_out)
*err_out = EILSEQ;
charset_conversion_reset(conv);
free(out);
return NULL;
}
out[out_used] = '\0';
return out;
}
bool charset_wire_init_sender(const char* spec) {
charset_wire_free();
if (!spec)
return true;
char* local;
char* remote;
if (charset_spec_parse(spec, &local, &remote) != 0)
return false;
void* conv = charset_conversion_open(local, remote);
free(local);
free(remote);
if (!conv)
return false;
g_wire_conv = (CharsetConversion*)conv;
return true;
}
bool charset_wire_init_receiver(const char* spec, const char* server_spec) {
charset_wire_free();
if (!spec)
return true;
char* local;
char* remote;
if (charset_spec_parse(spec, &local, &remote) != 0)
return false;
/* The wire charset is the client spec's REMOTE half; the local charset is
* the client spec's LOCAL half unless the server was itself started with
* --iconv naming a different local charset (the server halves above never
* travel, so the server's own flag is the only way its local charset can
* differ from what the client assumed). */
const char* wire = remote;
const char* target_local = local;
char* server_local = NULL;
char* server_remote = NULL;
if (server_spec) {
if (charset_spec_parse(server_spec, &server_local, &server_remote) != 0) {
free(local);
free(remote);
return false;
}
target_local = server_local;
}
void* conv = charset_conversion_open(wire, target_local);
free(server_local);
free(server_remote);
free(local);
free(remote);
if (!conv)
return false;
g_wire_conv = (CharsetConversion*)conv;
return true;
}
void charset_wire_free(void) {
if (g_wire_conv) {
charset_conversion_close(g_wire_conv);
g_wire_conv = NULL;
}
}
bool charset_wire_active(void) {
return g_wire_conv != NULL;
}
char* charset_wire_apply(const char* path) {
if (!g_wire_conv)
return str_dup(path);
return charset_convert(g_wire_conv, path, NULL);
}
static void charset_convert_failure_log(const char* path) {
char* escaped = output_escape(path, false);
log_message(LOG_LEVEL_ERROR, "--iconv: cannot convert file name '%s' to the target charset",
escaped ? escaped : "<unprintable>");
free(escaped);
}
bool send_wire_str(int file_descriptor, const char* local_path) {
if (!g_wire_conv)
return send_str(file_descriptor, local_path);
char* wire = charset_wire_apply(local_path);
if (!wire) {
charset_convert_failure_log(local_path);
return false;
}
bool ok = send_str(file_descriptor, wire);
free(wire);
return ok;
}
char* receive_wire_str(int file_descriptor) {
char* raw = receive_str(file_descriptor);
if (!raw)
return NULL;
if (!g_wire_conv)
return raw;
char* local = charset_convert(g_wire_conv, raw, NULL);
if (!local) {
charset_convert_failure_log(raw);
free(raw);
return NULL;
}
free(raw);
return local;
}
+85
View File
@@ -0,0 +1,85 @@
#ifndef CHARSET_H
#define CHARSET_H
#include <stdbool.h>
#include <stddef.h>
/* --iconv=CONVERT_SPEC file-name charset conversion (rsync compatibility).
*
* CONVERT_SPEC is "LOCAL[,REMOTE]": LOCAL is the charset of our own file
* names, REMOTE is the charset of the remote side's file names and defaults
* to LOCAL when the comma half is omitted. The sender converts every local
* path from LOCAL to REMOTE before it goes on the wire; the receiver converts
* every received path back from REMOTE to LOCAL. A NULL/disabled spec means
* identity with zero overhead (the common path never consults iconv).
*
* All helpers are friendly to the strict cold path: the wire conversion state
* is process-global (one direction per process -- a client only sends, a
* server only receives) and is initialized once, before any path is
* serialized, so conversion compiles to a single non-NULL check when disabled.
*/
/* Parse CONVERT_SPEC into malloc'd LOCAL and REMOTE charset names (caller
* frees both). REMOTE is a separate copy of LOCAL when no comma is present.
* Returns 0 on success, -1 on a malformed spec (empty halves / missing value /
* allocation failure); nothing is allocated on the -1 path. Both output
* pointers are REQUIRED (non-NULL). */
int charset_spec_parse(const char* spec, char** local_out, char** remote_out);
/* True when a CONVERT_SPEC is well-formed AND its charsets are usable for this
* feature: each pair opens in a probe iconv_open in BOTH directions (a sender
* converts local->remote, the receiver converts remote->local) and converting
* a representative ASCII name emits no embedded NUL byte (a UTF-16-style NUL
* emitter would be silently truncated by the C-string wire helpers). A typo'd
* charset name is therefore rejected at startup, not mid-run. NULL (iconv
* disabled) is always valid. */
bool charset_spec_valid(const char* spec);
/* Probe a concrete from->to conversion pair without keeping the descriptor:
* both charsets open AND a representative ASCII name converts with no embedded
* NUL. Used for direction-specific validation (e.g. the receiver's exact
* wire->local direction including a server-side charset override). */
bool charset_spec_valid_direction(const char* from_charset, const char* to_charset);
bool charset_pair_valid(const char* local, const char* remote);
/* One-shot conversion of a NUL-terminated input to a malloc'd NUL-terminated
* result, or NULL on failure. On failure *err_out (when non-NULL) receives
* the iconv errno (EILSEQ/EINVAL = the input is not representable in the
* target charset). The caller must free the result. */
char* charset_convert(const void* conversion, const char* in, int* err_out);
/* Open a conversion descriptor for direction from_charset -> to_charset.
* Returns NULL (errno = EINVAL) when a charset name is unsupported. Freed
* with charset_conversion_close. */
void* charset_conversion_open(const char* from_charset, const char* to_charset);
void charset_conversion_close(void* conversion);
/* Process-wide wire conversion. charset_wire_init_sender (client side) opens
* LOCAL->REMOTE; charset_wire_init_receiver (server side) opens
* wire(REMOTE)->server-local. server_spec is the server's own --iconv, whose
* LOCAL half may override the local charset the client assumed; NULL reuses
* the client spec's LOCAL half. Both return false on an unsupported spec.
* The state is freed with charset_wire_free. */
bool charset_wire_init_sender(const char* spec);
bool charset_wire_init_receiver(const char* spec, const char* server_spec);
void charset_wire_free(void);
bool charset_wire_active(void);
/* Pre-ack receiver-direction sanity (see charset_wire_init_receiver): true
* when the exact wire->server-local conversion the receiver will use (client
* spec's REMOTE half into the server's own LOCAL half, or the client's LOCAL
* half when the server has no --iconv) opens and produces NUL-free output. */
bool charset_wire_receiver_spec_valid(const char* spec, const char* server_spec);
/* Convert a path across the wire in the process direction. Returns a malloc'd
* string, or NULL when the name cannot be represented in the target charset. */
char* charset_wire_apply(const char* path);
/* Convenience wire string I/O: encode+send_str / receive_str+decode. Both
* return false/NULL (logging a clear --iconv error) on conversion failure, so
* an unconvertible path FAILS the transfer cleanly instead of silently sending
* a mangled name. */
bool send_wire_str(int file_descriptor, const char* local_path);
char* receive_wire_str(int file_descriptor);
#endif
+328 -17
View File
@@ -1,12 +1,170 @@
#include "checksum.h"
#include <fcntl.h>
#include <openssl/evp.h>
#include <string.h>
#include <strings.h>
#include <unistd.h>
/* delta.c owns the single XXH_IMPLEMENTATION that provides the xxHash symbols
* for the whole binary; this TU only needs the declarations. */
* for the whole binary; this TU only needs the declarations. The streaming
* state structs and XXH3_update are exposed only with XXH_STATIC_LINKING_ONLY. */
#define XXH_STATIC_LINKING_ONLY
#include <xxhash.h>
/* ---------------------------------------------------------------------------
* Self-contained MD4 (RFC 1320). OpenSSL's MD4 lives in the legacy provider
* and is not guaranteed present, so FastSync carries its own implementation to
* keep --checksum-choice=md4 working on every build.
* ------------------------------------------------------------------------- */
typedef struct {
uint32_t state[4];
uint64_t bit_count;
uint8_t buffer[64];
size_t buffer_len;
} Md4Ctx;
static uint32_t md4_rotl(uint32_t x, int n) {
return (x << n) | (x >> (32 - n));
}
static void md4_transform(uint32_t state[4], const uint8_t block[64]) {
uint32_t x[16];
for (int i = 0; i < 16; i++)
x[i] = (uint32_t)block[i * 4] | ((uint32_t)block[i * 4 + 1] << 8) |
((uint32_t)block[i * 4 + 2] << 16) | ((uint32_t)block[i * 4 + 3] << 24);
uint32_t a = state[0], b = state[1], c = state[2], d = state[3];
#define F(x, y, z) (((x) & (y)) | (~(x) & (z)))
#define G(x, y, z) (((x) & (y)) | ((x) & (z)) | ((y) & (z)))
#define H(x, y, z) ((x) ^ (y) ^ (z))
#define ROUND1(a, b, c, d, k, s) a = md4_rotl(a + F(b, c, d) + x[k], s)
#define ROUND2(a, b, c, d, k, s) a = md4_rotl(a + G(b, c, d) + x[k] + 0x5a827999u, s)
#define ROUND3(a, b, c, d, k, s) a = md4_rotl(a + H(b, c, d) + x[k] + 0x6ed9eba1u, s)
ROUND1(a, b, c, d, 0, 3);
ROUND1(d, a, b, c, 1, 7);
ROUND1(c, d, a, b, 2, 11);
ROUND1(b, c, d, a, 3, 19);
ROUND1(a, b, c, d, 4, 3);
ROUND1(d, a, b, c, 5, 7);
ROUND1(c, d, a, b, 6, 11);
ROUND1(b, c, d, a, 7, 19);
ROUND1(a, b, c, d, 8, 3);
ROUND1(d, a, b, c, 9, 7);
ROUND1(c, d, a, b, 10, 11);
ROUND1(b, c, d, a, 11, 19);
ROUND1(a, b, c, d, 12, 3);
ROUND1(d, a, b, c, 13, 7);
ROUND1(c, d, a, b, 14, 11);
ROUND1(b, c, d, a, 15, 19);
ROUND2(a, b, c, d, 0, 3);
ROUND2(d, a, b, c, 4, 5);
ROUND2(c, d, a, b, 8, 9);
ROUND2(b, c, d, a, 12, 13);
ROUND2(a, b, c, d, 1, 3);
ROUND2(d, a, b, c, 5, 5);
ROUND2(c, d, a, b, 9, 9);
ROUND2(b, c, d, a, 13, 13);
ROUND2(a, b, c, d, 2, 3);
ROUND2(d, a, b, c, 6, 5);
ROUND2(c, d, a, b, 10, 9);
ROUND2(b, c, d, a, 14, 13);
ROUND2(a, b, c, d, 3, 3);
ROUND2(d, a, b, c, 7, 5);
ROUND2(c, d, a, b, 11, 9);
ROUND2(b, c, d, a, 15, 13);
ROUND3(a, b, c, d, 0, 3);
ROUND3(d, a, b, c, 8, 9);
ROUND3(c, d, a, b, 4, 11);
ROUND3(b, c, d, a, 12, 15);
ROUND3(a, b, c, d, 2, 3);
ROUND3(d, a, b, c, 10, 9);
ROUND3(c, d, a, b, 6, 11);
ROUND3(b, c, d, a, 14, 15);
ROUND3(a, b, c, d, 1, 3);
ROUND3(d, a, b, c, 9, 9);
ROUND3(c, d, a, b, 5, 11);
ROUND3(b, c, d, a, 13, 15);
ROUND3(a, b, c, d, 3, 3);
ROUND3(d, a, b, c, 11, 9);
ROUND3(c, d, a, b, 7, 11);
ROUND3(b, c, d, a, 15, 15);
#undef F
#undef G
#undef H
#undef ROUND1
#undef ROUND2
#undef ROUND3
state[0] += a;
state[1] += b;
state[2] += c;
state[3] += d;
}
static void md4_init(Md4Ctx* ctx) {
ctx->state[0] = 0x67452301u;
ctx->state[1] = 0xefcdab89u;
ctx->state[2] = 0x98badcfeu;
ctx->state[3] = 0x10325476u;
ctx->bit_count = 0;
ctx->buffer_len = 0;
}
static void md4_update(Md4Ctx* ctx, const uint8_t* data, size_t len) {
ctx->bit_count += (uint64_t)len * 8;
while (len > 0) {
size_t space = sizeof(ctx->buffer) - ctx->buffer_len;
size_t take = len < space ? len : space;
memcpy(ctx->buffer + ctx->buffer_len, data, take);
ctx->buffer_len += take;
data += take;
len -= take;
if (ctx->buffer_len == sizeof(ctx->buffer)) {
md4_transform(ctx->state, ctx->buffer);
ctx->buffer_len = 0;
}
}
}
static void md4_final(Md4Ctx* ctx, uint8_t out[16]) {
uint64_t bit_count = ctx->bit_count;
uint8_t pad = 0x80;
md4_update(ctx, &pad, 1);
uint8_t zero = 0;
while (ctx->buffer_len != 56)
md4_update(ctx, &zero, 1);
uint8_t length_le[8];
for (int i = 0; i < 8; i++)
length_le[i] = (uint8_t)((bit_count >> (8 * i)) & 0xff);
md4_update(ctx, length_le, sizeof(length_le));
for (int i = 0; i < 4; i++) {
out[i * 4] = (uint8_t)(ctx->state[i] & 0xff);
out[i * 4 + 1] = (uint8_t)((ctx->state[i] >> 8) & 0xff);
out[i * 4 + 2] = (uint8_t)((ctx->state[i] >> 16) & 0xff);
out[i * 4 + 3] = (uint8_t)((ctx->state[i] >> 24) & 0xff);
}
}
/* One-shot EVP digest (md5/sha1). Returns false when OpenSSL refuses. */
static bool evp_digest(const EVP_MD* md, const void* data, size_t size, uint8_t* out,
size_t out_capacity, size_t* out_len) {
static const uint8_t empty = 0;
const void* input = data ? data : &empty;
unsigned int digest_len = 0;
if (EVP_Digest(input, size, out, &digest_len, md, NULL) != 1)
return false;
if (digest_len > out_capacity)
return false;
*out_len = digest_len;
return true;
}
bool checksum_digest(ChecksumAlgo algo, uint64_t seed, const void* data, size_t size, uint8_t* out,
size_t out_capacity, size_t* out_len) {
if (!out || !out_len || out_capacity < CHECKSUM_MAX_DIGEST_LEN)
@@ -14,39 +172,158 @@ bool checksum_digest(ChecksumAlgo algo, uint64_t seed, const void* data, size_t
if (data == NULL && size != 0)
return false;
if (algo == CHECKSUM_ALGO_XXH64) {
switch (algo) {
case CHECKSUM_ALGO_XXH64: {
uint64_t digest = XXH64(data, size, seed);
memcpy(out, &digest, sizeof(digest));
*out_len = sizeof(digest);
return true;
}
if (algo == CHECKSUM_ALGO_MD5) {
case CHECKSUM_ALGO_XXH3: {
uint64_t digest = XXH3_64bits_withSeed(data, size, seed);
memcpy(out, &digest, sizeof(digest));
*out_len = sizeof(digest);
return true;
}
case CHECKSUM_ALGO_XXH128: {
XXH128_hash_t digest = XXH3_128bits_withSeed(data, size, seed);
memcpy(out, &digest, sizeof(digest));
*out_len = sizeof(digest);
return true;
}
case CHECKSUM_ALGO_MD5:
/* md5 takes no seed; the caller's seed is deliberately ignored (documented
* in RSYNC_COMPAT.md). OpenSSL's one-shot EVP_Digest needs a non-NULL
* buffer even for an empty input, so map a NULL data + size==0 to an empty
* buffer. */
static const uint8_t empty = 0;
const void* input = data ? data : &empty;
unsigned int digest_len = 0;
if (EVP_Digest(input, size, out, &digest_len, EVP_md5(), NULL) != 1)
return false;
if (digest_len > out_capacity)
return false;
*out_len = digest_len;
* in RSYNC_COMPAT.md). */
return evp_digest(EVP_md5(), data, size, out, out_capacity, out_len);
case CHECKSUM_ALGO_MD4: {
Md4Ctx ctx;
md4_init(&ctx);
md4_update(&ctx, (const uint8_t*)data, size);
md4_final(&ctx, out);
*out_len = 16;
return true;
}
case CHECKSUM_ALGO_SHA1:
/* sha1 takes no seed; the caller's seed is deliberately ignored. */
return evp_digest(EVP_sha1(), data, size, out, out_capacity, out_len);
case CHECKSUM_ALGO_NONE:
/* No checksum requested: an empty digest is the successful result. */
*out_len = 0;
return true;
}
return false;
}
bool checksum_digest_file(ChecksumAlgo algo, uint64_t seed, const char* path, uint8_t* out,
size_t out_capacity, size_t* out_len) {
if (!path || !out || !out_len || out_capacity < CHECKSUM_MAX_DIGEST_LEN)
return false;
int fd = open(path, O_RDONLY | O_CLOEXEC);
if (fd < 0)
return false;
uint8_t buffer[64 * 1024];
bool ok = false;
if (algo == CHECKSUM_ALGO_MD5) {
EVP_MD_CTX* ctx = EVP_MD_CTX_new();
if (!ctx) {
close(fd);
return false;
}
unsigned int digest_len = 0;
if (EVP_DigestInit_ex(ctx, EVP_md5(), NULL) == 1) {
ok = true;
ssize_t got;
while ((got = read(fd, buffer, sizeof(buffer))) > 0) {
if (EVP_DigestUpdate(ctx, buffer, (size_t)got) != 1) {
ok = false;
break;
}
}
if (got < 0)
ok = false;
if (ok && EVP_DigestFinal_ex(ctx, out, &digest_len) == 1 && digest_len <= out_capacity)
*out_len = digest_len;
else
ok = false;
}
EVP_MD_CTX_free(ctx);
close(fd);
return ok;
}
XXH64_state_t xxh64;
XXH3_state_t* xxh3 = NULL;
if (algo == CHECKSUM_ALGO_XXH64) {
XXH64_reset(&xxh64, seed);
} else if (algo == CHECKSUM_ALGO_XXH3 || algo == CHECKSUM_ALGO_XXH128) {
xxh3 = XXH3_createState();
if (!xxh3) {
close(fd);
return false;
}
if (algo == CHECKSUM_ALGO_XXH3)
XXH3_64bits_reset_withSeed(xxh3, seed);
else
XXH3_128bits_reset_withSeed(xxh3, seed);
} else {
close(fd);
return false;
}
ok = true;
ssize_t got;
while ((got = read(fd, buffer, sizeof(buffer))) > 0) {
if (algo == CHECKSUM_ALGO_XXH64)
XXH64_update(&xxh64, buffer, (size_t)got);
else if (XXH3_64bits_update(xxh3, buffer, (size_t)got) == XXH_ERROR) {
ok = false;
break;
}
}
if (got < 0)
ok = false;
if (ok) {
if (algo == CHECKSUM_ALGO_XXH64) {
uint64_t digest = XXH64_digest(&xxh64);
memcpy(out, &digest, sizeof(digest));
*out_len = sizeof(digest);
} else if (algo == CHECKSUM_ALGO_XXH3) {
uint64_t digest = XXH3_64bits_digest(xxh3);
memcpy(out, &digest, sizeof(digest));
*out_len = sizeof(digest);
} else {
XXH128_hash_t digest = XXH3_128bits_digest(xxh3);
memcpy(out, &digest, sizeof(digest));
*out_len = sizeof(digest);
}
}
if (xxh3)
XXH3_freeState(xxh3);
close(fd);
return ok;
}
int checksum_algo_from_name(const char* name) {
if (!name)
return -1;
if (strcasecmp(name, "xxh64") == 0 || strcasecmp(name, "xxhash") == 0)
return (int)CHECKSUM_ALGO_XXH64;
if (strcasecmp(name, "xxh3") == 0)
return (int)CHECKSUM_ALGO_XXH3;
if (strcasecmp(name, "xxh128") == 0)
return (int)CHECKSUM_ALGO_XXH128;
if (strcasecmp(name, "md5") == 0)
return (int)CHECKSUM_ALGO_MD5;
if (strcasecmp(name, "md4") == 0)
return (int)CHECKSUM_ALGO_MD4;
if (strcasecmp(name, "sha1") == 0)
return (int)CHECKSUM_ALGO_SHA1;
if (strcasecmp(name, "none") == 0)
return (int)CHECKSUM_ALGO_NONE;
return -1;
}
@@ -54,22 +331,56 @@ const char* checksum_algo_name(ChecksumAlgo algo) {
switch (algo) {
case CHECKSUM_ALGO_XXH64:
return "xxh64";
case CHECKSUM_ALGO_XXH3:
return "xxh3";
case CHECKSUM_ALGO_XXH128:
return "xxh128";
case CHECKSUM_ALGO_MD5:
return "md5";
case CHECKSUM_ALGO_MD4:
return "md4";
case CHECKSUM_ALGO_SHA1:
return "sha1";
case CHECKSUM_ALGO_NONE:
return "none";
}
return "<unknown>";
}
bool checksum_algo_valid(int algo) {
return algo == (int)CHECKSUM_ALGO_XXH64 || algo == (int)CHECKSUM_ALGO_MD5;
return algo == (int)CHECKSUM_ALGO_XXH64 || algo == (int)CHECKSUM_ALGO_MD5 ||
algo == (int)CHECKSUM_ALGO_XXH3 || algo == (int)CHECKSUM_ALGO_XXH128 ||
algo == (int)CHECKSUM_ALGO_MD4 || algo == (int)CHECKSUM_ALGO_SHA1 ||
algo == (int)CHECKSUM_ALGO_NONE;
}
uint8_t checksum_digest_len(ChecksumAlgo algo) {
switch (algo) {
case CHECKSUM_ALGO_XXH64:
case CHECKSUM_ALGO_XXH3:
return 8;
case CHECKSUM_ALGO_XXH128:
case CHECKSUM_ALGO_MD5:
case CHECKSUM_ALGO_MD4:
return 16;
case CHECKSUM_ALGO_SHA1:
return 20;
case CHECKSUM_ALGO_NONE:
return 0;
}
return 0;
}
ChecksumAlgo checksum_negotiate_default(void) {
/* rsync 3.4.1 default preference order; every entry is compiled in, so this
* resolves to xxh128. */
static const ChecksumAlgo preference[] = {
CHECKSUM_ALGO_XXH128, CHECKSUM_ALGO_XXH3, CHECKSUM_ALGO_XXH64, CHECKSUM_ALGO_MD5,
CHECKSUM_ALGO_MD4, CHECKSUM_ALGO_SHA1, CHECKSUM_ALGO_NONE,
};
for (size_t i = 0; i < sizeof(preference) / sizeof(preference[0]); i++) {
if (checksum_algo_valid((int)preference[i]))
return preference[i];
}
return CHECKSUM_ALGO_XXH64;
}
+40 -10
View File
@@ -8,18 +8,35 @@
/* Whole-file content-digest algorithms selectable with --checksum-choice and
* seeded with --checksum-seed. The ids are the values actually placed on the
* wire (config frame), so they must be kept stable and validated on receive.
* CHECKSUM_ALGO_XXH64 == 0 is the default and is byte-for-byte what FastSync
* computed before these options existed (xxHash64 with seed 0). */
typedef enum { CHECKSUM_ALGO_XXH64 = 0, CHECKSUM_ALGO_MD5 = 1 } ChecksumAlgo;
* CHECKSUM_ALGO_XXH64 == 0 is the historical FastSync default and its numeric
* value is preserved. The full set mirrors the algorithms rsync 3.4.1 can be
* built with; every one of them is implemented here. */
typedef enum {
CHECKSUM_ALGO_XXH64 = 0,
CHECKSUM_ALGO_MD5 = 1,
CHECKSUM_ALGO_XXH3 = 2,
CHECKSUM_ALGO_XXH128 = 3,
CHECKSUM_ALGO_MD4 = 4,
CHECKSUM_ALGO_SHA1 = 5,
CHECKSUM_ALGO_NONE = 6
} ChecksumAlgo;
/* md5 digest is 16 bytes, the longest supported. */
#define CHECKSUM_MAX_DIGEST_LEN 16
/* FastSync's negotiated default (rsync 3.4.1 auto-negotiates xxh128 first).
* The wire default for Config->checksum_algo is this value. */
#define CHECKSUM_ALGO_DEFAULT CHECKSUM_ALGO_XXH128
/* sha1 digest is 20 bytes, the longest supported. */
#define CHECKSUM_MAX_DIGEST_LEN 20
/* Compute the whole-file digest of the first `size` bytes of `data`.
*
* - CHECKSUM_ALGO_XXH64: xxHash64(data, size, seed) (full 64-bit seed).
* - CHECKSUM_ALGO_MD5: md5(data, size) via OpenSSL EVP.
* md5 has no seed, so `seed` is ignored (documented).
* - CHECKSUM_ALGO_XXH3: XXH3_64bits_withSeed(data, size, seed).
* - CHECKSUM_ALGO_XXH128: XXH3_128bits_withSeed(data, size, seed).
* - CHECKSUM_ALGO_MD5: md5(data, size) via OpenSSL EVP (seed ignored).
* - CHECKSUM_ALGO_MD4: md4(data, size), self-contained RFC 1320 (seed ignored).
* - CHECKSUM_ALGO_SHA1: sha1(data, size) via OpenSSL EVP (seed ignored).
* - CHECKSUM_ALGO_NONE: no digest; *out_len is 0 and nothing is written.
* - `size == 0` hashes the empty input (plus its seed), not a NULL input.
*
* Writes up to `out_capacity` bytes into `out`, storing the digest length in
@@ -28,9 +45,16 @@ typedef enum { CHECKSUM_ALGO_XXH64 = 0, CHECKSUM_ALGO_MD5 = 1 } ChecksumAlgo;
bool checksum_digest(ChecksumAlgo algo, uint64_t seed, const void* data, size_t size, uint8_t* out,
size_t out_capacity, size_t* out_len);
/* Streaming whole-file digest: hash the contents of `path` without holding the
* whole file in memory. Same digest/capacity contract as checksum_digest.
* Returns false on open/read failure or an undersized buffer. */
bool checksum_digest_file(ChecksumAlgo algo, uint64_t seed, const char* path, uint8_t* out,
size_t out_capacity, size_t* out_len);
/* Resolve a --checksum-choice string (case-insensitive) to an algorithm id.
* Accepts "xxh64" and "xxhash" (both map to CHECKSUM_ALGO_XXH64, rsync's
* xxhash spelling) and "md5". Returns -1 for any unsupported name. */
* Accepts "xxh64"/"xxhash", "xxh3", "xxh128", "md5", "md4", "sha1", "none".
* "auto" is not an algorithm here; the caller resolves it to the negotiated
* default. Returns -1 for any unrecognized name. */
int checksum_algo_from_name(const char* name);
/* Canonical name of an algorithm (used in CLI error messages). */
@@ -39,7 +63,13 @@ const char* checksum_algo_name(ChecksumAlgo algo);
/* True when `algo` is a supported id (used by config receive validation). */
bool checksum_algo_valid(int algo);
/* Digest length in bytes for an algorithm (xxx64 = 8, md5 = 16). */
/* Digest length in bytes for an algorithm (xxh64/xxh3 = 8,
* md5/md4/xxh128 = 16, sha1 = 20, none = 0). */
uint8_t checksum_digest_len(ChecksumAlgo algo);
/* Pick the first algorithm from FastSync's compiled-in preference list that is
* supported on this build (rsync 3.4.1's `--version` order:
* xxh128 xxh3 xxh64 md5 md4 sha1 none). Used to resolve "auto". */
ChecksumAlgo checksum_negotiate_default(void);
#endif /* CHECKSUM_H */
+170 -77
View File
@@ -1,90 +1,183 @@
#include "chmod.h"
#include "file.h"
#include <stddef.h>
#include <string.h>
static bool parse_clause(mode_t* mode, const char* begin, const char* end) {
const char* p = begin;
unsigned who = 0;
while (p < end && strchr("ugoa", *p)) {
if (*p == 'a')
who = 7;
else
who |= *p == 'u' ? 1U : (*p == 'g' ? 2U : 4U);
p++;
}
if (who == 0)
who = 7;
if (p == end || (*p != '+' && *p != '-' && *p != '='))
return false;
char operation = *p++;
mode_t bits = 0;
while (p < end) {
mode_t bit;
switch (*p++) {
case 'r':
bit = 4;
break;
case 'w':
bit = 2;
break;
case 'x':
bit = 1;
break;
default:
return false;
}
bits |= bit;
}
for (unsigned class_index = 0; class_index < 3; class_index++) {
unsigned class_bit = 1U << class_index;
if (!(who & class_bit))
continue;
mode_t shift = (mode_t)((2U - class_index) * 3U);
mode_t mask = (mode_t)(7U << shift);
mode_t class_bits = (mode_t)(bits << shift);
if (operation == '+')
*mode |= class_bits;
else if (operation == '-')
*mode &= ~class_bits;
else
*mode = (*mode & ~mask) | class_bits;
}
return true;
}
/* rsync's --chmod parser (parse_chmod + tweak_mode). A single clause is
* applied as it is completed, so repeated clauses and repeated --chmod options
* (joined with commas by the CLI) accumulate exactly like rsync. The D/F
* selectors restrict a clause to directories/files; X adds execute only to
* directories or files that were already executable. */
#define CHMOD_BITS 07777
#define CHMOD_FLAG_X_KEEP (1U << 0)
#define CHMOD_FLAG_DIRS_ONLY (1U << 1)
#define CHMOD_FLAG_FILES_ONLY (1U << 2)
enum chmod_op { CHMOD_OP_ADD = 1, CHMOD_OP_SUB, CHMOD_OP_EQ, CHMOD_OP_SET };
enum chmod_state {
CHMOD_STATE_ERROR,
CHMOD_STATE_1ST_HALF,
CHMOD_STATE_2ND_HALF,
CHMOD_STATE_OCTAL
};
bool chmod_apply(mode_t mode, const char* spec, mode_t* result) {
if (!spec || !*spec || !result)
return false;
bool numeric = true;
size_t length = strlen(spec);
if (length > 4)
numeric = false;
for (size_t i = 0; i < length && numeric; i++)
numeric = spec[i] >= '0' && spec[i] <= '7';
if (numeric) {
if (length == 0 || length > 4)
return false;
mode_t parsed = 0;
for (size_t i = 0; i < length; i++)
parsed = (mode_t)((parsed << 3) | (spec[i] - '0'));
*result = parsed;
return true;
}
const mode_t nonperm = mode & ~(mode_t)CHMOD_BITS;
const bool initially_executable = (mode & 0111) != 0;
mode_t changed = mode;
const char* begin = spec;
while (*begin) {
const char* end = strchr(begin, ',');
if (!end)
end = begin + strlen(begin);
if (!parse_clause(&changed, begin, end))
return false;
if (*end == '\0')
int state = CHMOD_STATE_1ST_HALF;
unsigned where = 0;
int what = 0, op = 0, topbits = 0, topoct = 0, flags = 0;
const char* p = spec;
while (state != CHMOD_STATE_ERROR) {
if (*p == '\0' || *p == ',') {
int bits;
if (!op) {
state = CHMOD_STATE_ERROR;
break;
begin = end + 1;
if (!*begin)
return false;
}
*result = changed;
if (where)
bits = (int)(where * (unsigned)what);
else {
where = 0111;
bits = (int)((where * (unsigned)what) & ~(unsigned)file_process_umask());
}
int mode_and, mode_or;
switch (op) {
case CHMOD_OP_ADD:
mode_and = CHMOD_BITS;
mode_or = bits + topoct;
break;
case CHMOD_OP_SUB:
mode_and = CHMOD_BITS - bits - topoct;
mode_or = 0;
break;
case CHMOD_OP_EQ:
mode_and = CHMOD_BITS - (int)(where * 7U) - (topoct ? topbits : 0);
mode_or = bits + topoct;
break;
default:
mode_and = 0;
mode_or = bits;
break;
}
bool is_dir = S_ISDIR(nonperm);
if (!((flags & CHMOD_FLAG_DIRS_ONLY) && !is_dir) &&
!((flags & CHMOD_FLAG_FILES_ONLY) && is_dir)) {
changed &= (mode_t)mode_and;
if ((flags & CHMOD_FLAG_X_KEEP) && !initially_executable && !is_dir)
changed |= (mode_t)(mode_or & ~0111);
else
changed |= (mode_t)mode_or;
}
if (*p == '\0')
break;
p++;
state = CHMOD_STATE_1ST_HALF;
where = 0;
what = op = topoct = topbits = flags = 0;
continue;
}
switch (state) {
case CHMOD_STATE_1ST_HALF:
switch (*p) {
case 'D':
if (flags & CHMOD_FLAG_FILES_ONLY) {
state = CHMOD_STATE_ERROR;
break;
}
flags |= CHMOD_FLAG_DIRS_ONLY;
break;
case 'F':
if (flags & CHMOD_FLAG_DIRS_ONLY) {
state = CHMOD_STATE_ERROR;
break;
}
flags |= CHMOD_FLAG_FILES_ONLY;
break;
case 'u':
where |= 0100;
topbits |= 04000;
break;
case 'g':
where |= 0010;
topbits |= 02000;
break;
case 'o':
where |= 0001;
break;
case 'a':
where |= 0111;
break;
case '+':
op = CHMOD_OP_ADD;
state = CHMOD_STATE_2ND_HALF;
break;
case '-':
op = CHMOD_OP_SUB;
state = CHMOD_STATE_2ND_HALF;
break;
case '=':
op = CHMOD_OP_EQ;
state = CHMOD_STATE_2ND_HALF;
break;
default:
if (*p >= '0' && *p <= '7' && !where) {
op = CHMOD_OP_SET;
state = CHMOD_STATE_OCTAL;
where = 1;
what = *p - '0';
} else {
state = CHMOD_STATE_ERROR;
}
break;
}
break;
case CHMOD_STATE_2ND_HALF:
switch (*p) {
case 'r':
what |= 4;
break;
case 'w':
what |= 2;
break;
case 'X':
flags |= CHMOD_FLAG_X_KEEP;
/* fall through */
case 'x':
what |= 1;
break;
case 's':
if (topbits)
topoct |= topbits;
else
topoct = 04000;
break;
case 't':
topoct |= 01000;
break;
default:
state = CHMOD_STATE_ERROR;
break;
}
break;
default:
if (*p >= '0' && *p <= '7') {
what = what * 8 + (*p - '0');
if (what > CHMOD_BITS)
state = CHMOD_STATE_ERROR;
} else {
state = CHMOD_STATE_ERROR;
}
break;
}
p++;
}
if (state == CHMOD_STATE_ERROR)
return false;
*result = (changed & (mode_t)CHMOD_BITS) | nonperm;
return true;
}
+4 -1
View File
@@ -4,7 +4,10 @@
#include <stdbool.h>
#include <sys/stat.h>
/* Apply the supported rsync --chmod syntax to a permission mode. */
/* Apply rsync's --chmod syntax to a permission mode, including the D/F/X
* selectors and the s/t special bits. `mode` should carry the file type bits
* (S_IFDIR/S_IFREG) so D/F/X can be evaluated; the type bits are preserved in
* `result`. A spec may contain comma-separated clauses, which accumulate. */
bool chmod_apply(mode_t mode, const char* spec, mode_t* result);
#endif
+246 -78
View File
@@ -1,11 +1,13 @@
#include <stddef.h>
#include <stdint.h>
#include <limits.h>
#include <stdatomic.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include "array_list.h"
#include "charset.h"
#include "chunk.h"
#include "compression.h"
#include "data.h"
@@ -19,6 +21,33 @@
#define MAX_FILE_DATA_SIZE (64ULL * 1024 * 1024)
#define MAX_FILES_PER_CHUNK 65536U
/* Reserve `charge` against `session`'s connection budget. This mirrors the
static protocol_reserve_memory() in protocol.c: the receive-side call sites
only have the Data.owner pointer (a ProtocolSession*), and protocol.c is out
of scope for this fix, so the same atomic CAS accounting is reproduced here.
The matching release always goes through data_destroy()'s Data.owner path. */
static bool chunk_session_reserve(ProtocolSession* session, size_t charge) {
unsigned long long allocated = atomic_load(&session->total_allocated_bytes);
while (true) {
if (allocated > MAX_CONNECTION_MEMORY ||
(unsigned long long)charge > MAX_CONNECTION_MEMORY - allocated)
return false;
if (atomic_compare_exchange_weak(&session->total_allocated_bytes, &allocated,
allocated + (unsigned long long)charge))
return true;
}
}
bool data_charge_session(Data* data, ProtocolSession* session, size_t charge) {
if (!data || charge == 0 || session == NULL)
return true;
if (!chunk_session_reserve(session, charge))
return false;
data->owner = session;
data->protocol_charge = charge;
return true;
}
Chunk* chunk_create(File** items, int element_count) {
if (element_count < 0 || (element_count > 0 && items == NULL))
return NULL;
@@ -63,9 +92,22 @@ void chunk_destroy(void* item) {
free(chunk);
}
/* --iconv: a chunk blob carries wire-charset path/target bytes. Encode the
* sender-side path (a no-op copy when iconv is disabled) so the blob is in the
* same charset as every other wire string. */
static char* chunk_encode_wire(const char* path) {
if (!charset_wire_active())
return str_dup(path);
return charset_wire_apply(path);
}
static unsigned long long per_file_serialize_size(File* file, bool use_metadata) {
unsigned long long size = sizeof(size_t);
size_t path_len = strlen(file_wire_path(file));
char* wire_path = chunk_encode_wire(file_wire_path(file));
if (!wire_path)
return 0;
size_t path_len = strlen(wire_path);
free(wire_path);
unsigned long long metadata_size =
use_metadata ? sizeof(int) + (file->metadata ? FILE_METADATA_WIRE_SIZE : 0) : 0;
if ((unsigned long long)path_len > ULLONG_MAX - size)
@@ -74,16 +116,39 @@ static unsigned long long per_file_serialize_size(File* file, bool use_metadata)
if (metadata_size > ULLONG_MAX - size)
return 0;
size += metadata_size;
/* Entry type marker: 0 = regular file, 1 = explicit directory entry. */
/* Entry type marker: 0 = regular file, 1 = explicit directory entry,
2 = symlink entry (carries its target string), 3 = special/device node
(recreated by the receiver). */
if (sizeof(int) > ULLONG_MAX - size)
return 0;
size += sizeof(int);
/* A special node also carries its rdev major/minor. */
if (file->is_special) {
if (2 * sizeof(int32_t) > ULLONG_MAX - size)
return 0;
size += 2 * sizeof(int32_t);
}
if (sizeof(size_t) > ULLONG_MAX - size)
return 0;
size += sizeof(size_t);
if ((unsigned long long)file->data->size > ULLONG_MAX - size)
return 0;
return size + file->data->size;
size += file->data->size;
/* Symlink entries append the target string (length-prefixed). */
if (file->is_symlink) {
char* wire_target = chunk_encode_wire(file->symlink_target ? file->symlink_target : "");
if (!wire_target)
return 0;
size_t target_len = strlen(wire_target);
free(wire_target);
if (sizeof(size_t) > ULLONG_MAX - size)
return 0;
size += sizeof(size_t);
if ((unsigned long long)target_len > ULLONG_MAX - size)
return 0;
size += target_len;
}
return size;
}
Data* chunk_serialize(Chunk* chunk, bool use_metadata) {
@@ -109,17 +174,31 @@ Data* chunk_serialize(Chunk* chunk, bool use_metadata) {
char* data_pointer = data->data;
for (int i = 0; i < chunk->element_count; i++) {
File* file = chunk->items[i];
const char* wire_path = file_wire_path(file);
char* wire_path = chunk_encode_wire(file_wire_path(file));
if (wire_path == NULL) {
data_destroy(data);
return NULL;
}
size_t path_len = strlen(wire_path);
memcpy(data_pointer, &path_len, sizeof(size_t));
data_pointer += sizeof(size_t);
memcpy(data_pointer, wire_path, path_len);
data_pointer += path_len;
free(wire_path);
int entry_type = file->is_dir ? 1 : 0;
int entry_type = file->is_dir ? 1 : (file->is_symlink ? 2 : (file->is_special ? 3 : 0));
memcpy(data_pointer, &entry_type, sizeof(int));
data_pointer += sizeof(int);
if (file->is_special) {
int32_t special_major = file->rdev_major;
int32_t special_minor = file->rdev_minor;
memcpy(data_pointer, &special_major, sizeof(special_major));
data_pointer += sizeof(special_major);
memcpy(data_pointer, &special_minor, sizeof(special_minor));
data_pointer += sizeof(special_minor);
}
if (use_metadata)
metadata_to_buf(&data_pointer, file->metadata);
@@ -129,6 +208,21 @@ Data* chunk_serialize(Chunk* chunk, bool use_metadata) {
if (file_data_size > 0)
memcpy(data_pointer, file->data->data, file_data_size);
data_pointer += file_data_size;
if (file->is_symlink) {
char* wire_target = chunk_encode_wire(file->symlink_target ? file->symlink_target : "");
if (wire_target == NULL) {
data_destroy(data);
return NULL;
}
size_t target_len = strlen(wire_target);
memcpy(data_pointer, &target_len, sizeof(size_t));
data_pointer += sizeof(size_t);
if (target_len > 0)
memcpy(data_pointer, wire_target, target_len);
data_pointer += target_len;
free(wire_target);
}
}
return data;
}
@@ -141,17 +235,20 @@ Chunk* chunk_deserialize(Data* data, bool use_metadata) {
return NULL;
char* data_pointer = data->data;
size_t remaining_size = data->size;
/* The element currently being parsed is owned by `files` only after the
* array_list_add() at the end of the iteration; until then the error
* epilogue destroys it directly. Keeping this one pointer nulled after the
* hand-off makes the single cleanup path correct for every failure. */
File* file = NULL;
while (remaining_size > 0) {
if ((unsigned int)files->size >= MAX_FILES_PER_CHUNK) {
log_message(LOG_LEVEL_ERROR, "Chunk contains too many files");
array_list_delete(files);
return NULL;
goto error;
}
if (remaining_size < sizeof(size_t)) {
log_message(LOG_LEVEL_ERROR, "Invalid chunk format: not enough data for path length");
array_list_delete(files);
return NULL;
goto error;
}
size_t path_len;
@@ -161,95 +258,117 @@ Chunk* chunk_deserialize(Data* data, bool use_metadata) {
if (path_len > SIZE_MAX - 1 || remaining_size < path_len) {
log_message(LOG_LEVEL_ERROR, "Invalid chunk format: not enough data for path");
array_list_delete(files);
return NULL;
goto error;
}
if (path_len == SIZE_MAX) {
array_list_delete(files);
return NULL;
}
char* path = protocol_alloc(path_len + 1);
if (path == NULL) {
log_perror("Could not allocate memory for file path");
array_list_delete(files);
return NULL;
goto error;
}
memcpy(path, data_pointer, path_len);
path[path_len] = '\0';
if (memchr(path, '\0', path_len) != NULL) {
free(path);
array_list_delete(files);
return NULL;
goto error;
}
data_pointer += path_len;
remaining_size -= path_len;
if (path_len == 0 || has_path_traversal(path)) {
/* --iconv: the blob holds the wire charset; translate it to the receiver's
local charset before validation and creation so the destination gets the
local name. A name that cannot be decoded fails the file cleanly. */
if (charset_wire_active()) {
char* local_path = charset_wire_apply(path);
free(path);
array_list_delete(files);
return NULL;
if (local_path == NULL) {
log_message(LOG_LEVEL_ERROR,
"--iconv: received chunk file name cannot be converted to the local charset");
goto error;
}
path = local_path;
path_len = strlen(path);
}
File* file = file_create(path);
if (path_len == 0 || has_path_traversal(path)) {
free(path);
if (file == NULL) {
array_list_delete(files);
return NULL;
goto error;
}
file = file_create(path);
free(path);
if (file == NULL)
goto error;
if (remaining_size < sizeof(int)) {
log_message(LOG_LEVEL_ERROR, "Invalid chunk format: not enough data for entry type");
file_destroy(file);
array_list_delete(files);
return NULL;
goto error;
}
int entry_type;
memcpy(&entry_type, data_pointer, sizeof(int));
if (entry_type != 0 && entry_type != 1) {
if (entry_type != 0 && entry_type != 1 && entry_type != 2 && entry_type != 3) {
log_message(LOG_LEVEL_ERROR, "Invalid chunk format: bad entry type");
file_destroy(file);
array_list_delete(files);
return NULL;
goto error;
}
file->is_dir = entry_type == 1;
file->is_symlink = entry_type == 2;
file->is_special = entry_type == 3;
data_pointer += sizeof(int);
remaining_size -= sizeof(int);
if (file->is_special) {
if (remaining_size < 2 * (int32_t)sizeof(int32_t)) {
log_message(LOG_LEVEL_ERROR, "Invalid chunk format: not enough data for special rdev");
goto error;
}
int32_t special_major, special_minor;
memcpy(&special_major, data_pointer, sizeof(special_major));
data_pointer += sizeof(special_major);
memcpy(&special_minor, data_pointer, sizeof(special_minor));
data_pointer += sizeof(special_minor);
remaining_size -= 2 * sizeof(int32_t);
/* Reject an out-of-range/negative rdev here as a malformed chunk (the
same 0xffff / 0x00ffffff bounds file_special_rdev_valid uses), so a
bogus large-but-positive rdev is refused cleanly instead of being
deferred to the creation site where it would abort after the frame. */
if (special_major < 0 || special_minor < 0 || special_major > 0xffff ||
special_minor > 0x00ffffff) {
log_message(LOG_LEVEL_ERROR, "Invalid chunk format: out-of-range special rdev");
goto error;
}
file->rdev_major = special_major;
file->rdev_minor = special_minor;
}
if (use_metadata) {
if (remaining_size < sizeof(int)) {
log_message(LOG_LEVEL_ERROR, "Invalid chunk format: not enough data for metadata");
file_destroy(file);
array_list_delete(files);
return NULL;
goto error;
}
// Peek at present flag to determine total size needed before reading
/* Peek at the present flag to determine the total record size before
decoding. metadata_from_buf() independently bounds-checks every read
against remaining_size, so a short body can never over-read. */
int present_flag;
memcpy(&present_flag, data_pointer, sizeof(int));
if ((present_flag != 0 && present_flag != 1) ||
(present_flag == 1 && remaining_size < sizeof(int) + FILE_METADATA_WIRE_SIZE)) {
log_message(LOG_LEVEL_ERROR, "Invalid chunk format: not enough data for metadata body");
file_destroy(file);
array_list_delete(files);
return NULL;
goto error;
}
file->metadata = metadata_from_buf(&data_pointer);
remaining_size -= sizeof(int);
file->metadata = metadata_from_buf((const uint8_t*)data_pointer, remaining_size);
size_t metadata_consumed = sizeof(int);
if (present_flag == 1) {
if (file->metadata == NULL) {
file_destroy(file);
array_list_delete(files);
return NULL;
}
remaining_size -= FILE_METADATA_WIRE_SIZE;
if (file->metadata == NULL)
goto error;
metadata_consumed += FILE_METADATA_WIRE_SIZE;
}
data_pointer += metadata_consumed;
remaining_size -= metadata_consumed;
}
if (remaining_size < sizeof(size_t)) {
log_message(LOG_LEVEL_ERROR, "Invalid chunk format: not enough data for data size");
file_destroy(file);
array_list_delete(files);
return NULL;
goto error;
}
size_t file_data_size;
@@ -259,63 +378,103 @@ Chunk* chunk_deserialize(Data* data, bool use_metadata) {
if (remaining_size < file_data_size) {
log_message(LOG_LEVEL_ERROR, "Invalid chunk format: not enough data for file content");
file_destroy(file);
array_list_delete(files);
return NULL;
goto error;
}
// Reject individual file data larger than the maximum allowed size.
if (file_data_size > MAX_FILE_DATA_SIZE) {
log_message(LOG_LEVEL_ERROR, "File data size %zu exceeds maximum %llu", file_data_size,
(unsigned long long)MAX_FILE_DATA_SIZE);
file_destroy(file);
array_list_delete(files);
return NULL;
goto error;
}
size_t allocation_size = file_data_size > 0 ? file_data_size : 1;
void* file_data = protocol_alloc(allocation_size);
if (file_data == NULL) {
log_perror("Could not allocate memory for file data");
file_destroy(file);
array_list_delete(files);
return NULL;
goto error;
}
memcpy(file_data, data_pointer, file_data_size);
Data* replacement = data_create(file_data, file_data_size);
if (replacement == NULL) {
file_destroy(file);
array_list_delete(files);
return NULL;
if (replacement == NULL)
goto error;
/* Charge the retained per-file copy to the connection budget (when the
inbound chunk carries an owning session) so the queued copies are not
held outside MAX_CONNECTION_MEMORY (B6). A NULL owner (e.g. a local
batch apply) leaves the copy uncharged. */
if (!data_charge_session(replacement, data->owner, allocation_size)) {
log_message(LOG_LEVEL_ERROR, "Per-connection memory limit exceeded for chunk file data");
data_destroy(replacement);
goto error;
}
data_destroy(file->data);
file->data = replacement;
data_pointer += file_data_size;
remaining_size -= file_data_size;
if (!array_list_add(files, file)) {
file_destroy(file);
array_list_delete(files);
return NULL;
if (file->is_symlink) {
if (remaining_size < sizeof(size_t)) {
log_message(LOG_LEVEL_ERROR, "Invalid chunk format: not enough data for symlink target");
goto error;
}
size_t target_len;
memcpy(&target_len, data_pointer, sizeof(size_t));
data_pointer += sizeof(size_t);
remaining_size -= sizeof(size_t);
if (target_len == 0 || remaining_size < target_len) {
log_message(LOG_LEVEL_ERROR, "Invalid chunk format: bad symlink target");
goto error;
}
char* target = protocol_alloc(target_len + 1);
if (!target) {
log_perror("Could not allocate memory for symlink target");
goto error;
}
memcpy(target, data_pointer, target_len);
target[target_len] = '\0';
if (memchr(target, '\0', target_len) != NULL) {
free(target);
goto error;
}
/* The symlink target also rides the wire charset; decode it to the local
charset like the path (a target is a path). */
if (charset_wire_active()) {
char* local_target = charset_wire_apply(target);
free(target);
if (local_target == NULL) {
log_message(LOG_LEVEL_ERROR,
"--iconv: received chunk symlink target cannot be converted to the local "
"charset");
goto error;
}
target = local_target;
}
file->symlink_target = target;
data_pointer += target_len;
remaining_size -= target_len;
}
if (!array_list_add(files, file))
goto error;
file = NULL;
}
File** file_array = (File**)array_list_to_array(files);
if (files->size > 0 && file_array == NULL) {
array_list_delete(files);
return NULL;
}
if (files->size > 0 && file_array == NULL)
goto error;
Chunk* chunk = chunk_create(file_array, files->size);
free(file_array);
if (chunk == NULL) {
array_list_delete(files);
return NULL;
}
if (chunk == NULL)
goto error;
files->item_destroyer = NULL;
array_list_delete(files);
return chunk;
error:
if (file)
file_destroy(file);
array_list_delete(files);
return NULL;
}
Data* chunk_compress(Chunk* chunk, int compression_level, bool use_metadata) {
@@ -344,12 +503,21 @@ Chunk* receive_chunk_data(int fd, const Config* config) {
}
Data* data_to_process = chunk_data;
if (config->use_compression) {
/* Preserve the inbound session across decompression so the (larger)
decompressed chunk is charged to the same connection budget; the
compressed buffer's own charge is released by data_destroy below. */
ProtocolSession* owner = chunk_data->owner;
data_to_process = data_decompress_limited(chunk_data, MAX_CHUNK_SIZE);
data_destroy(chunk_data);
if (data_to_process == NULL) {
log_message(LOG_LEVEL_ERROR, "Failed to decompress chunk");
return NULL;
}
if (!data_charge_session(data_to_process, owner, data_to_process->size)) {
log_message(LOG_LEVEL_ERROR, "Per-connection memory limit exceeded for decompressed chunk");
data_destroy(data_to_process);
return NULL;
}
}
// Reject chunks larger than the maximum allowed size to prevent OOM.
+11
View File
@@ -23,4 +23,15 @@ Data* chunk_compress_with_threads(Chunk* chunk, int compression_level, bool use_
int compression_threads);
Chunk* receive_chunk_data(int fd, const Config* config);
/* Charge `charge` retained bytes of `data` against `session`'s per-connection
* budget (MAX_CONNECTION_MEMORY), mirroring the protocol layer's accounting, and
* record them on `data` so data_destroy() returns the charge through the
* Data.owner path. Returns false (leaving `data` uncharged) when the ceiling
* would be exceeded. A NULL/zero-size charge or a NULL session is a no-op
* success. The receive-side decompression and chunk-copy paths know the owning
* session only through the Data.owner of the buffer they are processing, so
* this is the entry point that lets them participate in the connection budget
* without a session handle (B6). */
bool data_charge_session(Data* data, ProtocolSession* session, size_t charge);
#endif
+555 -103
View File
@@ -2,138 +2,539 @@
#include "data.h"
#include "log.h"
#include "protocol.h"
#include <stdlib.h>
#include <limits.h>
#include <lz4.h>
#include <stdatomic.h>
#include <stdint.h>
#include <stdlib.h>
#include <string.h>
#include <strings.h>
#include <threads.h>
#include <unistd.h>
#include <zlib.h>
#include <zstd.h>
#define INITIAL_DECOMPRESS_BUF_SIZE (1024 * 1024)
#define MAX_DECOMPRESSED_SIZE (100ULL * 1024 * 1024) /* 100 MB hard ceiling */
static char* SKIP_COMPRESSION_EXTENSIONS[] = {".jpg", ".jpeg", ".png", ".gif", ".mp4", ".mkv",
".zip", ".gz", ".xz", ".zst", NULL};
/* rsync 3.4.1's built-in skip-compress suffix list (the `--skip-compress`
* defaults, in the man page's order). rsync stores it as space-separated
* "*.suffix" globs; FastSync matches the plain suffix after the final dot, so
* the leading "*." is omitted here. A user --skip-compress list replaces this
* default entirely (matching rsync). */
#define DEFAULT_SKIP_COMPRESS_SUFFIXES \
"3g2 3gp 7z aac ace apk avi bz2 deb dmg ear f4v flac flv gpg gz iso jar jpeg jpg lrz lz lz4 " \
"lzma " \
"lzo m1a m1v m2a m2ts m2v m4a m4b m4p m4r m4v mka mkv mov mp1 mp2 mp3 mp4 mpa mpeg mpg mpv mts " \
"odb odf odg odi odm odp ods odt oga ogg ogm ogv ogx opus otg oth otp ots ott oxt png qt rar " \
"rpm " \
"rz rzip spx squashfs sxc sxd sxg sxm sxw sz tbz tbz2 tgz tlz ts txz tzo vob war webm webp xz " \
"z " \
"zip zst"
bool compression_should_skip(const char* path) {
return compression_should_skip_with_suffixes(path, NULL, -1);
/* Self-describing compressed frames: the first byte is the CompressionAlgo id.
* zlib/lz4 store the uncompressed size as a little-endian uint32 after the
* codec byte so decompression can be exactly pre-sized and bounded. */
#define LZ4_SIZE_PREFIX_LEN 4
static _Atomic int g_compression_algo = COMPRESSION_ALGO_ZSTD;
/* Case-insensitive match of a bare suffix (no leading dot) against a
* space-separated suffix list. */
static bool suffix_in_list(const char* name, const char* list) {
size_t name_len = strlen(name);
while (*list) {
while (*list == ' ')
list++;
const char* start = list;
while (*list && *list != ' ')
list++;
size_t len = (size_t)(list - start);
if (len == name_len && strncasecmp(name, start, len) == 0)
return true;
}
return false;
}
bool compression_should_skip_with_suffixes(const char* path, char* const* suffixes, int count) {
if (!path)
return false;
const char* dot = strrchr(path, '.');
if (!dot)
if (!dot || dot[1] == '\0')
return false;
if (count < 0) {
suffixes = SKIP_COMPRESSION_EXTENSIONS;
count = 0;
while (SKIP_COMPRESSION_EXTENSIONS[count])
count++;
}
const char* name = dot + 1;
/* count < 0 (the user gave no --skip-compress) selects rsync's built-in
* default list; a non-negative count is the user's explicit list. */
if (count < 0)
return suffix_in_list(name, DEFAULT_SKIP_COMPRESS_SUFFIXES);
for (int i = 0; i < count; i++) {
if (strcasecmp(dot, suffixes[i]) == 0)
const char* suffix = suffixes[i];
if (suffix[0] == '.')
suffix++;
if (strcasecmp(name, suffix) == 0)
return true;
}
return false;
}
Data* data_compress(Data* data_to_compress, int compression_level) {
return data_compress_with_threads(data_to_compress, compression_level, 0);
int compression_algo_from_name(const char* name) {
if (!name)
return -1;
if (strcasecmp(name, "zstd") == 0)
return (int)COMPRESSION_ALGO_ZSTD;
if (strcasecmp(name, "lz4") == 0)
return (int)COMPRESSION_ALGO_LZ4;
if (strcasecmp(name, "zlib") == 0)
return (int)COMPRESSION_ALGO_ZLIB;
if (strcasecmp(name, "zlibx") == 0)
return (int)COMPRESSION_ALGO_ZLIBX;
if (strcasecmp(name, "none") == 0)
return (int)COMPRESSION_ALGO_NONE;
return -1;
}
Data* data_compress_with_threads(Data* data_to_compress, int compression_level,
const char* compression_algo_name(CompressionAlgo algo) {
switch (algo) {
case COMPRESSION_ALGO_NONE:
return "none";
case COMPRESSION_ALGO_ZSTD:
return "zstd";
case COMPRESSION_ALGO_LZ4:
return "lz4";
case COMPRESSION_ALGO_ZLIB:
return "zlib";
case COMPRESSION_ALGO_ZLIBX:
return "zlibx";
}
return "<unknown>";
}
bool compression_algo_valid(int algo) {
return algo == (int)COMPRESSION_ALGO_NONE || algo == (int)COMPRESSION_ALGO_ZSTD ||
algo == (int)COMPRESSION_ALGO_LZ4 || algo == (int)COMPRESSION_ALGO_ZLIB ||
algo == (int)COMPRESSION_ALGO_ZLIBX;
}
bool compression_algo_enabled(CompressionAlgo algo) {
return algo != COMPRESSION_ALGO_NONE;
}
CompressionAlgo compression_negotiate_default(void) {
/* rsync 3.4.1 default preference order; every entry is compiled in, so this
* resolves to zstd. */
static const CompressionAlgo preference[] = {
COMPRESSION_ALGO_ZSTD, COMPRESSION_ALGO_LZ4, COMPRESSION_ALGO_ZLIBX,
COMPRESSION_ALGO_ZLIB, COMPRESSION_ALGO_NONE,
};
for (size_t i = 0; i < sizeof(preference) / sizeof(preference[0]); i++) {
if (compression_algo_valid((int)preference[i]))
return preference[i];
}
return COMPRESSION_ALGO_ZSTD;
}
void compression_set_algo(CompressionAlgo algo) {
if (compression_algo_valid((int)algo))
atomic_store(&g_compression_algo, (int)algo);
}
CompressionAlgo compression_get_algo(void) {
return (CompressionAlgo)atomic_load(&g_compression_algo);
}
/* Per-thread cache of zstd contexts plus the grow-only compression scratch
* buffer. zstd contexts are stateful and not safe to share between threads,
* so each thread keeps its own (see compression_get_thread_ctx). The cache is
* stored in a C11 thread-specific storage slot whose destructor releases the
* contexts when the thread exits; this keeps LeakSanitizer clean for the
* short-lived sender/receiver/scanner worker threads without every worker
* entry point having to remember to call compression_free_thread_contexts().
* The main thread's slot is not torn down by tss at process exit, so an atexit
* hook releases it (and compression_free_thread_contexts allows eager
* release). */
typedef struct {
ZSTD_CCtx* cctx;
ZSTD_DCtx* dctx;
void* out_buf; /* reusable ZSTD_compressBound-sized output scratch */
size_t out_cap; /* bytes currently allocated for out_buf */
int level; /* compression level currently applied to cctx */
int workers; /* nbWorkers currently applied to cctx */
bool params_set;
bool cached; /* false when the TSS slot could not be used: caller owns */
} CompressionThreadCtx;
static once_flag compression_tls_once = ONCE_FLAG_INIT;
static tss_t compression_tls_key;
static bool compression_tls_ready;
static void compression_tls_make_key(void);
static void compression_ctx_free(CompressionThreadCtx* ctx) {
if (!ctx)
return;
if (ctx->cctx)
ZSTD_freeCCtx(ctx->cctx);
if (ctx->dctx)
ZSTD_freeDCtx(ctx->dctx);
free(ctx->out_buf);
free(ctx);
}
static void compression_tls_destructor(void* value) {
compression_ctx_free((CompressionThreadCtx*)value);
}
void compression_free_thread_contexts(void) {
call_once(&compression_tls_once, compression_tls_make_key);
if (!compression_tls_ready)
return;
CompressionThreadCtx* ctx = (CompressionThreadCtx*)tss_get(compression_tls_key);
if (!ctx)
return;
/* Clear the slot first so the thread-exit destructor cannot free it twice. */
tss_set(compression_tls_key, NULL);
compression_ctx_free(ctx);
}
static void compression_atexit_cleanup(void) {
compression_free_thread_contexts();
}
static void compression_tls_make_key(void) {
if (tss_create(&compression_tls_key, compression_tls_destructor) == thrd_success) {
compression_tls_ready = true;
atexit(compression_atexit_cleanup);
}
}
static CompressionThreadCtx* compression_get_thread_ctx(void) {
call_once(&compression_tls_once, compression_tls_make_key);
if (!compression_tls_ready) {
/* Extremely unlikely: fall back to an uncached context the caller frees. */
return (CompressionThreadCtx*)calloc(1, sizeof(CompressionThreadCtx));
}
CompressionThreadCtx* ctx = (CompressionThreadCtx*)tss_get(compression_tls_key);
if (ctx)
return ctx;
ctx = (CompressionThreadCtx*)calloc(1, sizeof(CompressionThreadCtx));
if (!ctx)
return NULL;
ctx->cached = true;
if (tss_set(compression_tls_key, ctx) != thrd_success)
ctx->cached = false;
return ctx;
}
/* Release an uncached context immediately; cached contexts are owned by the
* thread's TSS slot and freed on thread exit / compression_free_thread_contexts. */
static void compression_ctx_put(CompressionThreadCtx* ctx) {
if (ctx && !ctx->cached)
compression_ctx_free(ctx);
}
/* Build a frame consisting of a copy of `src` prefixed by `codec`. */
static Data* frame_with_codec(const void* src, size_t size, CompressionAlgo codec) {
if (size > SIZE_MAX - 1)
return NULL;
Data* out = data_create_empty(size + 1);
if (!out)
return NULL;
((uint8_t*)out->data)[0] = (uint8_t)codec;
if (size > 0)
memcpy((uint8_t*)out->data + 1, src, size);
out->size = size + 1;
return out;
}
static Data* zstd_compress(Data* in, int compression_level, int compression_threads) {
size_t dst_size = ZSTD_compressBound(in->size);
if (dst_size > SIZE_MAX - 1)
return NULL;
dst_size += 1; /* codec prefix */
CompressionThreadCtx* ctx = compression_get_thread_ctx();
if (ctx == NULL) {
log_message(LOG_LEVEL_ERROR, "Failed to allocate ZSTD compression context");
return NULL;
}
Data* compressed_data = NULL;
if (!ctx->cctx) {
ctx->cctx = ZSTD_createCCtx();
if (!ctx->cctx) {
log_message(LOG_LEVEL_ERROR, "Failed to create ZSTD compression context");
goto cleanup;
}
ctx->params_set = false;
}
/* Reset only the session: parameters (and any already-allocated zstd worker
* pool) stay attached to the context, so compressing the next file does not
* rebuild the pool. */
ZSTD_CCtx_reset(ctx->cctx, ZSTD_reset_session_only);
if (!ctx->params_set || ctx->level != compression_level) {
size_t zret = ZSTD_CCtx_setParameter(ctx->cctx, ZSTD_c_compressionLevel, compression_level);
if (ZSTD_isError(zret)) {
log_message(LOG_LEVEL_ERROR, "Failed to set compression level: %s", ZSTD_getErrorName(zret));
goto cleanup;
}
ctx->level = compression_level;
}
int available_threads = 0;
if (compression_threads > 0) {
long online_cpus = sysconf(_SC_NPROCESSORS_ONLN);
available_threads = online_cpus > 0 && online_cpus < compression_threads ? (int)online_cpus
: compression_threads;
}
if (!ctx->params_set || ctx->workers != available_threads) {
size_t zret = ZSTD_CCtx_setParameter(ctx->cctx, ZSTD_c_nbWorkers, available_threads);
if (ZSTD_isError(zret)) {
log_message(LOG_LEVEL_ERROR, "Failed to set compression threads: %s",
ZSTD_getErrorName(zret));
goto cleanup;
}
ctx->workers = available_threads;
}
ctx->params_set = true;
if (available_threads > 0) {
/* Streaming compression needs the source size before threaded mode can end a frame. */
size_t zret = ZSTD_CCtx_setPledgedSrcSize(ctx->cctx, in->size);
if (ZSTD_isError(zret)) {
log_message(LOG_LEVEL_ERROR, "Failed to set compression source size: %s",
ZSTD_getErrorName(zret));
goto cleanup;
}
}
if (ctx->out_cap < dst_size) {
void* grown = protocol_realloc(ctx->out_buf, dst_size);
if (grown == NULL) {
log_message(LOG_LEVEL_ERROR, "Failed to allocate compression buffer");
goto cleanup;
}
ctx->out_buf = grown;
ctx->out_cap = dst_size;
}
ZSTD_inBuffer input = {in->data, in->size, 0};
ZSTD_outBuffer output = {(uint8_t*)ctx->out_buf + 1, dst_size - 1, 0};
size_t ret;
do {
ret = ZSTD_compressStream2(ctx->cctx, &output, &input, ZSTD_e_end);
if (ZSTD_isError(ret)) {
log_message(LOG_LEVEL_ERROR, "Compression failed: %s", ZSTD_getErrorName(ret));
goto cleanup;
}
} while (ret > 0);
/* Hand off an exactly-sized copy; the scratch buffer stays cached so the next
* call does not reallocate a ZSTD_compressBound-sized block. */
compressed_data = data_create_empty(output.pos + 1);
if (compressed_data == NULL) {
log_message(LOG_LEVEL_ERROR, "Failed to allocate compressed data");
goto cleanup;
}
((uint8_t*)compressed_data->data)[0] = (uint8_t)COMPRESSION_ALGO_ZSTD;
if (output.pos > 0)
memcpy((uint8_t*)compressed_data->data + 1, (uint8_t*)ctx->out_buf + 1, output.pos);
compressed_data->size = output.pos + 1;
log_debug_message(LOG_DEBUG_UTIL, "Data succesfully compressed from %zu to %zu", in->size,
compressed_data->size);
cleanup:
compression_ctx_put(ctx);
return compressed_data;
}
static Data* lz4_compress(Data* in) {
int bound = LZ4_compressBound((int)in->size);
if (bound < 0 || in->size > (size_t)INT_MAX)
return NULL;
Data* out = data_create_empty((size_t)bound + 1 + LZ4_SIZE_PREFIX_LEN);
if (!out)
return NULL;
uint32_t raw_size = (uint32_t)in->size;
uint8_t* p = (uint8_t*)out->data;
p[0] = (uint8_t)COMPRESSION_ALGO_LZ4;
for (int i = 0; i < LZ4_SIZE_PREFIX_LEN; i++)
p[1 + i] = (uint8_t)((raw_size >> (8 * i)) & 0xff);
int written = 0;
if (in->size > 0) {
written = LZ4_compress_default((const char*)in->data, (char*)p + 1 + LZ4_SIZE_PREFIX_LEN,
(int)in->size, bound);
if (written <= 0) {
data_destroy(out);
return NULL;
}
}
out->size = (size_t)written + 1 + LZ4_SIZE_PREFIX_LEN;
return out;
}
static Data* zlib_compress(Data* in, CompressionAlgo algo, int compression_level) {
int level = compression_level;
if (level < 1)
level = Z_DEFAULT_COMPRESSION;
if (level > 9)
level = 9;
uLong bound = compressBound((uLong)in->size);
if (in->size > (size_t)ULONG_MAX)
return NULL;
Data* out = data_create_empty((size_t)bound + 1 + LZ4_SIZE_PREFIX_LEN);
if (!out)
return NULL;
uint32_t raw_size = (uint32_t)in->size;
uint8_t* p = (uint8_t*)out->data;
p[0] = (uint8_t)algo;
for (int i = 0; i < LZ4_SIZE_PREFIX_LEN; i++)
p[1 + i] = (uint8_t)((raw_size >> (8 * i)) & 0xff);
uLongf dest_len = bound;
int rc = compress2(p + 1 + LZ4_SIZE_PREFIX_LEN, &dest_len, (const Bytef*)in->data,
(uLong)in->size, level);
if (rc != Z_OK) {
data_destroy(out);
return NULL;
}
out->size = (size_t)dest_len + 1 + LZ4_SIZE_PREFIX_LEN;
return out;
}
Data* data_compress_codec(Data* data_to_compress, CompressionAlgo algo, int compression_level,
int compression_threads) {
if (!data_to_compress || (!data_to_compress->data && data_to_compress->size != 0) ||
compression_threads < 0 || compression_threads > COMPRESSION_MAX_THREADS)
return NULL;
if (!compression_algo_valid((int)algo))
return NULL;
log_message(LOG_LEVEL_DEBUG, "Starting to compress data");
size_t dst_size = ZSTD_compressBound(data_to_compress->size);
Data* compressed_data = data_create_empty(dst_size);
if (compressed_data == NULL)
return NULL;
ZSTD_CCtx* cctx = ZSTD_createCCtx();
if (!cctx) {
log_message(LOG_LEVEL_ERROR, "Failed to create ZSTD compression context");
data_destroy(compressed_data);
return NULL;
switch (algo) {
case COMPRESSION_ALGO_NONE:
return frame_with_codec(data_to_compress->data, data_to_compress->size, COMPRESSION_ALGO_NONE);
case COMPRESSION_ALGO_ZSTD:
return zstd_compress(data_to_compress, compression_level, compression_threads);
case COMPRESSION_ALGO_LZ4:
return lz4_compress(data_to_compress);
case COMPRESSION_ALGO_ZLIB:
case COMPRESSION_ALGO_ZLIBX:
return zlib_compress(data_to_compress, algo, compression_level);
}
size_t zret = ZSTD_CCtx_setParameter(cctx, ZSTD_c_compressionLevel, compression_level);
if (ZSTD_isError(zret)) {
log_message(LOG_LEVEL_ERROR, "Failed to set compression level: %s", ZSTD_getErrorName(zret));
ZSTD_freeCCtx(cctx);
data_destroy(compressed_data);
return NULL;
}
if (compression_threads > 0) {
long online_cpus = sysconf(_SC_NPROCESSORS_ONLN);
int available_threads = online_cpus > 0 && online_cpus < compression_threads
? (int)online_cpus
: compression_threads;
zret = ZSTD_CCtx_setParameter(cctx, ZSTD_c_nbWorkers, available_threads);
if (ZSTD_isError(zret)) {
log_message(LOG_LEVEL_ERROR, "Failed to set compression threads: %s",
ZSTD_getErrorName(zret));
ZSTD_freeCCtx(cctx);
data_destroy(compressed_data);
return NULL;
}
/* Streaming compression needs the source size before threaded mode can end a frame. */
zret = ZSTD_CCtx_setPledgedSrcSize(cctx, data_to_compress->size);
if (ZSTD_isError(zret)) {
log_message(LOG_LEVEL_ERROR, "Failed to set compression source size: %s",
ZSTD_getErrorName(zret));
ZSTD_freeCCtx(cctx);
data_destroy(compressed_data);
return NULL;
}
}
ZSTD_inBuffer input = {data_to_compress->data, data_to_compress->size, 0};
ZSTD_outBuffer output = {compressed_data->data, dst_size, 0};
size_t ret;
do {
ret = ZSTD_compressStream2(cctx, &output, &input, ZSTD_e_end);
if (ZSTD_isError(ret)) {
log_message(LOG_LEVEL_ERROR, "Compression failed: %s", ZSTD_getErrorName(ret));
ZSTD_freeCCtx(cctx);
data_destroy(compressed_data);
return NULL;
}
} while (ret > 0);
compressed_data->size = output.pos;
ZSTD_freeCCtx(cctx);
log_debug_message(LOG_DEBUG_UTIL, "Data succesfully compressed from %zu to %zu",
data_to_compress->size, compressed_data->size);
return compressed_data;
}
Data* data_decompress_limited(Data* compressed_data, size_t maximum_size) {
if (!compressed_data || (!compressed_data->data && compressed_data->size != 0) ||
maximum_size == 0)
Data* data_compress_with_threads(Data* data_to_compress, int compression_level,
int compression_threads) {
return data_compress_codec(data_to_compress, compression_get_algo(), compression_level,
compression_threads);
}
Data* data_compress(Data* data_to_compress, int compression_level) {
return data_compress_codec(data_to_compress, compression_get_algo(), compression_level, 0);
}
static Data* decompress_none(const Data* compressed_data, size_t maximum_size) {
size_t size = compressed_data->size - 1;
if (size > maximum_size)
return NULL;
Data* out = data_create_empty(size);
if (!out)
return NULL;
if (size > 0)
memcpy(out->data, (const uint8_t*)compressed_data->data + 1, size);
out->size = size;
return out;
}
/* Read the 4-byte little-endian raw size stored after the codec byte. */
static bool read_raw_size(const Data* in, uint32_t* raw_size) {
if (in->size < 1 + LZ4_SIZE_PREFIX_LEN)
return false;
const uint8_t* p = (const uint8_t*)in->data;
uint32_t v = 0;
for (int i = 0; i < LZ4_SIZE_PREFIX_LEN; i++)
v |= (uint32_t)p[1 + i] << (8 * i);
*raw_size = v;
return true;
}
static Data* lz4_decompress(Data* compressed_data, size_t maximum_size, size_t hard_limit) {
uint32_t raw_size = 0;
if (!read_raw_size(compressed_data, &raw_size))
return NULL;
if (raw_size > hard_limit || raw_size > maximum_size)
return NULL;
size_t comp_size = compressed_data->size - 1 - LZ4_SIZE_PREFIX_LEN;
Data* out = data_create_empty(raw_size);
if (!out)
return NULL;
if (raw_size == 0) {
out->size = 0;
return out;
}
int rc = LZ4_decompress_safe((const char*)compressed_data->data + 1 + LZ4_SIZE_PREFIX_LEN,
(char*)out->data, (int)comp_size, (int)raw_size);
if (rc < 0 || (uint32_t)rc != raw_size) {
log_message(LOG_LEVEL_ERROR, "LZ4 decompression failed");
data_destroy(out);
return NULL;
}
out->size = raw_size;
return out;
}
static Data* zlib_decompress(Data* compressed_data, size_t maximum_size, size_t hard_limit) {
uint32_t raw_size = 0;
if (!read_raw_size(compressed_data, &raw_size))
return NULL;
if (raw_size > hard_limit || raw_size > maximum_size)
return NULL;
size_t comp_size = compressed_data->size - 1 - LZ4_SIZE_PREFIX_LEN;
Data* out = data_create_empty(raw_size);
if (!out)
return NULL;
if (raw_size == 0) {
out->size = 0;
return out;
}
uLongf dest_len = raw_size;
int rc =
uncompress((Bytef*)out->data, &dest_len,
(const Bytef*)compressed_data->data + 1 + LZ4_SIZE_PREFIX_LEN, (uLong)comp_size);
if (rc != Z_OK || dest_len != raw_size) {
log_message(LOG_LEVEL_ERROR, "zlib decompression failed");
data_destroy(out);
return NULL;
}
out->size = raw_size;
return out;
}
static Data* zstd_decompress(Data* compressed_data, size_t maximum_size) {
/* The zstd frame starts after the codec byte. */
const void* frame = (const uint8_t*)compressed_data->data + 1;
size_t frame_size = compressed_data->size - 1;
log_debug_message(LOG_DEBUG_UTIL, "Start to decompress data");
unsigned long long dst_size =
ZSTD_getFrameContentSize(compressed_data->data, compressed_data->size);
if (ZSTD_isError(dst_size)) {
log_message(LOG_LEVEL_ERROR, "Failed to get decompressed size: %s",
ZSTD_getErrorName(dst_size));
unsigned long long dst_size = ZSTD_getFrameContentSize(frame, frame_size);
/* ZSTD_isError() is also true for ZSTD_CONTENTSIZE_ERROR and
* ZSTD_CONTENTSIZE_UNKNOWN (both are encoded near (size_t)-1), so test the
* sentinels explicitly instead of blanket-rejecting every error-ish value:
* only CONTENTSIZE_ERROR means an unreadable header, while CONTENTSIZE_UNKNOWN
* must reach the estimate fallback below. */
if (dst_size == ZSTD_CONTENTSIZE_ERROR) {
log_message(LOG_LEVEL_ERROR, "Failed to get decompressed size: invalid zstd frame");
return NULL;
}
// ZSTD_CONTENTSIZE_UNKNOWN (~2^64) can cause massive allocation;
// fall back to a conservative estimate (3x compressed size) when unknown.
if (dst_size == ZSTD_CONTENTSIZE_UNKNOWN) {
if (compressed_data->size > ULLONG_MAX / 3)
if (frame_size > ULLONG_MAX / 3)
return NULL;
dst_size = compressed_data->size * 3;
dst_size = frame_size * 3;
if (dst_size < INITIAL_DECOMPRESS_BUF_SIZE)
dst_size = INITIAL_DECOMPRESS_BUF_SIZE;
}
@@ -144,41 +545,51 @@ Data* data_decompress_limited(Data* compressed_data, size_t maximum_size) {
return NULL;
}
ZSTD_DCtx* dctx = ZSTD_createDCtx();
if (!dctx) {
log_message(LOG_LEVEL_ERROR, "Failed to create ZSTD decompression context");
CompressionThreadCtx* ctx = compression_get_thread_ctx();
if (ctx == NULL) {
log_message(LOG_LEVEL_ERROR, "Failed to allocate ZSTD decompression context");
return NULL;
}
Data* uncompressed_data = NULL;
if (!ctx->dctx) {
ctx->dctx = ZSTD_createDCtx();
if (!ctx->dctx) {
log_message(LOG_LEVEL_ERROR, "Failed to create ZSTD decompression context");
goto cleanup;
}
}
/* Reset only the session; decompression parameters are sticky. */
ZSTD_DCtx_reset(ctx->dctx, ZSTD_reset_session_only);
size_t buf_size = (dst_size > 0) ? (size_t)dst_size : INITIAL_DECOMPRESS_BUF_SIZE;
if (buf_size > maximum_size)
buf_size = maximum_size;
Data* uncompressed_data = data_create_empty(buf_size);
uncompressed_data = data_create_empty(buf_size);
if (!uncompressed_data) {
log_message(LOG_LEVEL_ERROR, "Failed to allocate decompression buffer");
ZSTD_freeDCtx(dctx);
return NULL;
goto cleanup;
}
ZSTD_inBuffer input = {compressed_data->data, compressed_data->size, 0};
ZSTD_inBuffer input = {frame, frame_size, 0};
ZSTD_outBuffer output = {uncompressed_data->data, buf_size, 0};
size_t ret;
do {
ret = ZSTD_decompressStream(dctx, &output, &input);
ret = ZSTD_decompressStream(ctx->dctx, &output, &input);
if (ZSTD_isError(ret)) {
log_message(LOG_LEVEL_ERROR, "Decompression failed: %s", ZSTD_getErrorName(ret));
ZSTD_freeDCtx(dctx);
data_destroy(uncompressed_data);
return NULL;
uncompressed_data = NULL;
goto cleanup;
}
if (ret > 0 && output.pos == output.size) {
if (buf_size >= hard_limit || buf_size > SIZE_MAX / 2) {
log_message(LOG_LEVEL_ERROR, "Decompressed data exceeds %llu bytes",
(unsigned long long)MAX_DECOMPRESSED_SIZE);
ZSTD_freeDCtx(dctx);
data_destroy(uncompressed_data);
return NULL;
uncompressed_data = NULL;
goto cleanup;
}
buf_size *= 2;
if (buf_size > hard_limit)
@@ -186,23 +597,64 @@ Data* data_decompress_limited(Data* compressed_data, size_t maximum_size) {
void* new_data = protocol_realloc(uncompressed_data->data, buf_size);
if (!new_data) {
log_message(LOG_LEVEL_ERROR, "Failed to grow decompression buffer");
ZSTD_freeDCtx(dctx);
data_destroy(uncompressed_data);
return NULL;
uncompressed_data = NULL;
goto cleanup;
}
uncompressed_data->data = new_data;
output.dst = new_data;
output.size = buf_size;
/* Re-attempt with the larger output buffer; the truncated-frame check
* below must not reject a complete frame that merely filled the previous
* buffer exactly. */
continue;
}
/* A positive hint with all input consumed means the frame is incomplete: a
* truncated stream would otherwise spin here forever (ZSTD_decompressStream
* keeps returning the same hint). Fail instead of burning CPU. */
if (ret != 0 && input.pos == input.size) {
log_message(LOG_LEVEL_ERROR,
"Truncated zstd frame: input exhausted with %zu bytes still expected", ret);
data_destroy(uncompressed_data);
uncompressed_data = NULL;
goto cleanup;
}
} while (ret > 0);
uncompressed_data->size = output.pos;
ZSTD_freeDCtx(dctx);
log_debug_message(LOG_DEBUG_UTIL, "Decompressed data successfully");
cleanup:
compression_ctx_put(ctx);
return uncompressed_data;
}
Data* data_decompress_limited(Data* compressed_data, size_t maximum_size) {
if (!compressed_data || (!compressed_data->data && compressed_data->size != 0) ||
maximum_size == 0)
return NULL;
if (compressed_data->size < 1)
return NULL;
unsigned long long hard_limit =
maximum_size < MAX_DECOMPRESSED_SIZE ? maximum_size : MAX_DECOMPRESSED_SIZE;
uint8_t codec = ((const uint8_t*)compressed_data->data)[0];
if (!compression_algo_valid(codec))
return NULL;
switch ((CompressionAlgo)codec) {
case COMPRESSION_ALGO_NONE:
return decompress_none(compressed_data, (size_t)hard_limit);
case COMPRESSION_ALGO_ZSTD:
return zstd_decompress(compressed_data, (size_t)hard_limit);
case COMPRESSION_ALGO_LZ4:
return lz4_decompress(compressed_data, maximum_size, (size_t)hard_limit);
case COMPRESSION_ALGO_ZLIB:
case COMPRESSION_ALGO_ZLIBX:
return zlib_decompress(compressed_data, maximum_size, (size_t)hard_limit);
}
return NULL;
}
Data* data_decompress(Data* compressed_data) {
return data_decompress_limited(compressed_data, MAX_DECOMPRESSED_SIZE);
}
+57 -2
View File
@@ -6,12 +6,67 @@
#define COMPRESSION_MAX_THREADS 64
/* Compression algorithms selectable with --compress-choice / -z. The ids are
* the values placed on the wire (Config->compression_algo), so they must be
* kept stable. NONE is "no compression"; ZSTD is the historical FastSync
* default and the negotiated "auto" choice. ZLIBX is rsync's zlib-without-
* matched-data variant: FastSync compresses only the delta/token bytes (it does
* not put matched file data in the compression stream), so its zlib codec is
* already the "x" form and zlib/zlibx share the same implementation, recorded
* under distinct ids. */
typedef enum {
COMPRESSION_ALGO_NONE = 0,
COMPRESSION_ALGO_ZSTD = 1,
COMPRESSION_ALGO_LZ4 = 2,
COMPRESSION_ALGO_ZLIB = 3,
COMPRESSION_ALGO_ZLIBX = 4
} CompressionAlgo;
/* Resolve a --compress-choice string (case-insensitive) to an algorithm id.
* Accepts "zstd", "lz4", "zlib", "zlibx", "none". "auto" is not an algorithm
* here; the caller resolves it to the negotiated default. Returns -1 for any
* unrecognized name. */
int compression_algo_from_name(const char* name);
const char* compression_algo_name(CompressionAlgo algo);
bool compression_algo_valid(int algo);
/* Pick the first algorithm from FastSync's compiled-in preference list
* (rsync 3.4.1's `--version` order: zstd lz4 zlibx zlib none). Resolves
* "auto". */
CompressionAlgo compression_negotiate_default(void);
/* True when the algorithm actually compresses (i.e. is not NONE). */
bool compression_algo_enabled(CompressionAlgo algo);
/* Select the process-wide codec used by the legacy wrappers below. Each
* process serves exactly one transfer config (the server forks per connection,
* the client configures itself before spawning transfer threads), so a
* process-global default is sufficient and constant for the lifetime of a
* transfer. Defaults to ZSTD when never set. Thread-safe. */
void compression_set_algo(CompressionAlgo algo);
CompressionAlgo compression_get_algo(void);
/* Codec-aware primitives. The compressed buffer is self-describing: its first
* byte is the CompressionAlgo id, so decompression never needs the codec passed
* separately (this keeps every existing Decompress call site source-compatible).
* `data_compress_codec` returns NULL on invalid input or an unsupported codec. */
Data* data_compress_codec(Data* data_to_compress, CompressionAlgo algo, int compression_level,
int compression_threads);
Data* data_decompress_limited(Data* compressed_data, size_t maximum_size);
/* Legacy zstd-default wrappers retained for existing callers/tests. */
Data* data_compress(Data* data_to_compress, int compression_level);
Data* data_compress_with_threads(Data* data_to_compress, int compression_level,
int compression_threads);
Data* data_decompress(Data* compressed_data);
Data* data_decompress_limited(Data* compressed_data, size_t maximum_size);
bool compression_should_skip(const char* path);
bool compression_should_skip_with_suffixes(const char* path, char* const* suffixes, int count);
/* Release the calling thread's cached zstd contexts (compressor, decompressor
* and scratch buffer). The cache is thread-local and is also released
* automatically when a worker thread exits (via a C11 tss destructor) and for
* the main thread at process exit; this explicit entry point exists so tests
* and long-lived callers can drop the cache deterministically. Safe to call
* when no context has been created, and idempotent. */
void compression_free_thread_contexts(void);
#endif
+1135 -391
View File
File diff suppressed because it is too large. Load diff
+997 -154
View File
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
+203
View File
@@ -0,0 +1,203 @@
#ifndef CREDENTIALS_H
#define CREDENTIALS_H
#include <stdbool.h>
#include <stddef.h>
#include <stdint.h>
#include <stdio.h>
/* Daemon password authentication (A7 remediation, protocol 2.19.0).
*
* FastSync authenticates a daemon connection with a SCRAM-SHA-256-style
* challenge/response handshake. The daemon stores only a salted PBKDF2
* verifier (never the password, and never a value that can be replayed as a
* bearer credential): the client proves knowledge of the password against a
* per-connection server nonce, and the server proves the same shared secret
* back. See credentials.c for the exact derivation.
*
* Server credential store format (--password-file and --early-input): one line
* per entry,
* user:$fastsync$1$pbkdf2-sha256$<iters>$<salt_b64>$<stored_key_b64>$<server_key_b64>
* with standard base64, a 16-byte salt and 32-byte keys, and iters in
* [CREDENTIAL_MIN_ITERS, CREDENTIAL_MAX_ITERS]. Blank lines and lines whose
* first non-space character is '#' or ';' are comments. The parser is STRICT:
* a malformed line fails the whole load so a typo can never silently change who
* may log in. A line holding the legacy (unsalted SHA-256 hex) secret is
* hard-rejected with an actionable "legacy" error; there is no auto-upgrade.
* Use `fastsync-server --hash-credentials` to generate new-format lines.
*
* Alongside the store, credentials_load maintains an exact-mode-0600
* `<store_path>.dummykey` sidecar holding the store-wide random dummy key. It
* is auto-created on first load and MUST be preserved across restarts: it makes
* the dummy challenge for an unknown user stable for the life of the store, so
* a daemon restart cannot be used as a username-enumeration oracle. A sidecar
* that is not an exact-mode-0600 regular file of exactly 32 bytes fails the load
* (fail closed); creation forces exact 0600 with fchmod (so a restrictive umask
* cannot leave the sidecar unreadable), and only a create/write/fsync/link or
* fchmod failure degrades to a transient per-run key with a warning. NOTE: the
* sidecar requires EXACT 0600, whereas the store / password files only reject
* group/other bits (a deliberate difference).
*
* Client --password-file format: the FIRST meaningful (non-comment, non-blank)
* line is `user:password`, holding the literal password. The client keeps it
* only for the duration of the handshake and wipes it at teardown; the file
* should be mode 0600 and readable only by its owner. */
/* Longest accepted credential-file line (excluding the trailing newline). */
#define CREDENTIAL_MAX_LINE 4096
/* Upper bound on a username in a credential file and on the wire. Kept well
* below MAX_STRING_SIZE so a wire username can never exhaust anything. */
#define CREDENTIAL_MAX_USER_LEN 256
/* Upper bound on a client-file password (before derivation). */
#define CREDENTIAL_MAX_PASSWORD_LEN 1024
/* SCRAM-SHA-256 parameters. Salt and client nonce sizes are fixed by the
* shared-auth-message framing; keys are always 32 bytes (SHA-256). */
#define CREDENTIAL_SALT_LEN 16
#define CREDENTIAL_NONCE_LEN 32
#define CREDENTIAL_KEY_LEN 32
#define CREDENTIAL_DEFAULT_ITERS 600000u
#define CREDENTIAL_MIN_ITERS 100000u
#define CREDENTIAL_MAX_ITERS 10000000u
/* Buffer size for the full AuthMessage (prefix + three length-prefixed fields).
* Worst case: 16 + 4 + 256 + 4 + 32 + 4 + 32. */
#define CREDENTIAL_AUTH_MESSAGE_MAX \
(16 + 4 + CREDENTIAL_MAX_USER_LEN + 4 + CREDENTIAL_NONCE_LEN + 4 + CREDENTIAL_NONCE_LEN)
typedef struct CredentialStore CredentialStore;
/* One resolved verifier. `found` is false for an unknown user or a user not on
* a module's auth list; the remaining fields then hold a deterministic dummy
* salt (HMAC of the store-wide dummy key over the username), the store-wide
* uniform iteration count (default for an empty store) and fixed dummy keys, so
* the server can run the same challenge/response math with no enumeration or
* timing oracle. */
typedef struct {
uint8_t salt[CREDENTIAL_SALT_LEN];
uint32_t iters;
uint8_t stored_key[CREDENTIAL_KEY_LEN];
uint8_t server_key[CREDENTIAL_KEY_LEN];
bool found;
} CredentialVerifier;
/* Load the daemon credential store.
*
* password_file and early_input_file are both NULL-or-path, matching the
* server's --password-file and --early-input options. A file that cannot be
* opened or that fails the strict grammar is a hard error (err filled, NULL
* returned) -- the daemon fails CLOSED rather than serving an auth-required
* module with a partial store. Both files may be NULL, which yields an empty
* store (every auth-required module then refuses connections). Every entry in
* the resulting store must agree on the iteration count; entries that disagree
* (within one file or across the two layered sources) are rejected. When both
* are given, the --early-input file is layered over --password-file: a duplicate
* username whose verifier matches is deduplicated; one whose verifier differs
* is an error (the two sources disagree), never a silent pick.
*
* The returned store is heap-owned; free it with credentials_free. */
CredentialStore* credentials_load(const char* password_file, const char* early_input_file,
char* err, size_t err_size);
/* Wipe every stored key/salt and free the store. */
void credentials_free(CredentialStore* store);
/* True when `user` is a single bounded token free of whitespace/control bytes
* (the rule applied to store users, client-file users and the module list). */
bool credentials_username_valid(const char* user);
/* Standard base64. encode writes NUL-terminated output to out (size out_sz).
* decode writes the raw bytes to out (capacity out_sz) and stores the length;
* the input must be a well-formed padded base64 string. Both return false on
* NULL arguments, a bad character/length, or insufficient output space. */
bool credentials_b64_encode(const uint8_t* in, size_t n, char* out, size_t out_sz);
bool credentials_b64_decode(const char* in, uint8_t* out, size_t out_sz, size_t* out_len);
/* Fill out[0..n) from the CSPRNG (RAND_bytes). Returns false on failure. */
bool credentials_random_bytes(uint8_t* out, size_t n);
/* Resolve `user` against the store AND the module's auth-user list. The list
* scan is a constant-time full-length comparison with no early break. On a
* miss, *out is filled with a dummy verifier (a deterministic per-username salt
* derived from the store's dummy key, the store-wide uniform iteration count,
* fixed dummy keys, found=false). Returns false on invalid arguments or an
* HMAC/crypto primitive failure. */
bool credentials_get_verifier(const CredentialStore* store, const char* user,
const char* const* module_users, int n, CredentialVerifier* out);
/* Derive the SCRAM keys from a plaintext password:
* K = PBKDF2-HMAC-SHA256(password, salt, iters, 32)
* ClientKey = HMAC-SHA256(K, "Client Key"); StoredKey = SHA256(ClientKey)
* ServerKey = HMAC-SHA256(K, "Server Key")
* Any of client_key/stored_key/server_key may be NULL when not needed.
* `iters` must lie in [CREDENTIAL_MIN_ITERS, CREDENTIAL_MAX_ITERS]. */
bool credentials_compute_keys(const char* password, const uint8_t salt[CREDENTIAL_SALT_LEN],
uint32_t iters, uint8_t client_key[CREDENTIAL_KEY_LEN],
uint8_t stored_key[CREDENTIAL_KEY_LEN],
uint8_t server_key[CREDENTIAL_KEY_LEN]);
/* Serialize the shared AuthMessage:
* "FastSync-Auth-v1" || be32(len(user)) || user
* || be32(32) || server_nonce
* || be32(32) || client_nonce
* out must hold at least CREDENTIAL_AUTH_MESSAGE_MAX bytes. *out_len receives
* the number of bytes written. */
bool credentials_build_auth_message(const char* user, const uint8_t* snonce, const uint8_t* cnonce,
uint8_t* out, size_t out_sz, size_t* out_len);
/* Client side: ClientProof = ClientKey XOR HMAC(StoredKey, AuthMessage), and
* the expected ServerSignature = HMAC(ServerKey, AuthMessage). */
bool credentials_client_proof(const uint8_t client_key[CREDENTIAL_KEY_LEN],
const uint8_t stored_key[CREDENTIAL_KEY_LEN],
const uint8_t server_key[CREDENTIAL_KEY_LEN], const uint8_t* auth_msg,
size_t msg_len, uint8_t proof[CREDENTIAL_KEY_LEN],
uint8_t server_sig[CREDENTIAL_KEY_LEN]);
/* Server side: recompute ClientSig' = HMAC(StoredKey, AuthMessage) and
* ClientKey' = proof XOR ClientSig', then accept iff v->found AND
* SHA256(ClientKey') equals StoredKey (constant-time over the 32-byte keys).
* Always computes server_sig_out = HMAC(ServerKey, AuthMessage). Returns the
* accept decision. */
bool credentials_verify_response(const CredentialVerifier* v, const char* user,
const uint8_t* snonce, const uint8_t* cnonce,
const uint8_t proof[CREDENTIAL_KEY_LEN],
uint8_t server_sig_out[CREDENTIAL_KEY_LEN]);
/* Derive a new-format store line for `user`/`password` and write it (without a
* trailing newline) into out. A random 16-byte salt is used. On failure err is
* filled. Used by --hash-credentials and by tests. */
bool credentials_hash_store_line(const char* user, const char* password, uint32_t iters, char* out,
size_t out_sz, char* err, size_t err_size);
/* Read `user:password` lines from `path` (the same no-group/other-bits check as
* the other secret files) and write one new-format store line per entry to
* `out`.
* Blank/comment lines are skipped; a malformed line fails the whole run.
* Returns 0 on success, -1 on error (err filled). Used by
* `--hash-credentials`. */
int credentials_hash_file(const char* path, uint32_t iters, FILE* out, char* err, size_t err_size);
/* Read the CLIENT-side secret file: the first meaningful line is
* `user:password` (the literal password). *user_out and *password_out are
* freshly allocated on success (password is plaintext -- the caller derives the
* proof and then burns/frees it); both are NULL on error. Returns 0 on
* success, -1 on failure (err filled: the path is named, never the credential
* itself). Only the line's trailing CR/LF are stripped: the password's bytes
* are otherwise preserved exactly, so a password with leading/trailing
* whitespace (after the ':') is kept usable. The username is trimmed of
* surrounding space/tabs. */
int credentials_read_secret_file(const char* path, char** user_out, char** password_out, char* err,
size_t err_size);
/* Constant-time equality over exactly len bytes. */
bool credentials_secure_equal(const char* a, const char* b, size_t len);
/* Overwrite secret[0..len) with zeros (best-effort wipe). */
void credentials_burn(char* secret, size_t len);
/* Number of entries currently in the store (tests/introspection). */
int credentials_store_size(const CredentialStore* store);
/* Whether the store contains an entry for `user` (tests/introspection). */
bool credentials_store_has(const CredentialStore* store, const char* user);
#endif
+794
View File
@@ -0,0 +1,794 @@
#include "daemon_conf.h"
#include "credentials.h"
#include "utils.h"
#include <arpa/inet.h>
#include <ctype.h>
#include <errno.h>
#include <limits.h>
#include <netinet/in.h>
#include <stdarg.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <strings.h>
/* ------------------------------------------------------------------ */
/* helpers */
/* ------------------------------------------------------------------ */
static void set_error(char* err, size_t err_size, const char* fmt, ...) {
if (!err || err_size == 0)
return;
va_list args;
va_start(args, fmt);
vsnprintf(err, err_size, fmt, args);
va_end(args);
}
/* Trim leading and trailing ASCII space/tab in place; returns the new start. */
static char* trim_ws(char* s) {
while (*s == ' ' || *s == '\t')
s++;
size_t len = strlen(s);
while (len > 0 && (s[len - 1] == ' ' || s[len - 1] == '\t'))
s[--len] = '\0';
return s;
}
/* Case-insensitive equality of a parsed key against a canonical key name. */
static bool key_equals(const char* key, const char* canonical) {
return strcasecmp(key, canonical) == 0;
}
static bool parse_bool_value(const char* value, bool* out) {
if (strcasecmp(value, "yes") == 0 || strcasecmp(value, "true") == 0 || strcmp(value, "1") == 0) {
*out = true;
return true;
}
if (strcasecmp(value, "no") == 0 || strcasecmp(value, "false") == 0 || strcmp(value, "0") == 0) {
*out = false;
return true;
}
return false;
}
/* Parse an IPv4/IPv6 CIDR "addr/prefix" into `bytes`/`*family`. Returns false
* for a malformed address, a missing/oversized prefix, or a prefix that does
* not fit the address family. */
static bool parse_cidr(const char* cidr, int* prefix_out, uint8_t* bytes, int* family_out) {
const char* slash = strchr(cidr, '/');
if (!slash)
return false;
size_t addr_len = (size_t)(slash - cidr);
if (addr_len == 0 || addr_len >= INET6_ADDRSTRLEN)
return false;
char addr[INET6_ADDRSTRLEN];
memcpy(addr, cidr, addr_len);
addr[addr_len] = '\0';
char* end = NULL;
long prefix = strtol(slash + 1, &end, 10);
if (end == slash + 1 || *end != '\0')
return false;
struct in_addr v4;
struct in6_addr v6;
if (inet_pton(AF_INET, addr, &v4) == 1) {
if (prefix < 0 || prefix > 32)
return false;
memcpy(bytes, &v4, sizeof(v4));
*prefix_out = (int)prefix;
*family_out = AF_INET;
return true;
}
if (inet_pton(AF_INET6, addr, &v6) == 1) {
if (prefix < 0 || prefix > 128)
return false;
memcpy(bytes, &v6, sizeof(v6));
*prefix_out = (int)prefix;
*family_out = AF_INET6;
return true;
}
return false;
}
/* A host pattern is valid when it is `*`, a valid IPv4/IPv6 literal, or a valid
* CIDR. Peer addresses reaching the matcher are always numeric, so hostname
* globs are rejected at parse time: accepting one would create a deny rule that
* silently never matches (fail-open). */
static bool host_pattern_valid(const char* pattern) {
if (!pattern || *pattern == '\0')
return false;
if (strcmp(pattern, "*") == 0)
return true;
if (strchr(pattern, '/')) {
uint8_t bytes[16];
int prefix;
int family;
return parse_cidr(pattern, &prefix, bytes, &family);
}
struct in_addr v4;
struct in6_addr v6;
return inet_pton(AF_INET, pattern, &v4) == 1 || inet_pton(AF_INET6, pattern, &v6) == 1;
}
/* Append every comma- and/or whitespace-separated host pattern in `value` to
* the heap-owned list (or replace the list when `replace` is set, which --dparam
* uses so an override can narrow access rather than only widen it). Returns
* false (err filled) on an invalid pattern or an allocation failure. */
static bool store_host_list(char*** list, int* count, const char* value, const char* key,
const char* module_name, bool replace, char* err, size_t err_size) {
if (replace) {
for (int i = 0; i < *count; i++)
free((*list)[i]);
free(*list);
*list = NULL;
*count = 0;
}
char* copy = str_dup(value);
if (!copy) {
if (module_name)
set_error(err, err_size, "out of memory parsing '%s' for module '%s'", key, module_name);
else
set_error(err, err_size, "out of memory parsing '%s'", key);
return false;
}
char* save = NULL;
int added = 0;
for (char* token = strtok_r(copy, ", \t", &save); token; token = strtok_r(NULL, ", \t", &save)) {
if (!host_pattern_valid(token)) {
if (module_name)
set_error(err, err_size, "module '%s': invalid host pattern '%s' in '%s'", module_name,
token, key);
else
set_error(err, err_size, "invalid host pattern '%s' in '%s'", token, key);
free(copy);
return false;
}
char** grown = realloc(*list, (size_t)(*count + 1) * sizeof(char*));
if (!grown) {
if (module_name)
set_error(err, err_size, "out of memory parsing '%s' for module '%s'", key, module_name);
else
set_error(err, err_size, "out of memory parsing '%s'", key);
free(copy);
return false;
}
*list = grown;
char* dup = str_dup(token);
if (!dup) {
if (module_name)
set_error(err, err_size, "out of memory parsing '%s' for module '%s'", key, module_name);
else
set_error(err, err_size, "out of memory parsing '%s'", key);
free(copy);
return false;
}
(*list)[(*count)++] = dup;
added++;
}
free(copy);
/* A present key with an empty (or separator-only) value would otherwise
* install a zero-length list, i.e. no ACL at all: a strict-parse config must
* never silently turn a restrictive directive into "allow everyone". */
if (added == 0) {
if (module_name)
set_error(err, err_size, "module '%s': '%s' must list at least one host pattern", module_name,
key);
else
set_error(err, err_size, "'%s' must list at least one host pattern", key);
return false;
}
return true;
}
/* Parse a `max connections` value: a positive integer (0/negative/garbage are
* rejected because they would silently disable the cap or admit nothing). */
static bool store_max_connections(int* slot, const char* value, const char* module_name, char* err,
size_t err_size) {
char* end = NULL;
errno = 0;
long n = strtol(value, &end, 10);
if (*value == '\0' || errno != 0 || *end != '\0' || n <= 0 || n > INT_MAX) {
if (module_name)
set_error(err, err_size,
"module '%s': invalid 'max connections' '%s' (must be a positive "
"integer)",
module_name, value);
else
set_error(err, err_size, "invalid 'max connections' '%s' (must be a positive integer)",
value);
return false;
}
*slot = (int)n;
return true;
}
/* Parse a non-negative concurrency cap where 0 means unlimited/disabled
* (per-module `max connections`, `max connections per host`,
* `auth lockout threshold`). Negative/garbage/oversized values are rejected. */
static bool store_optional_cap(int* slot, const char* value, int max_value, const char* key,
const char* module_name, char* err, size_t err_size) {
char* end = NULL;
errno = 0;
long n = strtol(value, &end, 10);
if (*value == '\0' || errno != 0 || *end != '\0' || n < 0 || n > max_value) {
if (module_name)
set_error(err, err_size, "module '%s': invalid '%s' '%s' (must be 0-%d)", module_name, key,
value, max_value);
else
set_error(err, err_size, "invalid '%s' '%s' (must be 0-%d)", key, value, max_value);
return false;
}
*slot = (int)n;
return true;
}
/* Parse an `auth failure delay` value: 0 (disabled) through the configured cap. */
static bool store_auth_failure_delay(int* slot, const char* value, char* err, size_t err_size) {
char* end = NULL;
errno = 0;
long n = strtol(value, &end, 10);
if (*value == '\0' || errno != 0 || *end != '\0' || n < 0 ||
n > DAEMON_CONF_MAX_AUTH_FAILURE_DELAY_MS) {
set_error(err, err_size, "invalid 'auth failure delay' '%s' (must be 0-%d milliseconds)", value,
DAEMON_CONF_MAX_AUTH_FAILURE_DELAY_MS);
return false;
}
*slot = (int)n;
return true;
}
bool daemon_module_name_valid(const char* name) {
if (!name || *name == '\0')
return false;
size_t len = strlen(name);
if (len > DAEMON_MAX_MODULE_NAME)
return false;
for (size_t i = 0; i < len; i++) {
unsigned char c = (unsigned char)name[i];
bool alnum = (c >= 'a' && c <= 'z') || (c >= 'A' && c <= 'Z') || (c >= '0' && c <= '9');
if (!alnum && c != '.' && c != '_' && c != '-')
return false;
}
return true;
}
DaemonConf* daemon_conf_create(void) {
DaemonConf* conf = calloc(1, sizeof(DaemonConf));
if (!conf)
return NULL;
conf->global.port = DAEMON_CONF_DEFAULT_PORT;
conf->global.max_connections = DAEMON_CONF_DEFAULT_MAX_CONNECTIONS;
conf->global.auth_failure_delay_ms = DAEMON_CONF_DEFAULT_AUTH_FAILURE_DELAY_MS;
conf->global.max_connections_per_host = DAEMON_CONF_DEFAULT_MAX_CONNECTIONS_PER_HOST;
conf->global.auth_lockout_threshold = DAEMON_CONF_DEFAULT_AUTH_LOCKOUT_THRESHOLD;
conf->global.auth_lockout_duration_sec = DAEMON_CONF_DEFAULT_AUTH_LOCKOUT_DURATION_SEC;
return conf;
}
/* Free a heap-owned pattern list of `count` entries. */
static void free_string_list(char** list, int count) {
for (int i = 0; i < count; i++)
free(list[i]);
free(list);
}
void daemon_conf_free(DaemonConf* conf) {
if (!conf)
return;
free(conf->global.motd_file);
free(conf->global.address);
free_string_list(conf->global.hosts_allow, conf->global.hosts_allow_count);
free_string_list(conf->global.hosts_deny, conf->global.hosts_deny_count);
for (int i = 0; i < conf->module_count; i++) {
DaemonModule* m = &conf->modules[i];
free(m->name);
free(m->path);
for (int j = 0; j < m->auth_user_count; j++)
free(m->auth_users[j]);
free(m->auth_users);
free_string_list(m->hosts_allow, m->hosts_allow_count);
free_string_list(m->hosts_deny, m->hosts_deny_count);
}
free(conf->modules);
free(conf);
}
const DaemonModule* daemon_conf_find_module(const DaemonConf* conf, const char* name) {
if (!conf || !name)
return NULL;
for (int i = 0; i < conf->module_count; i++) {
if (strcmp(conf->modules[i].name, name) == 0)
return &conf->modules[i];
}
return NULL;
}
/* Replace *slot with a str_dup of value; returns false on allocation failure. */
static bool store_string(char** slot, const char* value) {
char* dup = str_dup(value);
if (!dup)
return false;
free(*slot);
*slot = dup;
return true;
}
static bool store_port(int* slot, const char* value, char* err, size_t err_size) {
char* end;
errno = 0;
long p = strtol(value, &end, 10);
if (errno != 0 || *end != '\0' || *value == '\0' || p <= 0 || p > 65535) {
set_error(err, err_size, "invalid port '%s' (must be 1-65535)", value);
return false;
}
*slot = (int)p;
return true;
}
/* Apply a global scalar key/value. Keys are case-insensitive. Returns false
* (err filled) on an unknown key or an invalid value. */
static bool apply_global_key(DaemonConf* conf, char* key, const char* value, bool replace_hosts,
char* err, size_t err_size) {
if (key_equals(key, "port"))
return store_port(&conf->global.port, value, err, err_size);
if (key_equals(key, "motd file")) {
if (!store_string(&conf->global.motd_file, value)) {
set_error(err, err_size, "out of memory parsing 'motd file'");
return false;
}
return true;
}
if (key_equals(key, "address")) {
if (!store_string(&conf->global.address, value)) {
set_error(err, err_size, "out of memory parsing 'address'");
return false;
}
return true;
}
if (key_equals(key, "max connections"))
return store_max_connections(&conf->global.max_connections, value, NULL, err, err_size);
if (key_equals(key, "max connections per host"))
return store_optional_cap(&conf->global.max_connections_per_host, value,
DAEMON_CONF_MAX_CONCURRENCY_LIMIT, "max connections per host", NULL,
err, err_size);
if (key_equals(key, "auth failure delay"))
return store_auth_failure_delay(&conf->global.auth_failure_delay_ms, value, err, err_size);
if (key_equals(key, "auth lockout threshold"))
return store_optional_cap(&conf->global.auth_lockout_threshold, value,
DAEMON_CONF_MAX_CONCURRENCY_LIMIT, "auth lockout threshold", NULL,
err, err_size);
if (key_equals(key, "auth lockout duration"))
return store_optional_cap(&conf->global.auth_lockout_duration_sec, value,
DAEMON_CONF_MAX_AUTH_LOCKOUT_DURATION_SEC, "auth lockout duration",
NULL, err, err_size);
if (key_equals(key, "hosts allow"))
return store_host_list(&conf->global.hosts_allow, &conf->global.hosts_allow_count, value,
"hosts allow", NULL, replace_hosts, err, err_size);
if (key_equals(key, "hosts deny"))
return store_host_list(&conf->global.hosts_deny, &conf->global.hosts_deny_count, value,
"hosts deny", NULL, replace_hosts, err, err_size);
set_error(err, err_size, "unknown global key '%s'", key);
return false;
}
/* Apply a module key/value to the currently-open module. Returns false (err
* filled) on an unknown module key or an invalid value. */
static bool apply_module_key(DaemonModule* module, char* key, char* value, char* err,
size_t err_size) {
if (key_equals(key, "path")) {
if (*value == '\0') {
set_error(err, err_size, "module '%s': 'path' must not be empty", module->name);
return false;
}
if (!store_string(&module->path, value)) {
set_error(err, err_size, "out of memory parsing 'path' for module '%s'", module->name);
return false;
}
return true;
}
if (key_equals(key, "read only")) {
bool parsed;
if (!parse_bool_value(value, &parsed)) {
set_error(err, err_size,
"module '%s': 'read only' must be yes/no (or true/false/1/0), got '%s'",
module->name, value);
return false;
}
module->read_only = parsed;
return true;
}
if (key_equals(key, "client owner")) {
bool parsed;
if (!parse_bool_value(value, &parsed)) {
set_error(err, err_size,
"module '%s': 'client owner' must be yes/no (or true/false/1/0), got '%s'",
module->name, value);
return false;
}
module->client_owner = parsed;
return true;
}
if (key_equals(key, "auth users")) {
char* list = str_dup(value);
if (!list) {
set_error(err, err_size, "out of memory parsing 'auth users' for module '%s'", module->name);
return false;
}
char* save = NULL;
int added = 0;
for (char* token = strtok_r(list, ",", &save); token; token = strtok_r(NULL, ",", &save)) {
const char* user = trim_ws(token);
if (*user == '\0')
continue;
if (!credentials_username_valid(user)) {
set_error(err, err_size, "module '%s': invalid 'auth users' entry '%s'", module->name,
user);
free(list);
return false;
}
char** grown =
realloc(module->auth_users, (size_t)(module->auth_user_count + 1) * sizeof(char*));
if (!grown) {
free(list);
set_error(err, err_size, "out of memory parsing 'auth users' for module '%s'",
module->name);
return false;
}
module->auth_users = grown;
char* dup = str_dup(user);
if (!dup) {
free(list);
set_error(err, err_size, "out of memory parsing 'auth users' for module '%s'",
module->name);
return false;
}
module->auth_users[module->auth_user_count++] = dup;
added++;
}
free(list);
/* An empty/separator-only value must not silently disable authentication:
* the key's presence is an explicit request for an allow-list. */
if (added == 0) {
set_error(err, err_size, "module '%s': 'auth users' must list at least one user",
module->name);
return false;
}
return true;
}
if (key_equals(key, "max connections"))
return store_optional_cap(&module->max_connections, value, DAEMON_CONF_MAX_CONCURRENCY_LIMIT,
"max connections", module->name, err, err_size);
if (key_equals(key, "hosts allow"))
return store_host_list(&module->hosts_allow, &module->hosts_allow_count, value, "hosts allow",
module->name, false, err, err_size);
if (key_equals(key, "hosts deny"))
return store_host_list(&module->hosts_deny, &module->hosts_deny_count, value, "hosts deny",
module->name, false, err, err_size);
set_error(err, err_size, "unknown key '%s' in module '%s'", key, module->name);
return false;
}
static bool module_open_valid(const DaemonModule* module, char* err, size_t err_size) {
if (module->path == NULL) {
set_error(err, err_size, "module '%s' has no 'path'", module->name);
return false;
}
return true;
}
/* Validate a [section] header line body (text between the brackets) and set
* *name to the module name. Returns false on a malformed header. */
static bool parse_section_name(char* body, const char** name_out, char* err, size_t err_size) {
char* name = trim_ws(body);
if (!daemon_module_name_valid(name)) {
set_error(err, err_size, "invalid module name '%s' (must be 1-%d chars of [A-Za-z0-9._-])",
name, DAEMON_MAX_MODULE_NAME);
return false;
}
*name_out = name;
return true;
}
/* Open (or switch to) a module section. Closes any previously open module
* (validating it has a path) and appends the new one. */
static int open_module(DaemonConf* conf, int* current_module, const char* name, char* err,
size_t err_size) {
if (*current_module >= 0) {
if (!module_open_valid(&conf->modules[*current_module], err, err_size))
return -1;
}
if (daemon_conf_find_module(conf, name)) {
set_error(err, err_size, "duplicate module '%s'", name);
return -1;
}
if (conf->module_count >= DAEMON_CONF_MAX_MODULES) {
set_error(err, err_size, "too many modules (limit %d); module '%s' rejected",
DAEMON_CONF_MAX_MODULES, name);
return -1;
}
DaemonModule* grown =
realloc(conf->modules, (size_t)(conf->module_count + 1) * sizeof(DaemonModule));
if (!grown) {
set_error(err, err_size, "out of memory adding module '%s'", name);
return -1;
}
conf->modules = grown;
memset(&conf->modules[conf->module_count], 0, sizeof(DaemonModule));
conf->modules[conf->module_count].name = str_dup(name);
if (!conf->modules[conf->module_count].name) {
set_error(err, err_size, "out of memory adding module '%s'", name);
return -1;
}
conf->module_count++;
*current_module = conf->module_count - 1;
return 0;
}
/* Split a "key = value" line (value pointer returned in *value, pointing into
* line). Returns false when there is no '='. */
static bool split_key_value(char* line, char** key, char** value) {
char* eq = strchr(line, '=');
if (!eq)
return false;
*eq = '\0';
*key = trim_ws(line);
*value = trim_ws(eq + 1);
return true;
}
/* Strip one layer of surrounding double quotes from a trimmed value. A value
* that starts with '"' but does not end with '"' is an error. */
static bool unquote_value(char* value, char* err, size_t err_size) {
size_t len = strlen(value);
if (len == 0 || value[0] != '"')
return true;
if (len < 2 || value[len - 1] != '"') {
set_error(err, err_size, "unterminated quoted value");
return false;
}
memmove(value, value + 1, len - 2);
value[len - 2] = '\0';
return true;
}
DaemonConf* daemon_conf_load(const char* path, char* err, size_t err_size) {
if (err && err_size)
err[0] = '\0';
if (!path) {
set_error(err, err_size, "no daemon config path");
return NULL;
}
FILE* fp = fopen(path, "r");
if (!fp) {
set_error(err, err_size, "cannot open daemon config '%s': %s", path, strerror(errno));
return NULL;
}
DaemonConf* conf = daemon_conf_create();
if (!conf) {
fclose(fp);
set_error(err, err_size, "out of memory allocating daemon config");
return NULL;
}
int current_module = -1;
int line_no = 0;
char line[DAEMON_CONF_MAX_LINE + 2];
bool ok = true;
while (ok && fgets(line, sizeof(line), fp)) {
line_no++;
size_t len = strlen(line);
if (len == DAEMON_CONF_MAX_LINE + 1 && line[len - 1] != '\n') {
/* The read stopped at the buffer edge without a newline and there is
* more file to come: the line exceeds the bound. */
if (!feof(fp)) {
set_error(err, err_size, "line %d exceeds the %d-byte limit", line_no,
DAEMON_CONF_MAX_LINE);
ok = false;
break;
}
}
if (len > 0 && line[len - 1] == '\n')
line[--len] = '\0';
if (len > 0 && line[len - 1] == '\r')
line[--len] = '\0';
char* cursor = line;
while (*cursor == ' ' || *cursor == '\t')
cursor++;
if (*cursor == '\0' || *cursor == '#' || *cursor == ';')
continue; /* blank or comment line */
if (*cursor == '[') {
char* close = strchr(cursor, ']');
if (!close) {
set_error(err, err_size, "line %d: unterminated module header", line_no);
ok = false;
break;
}
*close = '\0';
char* trailing = close + 1;
const char* rest = trim_ws(trailing);
if (*rest != '\0') {
set_error(err, err_size, "line %d: unexpected text after module header", line_no);
ok = false;
break;
}
const char* name = NULL;
if (!parse_section_name(cursor + 1, &name, err, err_size)) {
ok = false;
break;
}
if (open_module(conf, &current_module, name, err, err_size) != 0) {
ok = false;
break;
}
continue;
}
char* key;
char* value;
if (!split_key_value(cursor, &key, &value)) {
set_error(err, err_size, "line %d: expected 'key = value'", line_no);
ok = false;
break;
}
if (*key == '\0') {
set_error(err, err_size, "line %d: empty key", line_no);
ok = false;
break;
}
if (!unquote_value(value, err, err_size)) {
ok = false;
break;
}
if (current_module >= 0) {
if (!apply_module_key(&conf->modules[current_module], key, value, err, err_size)) {
ok = false;
break;
}
} else {
if (!apply_global_key(conf, key, value, false, err, err_size)) {
ok = false;
break;
}
}
}
if (ok && ferror(fp)) {
set_error(err, err_size, "error reading daemon config '%s': %s", path, strerror(errno));
ok = false;
}
fclose(fp);
if (ok && current_module >= 0 &&
!module_open_valid(&conf->modules[current_module], err, err_size)) {
ok = false;
}
if (!ok) {
daemon_conf_free(conf);
return NULL;
}
return conf;
}
int daemon_conf_apply_dparam(DaemonConf* conf, const char* assignment, char* err, size_t err_size) {
if (err && err_size)
err[0] = '\0';
if (!conf || !assignment || *assignment == '\0') {
set_error(err, err_size, "--dparam requires a KEY=VALUE override");
return -1;
}
char* copy = str_dup(assignment);
if (!copy) {
set_error(err, err_size, "out of memory parsing --dparam");
return -1;
}
char* eq = strchr(copy, '=');
if (!eq) {
free(copy);
set_error(err, err_size, "--dparam '%s' has no '=' (expected KEY=VALUE)", assignment);
return -1;
}
*eq = '\0';
char* key = trim_ws(copy);
const char* value = trim_ws(eq + 1);
if (*key == '\0') {
free(copy);
set_error(err, err_size, "--dparam '%s' has an empty key", assignment);
return -1;
}
if (*value == '\0') {
free(copy);
set_error(err, err_size, "--dparam '%s' has an empty value", assignment);
return -1;
}
bool ok = apply_global_key(conf, key, value, true, err, err_size);
free(copy);
return ok ? 0 : -1;
}
/* Compare the first `prefix` bits of two 16-byte address buffers. */
static bool bit_prefix_match(const uint8_t* a, const uint8_t* b, int prefix) {
int whole = prefix / 8;
if (whole > 0 && memcmp(a, b, (size_t)whole) != 0)
return false;
int remainder = prefix % 8;
if (remainder == 0)
return true;
uint8_t mask = (uint8_t)(0xffu << (8 - remainder));
return (a[whole] & mask) == (b[whole] & mask);
}
/* Case-insensitive glob match used for hostname patterns. Falls back to the
* shared case-sensitive matcher when an operand is too long for the stack
* buffers. */
static bool host_glob_match(const char* pattern, const char* str) {
char pbuf[256];
char sbuf[256];
size_t plen = strlen(pattern);
size_t slen = strlen(str);
if (plen >= sizeof(pbuf) || slen >= sizeof(sbuf))
return glob_match(pattern, str);
for (size_t i = 0; i <= plen; i++)
pbuf[i] = (char)tolower((unsigned char)pattern[i]);
for (size_t i = 0; i <= slen; i++)
sbuf[i] = (char)tolower((unsigned char)str[i]);
return glob_match(pbuf, sbuf);
}
bool daemon_host_pattern_match(const char* pattern, const char* peer_ip) {
if (!pattern || *pattern == '\0' || !peer_ip || *peer_ip == '\0')
return false;
if (strcmp(pattern, "*") == 0)
return true;
if (strchr(pattern, '/')) {
uint8_t pattern_bytes[16];
uint8_t peer_bytes[16];
int prefix = 0;
int family = AF_UNSPEC;
if (!parse_cidr(pattern, &prefix, pattern_bytes, &family))
return false;
if (inet_pton(family, peer_ip, peer_bytes) != 1)
return false;
return bit_prefix_match(pattern_bytes, peer_bytes, prefix);
}
struct in_addr pattern_v4;
struct in_addr peer_v4;
if (inet_pton(AF_INET, pattern, &pattern_v4) == 1)
return inet_pton(AF_INET, peer_ip, &peer_v4) == 1 && pattern_v4.s_addr == peer_v4.s_addr;
struct in6_addr pattern_v6;
struct in6_addr peer_v6;
if (inet_pton(AF_INET6, pattern, &pattern_v6) == 1)
return inet_pton(AF_INET6, peer_ip, &peer_v6) == 1 &&
memcmp(&pattern_v6, &peer_v6, sizeof(pattern_v6)) == 0;
/* Not a literal: a hostname/glob pattern. */
return host_glob_match(pattern, peer_ip);
}
bool daemon_hosts_allowed(const char* peer_ip, char* const* allow, int allow_count,
char* const* deny, int deny_count) {
if (!peer_ip)
return false;
for (int i = 0; i < deny_count; i++) {
if (daemon_host_pattern_match(deny[i], peer_ip))
return false;
}
if (allow_count > 0) {
for (int i = 0; i < allow_count; i++) {
if (daemon_host_pattern_match(allow[i], peer_ip))
return true;
}
return false;
}
return true;
}
bool daemon_hosts_restricted(char* const* allow, int allow_count, char* const* deny,
int deny_count) {
(void)allow;
(void)deny;
return allow_count > 0 || deny_count > 0;
}
+182
View File
@@ -0,0 +1,182 @@
#ifndef DAEMON_CONF_H
#define DAEMON_CONF_H
#include <stdbool.h>
#include <stddef.h>
/* FastSync-native daemon configuration (a FastSync analog of rsyncd.conf).
*
* This is the config the fastsync-server --daemon listener consumes. It is
* line-based with an implicit global section followed by zero or more
* [module] sections. The full grammar is documented in RSYNC_COMPAT.md
* ("Daemon Mode") and summarized below; the parser lives entirely in
* daemon_conf.c so it can be unit tested without any socket code.
*
* The parser is STRICT: an unknown key, a malformed line, a value that does
* not parse, a module without a `path`, or a line longer than
* DAEMON_CONF_MAX_LINE all fail the whole load with a clear, line-numbered
* error instead of being silently ignored. This keeps a typo from silently
* changing what a module serves.
*/
/* A daemon module's configured root is used exactly like the standalone
* server's --destination-root: the daemon confines every connection that
* selects this module to this path (file_open_secure_parent /
* has_path_traversal / path_is_within all keep the existing confinement, just
* per-module). There is never any client-chosen root: a module path always
* stays confined. A daemon REFUSES every client-chosen ownership / super-user
* request by default -- --numeric-ids, --chown, --usermap/--groupmap,
* --fake-super, --copy-as and an explicit --super -- because there is no
* per-module opt-in unless the operator adds one. An operator opts a single
* module in with `client owner = yes` (DaemonModule.client_owner), which allows
* that client to choose ownership within that module's root (the standalone/SSH
* server honors such requests for its single operator-authorized root). The
* operator-level --no-super veto additionally forces super-user activities off
* for every daemon connection, even an opted-in module. See server_module_gate
* in server.c and RSYNC_COMPAT.md.
*
* `auth_users` is honored by Wave B daemon authentication: a module that
* declares auth users accepts a connection only when the presented username is
* on this list AND verifies against the daemon's credential store
* (--password-file / --early-input). An auth-required module with no usable
* store refuses (fail closed) rather than falling open; see server.c. Auth is
* never bypassed by ignoring the list. */
typedef struct DaemonModule {
char* name; /* module name, as the client requests it */
char* path; /* module root (daemon-side authorized root) */
bool read_only; /* `read only = yes/no`; default no */
bool client_owner; /* `client owner = yes/no`; default no. Per-module opt-in
that lets this module's clients choose ownership
(--numeric-ids/--chown/--usermap/--groupmap/--fake-super/
--copy-as) and request explicit --super super-user
activities. Without it the daemon refuses all of them. */
char** auth_users; /* `auth users = a,b`; Wave B credential list */
int auth_user_count;
/* `max connections = N` (optional per-module cap). 0 means unlimited. The
* per-connection child records the selected module in the shared registry
* (daemon_limits.c) once the config frame names it, so the cap is enforced
* across all forked children; the parent reclaims the slot on SIGCHLD. */
int max_connections;
char** hosts_allow; /* `hosts allow = a,b`; host access allow patterns */
int hosts_allow_count;
char** hosts_deny; /* `hosts deny = a,b`; host access deny patterns */
int hosts_deny_count;
} DaemonModule;
/* Global (pre-module) scalar keys. `motd file` is parsed and stored but has
* no wire effect yet (MOTD display is Wave C). */
typedef struct DaemonConfGlobals {
int port; /* `port`, default DAEMON_CONF_DEFAULT_PORT (873) */
char* motd_file; /* `motd file`, may be NULL */
char* address; /* `address` (optional bind address), may be NULL */
int max_connections; /* `max connections`, default
DAEMON_CONF_DEFAULT_MAX_CONNECTIONS (100) */
int auth_failure_delay_ms; /* `auth failure delay`, milliseconds; default
DAEMON_CONF_DEFAULT_AUTH_FAILURE_DELAY_MS */
int max_connections_per_host; /* `max connections per host`, concurrent cap per
source IP; default
DAEMON_CONF_DEFAULT_MAX_CONNECTIONS_PER_HOST (0 =
unlimited) */
int auth_lockout_threshold; /* `auth lockout threshold`, failed attempts from
one source before lockout; default
DAEMON_CONF_DEFAULT_AUTH_LOCKOUT_THRESHOLD (0
disables) */
int auth_lockout_duration_sec; /* `auth lockout duration`, seconds; default
DAEMON_CONF_DEFAULT_AUTH_LOCKOUT_DURATION_SEC
(0 disables) */
char** hosts_allow; /* `hosts allow`; global host access allow patterns */
int hosts_allow_count;
char** hosts_deny; /* `hosts deny`; global host access deny patterns */
int hosts_deny_count;
} DaemonConfGlobals;
typedef struct DaemonConf {
DaemonConfGlobals global;
DaemonModule* modules;
int module_count;
} DaemonConf;
#define DAEMON_CONF_DEFAULT_PORT 873
/* Default global connection cap when `max connections` is absent. Matches the
* historical hardcoded listener value. */
#define DAEMON_CONF_DEFAULT_MAX_CONNECTIONS 100
/* Default `auth failure delay` in milliseconds (0 disables the throttle). */
#define DAEMON_CONF_DEFAULT_AUTH_FAILURE_DELAY_MS 500
/* Default `max connections per host` (0 = unlimited). */
#define DAEMON_CONF_DEFAULT_MAX_CONNECTIONS_PER_HOST 0
/* Default cross-process auth lockout: 10 failed attempts from one source lock
* it out for 300 s (0 disables either knob). */
#define DAEMON_CONF_DEFAULT_AUTH_LOCKOUT_THRESHOLD 10
#define DAEMON_CONF_DEFAULT_AUTH_LOCKOUT_DURATION_SEC 300
/* Upper bound on a `max connections per host` or `auth lockout threshold`
* value, so a typo cannot size the shared registry absurdly. */
#define DAEMON_CONF_MAX_CONCURRENCY_LIMIT 1000000
/* Upper bound on `auth lockout duration` (7 days). */
#define DAEMON_CONF_MAX_AUTH_LOCKOUT_DURATION_SEC 604800
/* Largest accepted `auth failure delay`, so a typo cannot pin a connection
* child in nanosleep for an absurd time. */
/* Bounded well below the socket I/O timeout so a failed-auth child cannot hold
* a connection slot for long enough to amplify connection-cap exhaustion. */
#define DAEMON_CONF_MAX_AUTH_FAILURE_DELAY_MS 5000
/* Upper bound on the number of [module] sections, so the shared registry's
* per-module counter array stays fixed-size. The parser rejects the next
* section past this bound. */
#define DAEMON_CONF_MAX_MODULES 256
/* Longest accepted config line (excluding the trailing newline). Longer lines
* are rejected rather than buffered unboundedly. */
#define DAEMON_CONF_MAX_LINE 4096
/* Upper bound on a module name. Kept far below MAX_STRING_SIZE so a wire
* module name can never exhaust anything by being long. */
#define DAEMON_MAX_MODULE_NAME 200
/* Allocate an empty daemon config with defaulted globals (port 873, no
* modules, no motd/address). Never fails for an allocation failure; callers
* must still NULL-check. */
DaemonConf* daemon_conf_create(void);
/* Parse `path` into a freshly allocated DaemonConf. Returns NULL on any error
* and fills `err` (err_size bytes) with a clear, line-numbered message. The
* returned object is heap-owned; free it with daemon_conf_free. */
DaemonConf* daemon_conf_load(const char* path, char* err, size_t err_size);
void daemon_conf_free(DaemonConf* conf);
/* Case-sensitive exact module lookup by name. Returns the module or NULL.
* Module names are matched exactly (rsync semantics). */
const DaemonModule* daemon_conf_find_module(const DaemonConf* conf, const char* name);
/* Module-name syntax check: non-empty, at most DAEMON_MAX_MODULE_NAME chars,
* and only [A-Za-z0-9._-]. Used by the config parser, the client's
* host::module/path destination parser, and (implicitly) by the daemon lookup
* (a name that fails this can never match a parsed module). */
bool daemon_module_name_valid(const char* name);
/* Parse one --dparam=KEY=VALUE (or "--dparam KEY=VALUE") override string and
* apply it to the global keys only. Keys are case-insensitive and limited to
* the global keys defined by the grammar (port, motd file, address,
* max connections, max connections per host, auth failure delay,
* auth lockout threshold, auth lockout duration, hosts allow, hosts deny).
* Returns 0 on success, -1 on error (err filled). */
int daemon_conf_apply_dparam(DaemonConf* conf, const char* assignment, char* err, size_t err_size);
/* Host access-control matching (pure; no I/O). `daemon_host_pattern_match`
* matches one configured pattern against a numeric peer IP string. Supported
* patterns: `*` (match anything), an IPv4/IPv6 literal, an IPv4/IPv6 CIDR
* (`10.0.0.0/8`, `2001:db8::/32`), or a glob (`*.example.com`) evaluated with
* the same matcher as file globs; a glob only matches a peer string of the
* same shape, so a numeric peer never matches a hostname glob. */
bool daemon_host_pattern_match(const char* pattern, const char* peer_ip);
/* rsync-like combined decision over a deny list and an allow list: a matching
* deny rejects (deny takes precedence); otherwise, when any allow entries
* exist, a peer that matches none is rejected; with no allow entries every
* peer not denied is accepted. An empty/unset pair returns true. */
bool daemon_hosts_allowed(const char* peer_ip, char* const* allow, int allow_count,
char* const* deny, int deny_count);
/* True when at least one allow or deny pattern is configured (i.e. an
* unprovable peer must fail closed rather than being treated as unrestricted). */
bool daemon_hosts_restricted(char* const* allow, int allow_count, char* const* deny,
int deny_count);
#endif
+494
View File
@@ -0,0 +1,494 @@
#include "daemon_limits.h"
#include "daemon_conf.h"
#include "log.h"
#include <arpa/inet.h>
#include <netinet/in.h>
#include <stdatomic.h>
#include <stdint.h>
#include <stdlib.h>
#include <string.h>
#include <sys/mman.h>
#include <time.h>
/* The two module-count bounds must agree: the daemon config parser never
* produces more than DAEMON_CONF_MAX_MODULES modules, so the shared registry's
* per-module counter array is sized from the same bound. */
_Static_assert(DAEMON_LIMITS_MAX_MODULES == DAEMON_CONF_MAX_MODULES,
"daemon_limits module bound must match daemon_conf");
/* Slot lifecycle states (stored in slot_state). */
enum {
SLOT_FREE = 0,
SLOT_CLAIMED = 1,
SLOT_REGISTERED = 2,
};
/* The registry header lives at the base of the shared mapping; the pointer
* fields point at the arrays carved out of the same mapping. Absolute pointers
* remain valid in a forked child because fork() clones the address space and
* mapping, so parent and child observe the same virtual addresses. */
struct DaemonLimitRegistry {
int max_slots;
int module_count;
int host_slots; /* power of two; 1 when no per-source tracking is needed */
int per_host_cap;
int lockout_threshold;
int lockout_duration_sec;
size_t map_size;
_Atomic long long host_full_warn; /* last "table full" warning epoch */
_Atomic int* slot_state;
_Atomic int* slot_pid;
_Atomic int* slot_module;
_Atomic int* slot_host; /* per-source table bucket, or -1 */
_Atomic int* module_active;
_Atomic uint64_t* host_key; /* 0 == empty bucket */
_Atomic int* host_active;
_Atomic int* host_fail;
_Atomic long long* host_until; /* epoch seconds the lockout expires */
_Atomic long long* host_last_use; /* epoch seconds the bucket was last touched */
};
static size_t round_up(size_t n, size_t align) {
return (n + align - 1) & ~(align - 1);
}
static size_t next_pow2(size_t n) {
size_t p = 1;
while (p < n)
p <<= 1;
return p;
}
/* Parse a numeric IPv4/IPv6 peer string into family + raw bytes. */
static bool parse_peer_ip(const char* peer_ip, int* family, unsigned char* bytes) {
if (!peer_ip || *peer_ip == '\0')
return false;
struct in_addr v4;
if (inet_pton(AF_INET, peer_ip, &v4) == 1) {
memcpy(bytes, &v4, sizeof(v4));
*family = AF_INET;
return true;
}
struct in6_addr v6;
if (inet_pton(AF_INET6, peer_ip, &v6) == 1) {
memcpy(bytes, &v6, sizeof(v6));
*family = AF_INET6;
return true;
}
return false;
}
uint64_t daemon_limits_host_hash(const char* peer_ip, bool* ok) {
if (ok)
*ok = false;
unsigned char bytes[16];
int family = AF_UNSPEC;
if (!parse_peer_ip(peer_ip, &family, bytes))
return 0;
uint64_t hash = 14695981039346656037ULL ^ (uint64_t)(uint32_t)family;
size_t length = family == AF_INET ? 4 : 16;
for (size_t i = 0; i < length; i++) {
hash ^= bytes[i];
hash *= 1099511628211ULL;
}
if (hash == 0)
hash = 0x9e3779b97f4a7c15ULL;
if (ok)
*ok = true;
return hash;
}
/* True when the registry must maintain per-source buckets: either the per-host
* cap is configured, or the auth lockout is (threshold AND duration > 0). A
* lockout threshold without a duration is a no-op, so it must not size or intern
* the table. create(), register() and the lockout paths all agree on this. */
static bool registry_tracks_hosts(const DaemonLimitRegistry* registry) {
return registry->per_host_cap > 0 ||
(registry->lockout_threshold > 0 && registry->lockout_duration_sec > 0);
}
/* Find the bucket holding `peer_ip`, or -1 when it has no entry. Finding a
* bucket refreshes its last-use time so the eviction policy sees it as live. */
static int host_lookup(DaemonLimitRegistry* registry, const char* peer_ip) {
bool ok = false;
uint64_t key = daemon_limits_host_hash(peer_ip, &ok);
if (!ok)
return -1;
size_t mask = (size_t)registry->host_slots - 1;
size_t start = (size_t)(key & mask);
for (size_t i = 0; i < (size_t)registry->host_slots; i++) {
size_t idx = (start + i) & mask;
uint64_t current = atomic_load_explicit(&registry->host_key[idx], memory_order_acquire);
if (current == key) {
atomic_store_explicit(&registry->host_last_use[idx], (long long)time(NULL),
memory_order_relaxed);
return (int)idx;
}
if (current == 0)
return -1; /* no tombstones: an empty bucket ends the probe chain */
}
return -1;
}
/* A bucket with no live connection may be repurposed: immediately when its
* lockout deadline has already passed (the review's "expired" case), or after an
* idle window when it holds no pending lockout. A bucket with a future lockout
* deadline is retained so the lockout actually lasts its configured duration. */
static bool host_bucket_reclaimable(DaemonLimitRegistry* registry, size_t idx, long long now) {
if (atomic_load_explicit(&registry->host_active[idx], memory_order_relaxed) != 0)
return false;
long long until = atomic_load_explicit(&registry->host_until[idx], memory_order_relaxed);
if (until != 0)
return until <= now;
long long last_use = atomic_load_explicit(&registry->host_last_use[idx], memory_order_relaxed);
/* A bucket whose key is published but whose last_use has not yet been stamped
* (last_use == 0) must be treated as live: reclaiming it here would steal a
* bucket a racing child just claimed. The claim path also stamps last_use
* before publishing the key, so this window cannot persist. */
return last_use != 0 && now - last_use >= DAEMON_LIMITS_HOST_EVICT_IDLE_SEC;
}
/* Emit at most one "per-source table full" warning per
* DAEMON_LIMITS_HOST_FULL_WARN_SEC across all forked children. Called from a
* normal (non-signal) child path, so logging is safe here. */
static void host_warn_table_full(DaemonLimitRegistry* registry, long long now) {
long long last = atomic_load_explicit(&registry->host_full_warn, memory_order_relaxed);
if (last != 0 && now - last < DAEMON_LIMITS_HOST_FULL_WARN_SEC)
return;
if (atomic_compare_exchange_strong_explicit(&registry->host_full_warn, &last, now,
memory_order_relaxed, memory_order_relaxed)) {
log_message(LOG_LEVEL_WARNING,
"daemon: per-source registry is full (%d slots) and no bucket can be reclaimed; "
"'max connections per host' and the auth lockout are temporarily not enforced for "
"new sources (the per-module cap and host ACLs still apply)",
registry->host_slots);
}
}
/* Find or insert the bucket for `peer_ip`. Insertion is a lock-free CAS so two
* forked children racing on the same source converge on one bucket.
*
* When the probe finds no empty bucket it reclaims, via a key CAS, the first
* bucket that is reclaimable (expired lockout or idle, and no active
* connection) and resets its counters. This bounds the table's lifetime so it
* cannot fill permanently and stay fail-open. Returns -1 only when the address
* is unparseable or the table is genuinely full of live/locked buckets
* (callers fail open: the global/module caps and ACLs still apply). */
static int host_intern(DaemonLimitRegistry* registry, const char* peer_ip) {
bool ok = false;
uint64_t key = daemon_limits_host_hash(peer_ip, &ok);
if (!ok)
return -1;
long long now = (long long)time(NULL);
size_t mask = (size_t)registry->host_slots - 1;
size_t start = (size_t)(key & mask);
/* A couple of passes bound the work: the first normally claims/seeds a bucket;
* a lost eviction CAS retries once against the freshly observed table. */
for (int pass = 0; pass < 2; pass++) {
int evict = -1;
uint64_t evict_key = 0;
for (size_t i = 0; i < (size_t)registry->host_slots; i++) {
size_t idx = (start + i) & mask;
uint64_t current = atomic_load_explicit(&registry->host_key[idx], memory_order_acquire);
if (current == key) {
atomic_store_explicit(&registry->host_last_use[idx], now, memory_order_relaxed);
return (int)idx;
}
if (current == 0) {
/* Stamp last_use *before* publishing the key so a reclaimer racing the
* claim can never observe a claimed bucket with last_use == 0 and
* evict it. A pre-stamp is harmless if the CAS loses: the bucket is
* either still empty (never inspected for reclaim) or has just been
* taken by another source that wants a fresh timestamp anyway. */
atomic_store_explicit(&registry->host_last_use[idx], now, memory_order_relaxed);
uint64_t expected = 0;
if (atomic_compare_exchange_strong_explicit(&registry->host_key[idx], &expected, key,
memory_order_acq_rel, memory_order_acquire)) {
return (int)idx;
}
if (atomic_load_explicit(&registry->host_key[idx], memory_order_acquire) == key) {
return (int)idx;
}
continue; /* another child won this empty bucket; keep probing */
}
if (evict < 0 && host_bucket_reclaimable(registry, idx, now)) {
evict = (int)idx;
evict_key = current;
}
}
if (evict >= 0) {
/* Refresh the timestamp before the key changes hands so the reused bucket
* is not seen as immediately idle by a racing reclaimer. */
atomic_store_explicit(&registry->host_last_use[evict], now, memory_order_relaxed);
uint64_t expected = evict_key;
if (atomic_compare_exchange_strong_explicit(&registry->host_key[evict], &expected, key,
memory_order_acq_rel, memory_order_acquire)) {
/* The bucket now belongs to the new source; clear the evicted source's
* stale lockout/failure state. */
atomic_store_explicit(&registry->host_active[evict], 0, memory_order_relaxed);
atomic_store_explicit(&registry->host_fail[evict], 0, memory_order_relaxed);
atomic_store_explicit(&registry->host_until[evict], 0, memory_order_relaxed);
/* Two children can race to intern the same brand-new key into different
* eviction targets, leaving the table with duplicate buckets for `key`.
* Re-scan for the first (canonical) bucket holding `key`; when it
* precedes `evict`, drop our duplicate's occupancy and hand back the
* canonical bucket so per-source counts are not orphaned on the
* duplicate. The duplicate keeps its key, so no tombstone hole is
* created and probe chains stay intact; it ages out normally. */
for (size_t i = 0; i < (size_t)registry->host_slots; i++) {
size_t candidate = (start + i) & mask;
uint64_t found =
atomic_load_explicit(&registry->host_key[candidate], memory_order_acquire);
if (found == key) {
if (candidate != (size_t)evict) {
atomic_store_explicit(&registry->host_active[evict], 0, memory_order_relaxed);
return (int)candidate;
}
break;
}
if (found == 0)
break; /* the key is present at `evict`, so this cannot happen first */
}
return evict;
}
continue; /* lost the race; re-probe with fresh observations */
}
break; /* no free and no reclaimable bucket: genuinely full */
}
host_warn_table_full(registry, now);
return -1;
}
DaemonLimitRegistry* daemon_limits_create(int max_slots, int module_count, int per_host_cap,
int lockout_threshold, int lockout_duration_sec) {
if (max_slots < DAEMON_LIMITS_MIN_SLOTS)
max_slots = DAEMON_LIMITS_MIN_SLOTS;
if (max_slots > DAEMON_LIMITS_MAX_SLOTS)
max_slots = DAEMON_LIMITS_MAX_SLOTS;
if (module_count < 1)
module_count = 1;
if (module_count > DAEMON_LIMITS_MAX_MODULES)
module_count = DAEMON_LIMITS_MAX_MODULES;
if (per_host_cap < 0)
per_host_cap = 0;
if (lockout_threshold < 0)
lockout_threshold = 0;
if (lockout_duration_sec < 0)
lockout_duration_sec = 0;
bool need_hosts = per_host_cap > 0 || (lockout_threshold > 0 && lockout_duration_sec > 0);
int host_slots = 1;
if (need_hosts) {
size_t want = (size_t)max_slots * 4;
if (want < 64)
want = 64;
if (want > DAEMON_LIMITS_MAX_HOST_SLOTS)
want = DAEMON_LIMITS_MAX_HOST_SLOTS;
host_slots = (int)next_pow2(want);
}
size_t header = round_up(sizeof(DaemonLimitRegistry), 16);
size_t slot_bytes =
round_up((size_t)max_slots * sizeof(_Atomic int), 16) * 4; /* state,pid,module,host */
size_t module_bytes = round_up((size_t)module_count * sizeof(_Atomic int), 16);
size_t host_key_bytes = round_up((size_t)host_slots * sizeof(_Atomic uint64_t), 16);
size_t host_int_bytes = round_up((size_t)host_slots * sizeof(_Atomic int), 16) * 2;
size_t host_until_bytes = round_up((size_t)host_slots * sizeof(_Atomic long long), 16) * 2;
size_t total =
header + slot_bytes + module_bytes + host_key_bytes + host_int_bytes + host_until_bytes + 16;
void* map = mmap(NULL, total, PROT_READ | PROT_WRITE, MAP_SHARED | MAP_ANONYMOUS, -1, 0);
if (map == MAP_FAILED)
return NULL;
memset(map, 0, total);
DaemonLimitRegistry* registry = (DaemonLimitRegistry*)map;
registry->max_slots = max_slots;
registry->module_count = module_count;
registry->host_slots = host_slots;
registry->per_host_cap = per_host_cap;
registry->lockout_threshold = lockout_threshold;
registry->lockout_duration_sec = lockout_duration_sec;
registry->map_size = total;
unsigned char* cursor = (unsigned char*)map + header;
registry->slot_state = (atomic_int*)cursor;
cursor += (size_t)max_slots * sizeof(_Atomic int);
registry->slot_pid = (atomic_int*)cursor;
cursor += (size_t)max_slots * sizeof(_Atomic int);
registry->slot_module = (atomic_int*)cursor;
cursor += (size_t)max_slots * sizeof(_Atomic int);
registry->slot_host = (atomic_int*)cursor;
cursor += (size_t)max_slots * sizeof(_Atomic int);
registry->module_active = (atomic_int*)cursor;
cursor += (size_t)module_count * sizeof(_Atomic int);
cursor = (unsigned char*)round_up((size_t)(uintptr_t)cursor, 16);
registry->host_key = (_Atomic uint64_t*)cursor;
cursor += (size_t)host_slots * sizeof(_Atomic uint64_t);
registry->host_active = (atomic_int*)cursor;
cursor += (size_t)host_slots * sizeof(_Atomic int);
registry->host_fail = (atomic_int*)cursor;
cursor += (size_t)host_slots * sizeof(_Atomic int);
cursor = (unsigned char*)round_up((size_t)(uintptr_t)cursor, 16);
registry->host_until = (atomic_llong*)cursor;
cursor += (size_t)host_slots * sizeof(_Atomic long long);
registry->host_last_use = (atomic_llong*)cursor;
for (int i = 0; i < max_slots; i++) {
atomic_store(&registry->slot_module[i], -1);
atomic_store(&registry->slot_host[i], -1);
}
return registry;
}
void daemon_limits_destroy(DaemonLimitRegistry* registry) {
if (!registry)
return;
munmap(registry, registry->map_size);
}
int daemon_limits_claim_slot(DaemonLimitRegistry* registry) {
if (!registry)
return DAEMON_LIMITS_NO_SLOT;
for (int i = 0; i < registry->max_slots; i++) {
int expected = SLOT_FREE;
if (atomic_compare_exchange_strong(&registry->slot_state[i], &expected, SLOT_CLAIMED)) {
atomic_store(&registry->slot_pid[i], 0);
atomic_store(&registry->slot_module[i], -1);
atomic_store(&registry->slot_host[i], -1);
return i;
}
}
return DAEMON_LIMITS_NO_SLOT;
}
void daemon_limits_set_slot_pid(DaemonLimitRegistry* registry, int slot, long pid) {
if (!registry || slot < 0 || slot >= registry->max_slots)
return;
atomic_store(&registry->slot_pid[slot], (int)pid);
}
void daemon_limits_reclaim_slot(DaemonLimitRegistry* registry, int slot) {
if (!registry || slot < 0 || slot >= registry->max_slots)
return;
atomic_exchange_explicit(&registry->slot_state[slot], SLOT_FREE, memory_order_acq_rel);
atomic_store_explicit(&registry->slot_pid[slot], 0, memory_order_relaxed);
/* The module/host occupancy arrays are derived from the slot table; do not
* decrement here or a SIGKILL between a child's increment and its REGISTERED
* publish would leak a count. Callers that need the derived counts call
* daemon_limits_recompute. */
}
void daemon_limits_reclaim_pid(DaemonLimitRegistry* registry, long pid) {
if (!registry || pid <= 0)
return;
for (int i = 0; i < registry->max_slots; i++) {
if (atomic_load(&registry->slot_state[i]) == SLOT_FREE)
continue;
if (atomic_load(&registry->slot_pid[i]) == (int)pid) {
daemon_limits_reclaim_slot(registry, i);
return;
}
}
}
void daemon_limits_recompute(DaemonLimitRegistry* registry) {
if (!registry)
return;
/* Zero the derived arrays, then re-derive solely from the REGISTERED slots.
* A child that was SIGKILLed after incrementing a counter but before
* publishing REGISTERED is not counted, and its leaked increment is erased by
* the zeroing, so the leak cannot persist. */
for (int m = 0; m < registry->module_count; m++)
atomic_store_explicit(&registry->module_active[m], 0, memory_order_relaxed);
for (int h = 0; h < registry->host_slots; h++)
atomic_store_explicit(&registry->host_active[h], 0, memory_order_relaxed);
for (int i = 0; i < registry->max_slots; i++) {
if (atomic_load_explicit(&registry->slot_state[i], memory_order_acquire) != SLOT_REGISTERED)
continue;
int module = atomic_load_explicit(&registry->slot_module[i], memory_order_relaxed);
if (module >= 0 && module < registry->module_count)
atomic_fetch_add_explicit(&registry->module_active[module], 1, memory_order_relaxed);
int host = atomic_load_explicit(&registry->slot_host[i], memory_order_relaxed);
if (host >= 0 && host < registry->host_slots)
atomic_fetch_add_explicit(&registry->host_active[host], 1, memory_order_relaxed);
}
}
DaemonLimitResult daemon_limits_register(DaemonLimitRegistry* registry, int slot, int module_index,
const char* peer_ip, int module_cap) {
if (!registry || slot < 0 || slot >= registry->max_slots)
return DAEMON_LIMIT_UNAVAILABLE;
if (module_index < 0 || module_index >= registry->module_count)
return DAEMON_LIMIT_UNAVAILABLE;
if (atomic_load_explicit(&registry->slot_state[slot], memory_order_acquire) != SLOT_CLAIMED)
return DAEMON_LIMIT_UNAVAILABLE;
int host = -1;
if (registry_tracks_hosts(registry))
host = host_intern(registry, peer_ip);
int module_count = atomic_fetch_add(&registry->module_active[module_index], 1) + 1;
if (module_cap > 0 && module_count > module_cap) {
atomic_fetch_sub(&registry->module_active[module_index], 1);
return DAEMON_LIMIT_MODULE_FULL;
}
if (host >= 0) {
int host_count = atomic_fetch_add(&registry->host_active[host], 1) + 1;
if (registry->per_host_cap > 0 && host_count > registry->per_host_cap) {
atomic_fetch_sub(&registry->host_active[host], 1);
atomic_fetch_sub(&registry->module_active[module_index], 1);
return DAEMON_LIMIT_HOST_FULL;
}
}
atomic_store(&registry->slot_module[slot], module_index);
atomic_store(&registry->slot_host[slot], host);
atomic_store_explicit(&registry->slot_state[slot], SLOT_REGISTERED, memory_order_release);
return DAEMON_LIMIT_OK;
}
bool daemon_limits_auth_locked(DaemonLimitRegistry* registry, const char* peer_ip,
int* seconds_remaining) {
if (!registry || registry->lockout_threshold <= 0 || registry->lockout_duration_sec <= 0)
return false;
int bucket = host_lookup(registry, peer_ip);
if (bucket < 0)
return false;
long long until = atomic_load(&registry->host_until[bucket]);
long long now = (long long)time(NULL);
if (until > now) {
if (seconds_remaining)
*seconds_remaining = (int)(until - now);
return true;
}
if (until != 0) {
/* The previous lockout has expired: clear the stale counter so the source
* gets a fresh allowance. */
atomic_store(&registry->host_fail[bucket], 0);
atomic_store(&registry->host_until[bucket], 0);
}
return false;
}
void daemon_limits_auth_record_failure(DaemonLimitRegistry* registry, const char* peer_ip) {
if (!registry || registry->lockout_threshold <= 0 || registry->lockout_duration_sec <= 0)
return;
int bucket = host_intern(registry, peer_ip);
if (bucket < 0)
return;
int failures = atomic_fetch_add(&registry->host_fail[bucket], 1) + 1;
if (failures >= registry->lockout_threshold) {
long long now = (long long)time(NULL);
atomic_store(&registry->host_until[bucket], now + (long long)registry->lockout_duration_sec);
}
}
void daemon_limits_auth_record_success(DaemonLimitRegistry* registry, const char* peer_ip) {
if (!registry)
return;
int bucket = host_lookup(registry, peer_ip);
if (bucket < 0)
return;
atomic_store(&registry->host_fail[bucket], 0);
atomic_store(&registry->host_until[bucket], 0);
}
+147
View File
@@ -0,0 +1,147 @@
#ifndef DAEMON_LIMITS_H
#define DAEMON_LIMITS_H
#include <stdbool.h>
#include <stddef.h>
#include <stdint.h>
/* Cross-process daemon connection registry.
*
* The daemon listener forks ONE child per accepted connection, so any
* per-module / per-source accounting must live in state shared across the
* forked children. This module owns a fixed-size registry carved out of an
* anonymous shared mapping (mmap(MAP_SHARED | MAP_ANONYMOUS)) created by the
* accept-loop PARENT before it forks; every child inherits the mapping (and the
* pointer to it) across fork().
*
* Rules:
* - ONLY C11 atomics (atomic_*); never mtx_t/pthread locks, which can deadlock
* in a forked child if another thread held them at fork time.
* - No heap allocation after fork: the mapping is fixed-size and all access is
* atomic load/store/CAS over preallocated arrays.
*
* Slot lifecycle (the parent reclaims even when a child is SIGKILLed):
* FREE --(parent claim_slot)--> CLAIMED
* CLAIMED --(child register)--> REGISTERED
* any --(parent reclaim)--> FREE
* The child records its module index and per-source bucket into the slot before
* publishing REGISTERED; the parent's SIGCHLD handler matches the reaped pid to
* the slot and, when REGISTERED, decrements the module/per-source counters.
* A child killed before registering holds no counts, so reclaiming a CLAIMED
* slot only frees the slot.
*
* Per-source identity is the normalized numeric peer IP (IPv4-mapped IPv6 is
* already collapsed to IPv4 by utils_fd_peer_ip); it is interned into an
* open-addressed, linear-probing table keyed by a 64-bit hash. The same table
* also carries the cross-process auth-failure counter and lockout deadline.
*
* Per-source table lifetime: a bucket's key is never cleared back to empty (that
* would break every later probe chain that passed through it). Instead the
* table has a bounded-lifetime eviction policy: when no empty bucket exists, the
* first bucket that is reclaimable -- no active connection AND (its lockout
* deadline has passed OR it has been idle for
* DAEMON_LIMITS_HOST_EVICT_IDLE_SEC) -- is atomically repurposed for the new
* source via a CAS of its key, and its counters are reset. The table therefore
* cannot fill permanently, and a full table degrades to fail-open for the
* per-source cap/lockout of new sources (the per-module cap and host ACLs still
* apply) instead of staying fail-open forever. A rate-limited warning is logged
* on the fail-open path. The eviction race with a concurrent
* registration/reclaim on the same bucket is benign: it can at worst lose one
* source's counter (fail-open), never corrupt memory or the module caps.
*/
typedef struct DaemonLimitRegistry DaemonLimitRegistry;
/* Result of a per-connection admission check. */
typedef enum {
DAEMON_LIMIT_OK = 0, /* admitted; slot is now REGISTERED */
DAEMON_LIMIT_MODULE_FULL, /* module's `max connections` cap reached */
DAEMON_LIMIT_HOST_FULL, /* global `max connections per host` cap reached */
DAEMON_LIMIT_UNAVAILABLE, /* registry/slot unusable (caller fails open) */
} DaemonLimitResult;
/* Bounds for registry sizing. A slot is one concurrently live child. */
#define DAEMON_LIMITS_MIN_SLOTS 16
#define DAEMON_LIMITS_MAX_SLOTS 65536
#define DAEMON_LIMITS_MAX_HOST_SLOTS 65536
#define DAEMON_LIMITS_NO_SLOT (-1)
/* Upper bound on `module_count`, matching daemon_conf.h's DAEMON_CONF_MAX_MODULES
* (asserted in daemon_limits.c) so a caller can never size the per-module counter
* array larger than the config parser can produce. */
#define DAEMON_LIMITS_MAX_MODULES 256
/* Per-source table lifetime: a bucket with no active connection and no pending
* lockout is reclaimable once it has been idle this long, so a flood of distinct
* sources cannot pin the table full forever. A bucket whose lockout deadline
* has passed is reclaimable immediately (independent of this idle window). */
#define DAEMON_LIMITS_HOST_EVICT_IDLE_SEC 300
/* Minimum spacing between "per-source table is full" warnings, so a table-full
* attack cannot flood the log. */
#define DAEMON_LIMITS_HOST_FULL_WARN_SEC 60
/* Create the shared registry in the calling (parent) process. `max_slots` is
* the number of concurrently live children to track (clamped to
* [DAEMON_LIMITS_MIN_SLOTS, DAEMON_LIMITS_MAX_SLOTS]); `module_count` is the
* number of daemon modules (clamped to
* [1, DAEMON_LIMITS_MAX_MODULES]); `per_host_cap` and the lockout pair come
* from the daemon config (0 disables). Returns NULL on failure (e.g. mmap
* allocation); callers must degrade gracefully (global cap + ACLs still
* apply). */
DaemonLimitRegistry* daemon_limits_create(int max_slots, int module_count, int per_host_cap,
int lockout_threshold, int lockout_duration_sec);
/* Unmap the registry. Only the creating process may call this. */
void daemon_limits_destroy(DaemonLimitRegistry* registry);
/* Parent side: reserve a slot for the next fork. Returns the slot index or
* DAEMON_LIMITS_NO_SLOT when every slot is in use. */
int daemon_limits_claim_slot(DaemonLimitRegistry* registry);
/* Parent side: record the forked child's pid in a claimed slot. */
void daemon_limits_set_slot_pid(DaemonLimitRegistry* registry, int slot, long pid);
/* Parent side: release a slot. The slot becomes FREE; the module/per-source
* occupancy arrays are DERIVED state and are only refreshed by
* daemon_limits_recompute, which callers must invoke afterwards when they rely
* on the derived counts (the SIGCHLD handler batches one recompute for the whole
* reap). Idempotent. */
void daemon_limits_reclaim_slot(DaemonLimitRegistry* registry, int slot);
/* Parent SIGCHLD side: release the slot owned by `pid` (no-op when not found).
* Like reclaim_slot this does not touch the derived occupancy arrays; call
* daemon_limits_recompute after a batch of releases. */
void daemon_limits_reclaim_pid(DaemonLimitRegistry* registry, long pid);
/* Parent side (async-signal-safe; atomics only, no malloc/log): rebuild
* module_active[] / host_active[] from scratch by scanning the REGISTERED slots.
* The slot table is the single source of truth, so this self-heals any
* count leaked by a child that was SIGKILLed mid-registration (it zeroes the
* arrays and re-derives them). Bounded by max_slots + host_slots. A
* registration racing this call can be transiently undercounted until the next
* recompute, which can only relax a cap briefly -- never corrupt memory. */
void daemon_limits_recompute(DaemonLimitRegistry* registry);
/* Child side: admit the connection for `module_index` from `peer_ip`. Always
* tracks the module/per-source occupancy (so the parent's reclaim is
* symmetric); when `module_cap` > 0 it additionally enforces the per-module
* cap. A NULL/empty or non-numeric `peer_ip` skips the per-source track (the
* callers use that to exempt a trusted loopback peer from the per-host cap; the
* per-module cap still applies). Returns DAEMON_LIMIT_OK and publishes the
* slot, or a refusal reason. */
DaemonLimitResult daemon_limits_register(DaemonLimitRegistry* registry, int slot, int module_index,
const char* peer_ip, int module_cap);
/* Child side: true when `peer_ip` is currently locked out after too many failed
* authentications. `seconds_remaining` may be NULL. */
bool daemon_limits_auth_locked(DaemonLimitRegistry* registry, const char* peer_ip,
int* seconds_remaining);
/* Child side: count one failed authentication for `peer_ip`; once the threshold
* is reached the source is locked out for the configured duration. */
void daemon_limits_auth_record_failure(DaemonLimitRegistry* registry, const char* peer_ip);
/* Child side: clear the failure counter/lockout for a source that authenticated
* successfully (no-op when the source has no table entry). */
void daemon_limits_auth_record_success(DaemonLimitRegistry* registry, const char* peer_ip);
/* Pure helper: 64-bit FNV-1a hash of a numeric peer IP plus its family, used to
* index the per-source table. *ok is set false (and 0 returned) for a NULL or
* non-numeric address. Exposed for unit testing. */
uint64_t daemon_limits_host_hash(const char* peer_ip, bool* ok);
#endif
+7 -1
View File
@@ -23,6 +23,7 @@ Data* data_create_reserve(size_t size) {
d->data = NULL;
d->size = size;
d->protocol_charge = 0;
d->owner = NULL;
return d;
}
@@ -36,14 +37,19 @@ Data* data_create(void* data, size_t data_size) {
new_data->data = data;
new_data->size = data_size;
new_data->protocol_charge = 0;
new_data->owner = NULL;
return new_data;
}
void data_destroy(Data* data) {
if (data == NULL)
return;
if (data->protocol_charge != 0)
if (data->protocol_charge != 0) {
if (data->owner != NULL)
protocol_release_memory_for_session(data->owner, data->protocol_charge);
else
protocol_release_memory(data->protocol_charge);
}
free(data->data);
free(data);
}
+18
View File
@@ -3,11 +3,25 @@
#include <stdlib.h>
/* Forward declaration for the connection budget a received Data is charged
* against; defined in protocol.h (which includes this header). */
typedef struct ProtocolSession ProtocolSession;
typedef struct {
void* data;
size_t size;
/* Non-zero only for a buffer charged to the protocol connection budget. */
size_t protocol_charge;
/* Session whose budget `protocol_charge` was reserved from. When non-NULL,
* the charge is returned to this session directly, regardless of which
* session (if any) is bound to the destroying thread. owner is not
* guaranteed to be set whenever protocol_charge is non-zero: it is NULL for
* uncharged Data and for Data that has no recorded owner, in which case any
* charge falls back to the session bound at destroy time.
*
* Lifetime contract: a Data with a non-NULL owner must not outlive that
* ProtocolSession -- data_destroy dereferences owner to return the charge. */
ProtocolSession* owner;
} Data;
Data* data_create_empty(size_t data_size);
@@ -15,5 +29,9 @@ Data* data_create_reserve(size_t size);
Data* data_create(void* data, size_t data_size);
void data_destroy(Data* data);
void protocol_release_memory(size_t charge);
/* Release `charge` against `session` directly instead of the thread-local bound
* session. Used by data_destroy to honor Data.owner; `session` must outlive
* the Data whose charge is being returned. A NULL session is a no-op. */
void protocol_release_memory_for_session(ProtocolSession* session, size_t charge);
#endif
+14
View File
@@ -264,6 +264,20 @@ static bool delay_publish_entry(DelayUpdatesContext* context, const Config* conf
const StagedFileEntry* entry) {
if (!delay_publish_backup(context, config, entry))
return false;
/* --force: an incoming regular file/symlink may replace a destination
DIRECTORY (possibly non-empty). The immediate-install path handles this in
file_receive; a --delay-updates run stages elsewhere and only discovers the
blocking directory here, so clear it before the rename (rsync's
"could not make way for new regular file" without --force). */
if (config && config->force_delete && file_directory_exists_secure(entry->final_path)) {
if (!file_remove_tree_secure(entry->final_path)) {
char* escaped = output_escape(entry->final_path, false);
log_message(LOG_LEVEL_ERROR, "could not remove destination directory blocking '%s': %s",
escaped ? escaped : "<allocation failed>", strerror(errno));
free(escaped);
return false;
}
}
if (!file_rename_secure(entry->staged_path, entry->final_path)) {
if (errno == EXDEV) {
char* escaped = output_escape(entry->final_path, false);
+3 -2
View File
@@ -18,8 +18,9 @@ typedef struct {
/* Receiver-side --delay-updates staging registry. All successfully written
files land under a private staging directory inside the receive root and are
atomically renamed into their final destination only at the very end of the
transfer. A single PipelineContextReceiver has exactly one writer thread,
but the registry is still mutex-protected so the same object can be safely
transfer. A single receiver pipeline (see src/server/receiver_pipeline.h)
has exactly one writer thread, but the registry is still mutex-protected so
the same object can be safely
shared with the publish/cleanup phase that runs after the threads join. */
typedef struct DelayUpdatesContext {
char* root_directory; /* receive root the staging dir lives under */
+893
View File
@@ -0,0 +1,893 @@
#include "delete_plan.h"
#include "charset.h"
#include "delay_updates.h"
#include "file.h"
#include "log.h"
#include "utils.h"
#include <dirent.h>
#include <errno.h>
#include <fcntl.h>
#include <limits.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/stat.h>
#include <unistd.h>
/* Mirrors MAX_SERVER_DELETE_COUNT in file_receive.c: the server's hard bound on
* the number of entries one deletion commit may remove. A client
* --max-delete=NUM smaller than this replaces it for the run. */
#define DELETE_PLAN_SERVER_LIMIT 100000U
/* ------------------------------------------------------------------ */
/* Sender: plan builder */
/* ------------------------------------------------------------------ */
typedef struct PlanNode {
char* dir;
ArrayList* files; /* basenames kept directly in dir */
ArrayList* dirs; /* basenames of kept child directories */
bool sent;
struct PlanNode* hash_next;
} PlanNode;
struct DeletePlanSender {
PlanNode** buckets;
size_t capacity;
size_t count;
bool config_sent;
bool all_synced;
const ArrayList* synced_dirs;
/* Owned by the caller's synced_dirs list; non-NULL only for a general -R
transfer, where it is the destination prefix the delete walk is confined
to. NULL means the whole receive root (or a --files-from scope). */
const char* walk_root;
const ArrayList* protected_prefixes;
const ArrayList* size_skipped;
const ArrayList* missing_args;
size_t entries;
/* Transmitted FILE entries only. The caller's "empty scan" safety guard keys
off this (an I/O error that hid every file must refuse to delete even when
some directories were traversed), so directory keep entries do not count. */
size_t file_entries;
};
static size_t plan_hash(const char* key) {
size_t h = 5381;
for (const unsigned char* p = (const unsigned char*)key; *p; p++)
h = ((h << 5) + h) + *p;
return h;
}
static bool list_contains_str(const ArrayList* list, const char* value) {
if (!list)
return false;
for (int i = 0; i < list->size; i++) {
if (strcmp((const char*)list->items[i], value) == 0)
return true;
}
return false;
}
static bool list_add_str_unique(ArrayList* list, const char* value) {
if (!list || !value)
return false;
if (list_contains_str(list, value))
return true;
char* copy = str_dup(value);
if (!copy)
return false;
if (!array_list_add(list, copy)) {
free(copy);
return false;
}
return true;
}
DeletePlanSender* delete_plan_sender_create(void) {
DeletePlanSender* sender = calloc(1, sizeof(DeletePlanSender));
if (!sender)
return NULL;
sender->capacity = 64;
sender->buckets = calloc(sender->capacity, sizeof(PlanNode*));
if (!sender->buckets) {
free(sender);
return NULL;
}
sender->all_synced = true;
return sender;
}
static void plan_node_destroy(PlanNode* node) {
if (!node)
return;
free(node->dir);
array_list_delete(node->files);
array_list_delete(node->dirs);
free(node);
}
void delete_plan_sender_destroy(DeletePlanSender* sender) {
if (!sender)
return;
for (size_t i = 0; i < sender->capacity; i++) {
PlanNode* node = sender->buckets[i];
while (node) {
PlanNode* next = node->hash_next;
plan_node_destroy(node);
node = next;
}
}
free(sender->buckets);
free(sender);
}
static PlanNode* plan_find(const DeletePlanSender* sender, const char* dir) {
size_t index = plan_hash(dir) & (sender->capacity - 1);
for (PlanNode* node = sender->buckets[index]; node; node = node->hash_next) {
if (strcmp(node->dir, dir) == 0)
return node;
}
return NULL;
}
static bool plan_grow(DeletePlanSender* sender) {
size_t new_capacity = sender->capacity * 2;
PlanNode** buckets = calloc(new_capacity, sizeof(PlanNode*));
if (!buckets)
return false;
for (size_t i = 0; i < sender->capacity; i++) {
PlanNode* node = sender->buckets[i];
while (node) {
PlanNode* next = node->hash_next;
size_t index = plan_hash(node->dir) & (new_capacity - 1);
node->hash_next = buckets[index];
buckets[index] = node;
node = next;
}
}
free(sender->buckets);
sender->buckets = buckets;
sender->capacity = new_capacity;
return true;
}
static PlanNode* plan_ensure(DeletePlanSender* sender, const char* dir) {
PlanNode* node = plan_find(sender, dir);
if (node)
return node;
if (sender->count + 1 > sender->capacity * 3 / 4 && !plan_grow(sender))
return NULL;
node = calloc(1, sizeof(PlanNode));
if (!node)
return NULL;
node->dir = str_dup(dir);
node->files = array_list_create(free);
node->dirs = array_list_create(free);
if (!node->dir || !node->files || !node->dirs) {
plan_node_destroy(node);
return NULL;
}
size_t index = plan_hash(dir) & (sender->capacity - 1);
node->hash_next = sender->buckets[index];
sender->buckets[index] = node;
sender->count++;
return node;
}
static char* path_parent_dir(const char* path) {
const char* slash = strrchr(path, '/');
if (!slash)
return str_dup(".");
if (slash == path)
return str_dup(".");
size_t len = (size_t)(slash - path);
char* parent = malloc(len + 1);
if (!parent)
return NULL;
memcpy(parent, path, len);
parent[len] = '\0';
return parent;
}
static char* path_base_name(const char* path) {
const char* slash = strrchr(path, '/');
return str_dup(slash ? slash + 1 : path);
}
/* Copy `path`, stripping a leading '/' and any trailing '/'. */
static char* plan_clean_path(const char* path) {
while (*path == '/')
path++;
size_t len = strlen(path);
while (len > 0 && path[len - 1] == '/')
len--;
char* clean = malloc(len + 1);
if (!clean)
return NULL;
memcpy(clean, path, len);
clean[len] = '\0';
return clean;
}
static bool plan_ensure_ancestors(DeletePlanSender* sender, const char* dir) {
char* current = str_dup(dir);
if (!current)
return false;
bool ok = true;
while (strcmp(current, ".") != 0) {
char* parent = path_parent_dir(current);
char* base = path_base_name(current);
PlanNode* parent_node = parent ? plan_ensure(sender, parent) : NULL;
if (!parent || !base || !parent_node || !list_add_str_unique(parent_node->dirs, base)) {
ok = false;
free(parent);
free(base);
break;
}
free(base);
free(current);
current = parent;
}
free(current);
return ok;
}
bool delete_plan_sender_add(DeletePlanSender* sender, const char* path, bool is_dir) {
if (!sender || !path)
return false;
char* clean = plan_clean_path(path);
if (!clean)
return false;
if (*clean == '\0') {
free(clean);
return true;
}
char* parent = path_parent_dir(clean);
char* base = path_base_name(clean);
PlanNode* parent_node = parent ? plan_ensure(sender, parent) : NULL;
bool ok = parent && base && parent_node;
if (ok) {
if (is_dir) {
ok = list_add_str_unique(parent_node->dirs, base) && plan_ensure(sender, clean) != NULL;
} else {
ok = list_add_str_unique(parent_node->files, base);
}
}
if (ok)
ok = plan_ensure_ancestors(sender, parent);
if (ok) {
sender->entries++;
if (!is_dir)
sender->file_entries++;
}
free(clean);
free(parent);
free(base);
return ok;
}
void delete_plan_sender_finalize(DeletePlanSender* sender, const ArrayList* synced_dirs,
const char* walk_root) {
if (!sender)
return;
sender->synced_dirs = synced_dirs;
sender->all_synced = synced_dirs == NULL && walk_root == NULL;
sender->walk_root = walk_root;
}
bool delete_plan_sender_empty(const DeletePlanSender* sender) {
return !sender || sender->file_entries == 0;
}
void delete_plan_sender_set_config(DeletePlanSender* sender, const ArrayList* protected_prefixes,
const ArrayList* size_skipped, const ArrayList* missing_args) {
if (!sender)
return;
sender->protected_prefixes = protected_prefixes;
sender->size_skipped = size_skipped;
sender->missing_args = missing_args;
}
/* True when `dir` is `root` itself or a descendant of it (path-component
* aware, so "foo" does not match "foobar"). */
static bool path_at_or_under(const char* dir, const char* root) {
if (!dir || !root)
return false;
size_t n = strlen(root);
return strncmp(dir, root, n) == 0 && (dir[n] == '\0' || dir[n] == '/');
}
static bool plan_is_allowed(const DeletePlanSender* sender, const char* dir) {
if (sender->all_synced)
return true;
if (sender->walk_root)
return path_at_or_under(dir, sender->walk_root);
return list_contains_str(sender->synced_dirs, dir);
}
static int send_str_section(int fd, const ArrayList* list) {
int count = list ? list->size : 0;
if (!send_int(fd, count))
return -1;
for (int i = 0; i < count; i++) {
if (!send_wire_str(fd, (const char*)list->items[i]))
return -1;
}
return 0;
}
static int send_plan_node(int fd, DeletePlanSender* sender, PlanNode* node) {
if (!send_status(fd, STATUS_DELETE_PLAN))
return -1;
if (!send_int(fd, sender->config_sent ? 0 : 1))
return -1;
if (!sender->config_sent) {
if (send_str_section(fd, sender->protected_prefixes) != 0 ||
send_str_section(fd, sender->size_skipped) != 0 ||
send_str_section(fd, sender->missing_args) != 0)
return -1;
sender->config_sent = true;
}
if (!send_wire_str(fd, node->dir))
return -1;
if (send_str_section(fd, node->dirs) != 0 || send_str_section(fd, node->files) != 0)
return -1;
node->sent = true;
return 0;
}
static int send_prefix_plan(int fd, DeletePlanSender* sender, const char* dir) {
PlanNode* node = plan_find(sender, dir);
if (!node || node->sent)
return 0;
if (!plan_is_allowed(sender, dir))
return 0;
return send_plan_node(fd, sender, node);
}
int delete_plan_send_root(int fd, DeletePlanSender* sender) {
if (!sender)
return -1;
const char* root = sender->walk_root ? sender->walk_root : ".";
if (!plan_ensure(sender, root))
return -1;
return send_prefix_plan(fd, sender, root);
}
int delete_plan_send_for_path(int fd, DeletePlanSender* sender, const char* path, bool is_dir) {
if (!sender || !path)
return -1;
char* clean = plan_clean_path(path);
if (!clean)
return -1;
/* The walk root (the -R prefix, or ".") is sent up front by
delete_plan_send_root(); never emit the receive-root plan for a scoped -R
run, whose "." keep list would delete the prefix's siblings. */
int rc = sender->walk_root ? 0 : send_prefix_plan(fd, sender, ".");
if (rc == 0 && *clean != '\0') {
size_t len = strlen(clean);
size_t end = len;
if (!is_dir) {
const char* slash = strrchr(clean, '/');
end = slash ? (size_t)(slash - clean) : 0;
}
for (size_t i = 1; i <= end && rc == 0; i++) {
if (i == end || clean[i] == '/') {
char* prefix = malloc(i + 1);
if (!prefix) {
rc = -1;
break;
}
memcpy(prefix, clean, i);
prefix[i] = '\0';
rc = send_prefix_plan(fd, sender, prefix);
free(prefix);
}
}
}
free(clean);
return rc;
}
int delete_plan_send_remaining(int fd, DeletePlanSender* sender, const ArrayList* dirs) {
if (!sender || !dirs)
return 0;
for (int i = 0; i < dirs->size; i++) {
const char* dir = (const char*)dirs->items[i];
if (delete_plan_send_for_path(fd, sender, dir, true) != 0)
return -1;
}
return 0;
}
/* ------------------------------------------------------------------ */
/* Receiver: delete session */
/* ------------------------------------------------------------------ */
struct DeletePlanSession {
bool defer;
bool dry_run;
size_t max_delete;
size_t deleted;
size_t skipped;
bool limit_hit;
bool limit_logged;
bool config_seen;
bool missing_applied;
ArrayList* protected_prefixes;
ArrayList* size_skipped;
ArrayList* missing;
ArrayList* deferred;
};
DeletePlanSession* delete_plan_session_create(const Config* config) {
if (!config)
return NULL;
DeletePlanSession* session = calloc(1, sizeof(DeletePlanSession));
if (!session)
return NULL;
session->defer = config->delete_delay;
session->dry_run = config->dry_run;
bool user_limited =
config->max_delete >= 0 && (size_t)config->max_delete < DELETE_PLAN_SERVER_LIMIT;
session->max_delete =
user_limited ? (size_t)config->max_delete : (size_t)DELETE_PLAN_SERVER_LIMIT;
session->protected_prefixes = array_list_create(free);
session->size_skipped = array_list_create(free);
session->missing = array_list_create(free);
session->deferred = array_list_create(free);
if (!session->protected_prefixes || !session->size_skipped || !session->missing ||
!session->deferred) {
delete_plan_session_destroy(session);
return NULL;
}
return session;
}
void delete_plan_session_destroy(DeletePlanSession* session) {
if (!session)
return;
array_list_delete(session->protected_prefixes);
array_list_delete(session->size_skipped);
array_list_delete(session->missing);
array_list_delete(session->deferred);
free(session);
}
bool delete_plan_session_limit_reached(const DeletePlanSession* session) {
return session && session->limit_hit;
}
size_t delete_plan_session_deleted(const DeletePlanSession* session) {
return session ? session->deleted : 0;
}
/* True for a destination-relative path section entry (non-empty, relative,
* traversal-free). */
static bool valid_rel_path(const char* value) {
return value && value[0] != '\0' && value[0] != '/' && !has_path_traversal(value);
}
/* True for a single child name (non-empty, no slash, not "."/".."). */
static bool valid_name(const char* value) {
return value && value[0] != '\0' && strcmp(value, ".") != 0 && strcmp(value, "..") != 0 &&
strchr(value, '/') == NULL;
}
/* Read one count-prefixed section. `bytes` is the running per-frame budget,
* shared across every section of the frame so a hostile peer cannot retain more
* than MAX_MANIFEST_BYTES from one STATUS_DELETE_PLAN frame. */
static bool read_section(int fd, ArrayList* list, bool rel_path, size_t* bytes) {
int count;
if (!receive_int(fd, &count) || count < 0 || count > MAX_MANIFEST_ENTRIES)
return false;
for (int i = 0; i < count; i++) {
char* value = receive_wire_str(fd);
bool ok = value && (rel_path ? valid_rel_path(value) : valid_name(value));
if (ok) {
size_t entry_size = strlen(value) + sizeof(char*) + 16;
if (entry_size > MAX_MANIFEST_BYTES - *bytes) {
ok = false;
} else {
*bytes += entry_size;
ok = array_list_add(list, value);
}
}
if (!ok) {
free(value);
return false;
}
}
return true;
}
static int open_plan_dir(const Config* config, const char* dir) {
char* full = (strcmp(dir, ".") == 0) ? str_dup(config->receive_root_directory)
: path_cat(config->receive_root_directory, dir);
if (!full)
return -1;
int root_fd = utils_get_authorized_root_fd();
int fd = -1;
if (root_fd >= 0) {
if (utils_get_authorized_root_path())
fd = utils_open_authorized_destination(full);
else if (strcmp(dir, ".") == 0)
fd = dup(root_fd);
} else {
fd = open(full, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
}
free(full);
return fd;
}
typedef struct PlanSkips {
DeleteSkipEntry* entries;
int count;
} PlanSkips;
static bool build_plan_skips(const Config* config, const DeletePlanSession* session,
PlanSkips* out) {
out->entries = NULL;
out->count = 0;
int count = (config->delay_updates ? 1 : 0) + config->basis_count +
session->protected_prefixes->size + session->size_skipped->size;
if (count == 0)
return true;
out->entries = calloc((size_t)count, sizeof(DeleteSkipEntry));
if (!out->entries)
return false;
int idx = 0;
if (config->delay_updates) {
out->entries[idx].prefix = DELAY_UPDATES_STAGING_DIR;
out->entries[idx].top_level_only = true;
idx++;
}
for (int i = 0; i < config->basis_count; i++) {
out->entries[idx].prefix = config->basis_dirs[i].path;
out->entries[idx].top_level_only = false;
idx++;
}
for (int i = 0; i < session->protected_prefixes->size; i++) {
out->entries[idx].prefix = (const char*)session->protected_prefixes->items[i];
out->entries[idx].top_level_only = false;
idx++;
}
for (int i = 0; i < session->size_skipped->size; i++) {
out->entries[idx].prefix = (const char*)session->size_skipped->items[i];
out->entries[idx].top_level_only = false;
idx++;
}
out->count = idx;
return true;
}
static bool budget_available(const DeletePlanSession* session) {
return session->deleted < session->max_delete;
}
static void note_skipped(DeletePlanSession* session) {
session->limit_hit = true;
session->skipped++;
}
static void log_deleted(const char* rel) {
char* escaped = output_escape(rel, log_get_8_bit_output());
fprintf(stderr, " Deleted: %s\n", escaped ? escaped : "<allocation failed>");
free(escaped);
}
/* Append a snapshot path for --delete-delay. */
static bool defer_add(DeletePlanSession* session, const char* rel) {
char* copy = str_dup(rel);
if (!copy)
return false;
if (!array_list_add(session->deferred, copy)) {
free(copy);
return false;
}
session->deleted++;
return true;
}
/* Process the direct children of one directory. `keep_dirs`/`keep_files`
* (basenames) are the source entries that must be kept; NULL means every child
* is an extra (the forced path used inside a removed extra directory tree).
* `survives` reports that at least one child remains (kept, protected, or
* skipped by the budget). `force_now` removes even in --delete-delay mode
* (type conflicts must clear before the incoming data). */
static bool process_children(int dirfd, const char* dir_rel, const ArrayList* keep_dirs,
const ArrayList* keep_files, bool at_root, bool force_now,
const PlanSkips* skips, DeletePlanSession* session, bool* survives);
static bool process_extra_dir(int dirfd, const char* name, const char* child_rel, bool force_now,
const PlanSkips* skips, DeletePlanSession* session, bool* removed) {
*removed = false;
int childfd = openat(dirfd, name, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
if (childfd < 0) {
if (errno == ENOENT) {
*removed = true;
return true;
}
return false;
}
bool survives = false;
bool ok =
process_children(childfd, child_rel, NULL, NULL, false, force_now, skips, session, &survives);
close(childfd);
if (!ok)
return false;
if (survives)
return true;
if (!budget_available(session)) {
note_skipped(session);
return true;
}
if (session->defer && !force_now) {
if (!defer_add(session, child_rel))
return false;
*removed = true;
return true;
}
if (unlinkat(dirfd, name, AT_REMOVEDIR) == 0) {
session->deleted++;
log_deleted(child_rel);
*removed = true;
return true;
}
if (errno == ENOENT) {
*removed = true;
return true;
}
/* ENOTEMPTY/EEXIST: a protected entry the walker leaves behind survived, so
the directory stays; any other errno is a genuine failure. */
return errno == ENOTEMPTY || errno == EEXIST;
}
static bool process_extra_file(int dirfd, const char* name, const char* child_rel, bool force_now,
DeletePlanSession* session) {
if (!budget_available(session)) {
note_skipped(session);
return true;
}
if (session->defer && !force_now) {
return defer_add(session, child_rel);
}
if (unlinkat(dirfd, name, 0) == 0) {
session->deleted++;
log_deleted(child_rel);
} else if (errno != ENOENT) {
return false;
}
return true;
}
static bool process_children(int dirfd, const char* dir_rel, const ArrayList* keep_dirs,
const ArrayList* keep_files, bool at_root, bool force_now,
const PlanSkips* skips, DeletePlanSession* session, bool* survives) {
*survives = false;
int scanfd = openat(dirfd, ".", O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
if (scanfd < 0)
return false;
DIR* dir = fdopendir(scanfd);
if (!dir) {
close(scanfd);
return false;
}
bool operation_ok = true;
bool local_survives = false;
const struct dirent* entry;
while ((entry = readdir(dir)) != NULL) {
if (strcmp(entry->d_name, ".") == 0 || strcmp(entry->d_name, "..") == 0)
continue;
char* child_rel =
(strcmp(dir_rel, ".") == 0) ? str_dup(entry->d_name) : path_cat(dir_rel, entry->d_name);
if (!child_rel) {
operation_ok = false;
continue;
}
if (path_under_skip_prefix(child_rel, at_root, skips->entries, skips->count)) {
local_survives = true;
free(child_rel);
continue;
}
struct stat st;
if (fstatat(dirfd, entry->d_name, &st, AT_SYMLINK_NOFOLLOW) != 0) {
if (errno != ENOENT)
operation_ok = false;
free(child_rel);
continue;
}
bool is_dir = S_ISDIR(st.st_mode);
bool in_keep_dirs = is_dir && list_contains_str(keep_dirs, entry->d_name);
bool in_keep_files = !is_dir && list_contains_str(keep_files, entry->d_name);
if (in_keep_dirs) {
local_survives = true;
} else if (keep_dirs && !is_dir && list_contains_str(keep_dirs, entry->d_name)) {
/* Destination file blocks a source directory: clear it now, whatever the
delete timing, so the directory can be created. */
if (!process_extra_file(dirfd, entry->d_name, child_rel, true, session))
operation_ok = false;
} else if (in_keep_files) {
local_survives = true;
} else if (keep_files && is_dir && list_contains_str(keep_files, entry->d_name)) {
/* Destination directory blocks a source file: remove it now. */
bool removed = false;
if (!process_extra_dir(dirfd, entry->d_name, child_rel, true, skips, session, &removed))
operation_ok = false;
else if (!removed)
local_survives = true;
} else if (is_dir) {
bool removed = false;
if (!process_extra_dir(dirfd, entry->d_name, child_rel, force_now, skips, session, &removed))
operation_ok = false;
else if (!removed)
local_survives = true;
} else {
if (!process_extra_file(dirfd, entry->d_name, child_rel, force_now, session))
operation_ok = false;
}
free(child_rel);
}
closedir(dir);
*survives = local_survives;
return operation_ok;
}
static bool apply_plan_dir(DeletePlanSession* session, const Config* config, const char* dir,
const ArrayList* dirs, const ArrayList* files) {
int dirfd = open_plan_dir(config, dir);
if (dirfd < 0) {
/* An absent destination directory has nothing to delete. */
return errno == ENOENT || errno == ENOTDIR;
}
PlanSkips skips;
if (!build_plan_skips(config, session, &skips)) {
close(dirfd);
return false;
}
bool survives = false;
bool ok = process_children(dirfd, dir, dirs, files, strcmp(dir, ".") == 0, false, &skips, session,
&survives);
free(skips.entries);
close(dirfd);
if (!ok)
log_message(LOG_LEVEL_ERROR, "deletion failed while removing extraneous files");
return ok;
}
static bool apply_missing(DeletePlanSession* session, const Config* config) {
if (session->missing_applied)
return true;
session->missing_applied = true;
/* The server clears delete_missing_args when its --allow-delete policy is
off; never honor the client's exact-path requests then. */
if (!config->delete_missing_args || session->missing->size == 0)
return true;
DeleteManifest manifest = {
.keeps = NULL, .protected = NULL, .missing = session->missing, .dirs = NULL};
size_t remaining = budget_available(session) ? session->max_delete - session->deleted : 0;
size_t deleted = 0;
size_t skipped = 0;
bool limit = false;
bool ok = manifest_delete_missing_args_limited(config, &manifest, remaining, &deleted, &skipped,
&limit);
session->deleted += deleted;
session->skipped += skipped;
if (limit)
session->limit_hit = true;
return ok;
}
int delete_plan_session_receive(DeletePlanSession* session, const Config* config, int fd) {
if (!session || !config) {
send_status(fd, STATUS_ERROR);
return -1;
}
int has_config;
if (!receive_int(fd, &has_config) || (has_config != 0 && has_config != 1)) {
send_status(fd, STATUS_ERROR);
return -1;
}
size_t bytes = 0;
if (has_config) {
if (session->config_seen || !read_section(fd, session->protected_prefixes, true, &bytes) ||
!read_section(fd, session->size_skipped, true, &bytes) ||
!read_section(fd, session->missing, true, &bytes)) {
send_status(fd, STATUS_ERROR);
return -1;
}
session->config_seen = true;
}
char* dir = receive_wire_str(fd);
ArrayList* dirs = array_list_create(free);
ArrayList* files = array_list_create(free);
bool parsed = dir && (strcmp(dir, ".") == 0 || valid_rel_path(dir)) && dirs && files &&
read_section(fd, dirs, false, &bytes) && read_section(fd, files, false, &bytes);
if (!parsed) {
free(dir);
array_list_delete(dirs);
array_list_delete(files);
send_status(fd, STATUS_ERROR);
return -1;
}
bool enabled = config->use_delete || config->delete_missing_args;
bool ok = true;
if (!session->dry_run && enabled) {
if (!session->defer && !apply_missing(session, config))
ok = false;
if (ok && !apply_plan_dir(session, config, dir, dirs, files))
ok = false;
}
free(dir);
array_list_delete(dirs);
array_list_delete(files);
if (!ok) {
send_status(fd, STATUS_ERROR);
return -1;
}
if (session->limit_hit && !session->limit_logged) {
session->limit_logged = true;
log_message(LOG_LEVEL_WARNING, "Deletions stopped due to the delete limit (%zu skipped)",
session->skipped);
}
return 0;
}
/* Apply one snapshotted --delete-delay path (post-order: children precede their
* parent directory). */
static bool apply_deferred_path(DeletePlanSession* session, const Config* config, const char* rel) {
(void)session;
char* full = path_cat(config->receive_root_directory, rel);
if (!full)
return false;
char* leaf = NULL;
int parent_fd = file_open_secure_parent(full, &leaf, false);
free(full);
if (parent_fd < 0) {
free(leaf);
return errno == ENOENT || errno == ENOTDIR;
}
struct stat st;
if (fstatat(parent_fd, leaf, &st, AT_SYMLINK_NOFOLLOW) != 0) {
bool absent = errno == ENOENT;
close(parent_fd);
free(leaf);
return absent;
}
int rc;
if (S_ISDIR(st.st_mode))
rc = unlinkat(parent_fd, leaf, AT_REMOVEDIR);
else
rc = unlinkat(parent_fd, leaf, 0);
bool ok = rc == 0 || errno == ENOENT || errno == ENOTEMPTY || errno == EEXIST;
if (rc == 0)
log_deleted(rel);
close(parent_fd);
free(leaf);
return ok;
}
DeleteCommitResult delete_plan_session_commit(DeletePlanSession* session, const Config* config) {
if (!session || !config)
return DELETE_COMMIT_ERROR;
/* Central no-mutation guard (mirrors manifest_delete_all): a dry-run never
deletes. The receive path already skips plan application, but a hostile or
buggy peer could still reach the commit, so treat it as a no-op. */
if (session->dry_run)
return DELETE_COMMIT_OK;
bool ok = true;
if (session->defer) {
for (int i = 0; i < session->deferred->size && ok; i++)
ok = apply_deferred_path(session, config, (const char*)session->deferred->items[i]);
}
if (ok)
ok = apply_missing(session, config);
if (!ok)
return DELETE_COMMIT_ERROR;
if (session->limit_hit)
return DELETE_COMMIT_LIMIT_REACHED;
return DELETE_COMMIT_OK;
}
+83
View File
@@ -0,0 +1,83 @@
#ifndef DELETE_PLAN_H
#define DELETE_PLAN_H
#include "array_list.h"
#include "config.h"
#include "file_receive.h"
#include "protocol.h"
#include <stdbool.h>
/* Per-directory delete plans (protocol 2.24.0).
*
* rsync's --delete-during removes a directory's extras while the generator
* processes that directory, and --delete-delay records the deletion list during
* the scan but applies it only after a fully-successful transfer. FastSync has
* no per-directory generator pass; instead the sender streams one plan per
* source directory, in directory order, and the receiver applies it when it
* arrives (during) or snapshots its extras and commits them at the end (delay).
*
* The sender side builds a plan set from the path-only pre-scan (it needs every
* directory's complete direct-child list before the first data byte of that
* directory). The receiver side is a session that carries the global protected
* prefixes (filter-excluded and size-skipped source mirrors), the
* --delete-missing-args exact deletions, the shared --max-delete budget and,
* for --delete-delay, the snapshotted extras. */
/* ---- Sender: plan builder ---- */
typedef struct DeletePlanSender DeletePlanSender;
DeletePlanSender* delete_plan_sender_create(void);
void delete_plan_sender_destroy(DeletePlanSender* sender);
/* Record one transmitted entry. `path` is the destination-relative wire path;
* is_dir marks an explicit directory entry (--dirs, a -x mount point). */
bool delete_plan_sender_add(DeletePlanSender* sender, const char* path, bool is_dir);
/* Drop plans for directories outside `synced_dirs` (the --files-from
* synchronization scope; pass NULL when a full recursive transfer synchronized
* every directory). The receive root is the "." sentinel.
*
* `walk_root` scopes a general -R transfer: when non-NULL it is the
* reconstructed destination prefix the run actually transferred, and only the
* plan for that prefix (and directories below it) is ever transmitted, so the
* prefix's parent-directory siblings are never walked. Pass NULL for a plain
* recursive transfer and for --files-from. */
void delete_plan_sender_finalize(DeletePlanSender* sender, const ArrayList* synced_dirs,
const char* walk_root);
/* True when no transmitted FILE entry was recorded (an ambiguous empty scan).
Directory keep entries do not count, so an I/O error that hid every file
still refuses to delete. */
bool delete_plan_sender_empty(const DeletePlanSender* sender);
/* Attach the global config sections advertised on the first plan frame. */
void delete_plan_sender_set_config(DeletePlanSender* sender, const ArrayList* protected_prefixes,
const ArrayList* size_skipped, const ArrayList* missing_args);
/* Send the root plan (even before any data, so root extras are handled like
* rsync's first generator directory). Returns -1 on I/O error. */
int delete_plan_send_root(int fd, DeletePlanSender* sender);
/* Send the plans for every ancestor of `path` (root-first) and, when is_dir,
* for `path` itself; already-sent plans are skipped. */
int delete_plan_send_for_path(int fd, DeletePlanSender* sender, const char* path, bool is_dir);
/* Send the plan for every directory in `dirs` that has not been transmitted
* yet. Called after the data stream so an empty source directory's plan still
* clears its destination extras even though no file frame triggered it. */
int delete_plan_send_remaining(int fd, DeletePlanSender* sender, const ArrayList* dirs);
/* ---- Receiver: delete session ---- */
typedef struct DeletePlanSession DeletePlanSession;
DeletePlanSession* delete_plan_session_create(const Config* config);
void delete_plan_session_destroy(DeletePlanSession* session);
/* Read one STATUS_DELETE_PLAN frame (the leading status already consumed) and
* act on it. Returns 0 on success (including a dry-run/disabled no-op) and -1
* after signalling STATUS_ERROR on a malformed frame or a deletion failure. */
int delete_plan_session_receive(DeletePlanSession* session, const Config* config, int fd);
/* Apply the deferred snapshot (--delete-delay) and the missing-args deletions.
* Safe to call once; returns the commit outcome. */
DeleteCommitResult delete_plan_session_commit(DeletePlanSession* session, const Config* config);
/* True once the shared --max-delete budget stopped part of a deletion. */
bool delete_plan_session_limit_reached(const DeletePlanSession* session);
/* Number of destination entries the session's plans removed (or, for
--delete-delay, snapshotted for removal), for the end-of-transfer stats. */
size_t delete_plan_session_deleted(const DeletePlanSession* session);
#endif
+783 -96
View File
File diff suppressed because it is too large. Load diff
+103 -25
View File
@@ -22,13 +22,65 @@ bool file_load_data(File* file);
bool file_checksum(File* file, ChecksumAlgo algo, uint64_t seed, uint8_t* out, size_t out_capacity,
size_t* out_len);
size_t file_content_to_buffer(File* file);
FileMetadata* file_metadata_create(const struct stat* stats);
FileMetadata* file_metadata_create(const char* path, const struct stat* stats, bool capture_atime,
bool capture_crtime);
void file_metadata_destroy(void* metadata);
/* --open-noatime process-wide sender policy; see file.c. */
void file_set_open_noatime(bool enable);
bool file_get_open_noatime(void);
/* Capture the process umask ONCE, before any threads are created. Call this at
* the very top of main() in both entry points so the cached value is read while
* the process is still single-threaded: reading the umask needs a get+set round
* trip (umask(0); umask(old)), which would race against receiver threads
* creating files if it happened during the first write. Idempotent and safe to
* call more than once. */
void file_umask_capture(void);
/* Process-wide umask, captured once (thread-safe). Used to derive the mode of
* a brand-new destination like rsync: source_mode & 0777 & ~umask. Falls back
* to file_umask_capture() (behind pthread_once) if capture was never called. */
unsigned file_process_umask(void);
/* Open `path` read-only for transfer, honouring --open-noatime when set. */
int file_open_for_read(const char* path);
bool file_write_to_disk(const char* path, const void* data, unsigned long long data_size,
bool inplace, bool sparse);
/* A configured fd without a canonical identity deliberately rejects paths. */
bool file_set_authorized_root(int fd, const char* canonical_path);
/* Symlink trust-boundary helpers (Phase 4, symlink wave; rsync parity).
* --munge-links is a RECEIVER-side rewrite: rsync prefixes every stored symlink
* target with this marker, making the link unusable while the referenced
* directory does not exist. A SENDER receiving a munged source strips it back
* off before transmitting (so a munged tree round-trips through the receiver's
* re-munging). */
#define SYMLINK_MUNGE_PREFIX "/rsyncd-munged/"
char* file_symlink_munge(const char* target);
/* rsync 3.4.1 unsafe_symlink(): true when `target` escapes the transfer tree
* rooted at `link_path` (the symlink's transfer-relative path incl. its name).
* Absolute/empty targets and targets climbing above the transfer root (via
* "..") are unsafe, as are internal "/../" components and trailing "/..". */
bool file_symlink_unsafe(const char* target, const char* link_path);
/* True when a lexical target is relative and contains no ".." component, so it
* can never escape the receive root once created beneath it. */
bool file_symlink_target_contained(const char* target);
/* Strip a leading SYMLINK_MUNGE_PREFIX from `target` (mutable, in place);
* returns true when a marker was removed. */
bool file_symlink_unmunge(char* target);
/* Create a symlink at `path` -> `target`, confined below the authorized root
* (O_NOFOLLOW parent walk, symlinkat; the target is never followed). The link
* value is copied verbatim (rsync -l); only the placement path is confined.
* Returns false when a directory already occupies `path`. */
bool file_symlink_at_secure(const char* path, const char* target);
/* --keep-dirlinks (-K) receiver process-wide policy: allow an in-root existing
* symlink-to-directory to be followed as a directory. */
void file_set_keep_dirlinks(bool enable);
/* --trust-sender receiver process-wide policy (Phase 5). When set, the
* receiver trusts that the sender already produced a clean file list and skips
* its own redundant up-front re-validation of incoming paths (the empty/".."
* rejection and the escaping-symlink-target containment). The low-level
* fd-relative confinement primitives below are deliberately NOT disabled by
* this flag, so a hostile sender still cannot escape the authorized root. */
void file_set_trust_sender(bool enable);
bool file_get_trust_sender(void);
/* Secure path/filesystem primitives (symlink-safe, O_NOFOLLOW, root-confined). */
bool file_path_exists_secure(const char* path);
@@ -43,42 +95,68 @@ bool file_rename_secure(const char* old_path, const char* new_path);
regular file. See the .c for the exact success semantics. */
bool file_remove_tree_secure(const char* path);
/* Open a private 0700 directory (creating it on demand) that must live below
the authorized root. Used for the --temp-dir scratch directory and the
--delay-updates staging directory. */
the authorized root. Used for the --delay-updates staging directory. */
int file_open_private_dir(const char* dir_path);
/* Open an existing --temp-dir scratch directory as-is (absolute or relative;
no creation, no root confinement), matching rsync's --temp-dir handling. */
int file_open_temp_dir(const char* dir_path);
/* The file_to_disk_secure* variants write a temporary copy in the destination
directory and atomically rename it over `path`. temp_dir is an absolute,
root-confined scratch directory (already validated by the caller): when it
is non-NULL the temporary copy is instead created there (with a name unique
across the whole scratch directory) and atomically renamed into the
destination directory once fully written and fsynced. A rename across
filesystems (EXDEV) fails the write with an error; the file is never
silently copied into place. Pass NULL for the historical same-directory
behavior. --inplace writes never use temp_dir. */
directory and atomically rename it over `path`. temp_dir is a scratch
directory (an absolute path, or one the caller already resolved against the
destination root): when it is non-NULL the temporary copy is instead created
there (with a name unique across the whole scratch directory) and atomically
renamed into the destination directory once fully written and fsynced. When
that rename/link fails with EXDEV (the scratch dir is on another filesystem)
the write falls back to a non-atomic copy directly in the destination
directory, matching rsync. Pass NULL for the same-directory behavior.
--inplace writes never use temp_dir. */
bool file_to_disk_secure(const char* path, const void* data, unsigned long long data_size,
bool inplace, bool sparse, const FileMetadata* metadata,
bool preserve_executability, const char* temp_dir);
bool inplace, bool sparse, bool preallocate, const FileMetadata* metadata,
FileAttrPolicy policy, const char* temp_dir);
bool file_to_disk_secure_with_fsync(const char* path, const void* data,
unsigned long long data_size, bool inplace, bool sparse,
const FileMetadata* metadata, bool preserve_executability,
bool use_fsync, const char* temp_dir);
bool preallocate, const FileMetadata* metadata,
FileAttrPolicy policy, bool use_fsync, const char* temp_dir);
/* With update enabled, an existing newer destination is left untouched. The
check is descriptor-based for inplace writes; atomic replacement still has
an unavoidable final rename race without filesystem locking. */
bool file_to_disk_secure_update(const char* path, const void* data, unsigned long long data_size,
bool inplace, bool sparse, const FileMetadata* metadata,
bool preserve_executability, const char* temp_dir);
bool file_to_disk_secure_no_replace(const char* path, const void* data,
unsigned long long data_size, bool sparse,
const FileMetadata* metadata, bool preserve_executability,
bool inplace, bool sparse, bool preallocate,
const FileMetadata* metadata, FileAttrPolicy policy,
const char* temp_dir);
bool file_to_disk_secure_no_replace(const char* path, const void* data,
unsigned long long data_size, bool sparse, bool preallocate,
const FileMetadata* metadata, FileAttrPolicy policy,
const char* temp_dir);
/* Receiver write-path variant that also applies per-file xattrs (-X/-A) and the
* --fake-super stat xattr fd-relative before the final rename. `update` /
* `no_replace` / `use_fsync` mirror the plain wrappers above; `keep_partial`
* enables --partial best-effort retention of a failed write's temp. */
bool file_to_disk_secure_attrs(const char* path, const void* data, unsigned long long data_size,
bool inplace, bool sparse, bool preallocate,
const FileMetadata* metadata, FileAttrPolicy policy, bool update,
bool no_replace, bool use_fsync, const FileXattrList* xattrs,
bool fake_super, bool keep_partial, const char* temp_dir);
/* Atomic --link-dest install: replace `path` with a hard link to `basis_path`
(via a temp name + rename); fall back to a byte-identical local copy from
`data` when the link is impossible (EXDEV/EPERM/unsupported filesystem).
`metadata` is applied only on the copy fallback. */
`metadata` is applied only on the copy fallback. `preallocate` applies to
that copy fallback only (a hard-linked file shares the basis inode and is
never re-allocated). */
bool file_to_disk_secure_link(const char* path, const char* basis_path, const void* data,
unsigned long long data_size, const FileMetadata* metadata,
bool preserve_executability, bool use_fsync, const char* temp_dir);
unsigned long long data_size, bool preallocate,
const FileMetadata* metadata, FileAttrPolicy policy, bool use_fsync,
const char* temp_dir);
/* Like file_to_disk_secure_link, but the byte-copy fallback also applies the
* per-file xattrs (-X/-A) and --fake-super stat xattr (fd-relative). On a
* successful hard link no attributes are applied (the shared inode already
* carries the basis's). */
bool file_to_disk_secure_link_attrs(const char* path, const char* basis_path, const void* data,
unsigned long long data_size, bool preallocate,
const FileMetadata* metadata, FileAttrPolicy policy,
bool use_fsync, const FileXattrList* xattrs, bool fake_super,
const char* temp_dir);
#endif
+38
View File
@@ -0,0 +1,38 @@
#ifndef FILE_ATTR_H
#define FILE_ATTR_H
#include "config.h"
#include <stdbool.h>
#include <sys/stat.h>
/*
* Per-attribute receiver policy for applying a transmitted FileMetadata. This
* is the split-out replacement for the former single use_metadata bundle: each
* flag is applied independently, matching rsync's -p/-t/-o/-g/-E/-U semantics.
* `use_metadata` remains the transport/presence gate (whether the metadata frame
* travelled at all); this struct decides which attributes are ACTUALLY applied.
*
* It lives in its own header (rather than metadata.h) because xattr.h's
* fake_super_restore_fd() takes one and metadata.h <-> file_types.h form an
* include cycle that must not be entered from xattr.h.
*
* The mode leg is: perms wins over executability; an exec-bits-only change is
* made only when perms is off; when neither is set the receiver deliberately
* sets no source mode. file.c then substitutes the pre-existing destination
* mode for a brand-new destination with metadata it uses the sanitized
* source-mode-&-umask base (S_IWGRP|S_IWOTH cleared), and the fixed 0644
* default only when no metadata is available at all, so a no--p overwrite
* does not lose the destination's perms.
*/
typedef struct FileAttrPolicy {
bool perms; /* config->preserve_perms: apply the source mode bits */
bool times; /* config->preserve_times: apply the source mtime */
bool atimes; /* config->preserve_atimes (-U): apply the source atime */
bool executability; /* config->use_executability (-E): exec-bits-only mode */
} FileAttrPolicy;
/* Build the per-attribute policy from a connection's Config. A NULL config
* yields the all-off policy (no attribute application). */
FileAttrPolicy file_attr_policy_from_config(const Config* config);
#endif
+99 -22
View File
@@ -2,6 +2,7 @@
#include "log.h"
#include "utils.h"
#include <errno.h>
#include <limits.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
@@ -22,6 +23,8 @@ static void string_list_destroy(StringList* list) {
static bool string_list_add(StringList* list, const char* text) {
if (list->count == list->capacity) {
if (list->capacity > INT_MAX / 2)
return false;
int new_cap = list->capacity > 0 ? list->capacity * 2 : 16;
char** grown = realloc(list->items, (size_t)new_cap * sizeof(char*));
if (!grown)
@@ -51,11 +54,18 @@ static int normalize_entry(const char* raw, size_t len, bool strip_line_endings,
if (len == 0)
return 0;
if (raw[0] == '/') {
snprintf(err, err_size, "absolute path entries are not allowed: '%.*s'", (int)len, raw);
int print_len = len > (size_t)INT_MAX ? INT_MAX : (int)len;
snprintf(err, err_size, "absolute path entries are not allowed: '%.*s'", print_len, raw);
return -1;
}
/* Reject NUL bytes inside a token defensively. In NUL-delimited mode the
* delimiter itself is the final byte and is expected; in line mode any NUL is
* embedded garbage (strlen-based parsing would otherwise silently truncate). */
size_t scan_len = strip_line_endings ? len : len - 1;
if (memchr(raw, '\0', scan_len)) {
snprintf(err, err_size, "entry contains an embedded NUL byte");
return -1;
}
/* Reject NUL bytes inside a token defensively (NUL-delimited mode splits on
* them, so this only guards against embedded garbage). */
char* dup = malloc(len + 1);
if (!dup) {
snprintf(err, err_size, "memory allocation failed");
@@ -102,8 +112,27 @@ static int normalize_entry(const char* raw, size_t len, bool strip_line_endings,
return result;
}
/* Build the membership index over the exact entries only. `file_list_affects`
combines the exact/descendant lookups with a walk of the query's own ancestor
prefixes, so no ancestor prefix is ever materialized as a copy and the index
stays O(entry count) memory regardless of path depth. An empty entry (the
source root) sets whole_tree and short-circuits every query. */
static bool file_list_index_build(FileListSet* set, char* err, size_t err_size) {
if (!path_index_build(&set->index, (const char* const*)set->entries, (size_t)set->count)) {
snprintf(err, err_size, "memory allocation failed");
return false;
}
for (int i = 0; i < set->count; i++) {
if (set->entries[i][0] == '\0') {
set->whole_tree = true;
break;
}
}
return true;
}
static FileListSet* string_list_to_set(StringList* raw, char* err, size_t err_size) {
FileListSet* set = malloc(sizeof(FileListSet));
FileListSet* set = calloc(1, sizeof(FileListSet));
if (!set) {
snprintf(err, err_size, "memory allocation failed");
return NULL;
@@ -112,6 +141,10 @@ static FileListSet* string_list_to_set(StringList* raw, char* err, size_t err_si
set->entries = raw->items;
raw->items = NULL;
raw->count = 0;
if (!file_list_index_build(set, err, err_size)) {
file_list_destroy(set);
return NULL;
}
return set;
}
@@ -133,10 +166,20 @@ FileListSet* file_list_load(const char* path, bool null_separated, char* err, si
StringList raw = {0};
char* line = NULL;
size_t line_cap = 0;
ssize_t n;
bool ok = true;
char delim = null_separated ? '\0' : '\n';
while (ok && (n = getdelim(&line, &line_cap, delim, fp)) != -1) {
while (ok) {
ssize_t n = utils_getdelim_bounded(fp, &line, &line_cap, delim, UTILS_MAX_LINE_LEN);
if (n < 0) {
if (errno == EFBIG)
snprintf(err, err_size, "entry in file list exceeds %d bytes", (int)UTILS_MAX_LINE_LEN);
else
snprintf(err, err_size, "error reading file list: %s", strerror(errno));
ok = false;
break;
}
if (n == 0)
break;
int r = normalize_entry(line, (size_t)n, !null_separated, &raw, err, err_size);
if (r < 0) {
ok = false;
@@ -158,34 +201,68 @@ FileListSet* file_list_load(const char* path, bool null_separated, char* err, si
void file_list_destroy(FileListSet* set) {
if (!set)
return;
path_index_free(&set->index);
for (int i = 0; i < set->count; i++)
free(set->entries[i]);
free(set->entries);
free(set);
}
static bool path_has_prefix(const char* path, const char* prefix) {
size_t plen = strlen(prefix);
if (strncmp(path, prefix, plen) != 0)
return false;
return path[plen] == '/' || path[plen] == '\0';
}
bool file_list_affects(const FileListSet* set, const char* rel) {
if (!set)
return true;
if (!rel)
return false;
for (int i = 0; i < set->count; i++) {
const char* entry = set->entries[i];
if (entry[0] == '\0')
if (set->whole_tree)
return true; /* whole tree listed */
if (strcmp(rel, entry) == 0)
return true; /* the entry itself is listed */
if (path_has_prefix(rel, entry))
return true; /* rel lives under a listed directory */
if (path_has_prefix(entry, rel))
return true; /* rel is an ancestor directory of a listed entry */
/* An exact entry match means `rel` itself is listed. */
if (path_index_contains(&set->index, rel))
return true;
/* Otherwise `rel` is affected when a listed entry is an ancestor directory of
it; walk rel's own directory prefixes (which preserve path-boundary
semantics) and test each for an exact entry. No prefixes are stored. */
size_t len = strlen(rel);
while (len > 0) {
const char* slash = NULL;
for (size_t i = len; i-- > 0;) {
if (rel[i] == '/') {
slash = rel + i;
break;
}
}
if (!slash)
break;
len = (size_t)(slash - rel);
if (path_index_contains_n(&set->index, rel, len))
return true;
}
/* Finally `rel` is affected when it is an ancestor directory of a listed
entry (binary search for the first entry at or after `rel` + '/'). */
return path_index_has_descendant(&set->index, rel);
}
bool file_list_dir_in_scope(const FileListSet* set, const char* rel) {
if (!set || set->whole_tree)
return true;
if (!rel || rel[0] == '\0')
return false;
/* `rel` itself is listed, or one of its ancestor prefixes is an exact listed
directory (a listed prefix of a directory path is necessarily a
directory). */
size_t len = strlen(rel);
while (len > 0) {
const char* slash = NULL;
for (size_t i = len; i-- > 0;) {
if (rel[i] == '/') {
slash = rel + i;
break;
}
}
if (!slash)
break;
len = (size_t)(slash - rel);
if (path_index_contains_n(&set->index, rel, len))
return true;
}
return path_index_contains(&set->index, rel);
}
+20 -2
View File
@@ -1,6 +1,7 @@
#ifndef FILE_LIST_H
#define FILE_LIST_H
#include "utils.h"
#include <stdbool.h>
#include <stddef.h>
@@ -12,11 +13,18 @@
* of "." means the whole tree, absolute entries and ".." traversal are
* rejected at parse time. The set is immutable and shared read-only across
* scanner worker threads.
*/
*
* Membership is answered from `index`, built once at load time over the exact
* entries only: `index.exact` matches a listed path, the sorted view detects an
* ancestor directory of a listed entry, and `rel`'s own directory prefixes are
* matched against the exact set while descending. No ancestor prefix is stored
* as a separate string, so the index is O(entry count) memory however deep the
* paths are, and each query is O(path length) comparisons. */
typedef struct {
char** entries; /* normalized rel paths; "" means the whole tree */
int count;
PathIndex index;
bool whole_tree; /* an entry of "" lists the source root */
} FileListSet;
/* Load and validate a --files-from file. When `null_separated` (-0/--from0)
@@ -32,4 +40,14 @@ void file_list_destroy(FileListSet* set);
* this returns true, files are transferred only when it returns true. */
bool file_list_affects(const FileListSet* set, const char* rel);
/* True when the DIRECTORY `rel` (path relative to the source root) is inside a
* listed directory subtree: `rel` itself is a listed entry, or one of `rel`'s
* ancestor directory prefixes is an exact listed entry. Unlike
* file_list_affects this does NOT treat an ancestor of a listed entry as
* affected, so an implied parent directory of a listed file is not synchronized
* (rsync deletes nothing in it). With no set or a whole-tree set every
* directory is in scope. This is the delete-walker's "synchronized directory"
* predicate. */
bool file_list_dir_in_scope(const FileListSet* set, const char* rel);
#endif
+2227 -535
View File
File diff suppressed because it is too large. Load diff
+111 -6
View File
@@ -7,9 +7,72 @@
/* Server-side file receive/save path. */
/* Cumulative caps for the deferred directory-time accumulator. The sender may
* legitimately split a large tree across repeated STATUS_DIR_TIMES frames, so a
* per-frame bound is not enough: the receiver must bound the TOTAL it retains
* against a hostile sender. Mirror the delete-manifest limits
* (MAX_MANIFEST_ENTRIES / MAX_MANIFEST_BYTES): the entry count bounds the
* metadata array and the byte budget bounds the concatenated path strings. */
#define MAX_DIR_TIME_ENTRIES (1024 * 1024)
#define MAX_DIR_TIME_BYTES (16ULL * 1024 * 1024)
File* file_receive(const Config* config, int file_descriptor);
File* file_receive_directory(int file_descriptor);
File* file_receive_directory(int file_descriptor, const Config* config);
File* file_receive_dir_time(int file_descriptor, const Config* config);
File* file_receive_hardlink(int file_descriptor);
File* file_receive_symlink(int file_descriptor, const Config* config);
File* file_receive_special(int file_descriptor);
bool file_special_rdev_valid(int32_t major, int32_t minor, mode_t mode);
File* receive_incremental_check(int fd, const Config* config, bool* skipped);
/* Extended variant used by the receiver. `would_transfer` (may be NULL) is set
* true only on the server-contacting --dry-run path when the file is not up to
* date: the receiver has already sent STATUS_DRY_RUN_TRANSFER and returns NULL
* without storing anything. On that path `*skipped` is true for an up-to-date
* (STATUS_OK) file and both flags are false for a genuine error. */
File* receive_incremental_check_ex(int fd, const Config* config, bool* skipped,
bool* would_transfer);
/* P7 Wave D directory-time accumulator. The receiver collects the metadata of
* every directory it creates/receives (STATUS_MKDIR with metadata and/or the
* trailing STATUS_DIR_TIMES frame(s)) and applies the times only at the END of the
* transfer, after all children have been written and after the delete /
* --delay-updates phases have committed (writing or removing a child bumps the
* parent's mtime). -O/--omit-dir-times skips the application entirely. The
* list owns deep copies of the paths and metadata; freed on every path. */
typedef struct {
char** paths; /* owned, destination-relative wire paths */
FileMetadata* entries; /* owned, parallel to paths */
FileXattrList** xattrs; /* owned, parallel to paths; NULL when none */
size_t count;
size_t capacity;
size_t bytes; /* cumulative strlen of every retained path */
} DirTimeList;
/* Capture gate shared by the sender-side and receiver-side sinks: directory
* metadata is accumulated only when a directory attribute is requested
* (-p/--perms for directory modes, or -t/--times for directory mtimes with
* -O/--omit-dir-times not suppressing them) and metadata rides the wire. Kept
* here, next to the accumulator it guards, so both call sites express the same
* condition. */
bool dir_metadata_should_capture(const Config* config);
void dir_time_list_init(DirTimeList* list);
void dir_time_list_free(DirTimeList* list);
/* Deep-copy one directory's path + metadata (and, when non-NULL, its captured
* xattr/ACL block) into the list. Returns false on allocation failure OR when
* the cumulative entry/byte caps would be exceeded (the caller fails the
* transfer). */
bool dir_time_list_add(DirTimeList* list, const char* wire_path, const FileMetadata* metadata,
const FileXattrList* xattrs);
/* Apply every accumulated directory's metadata beneath `root_directory`,
* confined fd-relative: ownership through the negotiated identity policy,
* times (mtime, plus atime when -U captured one under -t), the mode (through
* --chmod when configured, under -p), and the captured xattrs/ACLs (under
* -X/-A). Best-effort per entry: an absent directory (an empty/pruned source
* dir that was deliberately not created) or a non-directory at the path is
* skipped QUIETLY, an unreachable one with a warning, and never fatal. */
void dir_metadata_list_apply(const DirTimeList* list, const char* root_directory,
const Config* config);
/* A received delete-manifest frame: the keep-set (`keeps`, destination-relative
paths the sender transferred/keeps) plus `protected`, destination-relative
@@ -25,11 +88,18 @@ typedef struct DeleteManifest {
ArrayList* keeps;
ArrayList* protected;
ArrayList* missing;
/* Destination-relative paths of the directories the sender synchronized for
this run. The extras walker only removes entries directly inside one of
these (the receive root is the "." sentinel); `--files-from` runs therefore
leave untransmitted directories and the unlisted parts of listed ones
alone, matching rsync's "delete only in synchronized directories". */
ArrayList* dirs;
} DeleteManifest;
void delete_manifest_free(DeleteManifest* manifest);
/* Read a delete-manifest frame: keep count + keeps, then protected count +
protected prefixes, then missing count + missing paths (self-delimiting; the
/* Read a delete-manifest frame (protocol 2.23.0): keep count + keeps, then
protected count + protected prefixes, then missing count + missing paths,
then synchronized-directory count + directory paths (self-delimiting; the
leading STATUS_MANIFEST code has been consumed). Returns an owned
DeleteManifest, or NULL after signalling STATUS_ERROR on a malformed frame. */
DeleteManifest* receive_manifest_entries(int fd);
@@ -48,11 +118,46 @@ bool manifest_delete_extras(const Config* config, DeleteManifest* manifest);
confinement or I/O error (the run then fails); tolerated per-path cases are
reported and skipped. */
bool manifest_delete_missing_args(const Config* config, DeleteManifest* manifest);
/* Budgeted form of manifest_delete_missing_args for the per-directory delete
session: each removed mirror draws from `max_delete` (SIZE_MAX = unlimited)
and the tallies are accumulated into `*deleted`/`*skipped`. `*limit_hit` is set
when the budget stopped the pass with entries left over. Returns false only
on a genuine deletion error. */
bool manifest_delete_missing_args_limited(const Config* config, DeleteManifest* manifest,
size_t max_delete, size_t* deleted, size_t* skipped,
bool* limit_hit);
/* Outcome of committing a delete manifest. LIMIT_REACHED reports rsync's
partial --max-delete result: the budget allowed some deletions and the rest
were skipped (the run still stores all file data but the client exits 25). */
typedef enum {
DELETE_COMMIT_OK = 0,
DELETE_COMMIT_LIMIT_REACHED,
DELETE_COMMIT_ERROR
} DeleteCommitResult;
/* Run every deletion family the manifest carries: the --delete-missing-args
exact-path deletions first (user requests are not blocked by exclusion
protection), then the ordinary extras walk when --delete is active. Returns
true when nothing to do or everything committed. */
bool manifest_delete_all(const Config* config, DeleteManifest* manifest);
protection), then the ordinary extras walk when --delete is active. Both
share one --max-delete budget. Returns DELETE_COMMIT_OK when nothing was to
do or everything committed, DELETE_COMMIT_LIMIT_REACHED when the budget
stopped part of the work, or DELETE_COMMIT_ERROR on a genuine failure. */
DeleteCommitResult manifest_delete_all(const Config* config, DeleteManifest* manifest);
/* Like manifest_delete_all, but reports how many destination entries the commit
removed (for the end-of-transfer wire stats). `deleted` may be NULL. */
DeleteCommitResult manifest_delete_all_counted(const Config* config, DeleteManifest* manifest,
size_t* deleted);
/* -n/--dry-run --delete would-delete reporting: walk the destination exactly as
the delete pass would and append (strdup'd) destination-relative paths that
WOULD be removed to `out`, without touching disk. Uses the same staging-dir,
basis-dir and protected-prefix skips as the real commit. Returns true on a
clean walk; `*count_out` receives the number of paths appended. */
bool manifest_would_delete_list(const Config* config, DeleteManifest* manifest, ArrayList* out,
size_t* count_out);
/* Convert one basis-directory path to the receive-root-relative protection
prefix the delete walker uses (NULL when it lies outside the root). Exposed
for unit tests of the root-of-"/" and normalization edge cases. */
char* file_receive_basis_delete_relative(const Config* config, const char* path);
/* Outcome of a single file_save_to_disk operation. The receiver needs to
distinguish "written" from "skipped" so --remove-source-files can be told
+46 -10
View File
@@ -10,23 +10,44 @@
#include <time.h>
#include <unistd.h>
#include "charset.h"
#include "compression.h"
#include "data.h"
#include "file.h"
#include "log.h"
#include "metadata.h"
#include "protocol.h"
#include "xattr.h"
/* Transmit a device/special node (--devices / --specials) as a STATUS_SPECIAL
* frame: the destination path, the metadata frame (whose mode's S_IFMT bits
* carry the node kind) and the device rdev major/minor. The receiver validates
* the kind and rdev and recreates the node (privilege-gating the mknod). */
bool file_send_special(const File* file, int file_descriptor, bool use_metadata) {
if (!file || !file_wire_path(file))
return false;
if (!send_status(file_descriptor, STATUS_SPECIAL))
return false;
if (!send_wire_str(file_descriptor, file_wire_path(file)))
return false;
if (use_metadata && !metadata_send(file_descriptor, file->metadata))
return false;
int32_t major = file->rdev_major;
int32_t minor = file->rdev_minor;
return send_n_data(file_descriptor, &major, sizeof(major)) &&
send_n_data(file_descriptor, &minor, sizeof(minor));
}
bool file_send_single_calls(File* file, int file_descriptor, bool use_metadata,
int compression_level, bool send_path) {
return file_send_single_calls_with_skip(file, file_descriptor, use_metadata, compression_level,
send_path, NULL, -1, 0);
send_path, NULL, -1, 0, false);
}
bool file_send_single_calls_with_skip(File* file, int file_descriptor, bool use_metadata,
int compression_level, bool send_path,
char* const* skip_suffixes, int skip_count,
int compression_threads) {
int compression_threads, bool send_xattrs) {
if (!file || !file->path || !file->data || (file->data->size != 0 && !file->data->data))
return false;
const Data* data_to_send = file->data;
@@ -41,7 +62,7 @@ bool file_send_single_calls_with_skip(File* file, int file_descriptor, bool use_
}
data_to_send = compressed_data;
}
if (send_path && !send_str(file_descriptor, file_wire_path(file))) {
if (send_path && !send_wire_str(file_descriptor, file_wire_path(file))) {
data_destroy(compressed_data);
return false;
}
@@ -49,6 +70,10 @@ bool file_send_single_calls_with_skip(File* file, int file_descriptor, bool use_
data_destroy(compressed_data);
return false;
}
if (send_xattrs && !xattr_send(file_descriptor, file ? file->xattrs : NULL)) {
data_destroy(compressed_data);
return false;
}
if (!send_data(file_descriptor, data_to_send)) {
data_destroy(compressed_data);
return false;
@@ -60,25 +85,27 @@ bool file_send_single_calls_with_skip(File* file, int file_descriptor, bool use_
bool file_send_sendfile(File* file, int file_descriptor, bool use_metadata, int compression_level,
bool send_path) {
return file_send_sendfile_with_skip(file, file_descriptor, use_metadata, compression_level,
send_path, NULL, -1, 0);
send_path, NULL, -1, 0, false);
}
bool file_send_sendfile_with_skip(File* file, int file_descriptor, bool use_metadata,
int compression_level, bool send_path, char* const* skip_suffixes,
int skip_count, int compression_threads) {
int skip_count, int compression_threads, bool send_xattrs) {
if (!file || !file->path || !file->data)
return false;
if (compression_level > 0)
return file_send_single_calls_with_skip(file, file_descriptor, use_metadata, compression_level,
send_path, skip_suffixes, skip_count,
compression_threads);
compression_threads, send_xattrs);
if (send_path && !send_str(file_descriptor, file_wire_path(file)))
if (send_path && !send_wire_str(file_descriptor, file_wire_path(file)))
return false;
if (use_metadata && !metadata_send(file_descriptor, file->metadata))
return false;
if (send_xattrs && !xattr_send(file_descriptor, file ? file->xattrs : NULL))
return false;
int fd = open(file->path, O_RDONLY);
int fd = file_open_for_read(file->path);
if (fd == -1) {
log_perror("Could not open file for sendfile");
return false;
@@ -116,10 +143,17 @@ bool file_send_sendfile_with_skip(File* file, int file_descriptor, bool use_meta
}
off_t offset = 0;
/* A non-positive --timeout disables the deadline: poll blocks until the
* socket is writable (rsync's --timeout=0 default). */
int io_timeout_sec = protocol_get_io_timeout_sec();
struct timespec deadline;
if (io_timeout_sec > 0) {
clock_gettime(CLOCK_MONOTONIC, &deadline);
deadline.tv_sec += 60;
deadline.tv_sec += io_timeout_sec;
}
while ((unsigned long long)offset < file_size) {
int timeout = -1;
if (io_timeout_sec > 0) {
struct timespec now;
clock_gettime(CLOCK_MONOTONIC, &now);
long long remaining = (long long)(deadline.tv_sec - now.tv_sec) * 1000LL +
@@ -128,8 +162,9 @@ bool file_send_sendfile_with_skip(File* file, int file_descriptor, bool use_meta
close(fd);
return false;
}
timeout = remaining > INT_MAX ? INT_MAX : (int)remaining;
}
struct pollfd pfd = {.fd = file_descriptor, .events = POLLOUT};
int timeout = remaining > INT_MAX ? INT_MAX : (int)remaining;
int polled = poll(&pfd, 1, timeout);
if (polled <= 0 || (pfd.revents & (POLLERR | POLLHUP | POLLNVAL))) {
close(fd);
@@ -147,6 +182,7 @@ bool file_send_sendfile_with_skip(File* file, int file_descriptor, bool use_meta
close(fd);
return false;
}
protocol_note_bytes_written((unsigned long long)sent);
}
close(fd);
+3 -2
View File
@@ -6,16 +6,17 @@
/* Client-side file send path. */
bool file_send_special(const File* file, int file_descriptor, bool use_metadata);
bool file_send_single_calls(File* file, int file_descriptor, bool use_metadata,
int compression_level, bool send_path);
bool file_send_single_calls_with_skip(File* file, int file_descriptor, bool use_metadata,
int compression_level, bool send_path,
char* const* skip_suffixes, int skip_count,
int compression_threads);
int compression_threads, bool send_xattrs);
bool file_send_sendfile(File* file, int file_descriptor, bool use_metadata, int compression_level,
bool send_path);
bool file_send_sendfile_with_skip(File* file, int file_descriptor, bool use_metadata,
int compression_level, bool send_path, char* const* skip_suffixes,
int skip_count, int compression_threads);
int skip_count, int compression_threads, bool send_xattrs);
#endif
+30 -172
View File
@@ -1,129 +1,7 @@
#include <errno.h>
#include <fcntl.h>
#include <libgen.h>
#include <limits.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <unistd.h>
#include "file_store.h"
#include "metadata.h"
#include "utils.h"
static int authorized_root_fd = -1;
static char* authorized_root_path;
static bool path_is_within_root(const char* root, const char* path) {
size_t root_length = strlen(root);
return strncmp(root, path, root_length) == 0 &&
(path[root_length] == '\0' || path[root_length] == '/');
}
bool file_store_set_authorized_root(int fd, const char* canonical_path) {
char* new_path = canonical_path ? str_dup(canonical_path) : NULL;
if (canonical_path && !new_path) {
authorized_root_fd = -1;
free(authorized_root_path);
authorized_root_path = NULL;
return false;
}
free(authorized_root_path);
authorized_root_path = new_path;
authorized_root_fd = fd;
return true;
}
int file_store_open_secure_parent(const char* path, char** leaf_out) {
char* copy = str_dup(path);
if (!copy)
return -1;
char* parent = dirname(copy);
const char* slash = strrchr(path, '/');
char* leaf = str_dup(slash ? slash + 1 : path);
if (!leaf) {
free(copy);
return -1;
}
int fd;
if (authorized_root_fd >= 0) {
if (!authorized_root_path || path[0] != '/' ||
!path_is_within_root(authorized_root_path, path)) {
free(copy);
free(leaf);
return -1;
}
fd = dup(authorized_root_fd);
if (fd < 0) {
free(copy);
free(leaf);
return -1;
}
size_t root_length = strlen(authorized_root_path);
char* relative = str_dup(path + root_length);
if (!relative) {
free(copy);
free(leaf);
close(fd);
return -1;
}
free(copy);
copy = relative;
parent = dirname(copy);
} else {
fd = (parent[0] == '/') ? open("/", O_RDONLY | O_DIRECTORY | O_CLOEXEC)
: open(".", O_RDONLY | O_DIRECTORY | O_CLOEXEC);
}
if (fd < 0) {
free(copy);
free(leaf);
return -1;
}
char* save = NULL;
char* component = strtok_r(parent, "/", &save);
while (component) {
if (strcmp(component, "..") == 0) {
close(fd);
free(copy);
free(leaf);
return -1;
}
if (strcmp(component, ".") != 0) {
int next = openat(fd, component, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
if (next < 0 && errno == ENOENT) {
if (mkdirat(fd, component, 0755) == 0 || errno == EEXIST)
next = openat(fd, component, O_RDONLY | O_DIRECTORY | O_NOFOLLOW | O_CLOEXEC);
}
if (next < 0) {
close(fd);
free(copy);
free(leaf);
return -1;
}
close(fd);
fd = next;
}
component = strtok_r(NULL, "/", &save);
}
free(copy);
*leaf_out = leaf;
return fd;
}
bool file_store_rename_secure(const char* old_path, const char* new_path) {
char *old_leaf = NULL, *new_leaf = NULL;
int old_parent = file_store_open_secure_parent(old_path, &old_leaf);
int new_parent = file_store_open_secure_parent(new_path, &new_leaf);
bool ok = old_parent >= 0 && new_parent >= 0 &&
renameat(old_parent, old_leaf, new_parent, new_leaf) == 0;
if (old_parent >= 0)
close(old_parent);
if (new_parent >= 0)
close(new_parent);
free(old_leaf);
free(new_leaf);
return ok;
}
static bool write_all(int fd, const void* data, unsigned long long size) {
const unsigned char* p = data;
@@ -139,60 +17,40 @@ static bool write_all(int fd, const void* data, unsigned long long size) {
return true;
}
bool file_store_write_secure(const char* path, const void* data, unsigned long long data_size,
bool inplace, bool sparse, const FileMetadata* metadata,
bool preserve_executability) {
char* leaf = NULL;
int dirfd = file_store_open_secure_parent(path, &leaf);
if (dirfd < 0)
/* A run of NUL bytes at least this long is emitted as a hole (lseek) rather
* than written, so the resulting file is genuinely sparse on the filesystem. */
#define SPARSE_HOLE_MIN 4096U
/* Sparse-aware writer (--sparse/-S). Walks `data`; any all-zero run of at
* least SPARSE_HOLE_MIN bytes is skipped with lseek(SEEK_CUR) so the block is
* never allocated (a real hole on the destination); every other byte is written
* normally. The file is pre-sized with ftruncate by the callers before this
* runs, so holes are guaranteed and the offset bookkeeping stays correct
* (each lseek advances the fd offset exactly as a write of that many bytes
* would). After the final run, ftruncate(size) guarantees the logical size is
* exactly `size` even when the tail was a hole. The full file image is in
* memory, so no wire change is needed. Returns false on I/O error. */
bool file_store_write_sparse(int fd, const unsigned char* data, unsigned long long size) {
unsigned long long i = 0;
while (i < size) {
if (data[i] == 0) {
unsigned long long run_start = i;
while (i < size && data[i] == 0)
i++;
unsigned long long run_len = i - run_start;
if (run_len >= SPARSE_HOLE_MIN) {
if (lseek(fd, (off_t)run_len, SEEK_CUR) < 0)
return false;
} else if (!write_all(fd, data + run_start, run_len)) {
return false;
int fd = -1;
bool ok = false;
if (inplace) {
fd = openat(dirfd, leaf, O_WRONLY | O_CREAT | O_TRUNC | O_CLOEXEC | O_NOFOLLOW, 0644);
if (fd >= 0) {
if (!sparse || data_size == 0 || ftruncate(fd, (off_t)data_size) == 0)
ok = write_all(fd, data, data_size);
if (ok && metadata)
ok = file_restore_metadata_fd(fd, metadata, preserve_executability);
}
} else {
int tmp_size = snprintf(NULL, 0, ".%s.tmp.%ld.%u", leaf, (long)getpid(), 99U);
if (tmp_size < 0) {
close(dirfd);
free(leaf);
unsigned long long run_start = i;
while (i < size && data[i] != 0)
i++;
if (!write_all(fd, data + run_start, i - run_start))
return false;
}
char* tmp = malloc((size_t)tmp_size + 1);
if (!tmp) {
close(dirfd);
free(leaf);
return false;
}
for (unsigned int i = 0; i < 100 && !ok; ++i) {
snprintf(tmp, (size_t)tmp_size + 1, ".%s.tmp.%ld.%u", leaf, (long)getpid(), i);
fd = openat(dirfd, tmp, O_WRONLY | O_CREAT | O_EXCL | O_CLOEXEC | O_NOFOLLOW, 0600);
if (fd < 0)
continue;
if (sparse && data_size > 0)
ok = ftruncate(fd, (off_t)data_size) == 0;
if (ok || (!sparse || data_size == 0))
ok = write_all(fd, data, data_size);
if (ok && metadata)
ok = file_restore_metadata_fd(fd, metadata, preserve_executability);
if (close(fd) != 0)
ok = false;
fd = -1;
if (ok && renameat(dirfd, tmp, dirfd, leaf) != 0)
ok = false;
if (!ok)
unlinkat(dirfd, tmp, 0);
}
free(tmp);
}
if (fd >= 0)
close(fd);
close(dirfd);
free(leaf);
return ok;
return ftruncate(fd, (off_t)size) == 0;
}
+7 -7
View File
@@ -1,14 +1,14 @@
#ifndef FILE_STORE_H
#define FILE_STORE_H
#include "file.h"
#include <stdbool.h>
bool file_store_set_authorized_root(int fd, const char* canonical_path);
int file_store_open_secure_parent(const char* path, char** leaf_out);
bool file_store_rename_secure(const char* old_path, const char* new_path);
bool file_store_write_secure(const char* path, const void* data, unsigned long long data_size,
bool inplace, bool sparse, const FileMetadata* metadata,
bool preserve_executability);
/* Sparse-aware write (--sparse/-S): every all-zero run of at least
* SPARSE_HOLE_MIN bytes is skipped with lseek(SEEK_CUR) so it becomes a real
* hole; every other byte is written. The caller pre-sizes the file with
* ftruncate; this function also ftruncate()s to `size` at the end so a trailing
* hole keeps the exact logical length. Shared by the file_store and file write
* paths. Returns false on write/lseek/ftruncate error. */
bool file_store_write_sparse(int fd, const unsigned char* data, unsigned long long size);
#endif
+62
View File
@@ -2,6 +2,8 @@
#define FILE_TYPES_H
#include "data.h"
#include "format.h"
#include "xattr.h"
#include <stdbool.h>
#include <sys/stat.h>
@@ -13,6 +15,18 @@ typedef struct {
gid_t gid;
time_t mtime_sec;
long mtime_nsec;
/* Optional access time (-U/--atimes) and creation/birth time (-N/--crtimes),
* appended for protocol 2.12.0. The SENDER sets the corresponding *_valid
* flag only when the preserve option is active (and, for crtime, only when
* the source platform exposed a birth time via statx STATX_BTIME). The wire
* always carries the fields and the flags; a false flag tells the receiver to
* ignore the value. */
bool atime_valid;
time_t atime_sec;
long atime_nsec;
bool crtime_valid;
time_t crtime_sec;
long crtime_nsec;
} FileMetadata;
typedef struct {
@@ -29,12 +43,60 @@ typedef struct {
/* True when this entry is an explicit directory entry (--dirs mode): the
* receiver creates the directory instead of writing a regular file. */
bool is_dir;
/* Receiver-only (P7 Wave D): this is a STATUS_DIR_TIMES entry. It carries a
* traversed source directory's metadata for DEFERRED application, but must
* NEVER create the directory: the scanner captures every traversed directory
* (including empty ones whose parents no child write created), so creation
* would resurrect the empty dirs that FastSync deliberately never transfers.
* file_save_to_disk_full short-circuits such an entry as FILE_SAVE_SKIPPED,
* and the sink still accumulates the metadata into its DirTimeList. */
bool dir_time_only;
/* Receiver-only, --link-dest: when set, install the destination entry as a
* hard link to this absolute (root-confined) path instead of writing
* `data`. The matching code has already verified the link target's content
* equals the incoming file, and `data` is kept as the cross-filesystem
* fallback (a local copy) if the hard link cannot be created. */
char* basis_link;
/* --hard-links (-H), sender + receiver wire state. link_group is a run-local
* id shared by every member of one source inode (0 = not part of a group).
* The FIRST member (link_first == true) carries its data on the wire and is
* written normally; every sibling (link_first == false) carries NO data and
* hardlink_target holds the first member's wire path so the receiver can link
* to (or copy from) the already-installed first member. */
int link_group;
bool link_first;
char* hardlink_target;
/* Symlink-type entry (-l/--links, or -k/--copy-dirlinks' keep-as-symlink
* branch). When true, `symlink_target` holds the (sender-munged, if
* --munge-links) target string that is carried on the wire; the receiver
* creates a symlink to (an unmunged) target instead of writing regular-file
* data. `data` is empty for a symlink entry. Sender + receiver state. */
bool is_symlink;
char* symlink_target;
/* Phase 4 special/devices: when `is_special` is true this entry is a device
* or special node to be RECREATED on the destination (mknod/mkfifo) rather
* than written from `data`. The concrete node kind is derived from the
* metadata mode's S_IFMT bits (receiver-validated), and rdev_major/minor
* carry the device major/minor numbers for char/block devices. CROSSES the
* wire (protocol 2.13.0). */
bool is_special;
int32_t rdev_major;
int32_t rdev_minor;
/* Phase-4 xattrs (-X/--xattrs, -A/--acls). Sender: captured from the source
* file when use_xattrs is set; transmitted in the per-file metadata frame.
* Receiver: parsed off the wire, attached here, and applied fd-relative on
* the written file. NULL/0 == the file carries no xattrs. */
FileXattrList* xattrs;
/* Sender-side output-parity state (never serialized): the receiver-reported
* pre-transfer destination snapshot for this entry, filled by the per-file
* STATUS_CHECK exchange when report_dest_info is set. `known` is false when
* no report was requested/received, in which case -i/--out-format treats the
* entry conservatively as newly created. */
OutputDestState dest_state;
/* Receiver-only wire-stats tally: the number of bytes reconstructed from the
* basis file (matched delta blocks) for this entry. 0 when the file was sent
* whole. Accumulated into ReceiverStats.matched_data by the receiver sink. */
unsigned long long matched_bytes;
} File;
/* The path that should be sent on the wire and used for the receiver-side
+616 -225
View File
@@ -1,162 +1,27 @@
#include "filter.h"
#include "log.h"
#include "utils.h"
#include <ctype.h>
#include <errno.h>
#include <limits.h>
#include <stdarg.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
/* ---- Single rule parsing ---- */
static bool rule_text_is_unsupported_word(const char* p, size_t len) {
static const char* const words[] = {"merge", "dir-merge", "hide", "show",
"protect", "risk", "clear"};
for (size_t i = 0; i < sizeof(words) / sizeof(words[0]); i++) {
size_t wl = strlen(words[i]);
if (len == wl && strncmp(p, words[i], wl) == 0)
return true;
}
return false;
/* Write a diagnostic message into the caller's optional buffer. A NULL `err`
* (or a zero size) is a no-op, so a caller that only needs the boolean status
* may pass NULL without the snprintf-on-NULL undefined behaviour. */
static void filter_set_error(char* err, size_t err_size, const char* fmt, ...) {
if (!err || err_size == 0)
return;
va_list ap;
va_start(ap, fmt);
vsnprintf(err, err_size, fmt, ap);
va_end(ap);
}
/* rsync include/exclude rule modifiers we do NOT implement. A rule whose +/- is
* immediately followed by one of these is rejected instead of being silently
* parsed as a literal pattern. */
static bool is_unsupported_rule_modifier(char c) {
return c == '!' || c == 'C' || c == 's' || c == 'r' || c == 'p' || c == 'x';
}
FilterRule* filter_rule_parse(const char* line, char* err, size_t err_size) {
if (err && err_size > 0)
err[0] = '\0';
if (!line)
return NULL;
char* text = str_dup(line);
if (!text) {
if (err)
snprintf(err, err_size, "memory allocation failed");
return NULL;
}
size_t len = strlen(text);
while (len > 0 && (text[len - 1] == '\n' || text[len - 1] == '\r'))
text[--len] = '\0';
const char* p = text;
while (*p == ' ' || *p == '\t')
p++;
if (*p == '\0') {
snprintf(err, err_size, "empty filter rule");
free(text);
return NULL;
}
FilterAction action = FILTER_ACTION_EXCLUDE;
if (*p == '+' || *p == '-') {
action = *p == '+' ? FILTER_ACTION_INCLUDE : FILTER_ACTION_EXCLUDE;
p++;
/* rsync attaches rule modifiers directly to the +/- (e.g. "-s foo"). Only
* the '/' anchor modifier is supported; anything else is a clear error
* rather than a silently-ignored literal. */
if (*p != ' ' && *p != '\t' && *p != '\0' && is_unsupported_rule_modifier(*p)) {
snprintf(err, err_size,
"filter rule modifier '%c' is not supported (only the '/' anchor after +/- "
"is implemented; put a space between +/- and the pattern)",
*p);
free(text);
return NULL;
}
while (*p == ' ' || *p == '\t')
p++;
} else {
/* ':' (dir-merge) and '.' (merge) are rsync filter-rule shorthands. At the
* start of a rule they mean "merge this file", so reject them instead of
* silently turning them into inert exclude patterns. */
if (*p == ':' || *p == '.' || *p == '!') {
snprintf(err, err_size,
"filter rule starting with '%c' is not supported (merge/dir-merge/list-clear "
"shorthands are not implemented; use +/- include/exclude rules)",
*p);
free(text);
return NULL;
}
const char* sp = p;
while (*sp != '\0' && *sp != ' ' && *sp != '\t')
sp++;
size_t word_len = (size_t)(sp - p);
if (rule_text_is_unsupported_word(p, word_len)) {
snprintf(err, err_size,
"'%.*s' filter directives are not supported (only +/- include/exclude rules "
"with an optional '/' anchor and trailing '/' dir marker)",
(int)word_len, p);
free(text);
return NULL;
}
if (word_len == strlen("include") && strncmp(p, "include", word_len) == 0) {
action = FILTER_ACTION_INCLUDE;
p = sp;
} else if (word_len == strlen("exclude") && strncmp(p, "exclude", word_len) == 0) {
action = FILTER_ACTION_EXCLUDE;
p = sp;
}
while (*p == ' ' || *p == '\t')
p++;
}
if (*p == '\0') {
snprintf(err, err_size, "filter rule has no pattern");
free(text);
return NULL;
}
/* A pattern beginning with '/' is anchored (either as "-/foo" or "- /foo"). */
bool anchored = false;
if (*p == '/') {
anchored = true;
p++;
while (*p == ' ' || *p == '\t')
p++;
}
if (*p == '\0') {
snprintf(err, err_size, "filter rule has no pattern after '/' anchor");
free(text);
return NULL;
}
/* Pattern runs to the end of the rule; a single trailing '/' marks dir-only. */
size_t pat_len = strlen(p);
bool dir_only = false;
if (pat_len > 1 && p[pat_len - 1] == '/') {
dir_only = true;
pat_len--;
} else if (pat_len == 1 && p[0] == '/') {
/* "//" anchored with nothing after: meaningless. */
snprintf(err, err_size, "filter rule has no pattern");
free(text);
return NULL;
}
FilterRule* rule = calloc(1, sizeof(FilterRule));
if (!rule) {
snprintf(err, err_size, "memory allocation failed");
free(text);
return NULL;
}
rule->pattern = malloc(pat_len + 1);
if (!rule->pattern) {
free(rule);
snprintf(err, err_size, "memory allocation failed");
free(text);
return NULL;
}
memcpy(rule->pattern, p, pat_len);
rule->pattern[pat_len] = '\0';
rule->action = action;
rule->anchored = anchored;
rule->dir_only = dir_only;
rule->owner = NULL;
free(text);
return rule;
}
/* ---- Ordered rule lists ---- */
void filter_rule_free(FilterRule* rule) {
if (!rule)
@@ -166,8 +31,6 @@ void filter_rule_free(FilterRule* rule) {
free(rule);
}
/* ---- Ordered rule lists ---- */
FilterRuleList* filter_rule_list_create(void) {
return calloc(1, sizeof(FilterRuleList));
}
@@ -176,6 +39,8 @@ bool filter_rule_list_add(FilterRuleList* list, FilterRule* rule) {
if (!list || !rule)
return false;
if (list->count == list->capacity) {
if (list->capacity > INT_MAX / 2)
return false;
int new_cap = list->capacity > 0 ? list->capacity * 2 : 8;
FilterRule** grown = realloc(list->items, (size_t)new_cap * sizeof(FilterRule*));
if (!grown)
@@ -187,28 +52,42 @@ bool filter_rule_list_add(FilterRuleList* list, FilterRule* rule) {
return true;
}
bool filter_rule_list_parse_append(FilterRuleList* list, const char* line, char* err,
size_t err_size) {
FilterRule* rule = filter_rule_parse(line, err, err_size);
if (!rule)
return false;
if (!filter_rule_list_add(list, rule)) {
filter_rule_free(rule);
snprintf(err, err_size, "memory allocation failed");
return false;
}
return true;
}
void filter_rule_list_free(FilterRuleList* list) {
if (!list)
return;
for (int i = 0; i < list->count; i++)
filter_rule_free(list->items[i]);
for (int i = 0; i < list->dir_merge_count; i++)
free(list->dir_merge_names[i]);
free(list->dir_merge_names);
free(list->items);
free(list);
}
/* Register a per-directory merge-file basename (for "dir-merge NAME"/": NAME"
* and -F's .rsync-filter). Duplicate names are ignored. */
bool filter_rule_list_add_dir_merge(FilterRuleList* list, const char* name) {
if (!list || !name || name[0] == '\0')
return false;
for (int i = 0; i < list->dir_merge_count; i++) {
if (strcmp(list->dir_merge_names[i], name) == 0)
return true;
}
if (list->dir_merge_count == list->dir_merge_capacity) {
int new_cap = list->dir_merge_capacity > 0 ? list->dir_merge_capacity * 2 : 4;
char** grown = realloc(list->dir_merge_names, (size_t)new_cap * sizeof(char*));
if (!grown)
return false;
list->dir_merge_names = grown;
list->dir_merge_capacity = new_cap;
}
char* dup = str_dup(name);
if (!dup)
return false;
list->dir_merge_names[list->dir_merge_count++] = dup;
return true;
}
static bool set_rule_owner(FilterRule* rule, const char* owner) {
char* dup = str_dup(owner ? owner : "");
if (!dup)
@@ -218,7 +97,315 @@ static bool set_rule_owner(FilterRule* rule, const char* owner) {
return true;
}
/* ---- CVS default excludes (-C) ---- */
/* ---- Rule parsing ---- */
/* A short rule prefix is a single character; a long rule name is alphabetic
* (with '-'). `is_short` distinguishes the modifier-attachment rules. */
typedef enum {
RULE_KIND_EXCLUDE,
RULE_KIND_INCLUDE,
RULE_KIND_HIDE,
RULE_KIND_SHOW,
RULE_KIND_PROTECT,
RULE_KIND_RISK,
RULE_KIND_MERGE,
RULE_KIND_DIR_MERGE,
RULE_KIND_CLEAR,
RULE_KIND_UNKNOWN,
} RuleKind;
static bool short_rule_char(char c, RuleKind* kind) {
switch (c) {
case '-':
*kind = RULE_KIND_EXCLUDE;
return true;
case '+':
*kind = RULE_KIND_INCLUDE;
return true;
case 'H':
*kind = RULE_KIND_HIDE;
return true;
case 'S':
*kind = RULE_KIND_SHOW;
return true;
case 'P':
*kind = RULE_KIND_PROTECT;
return true;
case 'R':
*kind = RULE_KIND_RISK;
return true;
case '.':
*kind = RULE_KIND_MERGE;
return true;
case ':':
*kind = RULE_KIND_DIR_MERGE;
return true;
case '!':
*kind = RULE_KIND_CLEAR;
return true;
default:
return false;
}
}
static bool long_rule_name(const char* name, size_t len, RuleKind* kind) {
struct {
const char* word;
RuleKind kind;
} table[] = {
{"exclude", RULE_KIND_EXCLUDE}, {"include", RULE_KIND_INCLUDE},
{"hide", RULE_KIND_HIDE}, {"show", RULE_KIND_SHOW},
{"protect", RULE_KIND_PROTECT}, {"risk", RULE_KIND_RISK},
{"merge", RULE_KIND_MERGE}, {"dir-merge", RULE_KIND_DIR_MERGE},
{"clear", RULE_KIND_CLEAR},
};
for (size_t i = 0; i < sizeof(table) / sizeof(table[0]); i++) {
if (strlen(table[i].word) == len && strncmp(name, table[i].word, len) == 0) {
*kind = table[i].kind;
return true;
}
}
return false;
}
static bool is_modifier_char(char c) {
return c == 's' || c == 'r' || c == 'p' || c == 'x' || c == '/' || c == '!' || c == 'C';
}
/* Parse "RULE[,MODIFIERS] [PATTERN]". On success `kind`, `sides`,
* `sides_explicit`, `negate`, `anchored_mod`, `perishable`, `xattr`,
* `cvs_inject` and the pattern span (`pat_start`/`pat_len`, possibly 0 for
* merge/clear) are filled. Returns true on success. */
static bool parse_rule_syntax(const char* text, RuleKind* kind, unsigned* sides,
bool* sides_explicit, bool* negate, bool* anchored_mod,
bool* perishable, bool* xattr, bool* cvs_inject,
const char** pat_start, size_t* pat_len) {
const char* p = text;
*sides = FILTER_SIDE_SENDER | FILTER_SIDE_RECEIVER;
*sides_explicit = false;
*negate = false;
*anchored_mod = false;
*perishable = false;
*xattr = false;
*cvs_inject = false;
*pat_start = NULL;
*pat_len = 0;
bool is_short = false;
if (short_rule_char(*p, kind)) {
is_short = true;
p++;
} else {
const char* name_start = p;
while (isalpha((unsigned char)*p) || *p == '-')
p++;
size_t name_len = (size_t)(p - name_start);
if (name_len == 0 || !long_rule_name(name_start, name_len, kind))
return false;
/* A long name must be followed by a separator, a comma or the end. */
if (*p != '\0' && *p != ',' && *p != ' ' && *p != '_')
return false;
}
/* Modifiers: long names require a comma; short names may attach directly.
Only commit a modifier run that terminates at a separator or the end, so a
pattern such as "*.tmp" written as "-*.tmp" is not mistaken for modifiers. */
const char* mod_start = p;
const char* mod_end = p;
if (*p == ',') {
p++;
mod_start = p;
while (is_modifier_char(*p))
p++;
mod_end = p;
} else if (is_short) {
const char* scan = p;
while (is_modifier_char(*scan))
scan++;
if (*scan == '\0' || *scan == ' ' || *scan == '_') {
mod_start = p;
mod_end = scan;
p = scan;
}
}
for (const char* m = mod_start; m < mod_end; m++) {
switch (*m) {
case 's':
*sides = FILTER_SIDE_SENDER;
*sides_explicit = true;
break;
case 'r':
*sides = FILTER_SIDE_RECEIVER;
*sides_explicit = true;
break;
case '!':
*negate = true;
break;
case '/':
*anchored_mod = true;
break;
case 'p':
*perishable = true;
break;
case 'x':
*xattr = true;
break;
case 'C':
*cvs_inject = true;
break;
default:
break;
}
}
/* A single space or underscore separates the rule/modifiers from the
pattern; further spaces/underscores belong to the pattern. */
const char* pat = p;
if (*pat == ' ' || *pat == '_')
pat++;
/* Trim a trailing newline/CR (the caller may pass a raw file line). */
*pat_start = pat;
*pat_len = strlen(pat);
while (*pat_len > 0 && (pat[*pat_len - 1] == '\n' || pat[*pat_len - 1] == '\r'))
(*pat_len)--;
return true;
}
FilterRule* filter_rule_parse(const char* line, const FilterParseOptions* opts, char* err,
size_t err_size) {
if (err && err_size > 0)
err[0] = '\0';
if (!line)
return NULL;
const char* p = line;
while (*p == ' ' || *p == '\t')
p++;
if (*p == '\0' || *p == '\n' || *p == '\r') {
filter_set_error(err, err_size, "empty filter rule");
return NULL;
}
RuleKind kind = RULE_KIND_UNKNOWN;
unsigned sides;
bool sides_explicit, negate, anchored_mod, perishable, xattr, cvs_inject;
const char* pat;
size_t pat_len;
if (!parse_rule_syntax(p, &kind, &sides, &sides_explicit, &negate, &anchored_mod, &perishable,
&xattr, &cvs_inject, &pat, &pat_len)) {
filter_set_error(err, err_size, "unrecognized filter rule syntax");
return NULL;
}
if (cvs_inject) {
/* The C modifier expands to the CVS defaults in place; the rule itself
carries no pattern and is handled by the caller. */
filter_set_error(err, err_size, "the C modifier is handled by the rule-list parser");
return NULL;
}
if (xattr) {
filter_set_error(err, err_size, "xattr-name filter rules (the x modifier) are not supported");
return NULL;
}
if (kind == RULE_KIND_MERGE || kind == RULE_KIND_DIR_MERGE) {
filter_set_error(err, err_size, "merge/dir-merge rules are handled by the rule-list parser");
return NULL;
}
if (kind == RULE_KIND_CLEAR) {
if (pat_len != 0) {
filter_set_error(err, err_size, "clear takes no pattern");
return NULL;
}
FilterRule* rule = calloc(1, sizeof(FilterRule));
if (!rule) {
filter_set_error(err, err_size, "memory allocation failed");
return NULL;
}
rule->action = FILTER_ACTION_NONE; /* clear marker: no pattern */
rule->sides = 0;
return rule;
}
FilterAction action;
switch (kind) {
case RULE_KIND_INCLUDE:
case RULE_KIND_SHOW:
case RULE_KIND_RISK:
action = FILTER_ACTION_INCLUDE;
break;
default:
action = FILTER_ACTION_EXCLUDE;
break;
}
if (kind == RULE_KIND_HIDE)
sides = FILTER_SIDE_SENDER;
else if (kind == RULE_KIND_SHOW)
sides = FILTER_SIDE_SENDER;
else if (kind == RULE_KIND_PROTECT)
sides = FILTER_SIDE_RECEIVER;
else if (kind == RULE_KIND_RISK)
sides = FILTER_SIDE_RECEIVER;
if (kind == RULE_KIND_HIDE || kind == RULE_KIND_SHOW || kind == RULE_KIND_PROTECT ||
kind == RULE_KIND_RISK)
sides_explicit = true;
/* --delete-excluded turns an unqualified (no explicit s/r) rule into a
sender-side-only rule, so it no longer protects the receiver. */
if (opts && opts->delete_excluded && !sides_explicit)
sides = FILTER_SIDE_SENDER;
if (pat_len == 0) {
filter_set_error(err, err_size, "filter rule has no pattern");
return NULL;
}
bool anchored = anchored_mod;
const char* pat_begin = pat;
if (*pat_begin == '/') {
anchored = true;
pat_begin++;
/* Drop the spaces that could follow the anchor in the "-/ foo" form. */
while (*pat_begin == ' ' || *pat_begin == '\t')
pat_begin++;
pat_len = strlen(pat_begin);
while (pat_len > 0 && (pat_begin[pat_len - 1] == '\n' || pat_begin[pat_len - 1] == '\r'))
pat_len--;
}
if (pat_len == 0) {
filter_set_error(err, err_size, "filter rule has no pattern after '/' anchor");
return NULL;
}
bool dir_only = false;
if (pat_len > 1 && pat_begin[pat_len - 1] == '/') {
dir_only = true;
pat_len--;
}
if (pat_len == 0) {
filter_set_error(err, err_size, "filter rule has no pattern");
return NULL;
}
FilterRule* rule = calloc(1, sizeof(FilterRule));
if (!rule) {
filter_set_error(err, err_size, "memory allocation failed");
return NULL;
}
rule->pattern = malloc(pat_len + 1);
if (!rule->pattern) {
free(rule);
filter_set_error(err, err_size, "memory allocation failed");
return NULL;
}
memcpy(rule->pattern, pat_begin, pat_len);
rule->pattern[pat_len] = '\0';
rule->action = action;
rule->sides = sides;
rule->anchored = anchored;
rule->dir_only = dir_only;
rule->negate = negate;
rule->perishable = perishable;
(void)xattr; /* xattr-name rules never match file/dir names; accepted/ignored */
return rule;
}
/* ---- CVS default excludes (-C and the C modifier) ---- */
typedef struct {
const char* pattern;
@@ -237,12 +424,13 @@ static const CvsDefaultRule CVS_DEFAULTS[] = {
{".svn/", true}, {".git/", true}, {".hg/", true}, {".bzr/", true},
};
static bool cvs_rule_list_append(FilterRuleList* list) {
static bool filter_list_append_cvs(FilterRuleList* list, unsigned sides) {
for (size_t i = 0; i < sizeof(CVS_DEFAULTS) / sizeof(CVS_DEFAULTS[0]); i++) {
FilterRule* rule = calloc(1, sizeof(FilterRule));
if (!rule)
return false;
rule->action = FILTER_ACTION_EXCLUDE;
rule->sides = sides;
rule->dir_only = CVS_DEFAULTS[i].dir_only;
size_t plen = strlen(CVS_DEFAULTS[i].pattern);
if (rule->dir_only && plen > 0 && CVS_DEFAULTS[i].pattern[plen - 1] == '/')
@@ -266,98 +454,261 @@ static bool cvs_rule_list_append(FilterRuleList* list) {
return true;
}
FilterRuleList* filter_base_build(const char* const* rule_texts, int rule_count, bool cvs_exclude,
#define FILTER_MAX_MERGE_DEPTH 16
static bool filter_list_parse_append_depth(FilterRuleList* list, const char* line,
const FilterParseOptions* opts, const char* base_dir,
int depth, char* err, size_t err_size);
/* Read a merge file and splice its rules into `list`. A relative path is
* resolved below `base_dir` when given, else used as-is (rsync resolves a
* command-line merge file relative to the current directory). */
static bool filter_list_merge_file(FilterRuleList* list, const char* name,
const FilterParseOptions* opts, const char* base_dir, int depth,
char* err, size_t err_size) {
if (name[0] == '\0') {
filter_set_error(err, err_size, "merge requires a filename");
return false;
}
char* path =
(base_dir && base_dir[0] && name[0] != '/') ? path_cat(base_dir, name) : str_dup(name);
if (!path) {
filter_set_error(err, err_size, "memory allocation failed");
return false;
}
FILE* fp = fopen(path, "r");
if (!fp) {
filter_set_error(err, err_size, "could not read merge file '%s': %s", path, strerror(errno));
free(path);
return false;
}
char* line = NULL;
size_t cap = 0;
bool ok = true;
while (true) {
ssize_t n = utils_getdelim_bounded(fp, &line, &cap, '\n', UTILS_MAX_LINE_LEN);
if (n < 0) {
filter_set_error(err, err_size, "error reading merge file '%s'", path);
ok = false;
break;
}
if (n == 0)
break;
const char* lp = line;
while (*lp == ' ' || *lp == '\t')
lp++;
if (*lp == '\0' || *lp == '\n' || *lp == '\r' || *lp == '#')
continue;
if (!filter_list_parse_append_depth(list, lp, opts, base_dir, depth + 1, err, err_size)) {
ok = false;
break;
}
}
free(line);
fclose(fp);
free(path);
return ok;
}
/* Parse one line and append/merge it into `list`. Handles clear, merge and
* dir-merge at the list level. */
static bool filter_list_parse_append_depth(FilterRuleList* list, const char* line,
const FilterParseOptions* opts, const char* base_dir,
int depth, char* err, size_t err_size) {
if (depth > FILTER_MAX_MERGE_DEPTH) {
filter_set_error(err, err_size, "merge files nested too deeply");
return false;
}
const char* p = line;
while (*p == ' ' || *p == '\t')
p++;
if (*p == '\0' || *p == '\n' || *p == '\r')
return true;
RuleKind kind = RULE_KIND_UNKNOWN;
unsigned sides;
bool sides_explicit, negate, anchored_mod, perishable, xattr, cvs_inject;
const char* pat;
size_t pat_len;
if (!parse_rule_syntax(p, &kind, &sides, &sides_explicit, &negate, &anchored_mod, &perishable,
&xattr, &cvs_inject, &pat, &pat_len)) {
filter_set_error(err, err_size, "unrecognized filter rule syntax: %s", p);
return false;
}
(void)sides_explicit;
(void)negate;
(void)anchored_mod;
(void)perishable;
(void)xattr;
if (cvs_inject) {
/* "C" injects the CVS defaults in place; no pattern is expected. */
return filter_list_append_cvs(list, sides);
}
if (kind == RULE_KIND_CLEAR) {
if (pat_len != 0) {
filter_set_error(err, err_size, "clear takes no pattern");
return false;
}
for (int i = 0; i < list->count; i++)
filter_rule_free(list->items[i]);
list->count = 0;
return true;
}
if (kind == RULE_KIND_MERGE) {
if (pat_len == 0) {
filter_set_error(err, err_size, "merge requires a filename");
return false;
}
char* name = malloc(pat_len + 1);
if (!name) {
filter_set_error(err, err_size, "memory allocation failed");
return false;
}
memcpy(name, pat, pat_len);
name[pat_len] = '\0';
bool ok = filter_list_merge_file(list, name, opts, base_dir, depth, err, err_size);
free(name);
return ok;
}
if (kind == RULE_KIND_DIR_MERGE) {
if (pat_len == 0) {
filter_set_error(err, err_size, "dir-merge requires a filename");
return false;
}
char* name = malloc(pat_len + 1);
if (!name) {
filter_set_error(err, err_size, "memory allocation failed");
return false;
}
memcpy(name, pat, pat_len);
name[pat_len] = '\0';
bool ok = filter_rule_list_add_dir_merge(list, name);
free(name);
if (!ok) {
filter_set_error(err, err_size, "memory allocation failed");
return false;
}
return true;
}
FilterRule* rule = filter_rule_parse(p, opts, err, err_size);
if (!rule)
return false;
if (!filter_rule_list_add(list, rule)) {
filter_rule_free(rule);
filter_set_error(err, err_size, "memory allocation failed");
return false;
}
return true;
}
bool filter_rule_list_parse_append(FilterRuleList* list, const char* line,
const FilterParseOptions* opts, const char* merge_base_dir,
char* err, size_t err_size) {
if (err && err_size > 0)
err[0] = '\0';
if (!list)
return false;
return filter_list_parse_append_depth(list, line, opts, merge_base_dir, 0, err, err_size);
}
FilterRuleList* filter_base_build(const char* const* rule_texts, int rule_count, bool cvs_exclude,
bool delete_excluded, char* err, size_t err_size) {
if (err && err_size > 0)
err[0] = '\0';
FilterRuleList* list = filter_rule_list_create();
if (!list) {
snprintf(err, err_size, "memory allocation failed");
filter_set_error(err, err_size, "memory allocation failed");
return NULL;
}
FilterParseOptions opts = {.delete_excluded = delete_excluded, .cvs_exclude = cvs_exclude};
for (int i = 0; i < rule_count; i++) {
if (!rule_texts || !rule_texts[i])
continue;
FilterRule* rule = filter_rule_parse(rule_texts[i], err, err_size);
if (!rule) {
if (!filter_rule_list_parse_append(list, rule_texts[i], &opts, NULL, err, err_size)) {
filter_rule_list_free(list);
return NULL;
}
if (!set_rule_owner(rule, "")) {
filter_rule_free(rule);
filter_rule_list_free(list);
snprintf(err, err_size, "memory allocation failed");
return NULL;
}
if (!filter_rule_list_add(list, rule)) {
filter_rule_free(rule);
if (cvs_exclude && !filter_list_append_cvs(list, FILTER_SIDE_SENDER | FILTER_SIDE_RECEIVER)) {
filter_rule_list_free(list);
snprintf(err, err_size, "memory allocation failed");
return NULL;
}
}
if (cvs_exclude && !cvs_rule_list_append(list)) {
filter_rule_list_free(list);
snprintf(err, err_size, "memory allocation failed");
filter_set_error(err, err_size, "memory allocation failed");
return NULL;
}
return list;
}
/* ---- Per-directory .rsync-filter files ---- */
/* ---- Per-directory merge files ---- */
FilterRuleList* filter_file_read(const char* dir_path, const char* owner_rel, bool* exists,
/* Undo the rules and dir-merge registrations that one merge file appended,
* leaving the caller's earlier content intact. A "clear" rule inside the file
* frees every rule, including the caller's; clamp to the surviving count so
* those already-freed rules are never resurrected and freed a second time. */
static void filter_file_rollback(FilterRuleList* list, int rules_before, int dir_merges_before) {
int first = rules_before < list->count ? rules_before : list->count;
for (int i = first; i < list->count; i++)
filter_rule_free(list->items[i]);
list->count = first;
for (int i = dir_merges_before; i < list->dir_merge_count; i++)
free(list->dir_merge_names[i]);
list->dir_merge_count = dir_merges_before;
}
bool filter_file_append(FilterRuleList* list, const char* dir_path, const char* name,
const char* owner_rel, const FilterParseOptions* opts, bool* exists,
char* err, size_t err_size) {
if (err && err_size > 0)
err[0] = '\0';
if (exists)
*exists = false;
char* filter_path = path_cat(dir_path, ".rsync-filter");
if (!list)
return false;
char* filter_path = path_cat(dir_path, name);
if (!filter_path) {
snprintf(err, err_size, "memory allocation failed");
return NULL;
filter_set_error(err, err_size, "memory allocation failed");
return false;
}
FILE* fp = fopen(filter_path, "r");
free(filter_path);
if (!fp) {
if (errno == ENOENT || errno == ENOTDIR)
return filter_rule_list_create();
log_message(LOG_LEVEL_WARNING, "Could not read .rsync-filter in %s: %s", dir_path,
strerror(errno));
return filter_rule_list_create();
return true;
char* escaped_dir = output_escape(dir_path, log_get_8_bit_output());
log_message(LOG_LEVEL_WARNING, "Could not read %s in %s: %s", name,
escaped_dir ? escaped_dir : "<allocation failed>", strerror(errno));
free(escaped_dir);
return true;
}
if (exists)
*exists = true;
FilterRuleList* list = filter_rule_list_create();
if (!list) {
fclose(fp);
snprintf(err, err_size, "memory allocation failed");
return NULL;
}
int rules_before = list->count;
int dir_merges_before = list->dir_merge_count;
char* line = NULL;
size_t line_cap = 0;
ssize_t n;
bool ok = true;
while ((n = getline(&line, &line_cap, fp)) != -1) {
while (true) {
ssize_t n = utils_getdelim_bounded(fp, &line, &line_cap, '\n', UTILS_MAX_LINE_LEN);
if (n < 0) {
if (errno == EFBIG) {
filter_set_error(err, err_size, "line in %s exceeds %d bytes", name,
(int)UTILS_MAX_LINE_LEN);
} else {
filter_set_error(err, err_size, "error reading %s: %s", name, strerror(errno));
}
ok = false;
break;
}
if (n == 0)
break;
const char* p = line;
while (*p == ' ' || *p == '\t')
p++;
if (*p == '\0' || *p == '\n' || *p == '\r' || *p == '#')
continue;
FilterRule* rule = filter_rule_parse(p, err, err_size);
if (!rule) {
ok = false;
break;
}
if (!set_rule_owner(rule, owner_rel)) {
filter_rule_free(rule);
snprintf(err, err_size, "memory allocation failed");
ok = false;
break;
}
if (!filter_rule_list_add(list, rule)) {
filter_rule_free(rule);
snprintf(err, err_size, "memory allocation failed");
/* Merge files inside a per-directory file resolve relative to that
directory. */
if (!filter_list_parse_append_depth(list, p, opts, dir_path, 0, err, err_size)) {
ok = false;
break;
}
@@ -365,12 +716,40 @@ FilterRuleList* filter_file_read(const char* dir_path, const char* owner_rel, bo
free(line);
fclose(fp);
if (!ok) {
filter_file_rollback(list, rules_before, dir_merges_before);
return false;
}
for (int i = rules_before; i < list->count; i++) {
if (!set_rule_owner(list->items[i], owner_rel)) {
filter_set_error(err, err_size, "memory allocation failed");
filter_file_rollback(list, rules_before, dir_merges_before);
return false;
}
}
return true;
}
FilterRuleList* filter_file_read_named(const char* dir_path, const char* name,
const char* owner_rel, const FilterParseOptions* opts,
bool* exists, char* err, size_t err_size) {
FilterRuleList* list = filter_rule_list_create();
if (!list) {
if (err && err_size > 0)
filter_set_error(err, err_size, "memory allocation failed");
return NULL;
}
if (!filter_file_append(list, dir_path, name, owner_rel, opts, exists, err, err_size)) {
filter_rule_list_free(list);
return NULL;
}
return list;
}
FilterRuleList* filter_file_read(const char* dir_path, const char* owner_rel, bool* exists,
char* err, size_t err_size) {
return filter_file_read_named(dir_path, ".rsync-filter", owner_rel, NULL, exists, err, err_size);
}
/* ---- Rule matching ---- */
/* Match a pattern that contains '/' (non-anchored) against the end of the
@@ -386,10 +765,10 @@ static bool glob_suffix_match(const char* pattern, const char* str) {
}
static FilterAction rule_matches(const FilterRule* rule, const char* rel_path, const char* leaf,
bool is_dir) {
bool is_dir, unsigned side) {
if (!rule || !rule->pattern)
return FILTER_ACTION_NONE;
if (rule->dir_only && !is_dir)
if (!(rule->sides & side))
return FILTER_ACTION_NONE;
/* A rule applies only to entries below its owner directory. */
const char* rel2 = rel_path;
@@ -404,24 +783,36 @@ static FilterAction rule_matches(const FilterRule* rule, const char* rel_path, c
if (rel2[0] == '\0')
return FILTER_ACTION_NONE;
bool matched;
if (rule->anchored) {
if (rule->dir_only && !is_dir)
matched = false;
else if (rule->anchored)
matched = glob_match(rule->pattern, rel2);
} else if (strchr(rule->pattern, '/') != NULL) {
else if (strchr(rule->pattern, '/') != NULL)
matched = glob_suffix_match(rule->pattern, rel2);
} else {
else
matched = glob_match(rule->pattern, leaf);
}
return matched ? rule->action : FILTER_ACTION_NONE;
if (rule->negate)
matched = !matched;
if (!matched)
return FILTER_ACTION_NONE;
if (side == FILTER_SIDE_RECEIVER)
return rule->action == FILTER_ACTION_EXCLUDE ? FILTER_ACTION_PROTECT : FILTER_ACTION_RISK;
return rule->action;
}
FilterAction filter_rules_apply(const FilterRuleList* list, const char* rel_path, const char* leaf,
bool is_dir) {
FilterAction filter_rules_apply_side(const FilterRuleList* list, const char* rel_path,
const char* leaf, bool is_dir, unsigned side) {
if (!list)
return FILTER_ACTION_NONE;
for (int i = 0; i < list->count; i++) {
FilterAction action = rule_matches(list->items[i], rel_path, leaf, is_dir);
FilterAction action = rule_matches(list->items[i], rel_path, leaf, is_dir, side);
if (action != FILTER_ACTION_NONE)
return action;
}
return FILTER_ACTION_NONE;
}
FilterAction filter_rules_apply(const FilterRuleList* list, const char* rel_path, const char* leaf,
bool is_dir) {
return filter_rules_apply_side(list, rel_path, leaf, is_dir, FILTER_SIDE_SENDER);
}
+88 -34
View File
@@ -4,38 +4,51 @@
#include <stdbool.h>
#include <stddef.h>
/* rsync-style filter rule engine (client-side file selection).
/* rsync-style filter rule engine (client-side file selection and the
* receiver-side protection set it feeds).
*
* Supported rule syntax (documented subset):
* [+|-] [anchored '/' prefix] pattern [trailing '/' for dir-only]
*
* "+ PATTERN" include rule (first match wins)
* "- PATTERN" exclude rule
* "PATTERN" implicit exclude rule (rsync default)
* "include PATTERN" / "exclude PATTERN" word forms
* leading '/' after the +/- anchors the pattern to its owner directory
* (the transfer root for command-line/-C rules, the directory that
* contains a .rsync-filter file for per-directory rules)
* a trailing '/' makes the rule match directories only
*
* Rejected explicitly (no silent no-ops): the rsync merge/dir-merge/list-clear
* shorthands written as a rule that starts with ':' or '.' or '!', the
* merge/dir-merge/hide/show/protect/risk/clear words, and every include/exclude
* rule modifier other than '/' (! C s r p x). The pattern must be separated
* from +/- by a space (or a single '/' anchor), exactly like rsync's
* "-s foo"/"-p ..." modifier syntax is refused.
* Rule syntax (see the rsync man page FILTER RULES section):
* RULE [PATTERN_OR_FILENAME]
* RULE,MODIFIERS [PATTERN_OR_FILENAME]
* Short RULE names may attach MODIFIERS directly ("-sr foo"); the long name
* form requires the comma. The pattern/filename is separated from the rule by
* one space or underscore. Rule names:
* exclude/- exclude (by default both sender-hide and receiver-protect)
* include/+ include (by default both sender-show and receiver-risk)
* hide/H sender-only exclude
* show/S sender-only include
* protect/P receiver-only exclude (protect from deletion)
* risk/R receiver-only include (allow deletion)
* merge/. read a client-side merge file for more rules
* dir-merge/: per-directory merge file (registered for the scanner)
* clear/! clear the current rule list (takes no argument)
* Modifiers: '/' absolute anchor, '!' negate match, 'C' inject CVS defaults,
* 's' sender side, 'r' receiver side, 'p' perishable, 'x' xattr name rule.
* A trailing '/' makes a pattern match directories only. A leading '/' anchors
* the pattern to its owner directory.
*/
typedef enum {
FILTER_ACTION_NONE = 0, /* no rule matched */
FILTER_ACTION_EXCLUDE = -1,
FILTER_ACTION_INCLUDE = 1
FILTER_ACTION_INCLUDE = 1,
/* Receiver-side-only verdicts: the entry is transferred but its destination
* mirror is protected from --delete (PROTECT) or explicitly left at risk
* (RISK). */
FILTER_ACTION_PROTECT = 2,
FILTER_ACTION_RISK = 3,
} FilterAction;
#define FILTER_SIDE_SENDER 1u
#define FILTER_SIDE_RECEIVER 2u
typedef struct {
FilterAction action;
FilterAction action; /* EXCLUDE or INCLUDE (the base pattern action) */
unsigned sides; /* FILTER_SIDE_SENDER | FILTER_SIDE_RECEIVER */
bool anchored; /* pattern anchored to the rule's owner directory */
bool dir_only; /* pattern had a trailing '/': matches directories only */
bool negate; /* '!' modifier: match succeeds when the pattern does not */
bool perishable; /* 'p' modifier (ignored in deleted directories) */
char* owner; /* owning directory rel path ("" == transfer root) */
char* pattern; /* cleaned glob pattern (no leading '/', no trailing '/') */
} FilterRule;
@@ -44,19 +57,41 @@ typedef struct {
FilterRule** items; /* owned array of rule pointers */
int count;
int capacity;
/* Per-directory merge-file basenames registered by "dir-merge NAME"/": NAME"
* or by -F (.rsync-filter). Owned strings; the scanner reads each name in
* every directory it traverses. */
char** dir_merge_names;
int dir_merge_count;
int dir_merge_capacity;
} FilterRuleList;
/* Context needed while parsing a rule list (merge files, --delete-excluded). */
typedef struct {
bool delete_excluded; /* --delete-excluded: default sides become sender-only */
bool cvs_exclude; /* -C: expand the CVS default excludes */
} FilterParseOptions;
/* Parse a single filter-rule line (no trailing newline required). Returns an
* owned rule, or NULL on unsupported/invalid syntax with a message in `err`. */
FilterRule* filter_rule_parse(const char* line, char* err, size_t err_size);
* owned rule, or NULL on unsupported/invalid syntax with a message in `err`.
* `opts` may be NULL (no merge expansion / no delete-excluded). */
FilterRule* filter_rule_parse(const char* line, const FilterParseOptions* opts, char* err,
size_t err_size);
void filter_rule_free(FilterRule* rule);
FilterRuleList* filter_rule_list_create(void);
/* Append a fully-parsed rule (takes ownership). Returns false on OOM. */
bool filter_rule_list_add(FilterRuleList* list, FilterRule* rule);
/* Parse `line` and append it. Returns false and fills `err` on bad syntax. */
bool filter_rule_list_parse_append(FilterRuleList* list, const char* line, char* err,
size_t err_size);
/* Register a per-directory merge-file basename (idempotent). Returns false on
* OOM. Used by the scanner to read custom "dir-merge" files. */
bool filter_rule_list_add_dir_merge(FilterRuleList* list, const char* name);
/* Parse `line` and append it. Handles "clear"/"!" (resets the list), "merge
* FILE"/". FILE" (splices the file's rules) and "dir-merge NAME"/": NAME"
* (registers a per-directory filename). Returns false and fills `err` on bad
* syntax or an unreadable merge file. `merge_base_dir` resolves a relative
* merge-file path (NULL means the process working directory). */
bool filter_rule_list_parse_append(FilterRuleList* list, const char* line,
const FilterParseOptions* opts, const char* merge_base_dir,
char* err, size_t err_size);
void filter_rule_list_free(FilterRuleList* list);
/* Build the command-line filter set: `rule_texts` (--filter=RULE in the order
@@ -64,19 +99,38 @@ void filter_rule_list_free(FilterRuleList* list);
* cvs_exclude is true. All rules are owned by "" (the transfer root).
* Returns NULL on unsupported rule text (message in `err`). */
FilterRuleList* filter_base_build(const char* const* rule_texts, int rule_count, bool cvs_exclude,
bool delete_excluded, char* err, size_t err_size);
/* Read "<dir_path>/<name>" and return its rules, each owned by `owner_rel`. A
* missing file yields an empty list with *exists=false; an unreadable file is
* treated as missing. Returns NULL only on parse or allocation failure
* (message in `err`). `opts` may be NULL. */
FilterRuleList* filter_file_read_named(const char* dir_path, const char* name,
const char* owner_rel, const FilterParseOptions* opts,
bool* exists, char* err, size_t err_size);
/* Append the rules of "<dir_path>/<name>" into an existing list (each owned by
* `owner_rel`). A missing file yields *exists=false and no error. Returns
* false only on parse/allocation failure (message in `err`). */
bool filter_file_append(FilterRuleList* list, const char* dir_path, const char* name,
const char* owner_rel, const FilterParseOptions* opts, bool* exists,
char* err, size_t err_size);
/* Read "<dir_path>/.rsync-filter" and return its rules, each owned by
* `owner_rel`. A missing file yields an empty list with *exists=false; an
* unreadable file is treated as missing. Returns NULL only on parse or
* allocation failure (message in `err`). */
/* filter_file_read_named with the default ".rsync-filter" name. */
FilterRuleList* filter_file_read(const char* dir_path, const char* owner_rel, bool* exists,
char* err, size_t err_size);
/* Evaluate an entry against one ordered rule list. Returns FILTER_ACTION_NONE
* when no rule matched, otherwise the first matching rule's action.
* `rel_path` is the entry's path relative to the transfer root ("" == root),
* `leaf` its final name, `is_dir` whether it is a directory. */
/* Evaluate an entry against one ordered rule list for one side. Returns
* FILTER_ACTION_NONE when no rule matched, otherwise the first matching rule's
* action (for the receiver side an EXCLUDE is reported as
* FILTER_ACTION_PROTECT and an INCLUDE as FILTER_ACTION_RISK). `rel_path` is
* the entry's path relative to the transfer root ("" == root), `leaf` its final
* name, `is_dir` whether it is a directory. */
FilterAction filter_rules_apply_side(const FilterRuleList* list, const char* rel_path,
const char* leaf, bool is_dir, unsigned side);
/* Sender-side convenience wrapper (kept for callers/tests that only need the
* transfer decision). */
FilterAction filter_rules_apply(const FilterRuleList* list, const char* rel_path, const char* leaf,
bool is_dir);
+129
View File
@@ -0,0 +1,129 @@
#include "format.h"
#include "protocol.h"
#include <stdio.h>
#include <string.h>
bool format_human_size_decimal(unsigned long long bytes, char* buffer, size_t buffer_size) {
if (!buffer || buffer_size == 0)
return false;
if (bytes < 1000ULL) {
int written = snprintf(buffer, buffer_size, "%llu", bytes);
return written >= 0 && (size_t)written < buffer_size;
}
static const char units[] = "KMGTPE";
double value = (double)bytes;
size_t divisions = 0;
while (value >= 1000.0 && divisions < sizeof(units) - 1) {
value /= 1000.0;
divisions++;
}
int written = snprintf(buffer, buffer_size, "%.2f%c", value, units[divisions - 1]);
return written >= 0 && (size_t)written < buffer_size;
}
bool format_big_num(unsigned long long value, bool human_readable, char* buffer,
size_t buffer_size) {
if (human_readable)
return format_human_size_decimal(value, buffer, buffer_size);
char digits[32];
int written = snprintf(digits, sizeof(digits), "%llu", value);
if (written < 0 || (size_t)written >= sizeof(digits))
return false;
size_t len = (size_t)written;
size_t separators = len > 1 ? (len - 1) / 3 : 0;
size_t total = len + separators;
if (total + 1 > buffer_size)
return false;
size_t out = total;
buffer[out] = '\0';
size_t digits_since_sep = 0;
for (size_t i = len; i > 0; i--) {
buffer[--out] = digits[i - 1];
digits_since_sep++;
if (digits_since_sep == 3 && i > 1) {
buffer[--out] = ',';
digits_since_sep = 0;
}
}
return true;
}
bool format_rsync_datetime(time_t when, bool dash, char* buffer, size_t buffer_size) {
if (!buffer || buffer_size == 0)
return false;
struct tm broken_down;
if (localtime_r(&when, &broken_down) == NULL)
return false;
const char* format = dash ? "%Y/%m/%d-%H:%M:%S" : "%Y/%m/%d %H:%M:%S";
return strftime(buffer, buffer_size, format, &broken_down) != 0;
}
bool format_dest_state_send(int fd, const OutputDestState* state) {
if (!state)
return false;
int32_t has_old = state->existed ? 1 : 0;
uint64_t size = (uint64_t)state->size;
int64_t mtime = (int64_t)state->mtime_sec;
int64_t mtime_nsec = state->mtime_nsec;
uint32_t mode = state->mode;
int32_t uid = state->uid;
int32_t gid = state->gid;
return send_n_data(fd, &has_old, sizeof(has_old)) && send_n_data(fd, &size, sizeof(size)) &&
send_n_data(fd, &mtime, sizeof(mtime)) &&
send_n_data(fd, &mtime_nsec, sizeof(mtime_nsec)) && send_n_data(fd, &mode, sizeof(mode)) &&
send_n_data(fd, &uid, sizeof(uid)) && send_n_data(fd, &gid, sizeof(gid));
}
bool format_dest_state_receive(int fd, OutputDestState* state) {
if (!state)
return false;
int32_t has_old = 0;
uint64_t size = 0;
int64_t mtime = 0;
int64_t mtime_nsec = 0;
uint32_t mode = 0;
int32_t uid = 0;
int32_t gid = 0;
if (!receive_n_data(fd, &has_old, sizeof(has_old)) || !receive_n_data(fd, &size, sizeof(size)) ||
!receive_n_data(fd, &mtime, sizeof(mtime)) ||
!receive_n_data(fd, &mtime_nsec, sizeof(mtime_nsec)) ||
!receive_n_data(fd, &mode, sizeof(mode)) || !receive_n_data(fd, &uid, sizeof(uid)) ||
!receive_n_data(fd, &gid, sizeof(gid)))
return false;
memset(state, 0, sizeof(*state));
state->known = true;
state->existed = has_old != 0;
state->size = size;
state->mtime_sec = mtime;
state->mtime_nsec = mtime_nsec;
state->mode = mode;
state->uid = uid;
state->gid = gid;
return true;
}
bool format_stats_send(int fd, const ReceiverStats* stats) {
if (!stats)
return false;
unsigned long long matched = stats->matched_data;
unsigned long long deleted = stats->deleted_files;
unsigned long long would = stats->would_delete_count;
return send_n_data(fd, &matched, sizeof(matched)) && send_n_data(fd, &deleted, sizeof(deleted)) &&
send_n_data(fd, &would, sizeof(would));
}
bool format_stats_receive(int fd, ReceiverStats* stats) {
if (!stats)
return false;
unsigned long long matched = 0;
unsigned long long deleted = 0;
unsigned long long would = 0;
if (!receive_n_data(fd, &matched, sizeof(matched)) ||
!receive_n_data(fd, &deleted, sizeof(deleted)) || !receive_n_data(fd, &would, sizeof(would)))
return false;
memset(stats, 0, sizeof(*stats));
stats->matched_data = matched;
stats->deleted_files = deleted;
stats->would_delete_count = would;
return true;
}
+76
View File
@@ -0,0 +1,76 @@
#ifndef FORMAT_H
#define FORMAT_H
#include <stdbool.h>
#include <stddef.h>
#include <stdint.h>
#include <time.h>
/* Low-level output-formatting primitives shared by the change-event model
* (change_list.c) and the transfer driver (client_send.c).
*
* The functions here are pure/string-level except for the STATUS_DEST_INFO
* codec, which lets the receiver report the pre-transfer destination entry so
* the sender can render rsync-accurate --itemize-changes / --out-format
* columns (see protocol.h). */
/* Pre-transfer destination snapshot, reported by the receiver when the wire
* config carries report_dest_info. `known` distinguishes "no report was
* requested/received" from "the destination did not exist" (`existed == false`
* with `known == true`). */
typedef struct {
bool known;
bool existed;
unsigned long long size;
long long mtime_sec;
long long mtime_nsec;
uint32_t mode;
int32_t uid;
int32_t gid;
} OutputDestState;
/* rsync's -h/--human-readable size (decimal, base 1000): integers below 1000
* print verbatim; larger values use the largest unit that keeps the value
* below 1000 (K/M/G/T/P/E) with exactly two decimals, so 1500000 -> "1.50M"
* and 999999 -> "1000.00K" (matching rsync's human_num). Returns false when
* the buffer is too small (nothing is written). */
bool format_human_size_decimal(unsigned long long bytes, char* buffer, size_t buffer_size);
/* rsync's general number formatting (big_num). When `human_readable` is true
* this is format_human_size_decimal; otherwise the integer is rendered with a
* ',' thousands separator every three digits (rsync's separator in the C
* locale). Returns false on an undersized buffer. */
bool format_big_num(unsigned long long value, bool human_readable, char* buffer,
size_t buffer_size);
/* rsync's %M/%t timestamp. When `dash` is true the separator between the date
* and the time is '-' (the %M form: "YYYY/MM/DD-HH:MM:SS"); otherwise it is a
* space (the %t form: "YYYY/MM/DD HH:MM:SS"). Local time. Returns false on a
* bad time or an undersized buffer. */
bool format_rsync_datetime(time_t when, bool dash, char* buffer, size_t buffer_size);
/* Fixed-width STATUS_DEST_INFO record codec (int32 has_old, uint64 size,
* int64 mtime, int64 mtime_nsec, uint32 mode, int32 uid, int32 gid). The
* status frame itself is sent/received by the caller. Returns false on I/O
* failure. */
bool format_dest_state_send(int fd, const OutputDestState* state);
bool format_dest_state_receive(int fd, OutputDestState* state);
/* End-of-transfer receiver counters reported through STATUS_STATS (protocol
* 2.25.0) when the wire config carries report_stats. `would_delete_count` is
* the number of destination-relative paths the receiver would have deleted in a
* -n/--dry-run --delete run; that many wire strings immediately follow the
* fixed record (sent/read by the caller). */
typedef struct {
unsigned long long matched_data;
unsigned long long deleted_files;
unsigned long long would_delete_count;
} ReceiverStats;
/* Fixed-width STATUS_STATS counter record. The status frame and the optional
* would-delete path list are sent/received by the caller. Returns false on I/O
* failure. */
bool format_stats_send(int fd, const ReceiverStats* stats);
bool format_stats_receive(int fd, ReceiverStats* stats);
#endif
+127
View File
@@ -0,0 +1,127 @@
#include "hardlink.h"
#include <stdint.h>
#include <stdlib.h>
#include <string.h>
#include "log.h"
#include "utils.h"
/* ---- Sender-side detection table ---- */
HardLinkTable* hardlink_table_create(void) {
HardLinkTable* table = calloc(1, sizeof(HardLinkTable));
if (!table)
return NULL;
if (mtx_init(&table->mutex, mtx_plain) != thrd_success) {
free(table);
return NULL;
}
table->next_gid = 1;
return table;
}
static void hardlink_item_destroy(HardLinkItem* item) {
if (!item)
return;
free(item->first_path);
item->first_path = NULL;
}
void hardlink_table_destroy(HardLinkTable* table) {
if (!table)
return;
for (size_t i = 0; i < table->count; i++)
hardlink_item_destroy(&table->items[i]);
free(table->items);
table->items = NULL;
table->count = 0;
table->capacity = 0;
mtx_destroy(&table->mutex);
free(table);
}
static HardLinkItem* hardlink_table_find_locked(HardLinkTable* table, dev_t dev, ino_t ino) {
for (size_t i = 0; i < table->count; i++) {
if (table->items[i].dev == dev && table->items[i].ino == ino)
return &table->items[i];
}
return NULL;
}
static bool hardlink_table_add_locked(HardLinkTable* table, dev_t dev, ino_t ino, const char* path,
int gid, HardLinkItem** out) {
if (table->count == table->capacity) {
size_t new_capacity = table->capacity == 0 ? 8 : table->capacity * 2;
if (new_capacity < table->capacity)
return false;
HardLinkItem* grown = realloc(table->items, new_capacity * sizeof(HardLinkItem));
if (!grown)
return false;
table->items = grown;
table->capacity = new_capacity;
}
HardLinkItem* item = &table->items[table->count];
char* dup = str_dup(path);
if (!dup)
return false;
memset(item, 0, sizeof(*item));
item->dev = dev;
item->ino = ino;
item->gid = gid;
item->first_path = dup;
table->count++;
*out = item;
return true;
}
bool hardlink_table_assign(HardLinkTable* table, const char* wire_path, dev_t dev, ino_t ino,
int* gid, bool* is_first, char** first_path_out) {
if (!table || !wire_path || !gid || !is_first || !first_path_out)
return false;
if (mtx_lock(&table->mutex) != thrd_success)
return false;
bool ok = true;
const HardLinkItem* item = hardlink_table_find_locked(table, dev, ino);
int next_gid;
if (item) {
*is_first = false;
char* dup = str_dup(item->first_path);
if (!dup) {
ok = false;
} else {
*gid = item->gid;
*first_path_out = dup;
}
next_gid = -1;
} else {
if (table->next_gid <= 0) {
ok = false;
next_gid = -1;
} else {
next_gid = table->next_gid;
HardLinkItem* created = NULL;
if (!hardlink_table_add_locked(table, dev, ino, wire_path, next_gid, &created)) {
ok = false;
} else {
char* dup = str_dup(wire_path);
if (!dup) {
hardlink_item_destroy(created);
table->count--;
ok = false;
} else {
*is_first = true;
*gid = next_gid;
*first_path_out = dup;
}
}
}
}
if (ok && next_gid > 0)
table->next_gid++;
mtx_unlock(&table->mutex);
if (!ok) {
log_message(LOG_LEVEL_ERROR, "memory allocation failed while detecting hard links");
}
return ok;
}
+66
View File
@@ -0,0 +1,66 @@
#ifndef HARDLINK_H
#define HARDLINK_H
#include <stdbool.h>
#include <stddef.h>
#include <sys/types.h>
#include <threads.h>
/*
* --hard-links / -H support.
*
* Sender side: a HardLinkTable detects regular files on the source that share
* an (st_dev, st_ino) identity (a `cp -al`-style hard-linked tree) and assigns
* each distinct inode a stable, run-local link-group id. The first member
* encountered carries the file data; every later member is marked as a sibling
* (no data payload) that the receiver creates as a hard link to the first
* member's destination file. Grouping is scoped by st_dev so inode reuse
* across different filesystems is never conflated. The table is mutex-guarded
* so the parallel (multi-threaded) scanner COULD share one instance across its
* worker threads; the first-thread-to-call designates the data-carrying member,
* which is safe because a hard-link group's members are byte-identical. (In
* practice the sender forces the sequential scanner whenever -H is on; the
* mutex guards the shared table for any path that supplies one.)
*
* ORDERING (why there is no receiver-side handshake): the receiver stores every
* file - including a hard-link group's first member - through a SINGLE writer
* thread draining a single FIFO queue driven by a single receive thread, so
* wire order == write order and every sibling is processed AFTER its group's
* first member. The sender additionally forces the sequential scanner with -H
* so the first-member frame always precedes its siblings on the wire. Sibling
* install therefore needs no present/wait registry: it hard-links to the first
* member (or copies it) knowing that path is already installed - or that, if
* the first member was skipped (already up to date), its destination still
* exists. This guarantee is REQUIRED; do not introduce a concurrent
* multi-writer receiver for -H without re-adding an ordering mechanism.
*/
typedef struct HardLinkItem {
dev_t dev;
ino_t ino;
int gid;
char* first_path; /* wire path of the group's data-carrying first member */
} HardLinkItem;
typedef struct HardLinkTable {
mtx_t mutex;
HardLinkItem* items;
size_t count;
size_t capacity;
int next_gid;
} HardLinkTable;
HardLinkTable* hardlink_table_create(void);
void hardlink_table_destroy(HardLinkTable* table);
/* Assign a link-group id to the regular file at `wire_path` with (dev, ino).
* On the first encounter the file becomes the group's first (data-carrying)
* member (*is_first = true) and a fresh gid is allocated. On a later member
* *is_first = false and *first_path_out is set to a malloc'd copy of the first
* member's wire path (the caller stores it and owns it; on the first member
* path the returned *first_path_out is a malloc'd copy of its own wire path).
* Returns false on allocation failure (transfer should abort). */
bool hardlink_table_assign(HardLinkTable* table, const char* wire_path, dev_t dev, ino_t ino,
int* gid, bool* is_first, char** first_path_out);
#endif
File diff suppressed because it is too large. Load diff
+162
View File
@@ -0,0 +1,162 @@
#ifndef IDENTITY_H
#define IDENTITY_H
#include "config.h"
#include <stdbool.h>
#include <stdint.h>
#include <sys/types.h>
/*
* Identity mapping: --numeric-ids / --usermap / --groupmap / --chown / --copy-as.
*
* FastSync transmits uid/gid numerically (int32 on the wire) and, by design,
* NEVER applies client-supplied ownership unless a user explicitly opts in with
* an identity flag below. This module is the controlled, opt-in,
* privilege-gated path for applying ownership on the receiver: the wire config
* is snapshotted once per connection via identity_set_active() and applied
* through an fd-relative fchown() in the receiver's metadata-restore path.
*
* Because only numeric ids cross the wire, name-based values are resolved to
* numbers at CLI parse time using the CLIENT (sender) machine's databases. On
* a shared-account source/destination this reproduces rsync's semantics; a
* genuinely different destination database is a documented divergence (see
* RSYNC_COMPAT.md).
*/
/* Parse one --usermap= / --groupmap= value (comma-separated FROM:TO rules,
* first match wins) into config->usermap / config->groupmap. is_group selects
* the group tables and name databases. Returns 0 on success, -1 on a
* malformed spec or an unresolvable name (never a silent no-op). */
int identity_parse_map(Config* config, const char* value, bool is_group);
/* Parse --chown=USER:GROUP. Supports USER:GROUP, USER (owner only), :GROUP
* (group only), '*' (current/root as appropriate) and numeric ids. Returns 0
* on success, -1 on a malformed spec / unresolvable name. */
int identity_parse_chown(Config* config, const char* value);
/* Parse --copy-as=USER[:GROUP] (P7 Wave E). USER is resolved with the same
* user-database rules as --chown (a name, @N/bare N numeric id, or '*' meaning
* the client's current euid); when ':GROUP' is present the group is resolved
* with the group database ('*' meaning the client's egid). When the group is
* omitted, the user's primary gid is used (getpwuid(uid)->pw_gid); if the
* resolved user is a numeric id with no local passwd entry, gid falls back to
* uid. On success sets copy_as_set/copy_as_uid/copy_as_gid and forces
* metadata transmission (ownership application needs the metadata path).
* Returns 0 on success, -1 on a malformed / empty / unresolvable spec (never a
* silent no-op). */
int identity_parse_copy_as(Config* config, const char* value);
/* True when a --copy-as request is active but the receiver is not permitted to
* perform the privileged ownership application it needs. This is the up-front
* refusal predicate: the server rejects the whole transfer at the config
* handshake rather than silently ignoring the requested ownership. It is a
* pure function of the config mode and the current effective uid (it does NOT
* read the active snapshot, so it is valid at the pre-STATUS_OK gate, before
* identity_set_active() has run). `super_mode` is the EFFECTIVE mode after any
* server-side policy veto. */
bool identity_copy_as_refused(const Config* config);
/* True when the CURRENT per-connection snapshot has a --copy-as active (i.e.
* identity_set_active() has run against a config with copy_as_set). The
* --fake-super owner replay consults this so a copy-as run never lets the
* recorded source owner overwrite the forced target owner. Reads the active
* snapshot, so call identity_set_active() first (the receiver does, before any
* write). */
bool identity_copy_as_active(void);
/* Receiver-side snapshot of the negotiated identity config. The server calls
* identity_set_active() once per connection (before any file write) using the
* config received over the wire; the snapshot is a deep copy so the caller may
* free its Config immediately. identity_clear_active() releases it.
*
* Returns true on success. On an allocation failure while deep-copying a
* requested usermap/groupmap it logs a LOG_LEVEL_ERROR, leaves the snapshot
* cleared (never a partial/wrong policy) and returns false; the caller must
* refuse the connection. */
bool identity_set_active(const Config* config);
void identity_clear_active(void);
/* True when any ownership-affecting identity option is present in the active
* snapshot. Ownership stays OFF ("do not apply") for every transfer that
* requests none of them, preserving FastSync's existing behavior. --super /
* --no-super alone does NOT enable ownership; an explicit identity flag
* (--numeric-ids / --chown / --usermap / --groupmap / --copy-as) or a
* preserve-source -o/--owner / -g/--group request is required. */
bool identity_active_enabled(void);
/* Per-side predicates over the ACTIVE per-connection snapshot (call
* identity_set_active() first). They mirror the owner_requested /
* group_requested conditions inside identity_resolve_targets() exactly, so
* callers that must apply only one side (e.g. the --fake-super owner replay)
* can pass (uid_t)-1 / (gid_t)-1 for the side that was NOT requested and leave
* it untouched. The owner side is requested by --copy-as, --chown USER,
* --numeric-ids, -o/--owner, or a non-empty --usermap; the group side by
* --copy-as, --chown :GROUP, --numeric-ids, -g/--group, or a non-empty
* --groupmap. */
bool identity_owner_requested(void);
bool identity_group_requested(void);
/* Pure, config-only predicate: true when the client requested ANY client-chosen
* ownership or super-user activity (--numeric-ids, --chown, --usermap/--groupmap,
* --copy-as, --fake-super, an explicit --super, or a preserve-source -o/-g).
* General awareness only; the daemon module gate uses the narrower
* identity_explicit_ownership_requested() below. Never reads the snapshot. */
bool identity_ownership_requested(const Config* config);
/* Pure, config-only predicate for the narrow set that lets the CLIENT choose an
* arbitrary owner/group: --numeric-ids, --chown, --usermap/--groupmap,
* --copy-as, --fake-super, or an explicit --super. Deliberately EXCLUDES a
* plain -o/--owner / -g/--group (or -a) preserve-source request, which the
* daemon gate handles by forcing super-user ownership activity off rather than
* refusing the whole transfer. Never reads the snapshot. */
bool identity_explicit_ownership_requested(const Config* config);
/* Apply the negotiated ownership to an already-written file descriptor.
* source_uid/source_gid are the transmitted numeric ids. Resolution order:
* --copy-as (highest priority, forces both ids), then a matching
* usermap/groupmap rule, then --chown, then --numeric-ids (raw), then a
* best-effort name lookup on the receiver's own databases (skipped when the
* transmitted id has no name on this system). Only calls fchown() when the
* result differs from the current value.
*
* Returns false ONLY when an active --copy-as ownership application failed: its
* forced ownership is REQUIRED, so the caller must treat the entry as failed
* rather than reporting success with the wrong owner. For every other identity
* policy an fchown EPERM/EACCES is logged and ignored and true is returned
* (rsync parity: the transfer must not abort). A no-op when no identity policy
* is active returns true. */
bool identity_apply_ownership(int fd, int32_t source_uid, int32_t source_gid);
/* Resolve the ownership that --fake-super should RECORD in the reserved xattr
* (rather than chown for real). A requested side (--copy-as / usermap /
* --chown / -o / -g, with --numeric-ids as the raw-id modifier) yields the
* resolved target; a side that was not requested keeps the transmitted source
* id. Must be called after identity_set_active(). */
void identity_resolve_storage_ids(int32_t source_uid, int32_t source_gid, uint32_t* out_uid,
uint32_t* out_gid);
/* P7 Wave D: the no-follow (symlink) counterpart. Resolves the same
* usermap/groupmap/chown/numeric-ids/copy-as policy but applies it with
* fchownat(..., AT_SYMLINK_NOFOLLOW) so a symlink's own ownership is changed
* without ever dereferencing it. A no-op unless an identity flag is active.
* The return value follows identity_apply_ownership(): false only when an
* active --copy-as application failed. */
bool identity_apply_ownership_link(int parent_fd, const char* leaf, int32_t source_uid,
int32_t source_gid);
/* Receiver-side wire validation of the resolved identity fields. */
bool identity_wire_valid(const Config* config);
/* P7 Wave E receiver-side permission gate for super-user activities (ownership
* application and char/block device-node creation). `privilege_super_permitted`
* consults the per-connection snapshot (call identity_set_active() first);
* `privilege_super_mode_permitted` is the pure mode predicate and is what
* callers holding a Config use (the config-frame gate, file_receive). Both
* return false only for SUPER_MODE_OFF; SUPER_MODE_ON and SUPER_MODE_AUTO (the
* default) permit a confined attempt, matching FastSync's historical
* best-effort behavior where an unprivileged attempt is refused by the kernel
* and skipped. Neither EVER elevates privileges. */
bool privilege_super_permitted(void);
bool privilege_super_mode_permitted(SuperMode mode);
#endif
+76 -26
View File
@@ -3,7 +3,9 @@
#include <stdbool.h>
#include <stdarg.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <threads.h>
#include <time.h>
static const char* log_level_strings[] = {"DEBUG", "INFO", "WARN", "ERROR"};
@@ -15,6 +17,18 @@ static FILE* log_fp = NULL;
static _Thread_local bool eight_bit_output;
static LogStderrMode stderr_mode = LOG_STDERR_ERRORS;
/* Serializes access to log_fp and makes each emitted line atomic: the
* timestamp prefix, formatted body, and trailing newline are written as one
* critical section so concurrent threads cannot interleave partial lines.
* Initialized lazily (matching the protocol.c bw_mutex idiom) because logging
* can happen before main() installs any synchronization. */
static mtx_t log_mutex;
static once_flag log_mutex_once = ONCE_FLAG_INIT;
static void log_mutex_init(void) {
mtx_init(&log_mutex, mtx_plain);
}
void set_log_level(LogLevel level) {
current_log_level = level;
}
@@ -27,6 +41,10 @@ uint32_t get_log_debug_flags(void) {
return current_debug_flags;
}
bool log_debug_enabled(LogDebugFlag flag) {
return current_log_level <= LOG_LEVEL_DEBUG && (current_debug_flags & flag) != 0;
}
void set_log_info_flags(uint32_t flags) {
info_flags = flags;
info_flags_explicit = true;
@@ -37,7 +55,10 @@ uint32_t get_log_info_flags(void) {
}
void log_set_file(FILE* fp) {
call_once(&log_mutex_once, log_mutex_init);
mtx_lock(&log_mutex);
log_fp = fp;
mtx_unlock(&log_mutex);
}
void log_set_8_bit_output(bool enabled) {
@@ -56,13 +77,48 @@ LogStderrMode log_get_stderr_mode(void) {
return stderr_mode;
}
static inline void write_message(FILE* dest_io, LogLevel log_level, struct tm t, const char* format,
/* Format one complete log line (timestamp prefix + body + newline) into a
* freshly allocated buffer. This is pure CPU/malloc work and must happen
* OUTSIDE the log mutex: the mutex only guards the log_fp pointer, so a
* stalled stderr/stdout pipe cannot block every logging thread. Returns NULL
* on allocation/formatting failure. */
static char* format_log_line(LogLevel log_level, const struct tm* t, const char* format,
va_list args) {
fprintf(dest_io, "%04d-%02d-%02d %02d:%02d:%02d [%s]: ", t.tm_year + 1900, t.tm_mon + 1,
t.tm_mday, t.tm_hour, t.tm_min, t.tm_sec, log_level_strings[log_level]);
char prefix[64];
int prefix_len = snprintf(
prefix, sizeof(prefix), "%04d-%02d-%02d %02d:%02d:%02d [%s]: ", t->tm_year + 1900,
t->tm_mon + 1, t->tm_mday, t->tm_hour, t->tm_min, t->tm_sec, log_level_strings[log_level]);
if (prefix_len < 0 || prefix_len >= (int)sizeof(prefix))
return NULL;
va_list copy;
va_copy(copy, args);
int body_len = vsnprintf(NULL, 0, format, copy);
va_end(copy);
if (body_len < 0)
return NULL;
size_t total = (size_t)prefix_len + (size_t)body_len;
char* line = malloc(total + 2); /* body bytes + '\n' + NUL */
if (!line)
return NULL;
memcpy(line, prefix, (size_t)prefix_len);
vsnprintf(line + prefix_len, (size_t)body_len + 1, format, args);
line[total] = '\n';
line[total + 1] = '\0';
return line;
}
vfprintf(dest_io, format, args);
fprintf(dest_io, "\n");
/* Write an already-formatted line to the console and, if configured, the log
* file. Only the log_fp pointer is read under the mutex (so log_set_file /
* config_delete cannot free it while it is in use); the single console fputs
* runs unlocked but is internally atomic per stdio stream. */
static void emit_log_line(FILE* console, const char* line) {
fputs(line, console);
call_once(&log_mutex_once, log_mutex_init);
mtx_lock(&log_mutex);
FILE* file = log_fp;
if (file)
fputs(line, file);
mtx_unlock(&log_mutex);
}
void log_message(LogLevel log_level, const char* format, ...) {
@@ -82,14 +138,12 @@ void log_message(LogLevel log_level, const char* format, ...) {
va_list args;
va_start(args, format);
write_message(dest_io, log_level, t, format, args);
char* line = format_log_line(log_level, &t, format, args);
va_end(args);
if (log_fp) {
va_start(args, format);
write_message(log_fp, log_level, t, format, args);
va_end(args);
}
if (!line)
return;
emit_log_line(dest_io, line);
free(line);
}
void log_debug_message(LogDebugFlag flag, const char* format, ...) {
@@ -103,14 +157,12 @@ void log_debug_message(LogDebugFlag flag, const char* format, ...) {
va_list args;
va_start(args, format);
write_message(stdout, LOG_LEVEL_DEBUG, t, format, args);
char* line = format_log_line(LOG_LEVEL_DEBUG, &t, format, args);
va_end(args);
if (log_fp) {
va_start(args, format);
write_message(log_fp, LOG_LEVEL_DEBUG, t, format, args);
va_end(args);
}
if (!line)
return;
emit_log_line(stdout, line);
free(line);
}
void log_info_message(LogInfoFlag flag, const char* format, ...) {
@@ -125,14 +177,12 @@ void log_info_message(LogInfoFlag flag, const char* format, ...) {
va_list args;
va_start(args, format);
write_message(stdout, LOG_LEVEL_INFO, t, format, args);
char* line = format_log_line(LOG_LEVEL_INFO, &t, format, args);
va_end(args);
if (log_fp) {
va_start(args, format);
write_message(log_fp, LOG_LEVEL_INFO, t, format, args);
va_end(args);
}
if (!line)
return;
emit_log_line(stdout, line);
free(line);
}
void log_perror(const char* context) {
+4
View File
@@ -29,6 +29,10 @@ void log_perror(const char* context);
void set_log_level(LogLevel level);
void set_log_debug_flags(uint32_t flags);
uint32_t get_log_debug_flags(void);
/* True when a log_debug_message() call with the same flag would actually emit:
* the debug log level is enabled AND the flag is selected. Hot paths use this
* to skip expensive message formatting/escaping when the line is filtered. */
bool log_debug_enabled(LogDebugFlag flag);
void log_debug_message(LogDebugFlag flag, const char* message, ...);
void set_log_info_flags(uint32_t flags);
uint32_t get_log_info_flags(void);
Loaded 100 of 201 files, more files were not shown because too many files have changed in this diff. Show more