Reclassify every rsync-compatibility row as parity / caveat / divergent
(replacing the misleading 143-OK / 0-divergence summary), and document the
protocol 2.23.0 behavior:
- Split the conflated `-M, --preserve` row: `-M` is `--remote-option`,
`--preserve` is the FastSync `-p`+`-t` alias.
- Fix `MAX_CONNECTION_MEMORY` (256 MiB, not 1 GB), `--rsync-path`
(client-only, never crosses the wire), and the `-p` mode behavior
(strict rsync parity; no masking).
- `--specials` now recreates sockets, so `-D` is real parity; fake-super
records the resolved owner and replays mode/time (never real-chowns).
- Document short options/clustering, checksum/compression choices, seed
randomization, timeout/max-alloc defaults, temp-dir confinement + EXDEV,
identity/map parity, verbatim symlinks, delete scoping, `--max-delete`
partial + exit 25, `--chmod`, output caveats, and server `--port`.
- Bump version refs to 2.23.0 and add the 2.23.0 CHANGELOG entry.
Docs-only; no source changes.
- --chmod no longer implies --preserve-perms; repeated --chmod options
accumulate, and D/F/X selectors plus s/t special bits are supported with
rsync's exact parse_chmod/tweak_mode semantics.
- Stop masking group/other write and setuid/setgid/sticky: -p copies the
source mode exactly, no-p new entries use source&~umask, directories keep
setgid/sticky, and special nodes follow the same rules.
- Apply ownership before mode on the fd path so a chown cannot clear the
setuid/setgid bits -p just restored (rsync order).
- Update unit and integration tests, including differential checks against
rsync 3.4.1.
Address review findings on feat/rsync-parity:
- confine --temp-dir below the receive root (reject absolute/.. like
backup-dir/partial-dir); keep EXDEV non-atomic fallback
- floor server session I/O deadlines at SERVER_IO_TIMEOUT_SEC (60s) and
install it on the socket layer at startup (slow-loris)
- charge each --delete-missing-args directory removal once and clamp the
extras-walk remaining budget so it can never underflow past --max-delete
- normalize --compress-choice=auto to zstd client-side and accept it on
receive so auto transfers no longer fail
- map received --max-alloc=0 to MAX_SERVER_ALLOC (receive path only)
- zero File.dest_state; include log-file-format in report_dest_info;
add STATUS_DELETE_LIMIT name; recognize --skip-compress as a
separate-value option; OOM-guard send_list_only root entry; drop the
dead -M= branch; record the bare relative protected prefix for -R
size-prunes in both scanners; refresh delete-manifest comment
- pin the rsync tarball sha256 and bump integrator image to v11
Tests: temp-dir rejection/relative/cross-device, server timeout floor,
delete-missing dir budget regression, compress-choice=auto e2e,
max-alloc=0 receive mapping, dest_state, report_dest_info modes,
skip-compress dash value, -M short forms, -R root size-prune mirror
protection (rsync 3.4.1 confirmed).
A non-empty --delete-missing-args directory removed under --force/--delete
now has its contents deleted entry-by-entry through the budgeted walker, so
every deleted file/dir counts toward --max-delete exactly like rsync (a
capped run leaves the remaining entries and exits 25).
- Scope the --delete extras walk to directories synchronized by the
transfer: add a synchronized-directory section to the delete manifest
(protocol 2.23.0) so --files-from subsets no longer delete untransmitted
paths outside listed directory subtrees (data-loss fix).
- Separate --max-size/--min-size prune protection from --delete-excluded so
size-pruned source mirrors survive (rsync parity).
- Unlink extraneous destination symlinks instead of skipping them.
- Make --max-delete partial (delete up to N, skip the rest) and exit 25;
accept negative values as unlimited.
- Draw --delete-missing-args deletions from the shared --max-delete budget.
- Honor --force during --delay-updates publication.
Add unit and integration regression tests; update the pinned config wire
golden and version strings for the 2.23.0 manifest/status additions.
#289 checksum/compression:
- -c/--checksum now implies the incremental content quick-check (without
implying -t), so an unchanged file is skipped like rsync.
- --checksum-choice/--cc accepts xxh64/xxhash, xxh3, xxh128, md5 and auto;
md4/sha1/none and the two-name form are rejected by name.
- --compress-choice/--zc rejects lz4/zlib/zlibx by name (zstd/none/auto kept).
- --checksum-seed=0 is randomized per transfer and sent on the wire.
- --skip-compress uses rsync 3.4.1's default suffix list; slash separators and
dot-less suffixes are accepted.
- add --no-whole-file.
#295 timeouts/alloc/temp-dir:
- --timeout default 0 (disabled), --contimeout default 60; 0 disables both,
plus --no-timeout/--no-contimeout.
- --max-alloc=0 means no allocation limit (was rejected).
- --temp-dir accepts any dir, requires it to exist, and falls back to a
non-atomic copy on EXDEV instead of aborting.
#296 connectivity/daemon:
- -M/--remote-option is rejected for daemon/TCP destinations (SSH-only).
- --trust-sender clarified as receiver-local; server-path tests added.
- --stop-at accepts rsync's full date form (y-m-dTh:m etc.).
Adds unit and integration coverage; no wire-field change, PROTOCOL_VERSION stays
2.22.0.
- #286: --numeric-ids is a mapping modifier only; it no longer activates
chown by itself (identity_active_enabled/owner/group predicates), and
--fake-super stores the resolved mapping instead of real-chowning.
- #286: apply owner/group to directories via the deferred directory
metadata path; capture+transmit+apply directory xattrs/ACLs (-aX/-aA),
including default ACLs, in STATUS_MKDIR/STATUS_DIR_TIMES.
- #294: --usermap/--groupmap support inclusive ranges, '*', empty FROM
(unnamed ids), and receiver-side TO name resolution; --chown mixing with
a same-side map is rejected like rsync.
- Protocol 2.22.0 -> 2.23.0 (map wire entry gains from_hi + to_name;
dir frames gain a bounded xattr block).
#287:
- --safe-links: keep safe in-tree links AS symlinks and skip unsafe
(absolute or ".."-escaping) ones, mirroring rsync's unsafe_symlink().
Skipped links are recorded as delete-protected so --delete does not
remove their destination mirror (no silent data loss).
- --copy-unsafe-links: preserve safe links as symlinks and dereference
only unsafe ones.
- --munge-links: receiver-side rewrite storing /rsyncd-munged/-prefixed
targets (rsync parity), replacing the no-op #SYMLINK sender prefix.
- -l: store the target verbatim, including absolute and ".." targets
(rsync -l parity); the old receiver containment silently dropped them.
#288:
- --specials: recreate unix-domain sockets via mknod(S_IFSOCK), which
Linux permits unprivileged; keep EEXIST/EPERM skip behavior.
- --copy-devices: copy a device's content into a regular file when
requested; skip unrequested non-regular entries like rsync's default.
#291:
- Compile --exclude/--include/--exclude-from/--include-from into the SAME
ordered rule list as --filter/-f (first match wins), so the common
`--include='*.txt' --exclude='*'` idiom and include-alone semantics match
rsync. The legacy per-kind scanner arrays are no longer applied.
- -x/--one-file-system emits the cross-device mount-point directory entry
(empty) instead of dropping it, in both the sequential and parallel scanners.
- Stop passing the legacy arrays to the scanner; document -f is --filter.
#292:
- New src/shared/format.c/.h: rsync "big_num" (comma-grouped integers) and
decimal -h human sizes, %M/%t timestamp, and the STATUS_DEST_INFO codec.
- Receiver answers each STATUS_CHECK with a pre-transfer destination snapshot
(new report_dest_info wire field + STATUS_DEST_INFO, PROTOCOL_VERSION
2.23.0) so the sender can render true itemize columns.
- Itemize now emits rsync-correct update/type chars and c/s/t/p/o/g columns
for files, dirs, symlinks and hard links, comparing size/time/perms/owner/
group against the reported destination.
- --out-format gains %i %n %f %l %b %M %t %o %p %B %U %G %L; %f is the
relative display path, %M the YYYY/MM/DD-HH:MM:SS form, %b the literal
bytes sent.
- --list-only prints transfer-relative names, directory entries and ls-style
grouped sizes.
- --stats prints rsync's multi-line block on stdout; -h uses decimal units.
Tests: unit tests for the filter ordering, format primitives, itemize
columns; integration + differential tests against real rsync 3.4.1 for
itemize/out-format/list-only/selection and -x. Golden wire len/hash and
protocol version strings updated for 2.23.0.
Implement rsync 3.4.1 client-CLI parity:
- cluster boolean shorts (-av, -aAX, -rlpt) and accept attached values
(-B1048576, -essh, -Mfoo); add the -r, -b, -L and -B short aliases
(-r is a faithful no-op since FastSync is always recursive)
- stop OPT_NOOP (-s/--secluded-args, -r/--recursive) from swallowing the
next argv
- add inline --opt=value for every value-taking long option, including
--exclude/--include/--exclude-from/--include-from/--log-file (#291)
- accept --port on the server CLI in addition to -p (#296)
- reject unknown flags naming the flag and stating it is unsupported
Unit tests cover clustering, attached/inline values, the OPT_NOOP
argument-consumption fix and rejected shorts.
Add POSIX ACL/xattr tooling (acl, attr), zstd/lz4/xxhash dev libs and build rsync 3.4.1 from source so drop-in parity tests can run inside CI. Bump all workflow/agent image references v10 -> v11.
The captured_config fixture assumed fault_dst already existed, relying on earlier tests in the same xdist worker creating it via _recover. Under --dist=load a worker can receive the capture test first, so the receiver rejected a missing destination root and the capture run failed. Create DEST_DIR in the autouse seeding fixture so test order/distribution cannot matter.
Split FastSync's single use_metadata bundle into four independent rsync-parity attributes: preserve_perms, preserve_times, preserve_owner, preserve_group. use_metadata is now a derived transport bit (config_derived_use_metadata).
CLI: real -p/--perms, -t/--times, -o/--owner, -g/--group plus --no-perms/--no-times/--no-owner/--no-group (short and long) and --no-preserve; -a is now rsync -rlptgoD; --preserve = -pt; -A implies -p; -X does not; --chmod implies -p; --usermap/--groupmap/--chown imply owner/group per side; --incremental/--delta still auto-preserve unless negated.
Receiver: per-attribute FileAttrPolicy gating for files, dirs (modes applied at end of transfer), symlinks and specials; rsync -E read-bit rule; new files get source_mode & ~umask sanitized (no group/other write); per-side identity resolution; deferred directory metadata; batch dir-metadata replay; daemon modules without 'client owner = yes' no longer refuse plain -a but force super off (no ownership) with a warning.
Wire: PROTOCOL_VERSION 2.21.0 -> 2.22.0 (four appended config bools, golden 653 / 95530566005420798). FileMetadata/chunk/batch framing unchanged. Docs/CHANGELOG/CMake updated to 2.22.0.
Bring README.md and RSYNC_COMPAT.md in line with the actual code/CLI and add
an automated guard so they cannot silently drift again.
Waves A-E:
- Correct stale compatibility claims: archive is `-rlptD` (owner/group are
opt-in via identity flags, not implied), and symlinks, hard links, xattrs,
ACLs and `--dirs` are implemented.
- Remove documented-but-nonexistent features: the six unread FASTSYNC_* env
vars, and `--client-cn` (server-only) from the client table.
- Repair the corrupted "Implementation Details" section (broken list numbering
and emphasis) and correct it against the source.
- Sync the client and server option tables with usage.c / server_cli.c, and
document server-contacting `--dry-run` (protocol 2.21.0).
- Hygiene: `# FastSync` heading, real build commands, consistent binary names,
runnable TLS examples, daemon module keys.
Also align the client `--help` / archive log wording and the RSYNC_COMPAT
archive rows with the opt-in ownership model, and add
tests/integration/test_readme_consistency.py (marked `ci`) asserting every
documented FASTSYNC_* var is read in src/ and every documented client/server
flag appears in the corresponding `--help`.
- generate_bench_data now writes exactly (1-random_ratio)*target bytes of
genuinely compressible repeated content instead of only the small fixed
STRUCTURED_FILES set; measured composition is reported and --dry-run prints
it for scaling checks
- verify each transfer against the source (paths/sizes/byte compare) before
recording timing; add --no-verify; failed runs are counted as invalid
- correct p50/p95 with linear-interpolation percentile (was int(len*0.95))
- tc/netem: run tc directly as root, else sudo; clear error when tc/iproute2
is missing or qdisc setup fails; netem_reset is always safe
- build into dedicated build-bench/ via --build-dir (Release), never reconfigure
the user's build/
- parse --configs with shlex.split
- add MB/s throughput column and throughput_mbps JSON field
- add --warm incremental mode: untimed full seed then measure add/change deltas
Follow-up to a237043 addressing three security/correctness re-review findings.
(1) MEDIUM: a server-contacting --dry-run with --compare-dest/--copy-dest/
--link-dest still read and hashed the basis file and compared it with the
client-supplied digest, a 1-bit content oracle. basis_match_find() gains a
hash_content parameter; the dry-run shortcut passes false and returns no
match without touching basis bytes, so an otherwise-matching entry is
reported as would-transfer. The real (non-dry-run) path is unchanged.
(2) LOW: xattr_capture_path() hardcoded preserve_acls=true, so the receiver's
hard-link copy fallback re-applied system.posix_acl_* even when -A was not
negotiated. The function now takes preserve_acls and members.* is
unaffected; scanner and receiver callers thread the negotiated flag.
(3) INFO: the --fsync --link-dest temp reopen now uses O_NONBLOCK and treats
a raced-in FIFO's ENXIO as a benign fsync-skip instead of blocking.
Tests: dry-run + basis unit test (asserts would-transfer, no content read) and
integration test; xattr capture ACL-filter test. Verified strict build, ASan,
clang-format, cppcheck, and the CI integration subset.
Re-review findings on the C3/C4 hardening branch:
- --stdio is the SSH transport whose remote argv is composed by the client
(including via --remote-option), so accepting --allow-super there let a
client defeat the C3 secure default for a root receiver. Reject it at CLI
parse time (standalone TCP only) and force the process-global flag off for
--stdio as defense in depth. Correct the help text and README/RSYNC_COMPAT:
the --stdio argv is client-composed, super stays off, and a forced command is
needed if the default must hold.
- daemon_conf: the per-module 'hosts allow'/'hosts deny' call sites passed
module_name and replace in the wrong order, so multiple lines replaced
instead of appended and the empty-value error omitted the module name. Pass
(module->name, false) like the global keys; add a unit test for two
per-module allow/deny lines appending.
- tls: read the client CN via ASN1_STRING_to_UTF8 so an exactly-required-length
name is accepted and only actual over-length CNs are rejected.
- client_cli: capture errno before output_escape() in
read_patterns_from_file() so an over-long line is still reported as
EFBIG instead of the (possibly malloc-clobbered) errno.
- file_list: guard string_list_add() capacity doubling against
overflow (capacity > INT_MAX / 2), matching filter_rule_list_add();
callers already surface the false as a memory-allocation error.
- compression: ZSTD_isError() is true for ZSTD_CONTENTSIZE_UNKNOWN,
which made the 3x unknown-size fallback dead code. Test the
CONTENTSIZE_ERROR/UNKNOWN sentinels explicitly so unknown-size frames
reach the estimate path (still bounded by the existing hard limit)
while invalid frames are rejected. Known-size frames and the 100 MB
ceiling/overflow checks are unchanged.
- tests: add an unknown-content-size-frame decompression test.
Tests: ./build/tests and ./build-asan/tests all pass (42/42);
clang-format + cppcheck clean.
- parse_ull_arg() rejects a leading '-'/'+' (strtoull would silently wrap
-1 to ULLONG_MAX) and --chunk-size/--delta-max enforce their upper bounds.
- Escape local untrusted paths before logging (client_send, scanner,
--filter rule, pattern-file reads) with output_escape(..., 8-bit mode).
- Read --exclude-from/--include-from through the bounded line reader.
- Open --log-file with O_NOFOLLOW|O_CLOEXEC, mode 0600, via open+fdopen;
create --write-batch with O_NOFOLLOW|O_CLOEXEC, mode 0600.
- Reject --dry-run together with --write-batch (dry-run must not write the
batch file), alongside the existing --read-batch/--only-write-batch rules.
Tests: signed/oversized numeric rejection, over-long pattern file, dry-run +
write-batch unit and integration coverage.
protocol_receive_n_data_until() aborted on a signal-interrupted plaintext
read (and on SSL_ERROR_SYSCALL with errno==EINTR); retry both, matching the
send path and protocol_read_status_until(). Also clamp each SSL_write() to
INT_MAX so a >INT_MAX size_t request can never truncate into a partial write.
Read per-directory filter files through the bounded reader, guard the rule
list's capacity doubling against INT_MAX/2 overflow, and escape the local
directory path before logging a read failure.
Read list files through utils_getdelim_bounded() so a single multi-gigabyte
line can no longer force unbounded allocation; over-long entries fail with a
clear error. Also add the documented memchr() NUL-byte check (excluding the
NUL delimiter in NUL-separated mode).
Tests: an over-long entry is rejected with an 'exceeds' diagnostic.
Replace the recursive glob matcher with an iterative O(pattern*string)
dynamic program. The old recursion explored exponentially many paths for
overlapping '*'/'**' wildcards (e.g. '*a*a*...*b' against a long run of
'a'), a CPU DoS reachable from --exclude/--include patterns and
.rsync-filter. A differential fuzz against the original matcher confirms
identical results. Doc: has_path_traversal() is a lexical '..' check only.
Add utils_getdelim_bounded(): a getdelim-style reader that never allocates
beyond UTILS_MAX_LINE_LEN, used to cap untrusted list/filter line reads.
Tests: pathological glob completes quickly; bounded reader returns EFBIG on
an over-long record.
data_decompress_limited() looped while ZSTD_decompressStream() returned a
positive hint. A truncated frame keeps returning that hint with all input
consumed, so a malformed/truncated payload spun forever (CPU DoS). Detect
input exhaustion with an incomplete frame and fail via the existing cleanup,
skipping the check when the output buffer merely needs to grow first.
Add a fork+alarm regression test that truncates a valid frame and asserts
decompression returns NULL promptly.
Address confirmed receiver security findings B1-B6:
B1 (HIGH): add O_NONBLOCK to the three receiver read-opens that opened an
existing destination/basis entry before the S_ISREG gate
(incremental_check_open_destination, basis_open_regular, hardlink_read_source)
so a client-planted FIFO can no longer block the receive thread forever while
the post-open type gate still rejects it.
B2 (HIGH/MED): --inplace now fstatat(AT_SYMLINK_NOFOLLOW)-probes the target and
refuses any existing non-regular entry, opens with O_NONBLOCK, and re-checks
S_ISREG on the opened fd. This stops a FIFO from hanging the open and stops a
char/block device from being written directly (bypassing --write-devices).
B3 (MED): under --dry-run the incremental quick-skip no longer reads/hashes the
destination file for --checksum/--delta; it decides from metadata only and
reports would-transfer when the comparison is inconclusive, closing the
read-only-module content-hash oracle.
B4 (LOW): xattr_name_appliable() now gates the two system.posix_acl_* names on
preserve_acls (--acls), not the derived use_xattrs (--xattrs OR --acls). The
receiver drops (never applies) ACL entries when -A was not negotiated while
keeping user.* working for -X.
B5 (INFO): receive_manifest_section() charges a per-entry overhead against
MAX_MANIFEST_BYTES and the aggregate entry count across all three sections is
capped at MAX_MANIFEST_ENTRIES.
B6 (MED): data_charge_session() reserves decompressed/chunk-copy bytes against
the owning ProtocolSession (MAX_CONNECTION_MEMORY) and records them on the Data
so data_destroy() releases them via the Data.owner path. Applied to the
whole-file/append/delta decompression sites and chunk_deserialize() per-file
copies; a missing session owner degrades to the previous uncharged behavior.
Tests: FIFO destination/basis non-hang (with alarm), --inplace FIFO/device
refusal, dry-run no-read oracle test plus updated metadata-only dry-run tests,
ACL-without--acls drop, manifest total-entry cap, and chunk session charging.
C2: --force is deletion authority (an incoming regular file may remove a
non-empty destination directory tree, and --delete-missing-args may
remove a non-empty directory mirror), but it was not masked by the
operator --allow-delete policy. The handler now clears
config->force_delete unless --allow-delete was given, exactly like
--delete and --delete-missing-args.
C3: a standalone TCP / --stdio server running as root defaulted to
SUPER_MODE_AUTO, so an untrusted client --devices/--write-devices/
--super could make it create device nodes, write raw devices, or apply
client-chosen ownership. A privileged standalone receiver now forces
SUPER_MODE_OFF unless the operator opts in with the new server-only
--allow-super flag. Non-root receivers are unchanged, and the daemon
path keeps its per-module `client owner = yes` gate. --allow-super is
rejected with --no-super or --daemon.
C6: tls_client_identity_allowed now rejects a CN whose reported length
reached the buffer bound, so a truncated over-long CN cannot be matched
by a required --client-cn prefix.
Tests: an integration regression proving --force cannot replace a
destination directory without --allow-delete; standalone-default tests
for --copy-as refusal and (root-only) skipped device creation; a CLI
unit test for the new flag. The integration shared_server fixture opts
in with --allow-super so the existing root-only ownership/device/copy-as
tests continue to exercise the opted-in configuration. README and
RSYNC_COMPAT document the flag and the force/delete gating.
secret_is_legacy_hex indexed s[0..63] without first checking the string
length, reading out of bounds for a shorter secret. Require
strlen(s) == 64 before scanning, and add a unit test that short and
63-hex-digit secrets are rejected as ordinary malformed verifiers (never
misreported as legacy).
C5: restrict the TLS 1.2 and below cipher list to ECDHE AEAD suites
(ECDHE+AESGCM:ECDHE+CHACHA20, minus NULL/eNULL/MD5/RC4/3DES) instead of
HIGH (which includes CBC), and set SSL_OP_CIPHER_SERVER_PREFERENCE so the
server's order decides the negotiated cipher. Client and server share
create_ssl_ctx, so both are updated.
C7: load the private key through an O_RDONLY|O_NOFOLLOW|O_CLOEXEC fd,
fstat that fd and validate owner/mode (now also rejecting group/other
execute bits), then load from the fd via BIO_new_fd. This removes the
stat-to-load TOCTOU race while keeping the exact-owner/0600 policy.
C8: verify an IP-literal client hostname against the certificate IP SAN
with X509_VERIFY_PARAM_set1_ip_asc instead of SSL_set1_host (a DNS
check), falling back to SSL_set1_host for real names.
Unit tests assert the server-preference option, the absence of CBC/RC4/
3DES suites, and that context creation still succeeds.
A present hosts allow/hosts deny/auth users key with an empty or
separator-only value produced a zero-length list, silently meaning no
ACL / no auth and contradicting the strict-parse contract.
store_host_list and the auth users parser now track how many entries a
present key actually added and fail the load with a clear error when it
is zero, so a restrictive directive can never silently become open.
Unit tests cover empty, whitespace-only and comma-only values.
A remote destination's user@host token is passed to ssh in option
position, so a host beginning with '-' (e.g. -oProxyCommand=...) was
parsed by ssh as an option, allowing arbitrary command execution.
- config_parse_ssh_dest now validates the user@host prefix and returns
-1 (with a clear logged error) for an empty host or a user/host that
starts with '-'; config_parse_transport_dest propagates the failure.
- transport_ssh.c's parse_remote_dest applies the same validation as
defense-in-depth, and ssh_build_client_argv inserts a '--'
end-of-options marker before the destination token.
- Unit tests cover -oProxyCommand=... / -prefixed hosts / empty host
rejection and the argv shape.
Extend _snapshot_tree to record mode, inode, xattrs, directories and
special nodes, and add coverage proving a server-contacting --dry-run
leaves the destination structurally identical for --delay-updates,
--backup, symlinks, hardlinks, FIFOs, and daemon modules (including a
read-only module). Add a regression test for the --read-batch --dry-run
refusal and for a missing/non-directory receive root failing a dry-run
exactly like a real run.
Fix stale version comments (2.20.0/633 -> 2.21.0/637) and RSYNC_COMPAT's
current --protocol value, and add a unit assertion that
--server-port/--port (and --server-host) set the dry-run routing bit.
receiver_save_file appended to context->outcomes for --remove-source-files
without the !dry_run guard the multithreaded pipeline has, so a hostile
dry-run client could grow outcomes unbounded (raw, uncharged realloc) and
force a per-frame ack. Guard the append on !dry_run.
--dry-run --read-batch=FILE still wrote to the destination because
batch_read_apply -> file_save_to_disk_full bypassed the per-caller
!dry_run guards. Guard file_save_to_disk_full and manifest_delete_all
directly (return SKIPPED/no-op) so every save/delete path is mutation-free
in dry-run, and keep the per-caller guards. Reject --dry-run combined with
--read-batch/--only-write-batch at CLI validation with a clear error (a
dry-run of a local batch apply is not meaningful).
--dry-run --server-host=H (or TLS / source-bind --address) silently ran the
client-side manifest even though a real run contacts the server. Add a
client-only, never-serialized server_host_set bit (alongside the existing
server_port_set) and extend dry_run_targets_server so every explicit remote
target contacts the receiver.
Also make incremental_check return the dry-run code (4) only when the
session actually requested dry-run; a stray STATUS_DRY_RUN_TRANSFER from a
hostile/buggy peer is now a logged protocol error (STATUS_ERROR) instead of
falling through to send file data and desync. Both normal send_single_file
callers handle rc == 4 explicitly as an abort.
A wire dry_run bit must not relax the destination-root precondition:
previously handlers skipped ensure_receive_root entirely in dry-run, so a
client could dry-run against a nonexistent/regular-file root a real session
rejects. Split the existence check (receive_root_exists, never creates)
from the create path and apply the precondition unconditionally: dry-run
runs the existence/directory check only, reports the failure, and creates
nothing (no --mkpath).
Also allow a `read only = yes` daemon module for a dry-run session (a
server-contacting dry-run IS a read-only wire operation) while still
refusing it for real writes, and update the stale read-only comments.
Address review/security findings in the 2.21.0 error-detail feature:
- Keepalive drain no longer erases the terminal detail: capture/clear is
skipped for STATUS_KEEPALIVE so the reason the peer just sent survives the
owed keepalive replies.
- Replace the capture path with a dedicated protocol_receive_error_detail:
the declared length is validated against MAX_ERROR_DETAIL_BYTES before any
allocation, over-cap bodies are drained through a fixed scratch buffer (so
the stream never desyncs), in-cap bodies read straight into the thread-local
detail buffer, and session->max_alloc is never raised. Lengths beyond
MAX_STRING_SIZE are treated as a fatal framing error.
- The detail body now honors the caller's deadline (timed/keepalive paths) and
polls the abort callback between drain chunks.
- Escape peer-controlled detail text with output_escape before logging it in
client_send.c and config.c.
- Clear io_error_detail in io_set_fds so a new connection on the same thread
cannot inherit a stale reason.
- Add unit tests for the keepalive-survival, over-cap drain, absurd-length
fatal framing, and deadline-clamped body read cases.
Today a server rejection sends a bare STATUS_ERROR and the reason only
reaches the server log, so the client cannot say why a transfer was
refused. Add an optional, bounded server->client error-detail frame:
- Status gains STATUS_ERROR_DETAIL appended LAST so existing wire
values are unchanged.
- send_error_detail(fd, msg) sends STATUS_ERROR_DETAIL followed by the
existing length-prefixed string primitive, slicing over-long messages
to MAX_ERROR_DETAIL_BYTES (4096).
- receive_status() (and the timed/keepalive status readers) always
consume the detail body and map the status back to STATUS_ERROR,
capturing the text into a thread-local buffer exposed by
protocol_last_error(); a bare STATUS_ERROR leaves it cleared. Every
existing call site keeps working and the stream cannot desync.
- Upgrade the daemon module gate / config validation (config.c), the
final transfer failure (server.c) and receiver-side path/node
validation (file_receive.c) to send a concrete reason; surface it on
the client in client_send.c/config.c.
- Bump PROTOCOL_VERSION to 2.21.0 (CMake VERSION, CHANGELOG, docs) and
update the pinned config wire golden hash / CLI-version tests.
- Add tests/test_protocol_error.c covering mapping+capture, the
over-long bound, bare-error clearing, and thread-locality.
The per-file STATUS_CHECK fast path was a single 534-line function that
was hard to review. Extract it into small static helpers called in order
by a short linear orchestrator:
- incremental_check_receive_request (receive/validate request frame)
- incremental_check_open_destination (secure open + stat)
- incremental_check_quick_skip (metadata/content skip decision)
- incremental_check_try_basis (compare/copy/link-dest)
- incremental_check_try_append_resume(--append tail resume)
- incremental_check_try_delta (block delta)
- incremental_check_try_fuzzy (--fuzzy basis)
- incremental_check_receive_full (STATUS_NEXT + whole file)
Pure refactor: the ordered sequence of wire operations
(send_status/send_n_data/receive_n_data/receive_status/receive_wire_str)
is byte-for-byte identical to the original, and every resource cleanup
is preserved (a unified idempotent cleanup replaces the duplicated
per-path close/free blocks). No functional changes.
--dry-run now handshakes with a remote/daemon receiver and reports what
WOULD transfer/skip based on receiver state, mutating nothing on either
side.
- Serialize Config.dry_run into the wire config frame and append
STATUS_DRY_RUN_TRANSFER to the status enum (no renumbering); bump
PROTOCOL_VERSION/CMake VERSION/CHANGELOG/golden wire to 2.21.0.
- Receiver: receive_incremental_check_ex runs the normal read-only
decision and answers STATUS_OK (skip) or STATUS_DRY_RUN_TRANSFER
(would transfer) with no basis materialization/append/delta/full
transfer. All mutation sites are guarded by !dry_run: file store,
manifest deletes, --mkpath root creation, --delay-updates staging,
publication, directory-time application, and outcome acks.
- Client: send_dry_run_remote connects, sends the config, checks each
regular file and prints the would-transfer set + trailer; no file data
or delete manifest is sent. Plain local destinations keep the
client-side manifest.
Stamp host_last_use before publishing a bucket key and treat an unstamped
(last_use == 0) bucket as live, so a just-claimed bucket can no longer be
stolen by a concurrent reclaimer.
After a successful eviction CAS, re-scan for the interned key and, when an
earlier bucket already holds it, zero the duplicate's active count and
return the canonical bucket, preventing orphaned per-host counts and cap
overshoot under full-table concurrency.
Add a message-carrying EXPECT_FAIL primitive and use it for the daemon-conf
buffer-overflow guard, and add a fork-based test that records auth failures
from forked children and asserts the parent observes the shared lockout.
The config_receive_{basis,skip,idmap}_count helpers wrote the
peer-controlled int through the Config member before range-checking it.
An over-cap basis_count therefore left config->basis_count huge while
config->basis_dirs was still NULL; config_receive()'s error path then
called config_delete(), whose basis loop dereferenced NULL and crashed
the daemon before authentication.
Read each count into a local, validate, and only then assign, leaving the
member untouched on failure. config_delete() also guards the basis loop
with the array pointer as defense in depth.
Add a regression test that feeds over-cap basis/idmap/skip counts and
asserts rejection without crashing, plus a direct config_delete() check
on the partial (count set, array NULL) state.
Address low-severity review findings on the X-macro config refactor:
1. The golden test only hashed config_send_wire_block(), so a
receive-side KIND that reads a different width/order could still
round-trip symmetrically. Add test_config_wire_golden_receive():
capture the same hash-pinned 633-byte frame and feed it through
config_receive(), asserting every field (config_wire_equal) plus the
derived use_delta/use_xattrs bits and representative bounded kinds.
Add test_config_wire_receive_bounds() for bounds the symmetric
round-trip cannot reach: an out-of-range BOOL (hand-built frame),
RAW_MAXALLOC zero, a malformed STR_MODULE, an over-cap
INT_IDMAPCOUNT, and an out-of-range INT_IDENTITY chown_uid.
2. golden_config_populate() set long runs of booleans to all-1, so an
adjacent swap within a run produced identical bytes. Alternate the
boolean values and make the fixture receiver-valid (chmod grammar
"u=rwx,go=rx" is the same 11 bytes; delta_max_file_size inside the
bound). Re-pin the golden: len stays 633, hash is now
9160991280011164139 (computed, not guessed).
3. Document in config.h and client_cli.c that the CLI option tables
remain hand-maintained and are deliberately not generated from the
wire-field X-macro (client-only fields, flag/alias/negation
semantics). No CLI-table rewrite.
PROTOCOL_VERSION stays "2.20.0"; src/shared/config.c is untouched and
the wire bytes are unchanged apart from the fixture's own new values.
Every client on loopback shares the 127.0.0.1 identity, so counting them
against 'max connections per host' or the default-on auth lockout lets one
local client deny service to all the others (and makes a shared-NAT/proxy
address a natural DoS vector for remote clients). Use
utils_fd_peer_is_local (fail-closed) in the daemon gate to exempt a
provably local peer from the per-source cap and the auth lockout while
keeping the per-module and global caps. Remote peers are unchanged.
Document the shared-NAT/proxy identity limitation and the loopback
exemption in README/RSYNC_COMPAT/CHANGELOG, update the integration test to
assert the exemption, and fix the README 'auth failure delay' cap (5000,
not 60000).
The per-source host table only grew: once its fixed open-addressed table
filled, host_intern returned -1 and the per-host cap plus the shared auth
lockout silently failed open forever. Add a bounded-lifetime eviction
policy: track a per-bucket last-use time and, when no empty bucket exists,
atomically repurpose the first bucket that has no active connection and
either has an expired lockout or has been idle, resetting its counters.
Warn (rate-limited) on the genuine fail-open path.
A child SIGKILLed mid-registration could also leak a module/host count
because the parent only decremented on a REGISTERED slot. Make the slot
table the source of truth: after the SIGCHLD reap the parent recomputes
module_active[]/host_active[] from the surviving REGISTERED slots (atomics
only, async-signal-safe) so any leaked increment is erased.
Also clamp module_count to DAEMON_LIMITS_MAX_MODULES and use one helper
for the sizing/register host-tracking condition (a lockout threshold with
duration 0 is a no-op and must not intern hosts).
Add a NULL guard to protocol_release_memory_for_session so it no-ops like
the sibling session setters. Correct the Data.owner doc comment, which
implied a non-zero protocol_charge always has an owner; document that
owner may be NULL for uncharged/ownerless Data, that any such charge
falls back to the bound session, and that a charged Data must not outlive
its owning session. Note the lifetime contract on the release API too.
Extend tests/test_protocol.c to cover destroying a charged Data with no
session bound (the other half of the original bug) and to assert that
data_create/data_create_reserve start with owner == NULL and
protocol_charge == 0.
- pr-review: replace invalid 'tea pr comment' with 'tea comment' (the
former is not a tea subcommand)
- integrator: drop stray '-M' from client examples (-M is now
--remote-option and requires an argument), use the canonical pytest
integration command, and bump the CI image tag to v10
- test-writer: build fuzz targets via -DENABLE_FUZZ=ON instead of
hand-rolled -fsanitize flags; fix the fuzz binary path
- cmake-expert: document -DSANITIZER=undefined, which is now live in
CMakeLists.txt
- README: add --allow-unauthenticated to the plain-TCP server example,
use --preserve for metadata (not -M), and use the canonical
integration command
- AGENTS.md: use the canonical integration command
Document on utils_get_authorized_root_path() that the returned pointer is
borrowed and invalidated by the next authorized-root setter, that the fd
and path are not read atomically (non-reentrant), and that the fd remains
caller-owned. Add a matching single-threaded/set-before-threads note at
the accessor definitions in utils.c.
In server.c, drop the redundant utils_set_authorized_root(-1, NULL) after
a failed utils_set_authorized_root(): the setter already fail-closes the
state on allocation failure. The following close(root_fd) is unchanged.
The agent and skill definitions had drifted badly from the current
codebase and tooling, repeating the same class of bug as the benchmark
tool (references to nonexistent scripts and invented flags):
- Replace the removed `python3 test.py` with the real integration
command (`python3 -m pytest tests/integration/ -n 4 --dist=load
-m "not setpriv"`) across agents and skills.
- Fix `feature-scout`'s fabricated CLI flag list (--host, --server-mode,
--use-* etc.) using the authoritative src/client/usage.c flags.
- Fix `perf-analyst` benchmark flags (-m -c -> -j -z) and point at
benchmark/bench.py instead of stale numbers.
- Correct `code-explainer` (no getopt_long; --sendfile not -f) and
version drift in the release skill (1.1.0 -> 2.20.0).
- Replace GitHub/`gh` workflows with Gitea/`tea` (PRs target dev; issues
via tea; branch strategy updated in all agents).
- Use the built-in `-DSANITIZER=address|thread` CMake option instead of
hand-rolled -fsanitize flags.
- Add `-p 8080 --allow-unauthenticated` to plain-TCP server examples.
- Merge the redundant security-screener into security-auditor; drop the
duplicate (16 agents remain).
Repo hygiene: gitignore `root/` and `test_partial_install_tmp/`, remove
the empty leftover trees, delete the tracked scratch scripts tmux.sh and
to_one_file.py, and note the compile_commands.json symlink in README.
test_config_wire_golden() serializes a fully-populated Config through
config_send_wire_block() and pins the exact frame to len=633 and FNV-1a
hash 6163263374908258816, captured from the pre-X-macro implementation.
Any field reorder, resize or codec change fails the test.
test_config_wire_roundtrip_all_fields() serializes/deserializes a defaults
Config and a fully-populated Config over a socketpair and compares every
serialized field. The comparison is itself generated from
CONFIG_WIRE_FIELDS (one CONFIG_CMP_<KIND> per table entry), so a new table
entry automatically extends coverage; it cannot fall out of sync. It
normalizes the receiver's NULL/"" canonicalization, the max_alloc server
clamp and the derived use_delta/use_xattrs bits.
Every Config field that crosses the wire was declared in up to six places
(struct member, config_set_defaults, send_*, receive_*, and the two CLI
option tables) and could drift silently. Add CONFIG_WIRE_FIELDS in
config.h: one ordered per-segment table where each serialized field is
declared once with its C type, default and wire codec (KIND).
config.h now expands the table to declare the struct members;
config_set_defaults() expands it to assign the defaults; and
config_send_wire_block()/config_receive() expand the per-segment lists to
emit/consume the frame. The per-segment function names, call order and
segment boundaries are preserved exactly.
Fields with genuinely custom logic keep dedicated helpers but are still
declared once in the table: the protocol-version handshake (HEADER), daemon
SCRAM auth (STR_REDACTED_AUTH), the daemon module name (STR_MODULE), the
repeated count+array blocks (BLOCK_SKIP_SUFFIXES/BLOCK_BASIS/BLOCK_IDMAP),
--copy-as presence/ids (COPY_AS_*), and the derived --delta / use_xattrs
bits (DERIVED_DELTA, BOOL_XATTR_DERIVE). The version field remains a
special header (validated before any other field is parsed) and is sent by
config_send_wire_block() explicitly.
No public field is renamed and PROTOCOL_VERSION stays "2.20.0". Because
the struct declaration order is no longer the wire order, the wire order is
now enforced solely by the table and by a byte-exact golden test
(follow-up commit). Add config_send_wire_block() so that test can hash the
frame body without the STATUS_OK handshake.
Wire the shared registry into the accept loop (parent claims a slot before
fork, blocks SIGCHLD across fork+pid publication, and reclaims the dead
child's slot from the SIGCHLD handler so per-module/per-source counts are
released even on SIGKILL). The connection child records the selected module
and normalized peer IP once the config frame names them: an over-cap module
or source is refused at the config gate with an audit log, and a source
that exceeded the auth-failure threshold is refused before a SCRAM
challenge (the counter is shared across children and cleared on success).
The existing global cap and host ACLs are untouched.
Add global keys `max connections per host` (default 0 = unlimited),
`auth lockout threshold` (default 10, 0 disables) and
`auth lockout duration` (default 300 s, 0 disables). Module
`max connections` now accepts 0 as unlimited. Bound the number of
[module] sections (DAEMON_CONF_MAX_MODULES) so the shared registry's
per-module counter array stays fixed-size; absent keys keep their
defaults so old configs still load.
The daemon forks one child per accepted connection, so per-module and
per-source accounting must live in state shared across the children. Add a
fixed-size registry carved from an anonymous shared mapping
(mmap(MAP_SHARED|MAP_ANONYMOUS)) created before the accept loop: a slot
lifecycle (FREE/CLAIMED/REGISTERED) with parent claim/reclaim and a
lock-free, open-addressed per-source table for the per-host occupancy and
the shared auth-failure counter. C11 atomics only; no pthread locks across
fork.
Unit tests cover slot exhaustion, the module/host caps, pid reclaim and
fork-shared visibility.
The benchmark tool used stale rsync-style spellings that map to
different FastSync options, so it never enabled the features it
claimed to measure:
-c -> --checksum (not compression)
-m -> --prune-empty-dirs (not multithreading)
-s -> --secluded-args, a no-op (not chunk serialization)
-f -> --filter, needs an argument (not sendfile)
Replace them with the real flags (-z, -j, --chunk-serialization,
--sendfile), force CMAKE_BUILD_TYPE=Release, route informational
output to stderr so --output json emits valid JSON, surface
client/rsync failures instead of silently dropping them, and widen
the results table for the longer config names. Update the benchmark
skill to match (correct flags, server invocation, and replace the
nonexistent test.py --full with benchmark/bench.py).
Data charged against a ProtocolSession kept only the charge amount, so
data_destroy released it from whatever session was thread-locally bound
at destroy time. Destroying a received Data on another thread, after the
session was unbound, or while a different session was bound leaked the
originating session's budget and underflowed the other's.
Add Data.owner, set it whenever protocol_receive_data_limited charges a
session, and have data_destroy release against that owner directly via
the newly-exported protocol_release_memory_for_session. Uncharged Data
(owner NULL) keeps the previous bound-session fallback.
Add a unit test proving a Data acquired on session A is released to A
even when unrelated session B is bound at destroy time.