Commit Graph
402 Commits
Author SHA1 Message Date
TapTap ea4ab661b4 fix(parity): rsync 3.4.1 symlink and special-node semantics (#287, #288)
#287:
- --safe-links: keep safe in-tree links AS symlinks and skip unsafe
  (absolute or ".."-escaping) ones, mirroring rsync's unsafe_symlink().
  Skipped links are recorded as delete-protected so --delete does not
  remove their destination mirror (no silent data loss).
- --copy-unsafe-links: preserve safe links as symlinks and dereference
  only unsafe ones.
- --munge-links: receiver-side rewrite storing /rsyncd-munged/-prefixed
  targets (rsync parity), replacing the no-op #SYMLINK sender prefix.
- -l: store the target verbatim, including absolute and ".." targets
  (rsync -l parity); the old receiver containment silently dropped them.

#288:
- --specials: recreate unix-domain sockets via mknod(S_IFSOCK), which
  Linux permits unprivileged; keep EEXIST/EPERM skip behavior.
- --copy-devices: copy a device's content into a regular file when
  requested; skip unrequested non-regular entries like rsync's default.
2026-09-15 21:57:48 +02:00
TapTap 23552e823d feat(cli): rsync short-option clustering and inline/attached values (#285)
Implement rsync 3.4.1 client-CLI parity:
- cluster boolean shorts (-av, -aAX, -rlpt) and accept attached values
  (-B1048576, -essh, -Mfoo); add the -r, -b, -L and -B short aliases
  (-r is a faithful no-op since FastSync is always recursive)
- stop OPT_NOOP (-s/--secluded-args, -r/--recursive) from swallowing the
  next argv
- add inline --opt=value for every value-taking long option, including
  --exclude/--include/--exclude-from/--include-from/--log-file (#291)
- accept --port on the server CLI in addition to -p (#296)
- reject unknown flags naming the flag and stating it is unsupported

Unit tests cover clustering, attached/inline values, the OPT_NOOP
argument-consumption fix and rejected shorts.
2026-09-15 21:12:53 +02:00
TapTap 93c1fc3c1f test: create fault-injection destination root in seeding fixture
CI / lint (push) Successful in 1m29s
CI / lint (pull_request) Successful in 1m29s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / sanitizers (undefined) (push) Successful in 1m4s
CI / sanitizers (address) (push) Successful in 1m10s
CI / fuzz-build (push) Successful in 39s
CI / coverage (push) Successful in 58s
CI / build-and-test (pull_request) Successful in 1m55s
CI / valgrind (push) Successful in 3m23s
CI / build-and-test (push) Successful in 5m11s
The captured_config fixture assumed fault_dst already existed, relying on earlier tests in the same xdist worker creating it via _recover. Under --dist=load a worker can receive the capture test first, so the receiver rejected a missing destination root and the capture run failed. Create DEST_DIR in the autouse seeding fixture so test order/distribution cannot matter.
2026-09-15 20:02:00 +02:00
TapTap 4815b1b281 test: account for group/other-write sanitization in new-dest mode expectation
CI / lint (push) Successful in 1m28s
CI / lint (pull_request) Successful in 1m28s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / sanitizers (undefined) (push) Successful in 1m1s
CI / sanitizers (address) (push) Successful in 1m6s
CI / fuzz-build (push) Successful in 38s
CI / coverage (push) Successful in 1m3s
CI / build-and-test (pull_request) Successful in 1m56s
CI / build-and-test (push) Failing after 4m50s
CI / valgrind (push) Successful in 3m23s
2026-09-15 19:41:12 +02:00
TapTap 34970b961c feat: per-attribute preservation flags -p/-t/-o/-g with --no-* negations (protocol 2.22.0)
Split FastSync's single use_metadata bundle into four independent rsync-parity attributes: preserve_perms, preserve_times, preserve_owner, preserve_group. use_metadata is now a derived transport bit (config_derived_use_metadata).

CLI: real -p/--perms, -t/--times, -o/--owner, -g/--group plus --no-perms/--no-times/--no-owner/--no-group (short and long) and --no-preserve; -a is now rsync -rlptgoD; --preserve = -pt; -A implies -p; -X does not; --chmod implies -p; --usermap/--groupmap/--chown imply owner/group per side; --incremental/--delta still auto-preserve unless negated.

Receiver: per-attribute FileAttrPolicy gating for files, dirs (modes applied at end of transfer), symlinks and specials; rsync -E read-bit rule; new files get source_mode & ~umask sanitized (no group/other write); per-side identity resolution; deferred directory metadata; batch dir-metadata replay; daemon modules without 'client owner = yes' no longer refuse plain -a but force super off (no ownership) with a warning.

Wire: PROTOCOL_VERSION 2.21.0 -> 2.22.0 (four appended config bools, golden 653 / 95530566005420798). FileMetadata/chunk/batch framing unchanged. Docs/CHANGELOG/CMake updated to 2.22.0.
2026-09-15 19:32:02 +02:00
TapTap 09c384d7d0 Merge branch 'docs/readme-refresh' into dev
CI / lint (push) Successful in 1m25s
CI / lint (pull_request) Successful in 1m24s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / sanitizers (undefined) (push) Successful in 1m1s
CI / sanitizers (address) (push) Successful in 1m7s
CI / fuzz-build (push) Successful in 36s
CI / coverage (push) Successful in 56s
CI / build-and-test (pull_request) Successful in 1m54s
CI / valgrind (push) Successful in 3m18s
CI / build-and-test (push) Successful in 5m34s
# Conflicts:
#	README.md
2026-09-14 18:52:54 +02:00
TapTap 81ad313ee5 docs: refresh README against implementation and guard against drift
Bring README.md and RSYNC_COMPAT.md in line with the actual code/CLI and add
an automated guard so they cannot silently drift again.

Waves A-E:
- Correct stale compatibility claims: archive is `-rlptD` (owner/group are
  opt-in via identity flags, not implied), and symlinks, hard links, xattrs,
  ACLs and `--dirs` are implemented.
- Remove documented-but-nonexistent features: the six unread FASTSYNC_* env
  vars, and `--client-cn` (server-only) from the client table.
- Repair the corrupted "Implementation Details" section (broken list numbering
  and emphasis) and correct it against the source.
- Sync the client and server option tables with usage.c / server_cli.c, and
  document server-contacting `--dry-run` (protocol 2.21.0).
- Hygiene: `# FastSync` heading, real build commands, consistent binary names,
  runnable TLS examples, daemon module keys.

Also align the client `--help` / archive log wording and the RSYNC_COMPAT
archive rows with the opt-in ownership model, and add
tests/integration/test_readme_consistency.py (marked `ci`) asserting every
documented FASTSYNC_* var is read in src/ and every documented client/server
flag appears in the corresponding `--help`.
2026-09-14 18:48:42 +02:00
TapTap 7badac7f97 test(compression): free inherited Data in forked truncation test (valgrind) 2026-09-14 17:54:20 +02:00
TapTap e79d2b47b0 Merge branch 'fix/sec-server' into fix/sec-integration 2026-09-14 17:27:38 +02:00
TapTap 7c24a365cf Merge branch 'fix/sec-receiver' into fix/sec-integration 2026-09-14 17:27:38 +02:00
TapTap 5893de4a34 fix(receiver): close re-review findings — dry-run basis oracle, ACL capture, fsync reopen
Follow-up to a237043 addressing three security/correctness re-review findings.

(1) MEDIUM: a server-contacting --dry-run with --compare-dest/--copy-dest/
    --link-dest still read and hashed the basis file and compared it with the
    client-supplied digest, a 1-bit content oracle. basis_match_find() gains a
    hash_content parameter; the dry-run shortcut passes false and returns no
    match without touching basis bytes, so an otherwise-matching entry is
    reported as would-transfer. The real (non-dry-run) path is unchanged.

(2) LOW: xattr_capture_path() hardcoded preserve_acls=true, so the receiver's
    hard-link copy fallback re-applied system.posix_acl_* even when -A was not
    negotiated. The function now takes preserve_acls and members.* is
    unaffected; scanner and receiver callers thread the negotiated flag.

(3) INFO: the --fsync --link-dest temp reopen now uses O_NONBLOCK and treats
    a raced-in FIFO's ENXIO as a benign fsync-skip instead of blocking.

Tests: dry-run + basis unit test (asserts would-transfer, no content read) and
integration test; xattr capture ACL-filter test. Verified strict build, ASan,
clang-format, cppcheck, and the CI integration subset.
2026-09-14 17:19:27 +02:00
TapTap 825ba69753 fix(server): reject --allow-super with --stdio, fix module host-list append
Re-review findings on the C3/C4 hardening branch:

- --stdio is the SSH transport whose remote argv is composed by the client
  (including via --remote-option), so accepting --allow-super there let a
  client defeat the C3 secure default for a root receiver.  Reject it at CLI
  parse time (standalone TCP only) and force the process-global flag off for
  --stdio as defense in depth.  Correct the help text and README/RSYNC_COMPAT:
  the --stdio argv is client-composed, super stays off, and a forced command is
  needed if the default must hold.
- daemon_conf: the per-module 'hosts allow'/'hosts deny' call sites passed
  module_name and replace in the wrong order, so multiple lines replaced
  instead of appended and the empty-value error omitted the module name.  Pass
  (module->name, false) like the global keys; add a unit test for two
  per-module allow/deny lines appending.
- tls: read the client CN via ASN1_STRING_to_UTF8 so an exactly-required-length
  name is accepted and only actual over-length CNs are rejected.
2026-09-14 17:16:01 +02:00
TapTap 34abaadb9a fix: address low/informational sec-parser follow-ups
- client_cli: capture errno before output_escape() in
  read_patterns_from_file() so an over-long line is still reported as
  EFBIG instead of the (possibly malloc-clobbered) errno.
- file_list: guard string_list_add() capacity doubling against
  overflow (capacity > INT_MAX / 2), matching filter_rule_list_add();
  callers already surface the false as a memory-allocation error.
- compression: ZSTD_isError() is true for ZSTD_CONTENTSIZE_UNKNOWN,
  which made the 3x unknown-size fallback dead code.  Test the
  CONTENTSIZE_ERROR/UNKNOWN sentinels explicitly so unknown-size frames
  reach the estimate path (still bounded by the existing hard limit)
  while invalid frames are rejected.  Known-size frames and the 100 MB
  ceiling/overflow checks are unchanged.
- tests: add an unknown-content-size-frame decompression test.

Tests: ./build/tests and ./build-asan/tests all pass (42/42);
clang-format + cppcheck clean.
2026-09-14 17:11:02 +02:00
TapTap 10c4ffebdf fix(client): harden CLI args, log escaping, and local artifact opens
- parse_ull_arg() rejects a leading '-'/'+' (strtoull would silently wrap
  -1 to ULLONG_MAX) and --chunk-size/--delta-max enforce their upper bounds.
- Escape local untrusted paths before logging (client_send, scanner,
  --filter rule, pattern-file reads) with output_escape(..., 8-bit mode).
- Read --exclude-from/--include-from through the bounded line reader.
- Open --log-file with O_NOFOLLOW|O_CLOEXEC, mode 0600, via open+fdopen;
  create --write-batch with O_NOFOLLOW|O_CLOEXEC, mode 0600.
- Reject --dry-run together with --write-batch (dry-run must not write the
  batch file), alongside the existing --read-batch/--only-write-batch rules.

Tests: signed/oversized numeric rejection, over-long pattern file, dry-run +
write-batch unit and integration coverage.
2026-09-14 16:39:19 +02:00
TapTap 1a26bde2d4 fix(file_list): bound entry length and reject embedded NUL bytes
Read list files through utils_getdelim_bounded() so a single multi-gigabyte
line can no longer force unbounded allocation; over-long entries fail with a
clear error.  Also add the documented memchr() NUL-byte check (excluding the
NUL delimiter in NUL-separated mode).

Tests: an over-long entry is rejected with an 'exceeds' diagnostic.
2026-09-14 16:39:14 +02:00
TapTap 58a28334b8 fix(utils): bound glob matching and line reads
Replace the recursive glob matcher with an iterative O(pattern*string)
dynamic program.  The old recursion explored exponentially many paths for
overlapping '*'/'**' wildcards (e.g. '*a*a*...*b' against a long run of
'a'), a CPU DoS reachable from --exclude/--include patterns and
.rsync-filter.  A differential fuzz against the original matcher confirms
identical results.  Doc: has_path_traversal() is a lexical '..' check only.

Add utils_getdelim_bounded(): a getdelim-style reader that never allocates
beyond UTILS_MAX_LINE_LEN, used to cap untrusted list/filter line reads.

Tests: pathological glob completes quickly; bounded reader returns EFBIG on
an over-long record.
2026-09-14 16:39:10 +02:00
TapTap e48f19ee2b fix(compression): fail truncated zstd frames instead of spinning
data_decompress_limited() looped while ZSTD_decompressStream() returned a
positive hint.  A truncated frame keeps returning that hint with all input
consumed, so a malformed/truncated payload spun forever (CPU DoS).  Detect
input exhaustion with an incomplete frame and fail via the existing cleanup,
skipping the check when the output buffer merely needs to grow first.

Add a fork+alarm regression test that truncates a valid frame and asserts
decompression returns NULL promptly.
2026-09-14 16:39:05 +02:00
TapTap a2370433b2 fix(receiver): non-blocking receiver opens, inplace type gate, dry-run/B4/B5/B6
Address confirmed receiver security findings B1-B6:

B1 (HIGH): add O_NONBLOCK to the three receiver read-opens that opened an
existing destination/basis entry before the S_ISREG gate
(incremental_check_open_destination, basis_open_regular, hardlink_read_source)
so a client-planted FIFO can no longer block the receive thread forever while
the post-open type gate still rejects it.

B2 (HIGH/MED): --inplace now fstatat(AT_SYMLINK_NOFOLLOW)-probes the target and
refuses any existing non-regular entry, opens with O_NONBLOCK, and re-checks
S_ISREG on the opened fd.  This stops a FIFO from hanging the open and stops a
char/block device from being written directly (bypassing --write-devices).

B3 (MED): under --dry-run the incremental quick-skip no longer reads/hashes the
destination file for --checksum/--delta; it decides from metadata only and
reports would-transfer when the comparison is inconclusive, closing the
read-only-module content-hash oracle.

B4 (LOW): xattr_name_appliable() now gates the two system.posix_acl_* names on
preserve_acls (--acls), not the derived use_xattrs (--xattrs OR --acls).  The
receiver drops (never applies) ACL entries when -A was not negotiated while
keeping user.* working for -X.

B5 (INFO): receive_manifest_section() charges a per-entry overhead against
MAX_MANIFEST_BYTES and the aggregate entry count across all three sections is
capped at MAX_MANIFEST_ENTRIES.

B6 (MED): data_charge_session() reserves decompressed/chunk-copy bytes against
the owning ProtocolSession (MAX_CONNECTION_MEMORY) and records them on the Data
so data_destroy() releases them via the Data.owner path.  Applied to the
whole-file/append/delta decompression sites and chunk_deserialize() per-file
copies; a missing session owner degrades to the previous uncharged behavior.

Tests: FIFO destination/basis non-hang (with alarm), --inplace FIFO/device
refusal, dry-run no-read oracle test plus updated metadata-only dry-run tests,
ACL-without--acls drop, manifest total-entry cap, and chunk session charging.
2026-09-14 16:19:26 +02:00
TapTap 9da5a0a9ed fix(server): gate --force by --allow-delete and secure root super default
C2: --force is deletion authority (an incoming regular file may remove a
non-empty destination directory tree, and --delete-missing-args may
remove a non-empty directory mirror), but it was not masked by the
operator --allow-delete policy.  The handler now clears
config->force_delete unless --allow-delete was given, exactly like
--delete and --delete-missing-args.

C3: a standalone TCP / --stdio server running as root defaulted to
SUPER_MODE_AUTO, so an untrusted client --devices/--write-devices/
--super could make it create device nodes, write raw devices, or apply
client-chosen ownership.  A privileged standalone receiver now forces
SUPER_MODE_OFF unless the operator opts in with the new server-only
--allow-super flag.  Non-root receivers are unchanged, and the daemon
path keeps its per-module `client owner = yes` gate.  --allow-super is
rejected with --no-super or --daemon.

C6: tls_client_identity_allowed now rejects a CN whose reported length
reached the buffer bound, so a truncated over-long CN cannot be matched
by a required --client-cn prefix.

Tests: an integration regression proving --force cannot replace a
destination directory without --allow-delete; standalone-default tests
for --copy-as refusal and (root-only) skipped device creation; a CLI
unit test for the new flag.  The integration shared_server fixture opts
in with --allow-super so the existing root-only ownership/device/copy-as
tests continue to exercise the opted-in configuration.  README and
RSYNC_COMPAT document the flag and the force/delete gating.
2026-09-14 16:09:44 +02:00
TapTap 80c1ff321c fix(credentials): length-check before legacy-hex scan (C9)
secret_is_legacy_hex indexed s[0..63] without first checking the string
length, reading out of bounds for a shorter secret.  Require
strlen(s) == 64 before scanning, and add a unit test that short and
63-hex-digit secrets are rejected as ordinary malformed verifiers (never
misreported as legacy).
2026-09-14 16:09:25 +02:00
TapTap 551c187005 fix(tls): AEAD-only 1.2 suites, server preference, TOCTOU key load, IP SAN
C5: restrict the TLS 1.2 and below cipher list to ECDHE AEAD suites
(ECDHE+AESGCM:ECDHE+CHACHA20, minus NULL/eNULL/MD5/RC4/3DES) instead of
HIGH (which includes CBC), and set SSL_OP_CIPHER_SERVER_PREFERENCE so the
server's order decides the negotiated cipher.  Client and server share
create_ssl_ctx, so both are updated.

C7: load the private key through an O_RDONLY|O_NOFOLLOW|O_CLOEXEC fd,
fstat that fd and validate owner/mode (now also rejecting group/other
execute bits), then load from the fd via BIO_new_fd.  This removes the
stat-to-load TOCTOU race while keeping the exact-owner/0600 policy.

C8: verify an IP-literal client hostname against the certificate IP SAN
with X509_VERIFY_PARAM_set1_ip_asc instead of SSL_set1_host (a DNS
check), falling back to SSL_set1_host for real names.

Unit tests assert the server-preference option, the absence of CBC/RC4/
3DES suites, and that context creation still succeeds.
2026-09-14 16:09:21 +02:00
TapTap 0d6c1f784f fix(daemon-conf): reject empty hosts/auth allow-lists (C4)
A present hosts allow/hosts deny/auth users key with an empty or
separator-only value produced a zero-length list, silently meaning no
ACL / no auth and contradicting the strict-parse contract.

store_host_list and the auth users parser now track how many entries a
present key actually added and fail the load with a clear error when it
is zero, so a restrictive directive can never silently become open.
Unit tests cover empty, whitespace-only and comma-only values.
2026-09-14 16:09:16 +02:00
TapTap f75a69f96a fix(ssh): reject option-injection destinations (C1)
A remote destination's user@host token is passed to ssh in option
position, so a host beginning with '-' (e.g. -oProxyCommand=...) was
parsed by ssh as an option, allowing arbitrary command execution.

- config_parse_ssh_dest now validates the user@host prefix and returns
  -1 (with a clear logged error) for an empty host or a user/host that
  starts with '-'; config_parse_transport_dest propagates the failure.
- transport_ssh.c's parse_remote_dest applies the same validation as
  defense-in-depth, and ssh_build_client_argv inserts a '--'
  end-of-options marker before the destination token.
- Unit tests cover -oProxyCommand=... / -prefixed hosts / empty host
  rejection and the argv shape.
2026-09-14 16:08:56 +02:00
TapTap 5d39619a8a Merge branch 'feat/w9-dryrun' into fix/w9-integration 2026-09-13 13:27:32 +02:00
TapTap 6269ae54e5 test(dry-run): strengthen no-mutation coverage and refresh docs
Extend _snapshot_tree to record mode, inode, xattrs, directories and
special nodes, and add coverage proving a server-contacting --dry-run
leaves the destination structurally identical for --delay-updates,
--backup, symlinks, hardlinks, FIFOs, and daemon modules (including a
read-only module).  Add a regression test for the --read-batch --dry-run
refusal and for a missing/non-directory receive root failing a dry-run
exactly like a real run.

Fix stale version comments (2.20.0/633 -> 2.21.0/637) and RSYNC_COMPAT's
current --protocol value, and add a unit assertion that
--server-port/--port (and --server-host) set the dry-run routing bit.
2026-09-13 12:57:37 +02:00
TapTap 334fc5b3e8 fix(protocol): harden STATUS_ERROR_DETAIL receive path
Address review/security findings in the 2.21.0 error-detail feature:

- Keepalive drain no longer erases the terminal detail: capture/clear is
  skipped for STATUS_KEEPALIVE so the reason the peer just sent survives the
  owed keepalive replies.
- Replace the capture path with a dedicated protocol_receive_error_detail:
  the declared length is validated against MAX_ERROR_DETAIL_BYTES before any
  allocation, over-cap bodies are drained through a fixed scratch buffer (so
  the stream never desyncs), in-cap bodies read straight into the thread-local
  detail buffer, and session->max_alloc is never raised.  Lengths beyond
  MAX_STRING_SIZE are treated as a fatal framing error.
- The detail body now honors the caller's deadline (timed/keepalive paths) and
  polls the abort callback between drain chunks.
- Escape peer-controlled detail text with output_escape before logging it in
  client_send.c and config.c.
- Clear io_error_detail in io_set_fds so a new connection on the same thread
  cannot inherit a stale reason.
- Add unit tests for the keepalive-survival, over-cap drain, absurd-length
  fatal framing, and deadline-clamped body read cases.
2026-09-13 12:40:03 +02:00
TapTap 88aee6ce94 feat(protocol): add optional STATUS_ERROR_DETAIL rejection reason (2.21.0)
Today a server rejection sends a bare STATUS_ERROR and the reason only
reaches the server log, so the client cannot say why a transfer was
refused.  Add an optional, bounded server->client error-detail frame:

  - Status gains STATUS_ERROR_DETAIL appended LAST so existing wire
    values are unchanged.
  - send_error_detail(fd, msg) sends STATUS_ERROR_DETAIL followed by the
    existing length-prefixed string primitive, slicing over-long messages
    to MAX_ERROR_DETAIL_BYTES (4096).
  - receive_status() (and the timed/keepalive status readers) always
    consume the detail body and map the status back to STATUS_ERROR,
    capturing the text into a thread-local buffer exposed by
    protocol_last_error(); a bare STATUS_ERROR leaves it cleared.  Every
    existing call site keeps working and the stream cannot desync.
  - Upgrade the daemon module gate / config validation (config.c), the
    final transfer failure (server.c) and receiver-side path/node
    validation (file_receive.c) to send a concrete reason; surface it on
    the client in client_send.c/config.c.
  - Bump PROTOCOL_VERSION to 2.21.0 (CMake VERSION, CHANGELOG, docs) and
    update the pinned config wire golden hash / CLI-version tests.
  - Add tests/test_protocol_error.c covering mapping+capture, the
    over-long bound, bare-error clearing, and thread-locality.
2026-09-13 12:19:46 +02:00
TapTap 6f974eff19 feat(dry-run): server-contacting --dry-run (protocol 2.21.0)
--dry-run now handshakes with a remote/daemon receiver and reports what
WOULD transfer/skip based on receiver state, mutating nothing on either
side.

- Serialize Config.dry_run into the wire config frame and append
  STATUS_DRY_RUN_TRANSFER to the status enum (no renumbering); bump
  PROTOCOL_VERSION/CMake VERSION/CHANGELOG/golden wire to 2.21.0.
- Receiver: receive_incremental_check_ex runs the normal read-only
  decision and answers STATUS_OK (skip) or STATUS_DRY_RUN_TRANSFER
  (would transfer) with no basis materialization/append/delta/full
  transfer.  All mutation sites are guarded by !dry_run: file store,
  manifest deletes, --mkpath root creation, --delay-updates staging,
  publication, directory-time application, and outcome acks.
- Client: send_dry_run_remote connects, sends the config, checks each
  regular file and prints the would-transfer set + trailer; no file data
  or delete manifest is sent.  Plain local destinations keep the
  client-side manifest.
2026-09-13 11:56:05 +02:00
TapTap eb71b29d1c Merge branch 'fix/w8-authroot' into fix/w8-integration 2026-09-13 11:13:40 +02:00
TapTap 4d5befedfe Merge branch 'fix/w8-charge' into fix/w8-integration 2026-09-13 11:13:40 +02:00
TapTap f00844cf9a Merge branch 'fix/w8-daemonlim' into fix/w8-integration 2026-09-13 11:13:40 +02:00
TapTap 2ec17e821c fix(daemon): harden bounded per-source registry races
Stamp host_last_use before publishing a bucket key and treat an unstamped
(last_use == 0) bucket as live, so a just-claimed bucket can no longer be
stolen by a concurrent reclaimer.

After a successful eviction CAS, re-scan for the interned key and, when an
earlier bucket already holds it, zero the duplicate's active count and
return the canonical bucket, preventing orphaned per-host counts and cap
overshoot under full-table concurrency.

Add a message-carrying EXPECT_FAIL primitive and use it for the daemon-conf
buffer-overflow guard, and add a fork-based test that records auth failures
from forked children and asserts the parent observes the shared lockout.
2026-09-13 11:11:23 +02:00
TapTap 6d47d93fd7 fix(config): validate received counts before publishing them
The config_receive_{basis,skip,idmap}_count helpers wrote the
peer-controlled int through the Config member before range-checking it.
An over-cap basis_count therefore left config->basis_count huge while
config->basis_dirs was still NULL; config_receive()'s error path then
called config_delete(), whose basis loop dereferenced NULL and crashed
the daemon before authentication.

Read each count into a local, validate, and only then assign, leaving the
member untouched on failure.  config_delete() also guards the basis loop
with the array pointer as defense in depth.

Add a regression test that feeds over-cap basis/idmap/skip counts and
asserts rejection without crashing, plus a direct config_delete() check
on the partial (count set, array NULL) state.
2026-09-13 11:05:57 +02:00
TapTap 6ea966781f test(config): add receive-side golden oracle and sharpen fixture
Address low-severity review findings on the X-macro config refactor:

1. The golden test only hashed config_send_wire_block(), so a
   receive-side KIND that reads a different width/order could still
   round-trip symmetrically.  Add test_config_wire_golden_receive():
   capture the same hash-pinned 633-byte frame and feed it through
   config_receive(), asserting every field (config_wire_equal) plus the
   derived use_delta/use_xattrs bits and representative bounded kinds.
   Add test_config_wire_receive_bounds() for bounds the symmetric
   round-trip cannot reach: an out-of-range BOOL (hand-built frame),
   RAW_MAXALLOC zero, a malformed STR_MODULE, an over-cap
   INT_IDMAPCOUNT, and an out-of-range INT_IDENTITY chown_uid.

2. golden_config_populate() set long runs of booleans to all-1, so an
   adjacent swap within a run produced identical bytes.  Alternate the
   boolean values and make the fixture receiver-valid (chmod grammar
   "u=rwx,go=rx" is the same 11 bytes; delta_max_file_size inside the
   bound).  Re-pin the golden: len stays 633, hash is now
   9160991280011164139 (computed, not guessed).

3. Document in config.h and client_cli.c that the CLI option tables
   remain hand-maintained and are deliberately not generated from the
   wire-field X-macro (client-only fields, flag/alias/negation
   semantics).  No CLI-table rewrite.

PROTOCOL_VERSION stays "2.20.0"; src/shared/config.c is untouched and
the wire bytes are unchanged apart from the fixture's own new values.
2026-09-13 10:56:34 +02:00
TapTap 0a7f5faea6 test: replace strcat with a bounds-checked append in daemon-conf test 2026-09-13 10:51:01 +02:00
TapTap 25909110ac fix(daemon): exempt trusted loopback peers from per-host limits
Every client on loopback shares the 127.0.0.1 identity, so counting them
against 'max connections per host' or the default-on auth lockout lets one
local client deny service to all the others (and makes a shared-NAT/proxy
address a natural DoS vector for remote clients).  Use
utils_fd_peer_is_local (fail-closed) in the daemon gate to exempt a
provably local peer from the per-source cap and the auth lockout while
keeping the per-module and global caps.  Remote peers are unchanged.

Document the shared-NAT/proxy identity limitation and the loopback
exemption in README/RSYNC_COMPAT/CHANGELOG, update the integration test to
assert the exemption, and fix the README 'auth failure delay' cap (5000,
not 60000).
2026-09-13 10:50:58 +02:00
TapTap bd43448af2 fix(daemon): bound per-source table lifetime and recompute occupancy
The per-source host table only grew: once its fixed open-addressed table
filled, host_intern returned -1 and the per-host cap plus the shared auth
lockout silently failed open forever.  Add a bounded-lifetime eviction
policy: track a per-bucket last-use time and, when no empty bucket exists,
atomically repurpose the first bucket that has no active connection and
either has an expired lockout or has been idle, resetting its counters.
Warn (rate-limited) on the genuine fail-open path.

A child SIGKILLed mid-registration could also leak a module/host count
because the parent only decremented on a REGISTERED slot.  Make the slot
table the source of truth: after the SIGCHLD reap the parent recomputes
module_active[]/host_active[] from the surviving REGISTERED slots (atomics
only, async-signal-safe) so any leaked increment is erased.

Also clamp module_count to DAEMON_LIMITS_MAX_MODULES and use one helper
for the sizing/register host-tracking condition (a lockout threshold with
duration 0 is a no-op and must not intern hosts).
2026-09-13 10:50:53 +02:00
TapTap 18d1b84246 refactor(protocol): guard session release, clarify Data.owner contract
Add a NULL guard to protocol_release_memory_for_session so it no-ops like
the sibling session setters.  Correct the Data.owner doc comment, which
implied a non-zero protocol_charge always has an owner; document that
owner may be NULL for uncharged/ownerless Data, that any such charge
falls back to the bound session, and that a charged Data must not outlive
its owning session.  Note the lifetime contract on the release API too.

Extend tests/test_protocol.c to cover destroying a charged Data with no
session bound (the other half of the original bug) and to assert that
data_create/data_create_reserve start with owner == NULL and
protocol_charge == 0.
2026-09-13 10:42:45 +02:00
TapTap 4e918a1b69 test(config): pin wire bytes and round-trip every field
test_config_wire_golden() serializes a fully-populated Config through
config_send_wire_block() and pins the exact frame to len=633 and FNV-1a
hash 6163263374908258816, captured from the pre-X-macro implementation.
Any field reorder, resize or codec change fails the test.

test_config_wire_roundtrip_all_fields() serializes/deserializes a defaults
Config and a fully-populated Config over a socketpair and compares every
serialized field.  The comparison is itself generated from
CONFIG_WIRE_FIELDS (one CONFIG_CMP_<KIND> per table entry), so a new table
entry automatically extends coverage; it cannot fall out of sync.  It
normalizes the receiver's NULL/"" canonicalization, the max_alloc server
clamp and the derived use_delta/use_xattrs bits.
2026-09-13 10:28:31 +02:00
TapTap 4c17122b00 feat(daemon): enforce per-module/per-host caps and shared auth lockout
Wire the shared registry into the accept loop (parent claims a slot before
fork, blocks SIGCHLD across fork+pid publication, and reclaims the dead
child's slot from the SIGCHLD handler so per-module/per-source counts are
released even on SIGKILL). The connection child records the selected module
and normalized peer IP once the config frame names them: an over-cap module
or source is refused at the config gate with an audit log, and a source
that exceeded the auth-failure threshold is refused before a SCRAM
challenge (the counter is shared across children and cleared on success).
The existing global cap and host ACLs are untouched.
2026-09-13 10:24:05 +02:00
TapTap 0abaa62193 feat(daemon): parse per-host cap and auth lockout config keys
Add global keys `max connections per host` (default 0 = unlimited),
`auth lockout threshold` (default 10, 0 disables) and
`auth lockout duration` (default 300 s, 0 disables). Module
`max connections` now accepts 0 as unlimited. Bound the number of
[module] sections (DAEMON_CONF_MAX_MODULES) so the shared registry's
per-module counter array stays fixed-size; absent keys keep their
defaults so old configs still load.
2026-09-13 10:24:01 +02:00
TapTap 5334397b81 feat(daemon): add shared cross-process connection registry
The daemon forks one child per accepted connection, so per-module and
per-source accounting must live in state shared across the children. Add a
fixed-size registry carved from an anonymous shared mapping
(mmap(MAP_SHARED|MAP_ANONYMOUS)) created before the accept loop: a slot
lifecycle (FREE/CLAIMED/REGISTERED) with parent claim/reclaim and a
lock-free, open-addressed per-source table for the per-host occupancy and
the shared auth-failure counter. C11 atomics only; no pthread locks across
fork.

Unit tests cover slot exhaustion, the module/host caps, pid reclaim and
fork-shared visibility.
2026-09-13 10:23:58 +02:00
TapTap 3260a39ab4 refactor(shared): single owner for authorized_root state 2026-09-13 10:06:04 +02:00
TapTap 5d3c43305e fix(protocol): release Data charge to its owning session
Data charged against a ProtocolSession kept only the charge amount, so
data_destroy released it from whatever session was thread-locally bound
at destroy time. Destroying a received Data on another thread, after the
session was unbound, or while a different session was bound leaked the
originating session's budget and underflowed the other's.

Add Data.owner, set it whenever protocol_receive_data_limited charges a
session, and have data_destroy release against that owner directly via
the newly-exported protocol_release_memory_for_session. Uncharged Data
(owner NULL) keeps the previous bound-session fallback.

Add a unit test proving a Data acquired on session A is released to A
even when unrelated session B is bound at destroy time.
2026-09-13 10:05:38 +02:00
TapTap 3499baf80b build: explicit CMake targets; move receiver pipeline out of shared 2026-09-13 07:20:28 +02:00
TapTap c2df0347ef fix(client,protocol): EINTR-safe sends, armed abort, keepalive drain grace, TLS WANT_WRITE 2026-09-13 06:59:02 +02:00
TapTap dd44537b44 Merge branch 'fix/w6-tests' into fix/w6-integration 2026-09-13 06:25:05 +02:00
TapTap 2854a9d149 test: fuzz manifest/protocol/xattr, hardlink unit, fault injection 2026-09-13 06:24:46 +02:00
TapTap 1fa2fbd266 feat(client): --port alias, --threads=N, graceful abort, keepalive 2026-09-13 06:22:28 +02:00
TapTap dcc78c14c5 docs,fuzz: fix ownership/alloc comments; fuzz chunk metadata path 2026-09-13 05:48:27 +02:00