A present hosts allow/hosts deny/auth users key with an empty or
separator-only value produced a zero-length list, silently meaning no
ACL / no auth and contradicting the strict-parse contract.
store_host_list and the auth users parser now track how many entries a
present key actually added and fail the load with a clear error when it
is zero, so a restrictive directive can never silently become open.
Unit tests cover empty, whitespace-only and comma-only values.
A remote destination's user@host token is passed to ssh in option
position, so a host beginning with '-' (e.g. -oProxyCommand=...) was
parsed by ssh as an option, allowing arbitrary command execution.
- config_parse_ssh_dest now validates the user@host prefix and returns
-1 (with a clear logged error) for an empty host or a user/host that
starts with '-'; config_parse_transport_dest propagates the failure.
- transport_ssh.c's parse_remote_dest applies the same validation as
defense-in-depth, and ssh_build_client_argv inserts a '--'
end-of-options marker before the destination token.
- Unit tests cover -oProxyCommand=... / -prefixed hosts / empty host
rejection and the argv shape.
Extend _snapshot_tree to record mode, inode, xattrs, directories and
special nodes, and add coverage proving a server-contacting --dry-run
leaves the destination structurally identical for --delay-updates,
--backup, symlinks, hardlinks, FIFOs, and daemon modules (including a
read-only module). Add a regression test for the --read-batch --dry-run
refusal and for a missing/non-directory receive root failing a dry-run
exactly like a real run.
Fix stale version comments (2.20.0/633 -> 2.21.0/637) and RSYNC_COMPAT's
current --protocol value, and add a unit assertion that
--server-port/--port (and --server-host) set the dry-run routing bit.
Address review/security findings in the 2.21.0 error-detail feature:
- Keepalive drain no longer erases the terminal detail: capture/clear is
skipped for STATUS_KEEPALIVE so the reason the peer just sent survives the
owed keepalive replies.
- Replace the capture path with a dedicated protocol_receive_error_detail:
the declared length is validated against MAX_ERROR_DETAIL_BYTES before any
allocation, over-cap bodies are drained through a fixed scratch buffer (so
the stream never desyncs), in-cap bodies read straight into the thread-local
detail buffer, and session->max_alloc is never raised. Lengths beyond
MAX_STRING_SIZE are treated as a fatal framing error.
- The detail body now honors the caller's deadline (timed/keepalive paths) and
polls the abort callback between drain chunks.
- Escape peer-controlled detail text with output_escape before logging it in
client_send.c and config.c.
- Clear io_error_detail in io_set_fds so a new connection on the same thread
cannot inherit a stale reason.
- Add unit tests for the keepalive-survival, over-cap drain, absurd-length
fatal framing, and deadline-clamped body read cases.
Today a server rejection sends a bare STATUS_ERROR and the reason only
reaches the server log, so the client cannot say why a transfer was
refused. Add an optional, bounded server->client error-detail frame:
- Status gains STATUS_ERROR_DETAIL appended LAST so existing wire
values are unchanged.
- send_error_detail(fd, msg) sends STATUS_ERROR_DETAIL followed by the
existing length-prefixed string primitive, slicing over-long messages
to MAX_ERROR_DETAIL_BYTES (4096).
- receive_status() (and the timed/keepalive status readers) always
consume the detail body and map the status back to STATUS_ERROR,
capturing the text into a thread-local buffer exposed by
protocol_last_error(); a bare STATUS_ERROR leaves it cleared. Every
existing call site keeps working and the stream cannot desync.
- Upgrade the daemon module gate / config validation (config.c), the
final transfer failure (server.c) and receiver-side path/node
validation (file_receive.c) to send a concrete reason; surface it on
the client in client_send.c/config.c.
- Bump PROTOCOL_VERSION to 2.21.0 (CMake VERSION, CHANGELOG, docs) and
update the pinned config wire golden hash / CLI-version tests.
- Add tests/test_protocol_error.c covering mapping+capture, the
over-long bound, bare-error clearing, and thread-locality.
--dry-run now handshakes with a remote/daemon receiver and reports what
WOULD transfer/skip based on receiver state, mutating nothing on either
side.
- Serialize Config.dry_run into the wire config frame and append
STATUS_DRY_RUN_TRANSFER to the status enum (no renumbering); bump
PROTOCOL_VERSION/CMake VERSION/CHANGELOG/golden wire to 2.21.0.
- Receiver: receive_incremental_check_ex runs the normal read-only
decision and answers STATUS_OK (skip) or STATUS_DRY_RUN_TRANSFER
(would transfer) with no basis materialization/append/delta/full
transfer. All mutation sites are guarded by !dry_run: file store,
manifest deletes, --mkpath root creation, --delay-updates staging,
publication, directory-time application, and outcome acks.
- Client: send_dry_run_remote connects, sends the config, checks each
regular file and prints the would-transfer set + trailer; no file data
or delete manifest is sent. Plain local destinations keep the
client-side manifest.
Stamp host_last_use before publishing a bucket key and treat an unstamped
(last_use == 0) bucket as live, so a just-claimed bucket can no longer be
stolen by a concurrent reclaimer.
After a successful eviction CAS, re-scan for the interned key and, when an
earlier bucket already holds it, zero the duplicate's active count and
return the canonical bucket, preventing orphaned per-host counts and cap
overshoot under full-table concurrency.
Add a message-carrying EXPECT_FAIL primitive and use it for the daemon-conf
buffer-overflow guard, and add a fork-based test that records auth failures
from forked children and asserts the parent observes the shared lockout.
The config_receive_{basis,skip,idmap}_count helpers wrote the
peer-controlled int through the Config member before range-checking it.
An over-cap basis_count therefore left config->basis_count huge while
config->basis_dirs was still NULL; config_receive()'s error path then
called config_delete(), whose basis loop dereferenced NULL and crashed
the daemon before authentication.
Read each count into a local, validate, and only then assign, leaving the
member untouched on failure. config_delete() also guards the basis loop
with the array pointer as defense in depth.
Add a regression test that feeds over-cap basis/idmap/skip counts and
asserts rejection without crashing, plus a direct config_delete() check
on the partial (count set, array NULL) state.
Address low-severity review findings on the X-macro config refactor:
1. The golden test only hashed config_send_wire_block(), so a
receive-side KIND that reads a different width/order could still
round-trip symmetrically. Add test_config_wire_golden_receive():
capture the same hash-pinned 633-byte frame and feed it through
config_receive(), asserting every field (config_wire_equal) plus the
derived use_delta/use_xattrs bits and representative bounded kinds.
Add test_config_wire_receive_bounds() for bounds the symmetric
round-trip cannot reach: an out-of-range BOOL (hand-built frame),
RAW_MAXALLOC zero, a malformed STR_MODULE, an over-cap
INT_IDMAPCOUNT, and an out-of-range INT_IDENTITY chown_uid.
2. golden_config_populate() set long runs of booleans to all-1, so an
adjacent swap within a run produced identical bytes. Alternate the
boolean values and make the fixture receiver-valid (chmod grammar
"u=rwx,go=rx" is the same 11 bytes; delta_max_file_size inside the
bound). Re-pin the golden: len stays 633, hash is now
9160991280011164139 (computed, not guessed).
3. Document in config.h and client_cli.c that the CLI option tables
remain hand-maintained and are deliberately not generated from the
wire-field X-macro (client-only fields, flag/alias/negation
semantics). No CLI-table rewrite.
PROTOCOL_VERSION stays "2.20.0"; src/shared/config.c is untouched and
the wire bytes are unchanged apart from the fixture's own new values.
Every client on loopback shares the 127.0.0.1 identity, so counting them
against 'max connections per host' or the default-on auth lockout lets one
local client deny service to all the others (and makes a shared-NAT/proxy
address a natural DoS vector for remote clients). Use
utils_fd_peer_is_local (fail-closed) in the daemon gate to exempt a
provably local peer from the per-source cap and the auth lockout while
keeping the per-module and global caps. Remote peers are unchanged.
Document the shared-NAT/proxy identity limitation and the loopback
exemption in README/RSYNC_COMPAT/CHANGELOG, update the integration test to
assert the exemption, and fix the README 'auth failure delay' cap (5000,
not 60000).
The per-source host table only grew: once its fixed open-addressed table
filled, host_intern returned -1 and the per-host cap plus the shared auth
lockout silently failed open forever. Add a bounded-lifetime eviction
policy: track a per-bucket last-use time and, when no empty bucket exists,
atomically repurpose the first bucket that has no active connection and
either has an expired lockout or has been idle, resetting its counters.
Warn (rate-limited) on the genuine fail-open path.
A child SIGKILLed mid-registration could also leak a module/host count
because the parent only decremented on a REGISTERED slot. Make the slot
table the source of truth: after the SIGCHLD reap the parent recomputes
module_active[]/host_active[] from the surviving REGISTERED slots (atomics
only, async-signal-safe) so any leaked increment is erased.
Also clamp module_count to DAEMON_LIMITS_MAX_MODULES and use one helper
for the sizing/register host-tracking condition (a lockout threshold with
duration 0 is a no-op and must not intern hosts).
Add a NULL guard to protocol_release_memory_for_session so it no-ops like
the sibling session setters. Correct the Data.owner doc comment, which
implied a non-zero protocol_charge always has an owner; document that
owner may be NULL for uncharged/ownerless Data, that any such charge
falls back to the bound session, and that a charged Data must not outlive
its owning session. Note the lifetime contract on the release API too.
Extend tests/test_protocol.c to cover destroying a charged Data with no
session bound (the other half of the original bug) and to assert that
data_create/data_create_reserve start with owner == NULL and
protocol_charge == 0.
test_config_wire_golden() serializes a fully-populated Config through
config_send_wire_block() and pins the exact frame to len=633 and FNV-1a
hash 6163263374908258816, captured from the pre-X-macro implementation.
Any field reorder, resize or codec change fails the test.
test_config_wire_roundtrip_all_fields() serializes/deserializes a defaults
Config and a fully-populated Config over a socketpair and compares every
serialized field. The comparison is itself generated from
CONFIG_WIRE_FIELDS (one CONFIG_CMP_<KIND> per table entry), so a new table
entry automatically extends coverage; it cannot fall out of sync. It
normalizes the receiver's NULL/"" canonicalization, the max_alloc server
clamp and the derived use_delta/use_xattrs bits.
Wire the shared registry into the accept loop (parent claims a slot before
fork, blocks SIGCHLD across fork+pid publication, and reclaims the dead
child's slot from the SIGCHLD handler so per-module/per-source counts are
released even on SIGKILL). The connection child records the selected module
and normalized peer IP once the config frame names them: an over-cap module
or source is refused at the config gate with an audit log, and a source
that exceeded the auth-failure threshold is refused before a SCRAM
challenge (the counter is shared across children and cleared on success).
The existing global cap and host ACLs are untouched.
Add global keys `max connections per host` (default 0 = unlimited),
`auth lockout threshold` (default 10, 0 disables) and
`auth lockout duration` (default 300 s, 0 disables). Module
`max connections` now accepts 0 as unlimited. Bound the number of
[module] sections (DAEMON_CONF_MAX_MODULES) so the shared registry's
per-module counter array stays fixed-size; absent keys keep their
defaults so old configs still load.
The daemon forks one child per accepted connection, so per-module and
per-source accounting must live in state shared across the children. Add a
fixed-size registry carved from an anonymous shared mapping
(mmap(MAP_SHARED|MAP_ANONYMOUS)) created before the accept loop: a slot
lifecycle (FREE/CLAIMED/REGISTERED) with parent claim/reclaim and a
lock-free, open-addressed per-source table for the per-host occupancy and
the shared auth-failure counter. C11 atomics only; no pthread locks across
fork.
Unit tests cover slot exhaustion, the module/host caps, pid reclaim and
fork-shared visibility.
Data charged against a ProtocolSession kept only the charge amount, so
data_destroy released it from whatever session was thread-locally bound
at destroy time. Destroying a received Data on another thread, after the
session was unbound, or while a different session was bound leaked the
originating session's budget and underflowed the other's.
Add Data.owner, set it whenever protocol_receive_data_limited charges a
session, and have data_destroy release against that owner directly via
the newly-exported protocol_release_memory_for_session. Uncharged Data
(owner NULL) keeps the previous bound-session fallback.
Add a unit test proving a Data acquired on session A is released to A
even when unrelated session B is bound at destroy time.